Cryptocurrency anti-money laundering method based on semi-supervised learning

By adopting the SSL-AML method based on semi-supervised learning in cryptocurrency money laundering detection, using meta-path modeling heterogeneous graphs and learnable data augmentation technology, the scarcity, class imbalance and heterogeneity of labeled data is solved, and efficient cryptocurrency money laundering detection is achieved.

CN120047147APending Publication Date: 2025-05-27BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411881945.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing cryptocurrency money laundering detection methods have high dependence on labeled data and poor generalization capabilities, and face scarce labeled data, serious imbalance and heterogeneity problems, resulting in poor detection results.

Method used

A cryptocurrency anti-money laundering method based on semi-supervised learning is proposed. A semi-supervised learning framework is constructed through meta-path modeling heterogeneous graphs and two stages of pre-training and consistency training. This method utilizes similarity-based mechanisms to dynamically generate high-quality pseudo-labels, and enhances the robustness and expression of the optimization model through learning data, alleviating high imbalance in the dataset.

Benefits of technology

It effectively solves the problems of scarcity, class imbalance and heterogeneity of labeled data in cryptocurrency money laundering detection, improves the robustness and expression ability of the model, achieves a higher detection accuracy and recall rate, which is better than the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047147A_ABST
    Figure CN120047147A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised learning-based cryptocurrency anti-money laundering method, which realizes effective detection of cryptocurrency money laundering behaviors. According to the method, encrypted currency transaction data is constructed into a heterogeneous graph network based on meta-path representation, and a semi-supervised learning framework combining pre-training and consistency training is adopted to detect the transaction data. In the pre-training stage, a similarity-based mechanism is utilized to dynamically generate a selection threshold value of a high-quality pseudo tag, so that selection of the pseudo tag is more flexible and effective; learnable data enhancement is introduced in the consistency training stage, the model is optimized from the two aspects of consistency and diversity, class-unbalanced transaction data is processed through data homogeneity distribution, and detection of cryptocurrency laundering behaviors is achieved. Experimental results show that the method can significantly improve various performance indexes of cryptocurrency money laundering detection, realizes accurate detection of cryptocurrency money laundering activities, and effectively guarantees security of cryptocurrency transactions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of cryptocurrency security, and specifically relates to an anti-money laundering method for cryptocurrencies based on semi-supervised learning. Background Art

[0002] In the digital age, the rise of cryptocurrencies has brought changes to global financial transactions. They offer decentralized and anonymous ways of payment and value transfer. However, due to their technical characteristics, cryptocurrencies are often used in illegal activities such as money laundering. According to a report released by Chainalysis, despite market fluctuations, illegal activities related to cryptocurrencies, especially money laundering activities, are still on the rise.

[0003] Money laundering activities take advantage of the anonymity and global reach of cryptocurrencies to convert illegally obtained funds into legal assets, thus providing criminals with a means to evade legal sanctions. Such behavior not only disrupts the financial order but also fuels more serious criminal activities such as cybercrime, drug trafficking, and terrorist financing. Of particular concern is the continuous development of cryptocurrency money laundering techniques.

[0004] From traditional mixing services to complex decentralized finance protocols, criminals are becoming increasingly sophisticated, using technological advancements to hide the sources and destinations of illegal funds. For example, the rise of dark web markets, the internationalization of illegal cryptocurrency services, and the emergence of new underground money laundering services have all increased the difficulty of tracking and combating money laundering activities. Against this background, detecting cryptocurrency money laundering behavior has become particularly important, and there is an urgent need to propose an effective detection method to help law enforcement agencies promptly identify and combat money laundering activities and prevent the recycling of illegal funds.

[0005] Traditional methods for checking cryptocurrency money laundering using supervised learning or unsupervised learning have problems such as high dependence on labeled data and poor generalization ability.

[0006] In addition, due to the following characteristics of the money laundering dataset, money laundering remains a complex and difficult task.

[0007] · Scarce labeled data. Among the 203,769 nodes in the dataset, approximately 77% of the nodes are unlabeled. If the dataset is divided by the time step and the unlabeled ratio for each time step is calculated, this ratio can reach as high as 34.59. The scarcity of labeled data severely limits the supervised learning algorithms in AML.

[0008] ·Severe class imbalance. Even with a limited amount of labeled data, there is still a severe class imbalance. Specifically, in the Elliptic dataset, only 2% of the nodes are labeled as illegal, while 21% of the nodes are labeled as legal. If the ratio for each time step is calculated, it can reach 355. Such a high class imbalance poses a huge challenge to AML.

[0009] ·Heterogeneity. Elliptic++ is a variant of the original Elliptic dataset. When Elliptic++(Transaction) and Elliptic++(Actor) are used in combination, it contains different types of nodes. Therefore, the transaction dataset becomes a heterogeneous graph. Due to the complexity of different nodes and edges, the algorithm needs to aggregate different node features while maintaining structural integrity, and needs to distinguish the different semantics and importance of different edge types.

[0010] To solve the above problems existing in the current cryptocurrency money laundering detection, the present invention proposes a cryptocurrency anti-money laundering method based on semi-supervised learning, SSL-AML (Semi-Supervised Learning framework for Anti-Money Laundering). SSL-AML models the heterogeneous graph based on meta-paths and constructs a semi-supervised learning framework including a pre-training stage and a consistency training stage to achieve the detection of cryptocurrency money laundering. In the pre-training stage, a similarity-based mechanism instead of a fixed confidence is used to dynamically generate a selection threshold for high-quality pseudo-labels, making the selection of pseudo-labels more flexible and effective. In the consistency training stage, learnable data augmentation is introduced and optimized from two aspects of consistency and diversity to improve the robustness and expressive ability of the model. In addition, without changing the original data distribution, homomorphic distribution is adopted to alleviate the high imbalance in the dataset. Summary of the Invention

[0011] The present invention proposes a cryptocurrency anti-money laundering method based on semi-supervised learning, and its design process is as follows: First, the cryptocurrency transaction data is constructed into a homogeneous graph or a heterogeneous graph. Then, in the pre-training stage, a label generator and a label evaluator are used to screen the pseudo-labels generated by the student model with a dynamic threshold to obtain high-quality pseudo-labeled data. Finally, the labeled data and the high-quality pseudo-labeled data are used for consistency training, and the relevant loss values are calculated and the node classification effect of the model is optimized. The detailed steps are as follows:

[0012] (1) Graph network construction and meta-path aggregation based on homomorphic distribution: The transactions or wallet addresses are used as nodes, and the directed edges are used as the direction of fund flow to construct a graph network. For heterogeneous transaction data, the method of meta-paths is also introduced to construct a heterogeneous transaction network, and then the meta-path information is aggregated based on the concept of homomorphic distribution.

[0013] 1) Heterogeneous graphs and homogeneous graphs: Graph is defined as where V and E represent the node set and the edge set respectively. For each node v ∈ V, there exists a mapping function τ(v): V → A that maps the node to its type. Similarly, each edge e ∈ E is associated with a mapping function ξ(e): E → R for the associated type. Here, A and R represent the predefined sets of node types and edge types respectively. Graph is a heterogeneous graph when |A| + |R| > 2, otherwise it is a homogeneous graph. In the heterogeneous cryptocurrency transaction graph, the node set V can be divided into two subsets: the node set V TX of all cryptocurrency transaction entities and the node set V AD .

[0014] 2) Meta-path construction of heterogeneous graphs: Meta-paths are constructed to capture the feature information between different nodes, and their definitions are as follows: The meta-path on the cryptocurrency transaction graph is represented as a triple (τ(i), ξ(e), τ(j)), which represents the heterogeneous relationship from the starting node i to the target node j, and the connecting edge is e. τ(·) represents the type category of the corresponding node, and ξ(·) represents the type of the edge e. Meta-paths are constructed, and three types of meta-paths are defined according to the types of wallet address nodes. Meta-path A: T X → Add ← T X , Meta-path B: T X ← Add → T X and Meta-path C: T X → Add → T X . Meta-path A corresponds to only input address nodes, Meta-path B corresponds to only output address nodes, and Meta-path C corresponds to input address nodes and output from the address nodes.

[0015] 3) Meta-path aggregation based on homomorphic distribution: In class-imbalanced datasets, the homomorphic distribution patterns of normal nodes and abnormal nodes are very different. Specifically, normal nodes tend to form connections with other normal neighbors, resulting in high homogeneity. On the contrary, abnormal nodes are often surrounded by a large proportion of normal neighbors, resulting in reduced homogeneity. Therefore, the homogeneous distribution is considered an effective method for distinguishing normal and abnormal nodes. The meta-path aggregation in the present invention uses a homogeneity-based encoder to distinguish legal and illegal nodes. Given a meta-path instance P(v, u), which connects the target node v and the neighbor The embedding representation of node v at the l-th layer is calculated as:

[0016]

[0017] In the formula, θ his the parameter set of MLP(·). Here, MLP(·) is used to calculate the homomorphism between two nodes, thereby obtaining the homomorphic representation of the edge. After that, the homomorphisms of all connected edges are aggregated to describe the overall context homomorphic distribution of node v, which indicates the possibility that it is a normal node and helps to improve subsequent classification. In this way, the class imbalance problem is solved without changing the original distribution of the dataset. For simplicity, h v is used to represent the embedding of node v below. In addition, a simplified graph neural network (GNN) is used, where the prediction p v ∈R 2 is defined as:

[0018] p v = F(h v ; θ p ) = softmax(W p h v + b p ) (2)

[0019] where are the parameters. Therefore, the corresponding predicted label y v is calculated as follows:

[0020] y v [argmax p v = 1(3)

[0021] Since there is a direct correlation between p v and y v , for simplicity, the prediction is sometimes referred to as the label.

[0022] (2) Pretraining stage: In this stage, the evaluator is pre-trained to provide a threshold for selecting high-quality pseudo-labels. To improve the generalization and flexibility of pseudo-label selection, a similarity-based evaluation mechanism is used to evaluate the pseudo-labels. The evaluation module consists of a label generator and a label evaluator, and pre-trains the teacher model f t (·) to select pseudo-labels.

[0023] Label generator: The generator G extracts the semantic information of its ground truth label v from the embedding vector h of the labeled node v. It is calculated as:

[0024] G(h v ; θ G ) = MLP(EMB(h v )) (4)

[0025] where θ G is the parameter set of G. The output of the generator G is not the student model f SThe pseudo-labels used in (·) are thus called false labels to show the difference. At the same time, the similarity between the pseudo-label p′ v and the real label p v is defined as the scaled cosine similarity:

[0026]

[0027] Label evaluator: The evaluator E receives the node with the embedding h v and its pseudo-label p′ v as inputs, extracts the semantic information of p v from h v , and calculates the semantic correlation between p v and p′ v to evaluate the similarity:

[0028] E(h v , p′ v ; θ E ) = Sigmoid(MLP(ATT(EMB(h v ), EMB(p′ v )))) (6)

[0029] where θ E is the parameter set of E.

[0030] Pretraining process: Before the student model fS(·) classifies the nodes using the selected pseudo-labels, the teacher model f t (·) is pre-trained to find a reasonable selection threshold. The auxiliary loss L aux is calculated as:

[0031] L aux = L G + L E (7)

[0032]

[0033] where B R is the size of the randomly subsampled batch. To prevent interference during the training process, the parameters of E, i.e., θ E , are frozen when updating G, and vice versa. Equation 8 is the loss of G, indicating that the generator is trained to make the semantic correlation between h v and its false label p′ v = G(h v ) approximate to 1, representing the maximum correlation. Similarly, Equation 9 is the loss of E, indicating that the evaluator is trained to make the estimated similarity E(h v , G(h v )) approximate to the actual similarity S(p v , G(hv Therefore, the similarity estimated by E can be used to evaluate the pseudo-labels after pre-training.

[0034] 1) High-quality nodes: Given an unlabeled node set V u and a labeled node set V l , the high-quality node set V hq can be defined as:

[0035] V hq = {v | v ∈ V u ∧ E(h v , G(h v )) ≥ τ} (10)

[0036] where τ is a threshold, set as the mean of the training batch E(h v , G(h v ))

[0037] (3) Consistency training stage: Existing graph data augmentation techniques are difficult to learn and difficult to control the degree of augmentation, which hinders subsequent learning tasks. Therefore, a learnable data augmentation method is used to further improve the classification in cryptocurrency trading data. The specific steps are as follows:

[0038] 1) Given a high-quality node v ∈ V hq , sharpen its embedding h v by masking some less important dimensions, expressed as:

[0039]

[0040] where is the sharpened synthetic embedding, θ lda is the parameter set of learnable data augmentation. ATT(·) is the attention mechanism for calculating the importance of each dimension of h v , such as the linear transformation Wh v + b. MASK(·) is the masking operation. According to the attention calculation of h v , the dimensions with smaller values are set to zero, and the dimensions with larger values are set to 1. The parameter θ mask controls the threshold.

[0041] 2) To optimize the learnable data augmentation, diversity and consistency metrics are used to design the loss function. The diversity metric emphasizes the difference between the original h v and the augmented , so the corresponding loss L d is denoted as:

[0042]

[0043] On the other hand, the consistency metric emphasizes correctness, i.e., enhancing the representation of the predicted label and the approximation degree with the original pseudo-label y″ v . Therefore, the corresponding loss L c is denoted as:

[0044]

[0045] where is the predicted value of v [argmaxpv]=1.

[0046] Therefore, this loss function allows injecting noisy data to obtain data diversity while ensuring consistency with the original data, improving the quality of data augmentation. Supervised and unsupervised loss calculations. Given a set of labeled nodes V l , a set of high-quality nodes V hq , the total loss L total is defined as the sum of the supervised loss L s and the unsupervised loss L u , denoted as:

[0047] L total =L s +L u (14)

[0048]

[0049] The pseudo-code of the training algorithm is as follows:

[0050] Table 1 Pseudo-code of the SSL-AML algorithm

[0051]

[0052]

[0053] The creativity of the present invention is mainly reflected in:

[0054] (1) The present invention proposes a semi-supervised learning framework to alleviate the scarcity of labeled data. This framework evaluates pseudo-labels based on similarity rather than confidence, overcomes the limitation of the confidence-based method that highly depends on manual design, and makes the selection of pseudo-labels more flexible and effective. In addition, learnable data augmentation and its corresponding training paradigm are introduced, which retain the diversity of the data space while pursuing consistency, thereby improving the robustness and overall performance of the model.

[0055] (2) The present invention constructs a simplified GNN with three transaction meta - paths to describe the complex semantic relationship between transactions and addresses, allowing SSL - AML to process heterogeneous cryptocurrency data.

[0056] (3) The present invention uses the homomorphic distribution of target nodes to distinguish legal and illegal nodes, thus solving the class - imbalance problem without changing the data distribution.

[0057] (4) The present invention conducts comprehensive experiments to evaluate SSL - AML on the Elliptic and Elliptic++ datasets. The experimental results show that the accuracy of this method is 95.7%, the recall rate is 69.9%, the F1 score is 80.6%, and the precision is 97.8%, all of which are better than the comparison baselines of homogeneous or heterogeneous datasets. Brief Description of the Drawings

[0058] Figure 1 is the semi - supervised learning cryptocurrency money - laundering detection framework diagram of the present invention. Detailed Description of the Invention

[0059] The present invention designs a cryptocurrency anti - money - laundering method based on semi - supervised learning; this method constructs a heterogeneous graph network through a meta - path - based method, optimizes the pseudo - label screening process based on pre - training technology to obtain high - quality pseudo - labeled nodes, and then classifies abnormal nodes through a consistency training process, successfully solving the problems of scarce labeled data, class imbalance, and heterogeneity in cryptocurrency money - laundering detection.

[0060] The experimental data comes from Elliptic and Elliptic++, and the specific data is shown in Table 1. When Elliptic++(Transaction) and Elliptic++(Actor) are used alone, they are homogeneous graph datasets, and when they are used together, they form a heterogeneous graph dataset, as shown in Table 2.

[0061] Table 2 Homogeneous Graph Datasets

[0062]

[0063]

[0064] Table 3 Heterogeneous Dataset Elliptic++(Actor + Transaction)

[0065]

[0066] The present invention adopts the following technical methods and implementation steps:

[0067] First, construct the cryptocurrency trading data into a homogeneous graph or a heterogeneous graph. Then, in the pre-training stage, use a label generator and a label evaluator to screen the pseudo-labels generated by the student model with a dynamic threshold to obtain high-quality pseudo-labeled data. Finally, use the labeled data and the high-quality pseudo-labeled data for consistency training, calculate the relevant loss value, and optimize the node classification effect of the model. The detailed steps are as follows: (1) Graph network construction and meta-path aggregation based on homomorphic distribution: Use the DGL third-party library to construct a graph network with transactions or wallet addresses as nodes and directed edges as the direction of fund flow. For heterogeneous transaction data, a method of introducing 3 meta-paths is also used to construct a heterogeneous transaction network, and then aggregate the meta-path information based on the concept of homomorphic distribution.

[0068] 1) Meta-path construction of heterogeneous graph: Construct meta-paths to capture the feature information between different nodes, and its definition is as follows: The meta-path on the cryptocurrency trading graph is represented as a triple (τ(i), ξ(e), τ(j)), which represents the heterogeneous relationship from the starting node i to the target node j, and the connecting edge is e. τ(·) represents the type category of the corresponding node, and ξ(·) represents the type of the edge e. Construct meta-paths, and three types of meta-paths are defined according to the type of wallet address nodes. Meta-path A: T X →Add←T X 、Meta-path B: T X ←Add→T X and Meta-path C: T X →Add→T X . Meta-path A corresponds to only input address nodes, meta-path B corresponds to only output address nodes, and meta-path C corresponds to input address nodes and output from the address node.

[0069] 2) Meta-path aggregation based on homomorphic distribution: In a class-imbalanced dataset, the homomorphic distribution patterns of normal nodes and abnormal nodes are very different. Specifically, normal nodes tend to form connections with other normal neighbors, resulting in high homogeneity. On the contrary, abnormal nodes are often surrounded by a large proportion of normal neighbors, resulting in reduced homogeneity. Therefore, the homogeneity distribution is considered an effective method to distinguish normal and abnormal nodes. The meta-path aggregation in the present invention uses a homogeneity-based encoder to distinguish legal and illegal nodes. Given a meta-path instance P(v, u), which connects the target node v and the meta-path-based neighbor The embedding representation of node v at the l-th layer is calculated as:

[0070]

[0071] In the formula, θ his the parameter set of MLP(·). Here, MLP(·) is used to calculate the homomorphism between two nodes, thereby obtaining the homomorphic representation of the edge. After that, the homomorphisms of all connected edges are aggregated to describe the overall context homomorphic distribution of node v, which indicates the possibility that it is a normal node and helps to improve subsequent classification. In this way, the class imbalance problem is solved without changing the original distribution of the dataset. For simplicity, h v is used to represent the embedding of node v below. In addition, a simplified graph neural network (GNN) is used in this patent, where the prediction p v ∈R 2 of node v is defined as:

[0072] p v = F(h v ; θ p ) = softmax(W p h v + b p ) (2)

[0073] where are parameters. Therefore, the corresponding predicted label y v is calculated as follows:

[0074] y v [argmax p v = 1 (3)

[0075] Since there is a direct correlation between p v and y v , for simplicity, the prediction is sometimes referred to as the label.

[0076] (2) Pretraining stage: In this stage, the evaluator is pre-trained to provide a threshold for selecting high-quality pseudo-labels. To improve the generalization and flexibility of pseudo-label selection, a similarity-based evaluation mechanism is used to evaluate the pseudo-labels. The evaluation module consists of a label generator and a label evaluator, and pre-trains the teacher model f t(·) to select pseudo-labels.

[0077] 1) Label generator: The generator G extracts the semantic information of its ground truth label v from the embedding vector h of the labeled node v. It is calculated as:

[0078] G(h v ; θ G ) = MLP(EMB(h v )) (4)

[0079] where θ G is the parameter set of G. The output of the generator G is not the student model f SThe pseudo-labels used in (·) are thus called false labels to show the difference. At the same time, the similarity between the pseudo-label p′ v and the real label p v is defined as the scaled cosine similarity:

[0080]

[0081] 2) Label evaluator: The evaluator E receives the nodes of the embedding h v and their pseudo-labels p′ v as inputs, extracts the semantic information of p v from h v , and calculates the semantic correlation between p v and p′ v to evaluate the similarity:

[0082] E(h v , p′ v ; θ E ) = Sigmoid(MLP(ATT(EMB(h v ), EMB(p′ v )))) (6)

[0083] where θ E is the parameter set of E.

[0084] 3) Pre-training process: Before the student model f S (·) classifies the nodes using the selected pseudo-labels, the teacher model f t (·) is pre-trained to find a reasonable selection threshold. The auxiliary loss L aux is calculated as:

[0085] L aux = L G + L E (7)

[0086]

[0087] where B R is the size of the randomly subsampled batch. To prevent interference during the training process, the parameters of E, namely θ E , are frozen when updating G, and vice versa. Equation 8 is the loss of G, indicating that the generator is trained to make the semantic correlation between h v and its false label p′ v = G(h v ) approximate to 1, indicating the maximum correlation. Similarly, Equation 9 is the loss of E, indicating that the evaluator is trained to make the estimated similarity E(h v , G(h v )) approximate to the actual similarity S(pv ,G(h v ))。Therefore, the similarity estimated by E can be used to evaluate the pre-trained pseudo-labels.

[0088] 4) High-quality nodes: Given an unlabeled node set V u and a labeled node set V l , the high-quality node V hq can be defined as:

[0089] V hq ={v|v∈V u ∧E(h v , G(h v ))≥τ} (10)

[0090] where τ is the threshold, set as the mean of the training batch E(h v , G(h v ). The hyperparameter settings are shown in Table 3.

[0091] Table 4 Hyperparameter settings

[0092]

[0093] (3) Consistency training stage: Existing graph data augmentation techniques are difficult to learn and difficult to control the degree of augmentation, hindering subsequent learning tasks. Therefore, a learnable data augmentation method is used to further improve the classification in cryptocurrency transaction data.

[0094] 1) Given a high-quality node v∈V hq , sharpen its embedding h v by masking some less important dimensions, expressed as:

[0095]

[0096] where is the sharpened synthetic embedding, and θ lda is the parameter set of learnable data augmentation. Among them, ATT(·) is the attention mechanism for calculating the importance of each dimension of h v , such as the linear transformation Wh v +b. MASK(·) is the masking operation, which sets the dimensions with smaller values to zero and the dimensions with larger values to 1 according to the attention calculation of h v , and the parameter θ msk controls the threshold.

[0097] 2) To optimize the learnable data augmentation, diversity and consistency metrics are used to design the loss function. The diversity metric emphasizes the difference between the original h v and the augmented , so the corresponding loss Ld Denoted as:

[0098]

[0099] On the other hand, the consistency metric emphasizes correctness, that is, enhancing the representation of the predicted label with the original pseudo-label y″ v of the approximation degree. Therefore, the corresponding loss L c Denoted as:

[0100]

[0101] is the predicted value of, y″ v [argmaxpv] = 1.

[0102] 3) Therefore, this loss function allows injecting noisy data to obtain data diversity while ensuring consistency with the original data and improving the quality of data augmentation. Supervised and unsupervised loss calculations. Given a set of labeled nodes V l , a set of high-quality nodes V hq , the total loss L total is defined as the sum of the supervised loss L s and the unsupervised loss L u , denoted as:

[0103] L total = L s + L u (14)

[0104]

[0105] The detection results are shown in Tables 4 and 5. Among them, the four indicators of Precision (Pre), Recall (Rec), F1score (F1), and Accuracy (Acc) are all expressed in percentage values. The optimal value is in bold, and the sub-optimal value is underlined. The formula for the indicator is as follows:

[0106]

[0107] Where Tp represents the number of nodes whose true label is licit and the model finally predicts licit nodes; Fn represents the number of nodes whose true label is legal and the model finally predicts illegal nodes; Fp represents the number of nodes whose true label is illegal and the model finally predicts legal nodes; Tn represents the number of nodes whose true label is illegal and the model finally predicts illegal nodes.

[0108] Table 5 Results of the homogeneous graph dataset

[0109]

[0110] As shown in Table 4, SSL-AML has significant advantages compared to the baselines, especially in the Elliptic and Elliptic++ (Transaction) datasets. It achieved the best performance in 9 out of 12 cases of 4 metrics on 3 datasets, reaching an accuracy of 95.7%, a recall of 69.9%, an F1-score of 80.6%, and an accuracy of 97.8% respectively. In contrast, the other 11 baselines did not have the best performance or only had the best performance in 1 case. Compared with the EvolveGCN method that achieved multiple sub-optimal performances, the method had an 8.4%, 9.4%, and 4.5% higher F1-score on the three datasets respectively. In addition, the method performed particularly well in terms of the F1-score, being the best in all three datasets, with an average increase of 7.1% across the three datasets compared to all baselines. Graph-based methods are very suitable for the AML task. According to Table 4, traditional machine learning-based methods such as LR or LSTM are significantly inferior to graph-based methods including SSL-AML, especially in terms of accuracy, indicating that graph-based methods are more suitable for the AML task because of their strong ability to capture complex structures and semantic information in transaction networks. The scale of the dataset has a significant impact on the detection results. The overall performance of almost all methods decreased significantly on Elliptic++ (Actor). The reason is that Elliptic++ (Actor) is very large in scale. It is the largest dataset among the three, with 822,942 nodes and 2,868,964 edges, posing a great challenge to the detection model.

[0111] Table 6 Results of Heterogeneous Graph Datasets

[0112]

[0113] As shown in Table 5, experiments were conducted using the heterogeneous graph dataset combined with Elliptic++ (Transaction) and Elliptic++ (Actor). SSL-AML outperformed all baselines in terms of the F1-score. The F1-score of SSL-AML reached 63.4%, with an average increase of 37.6% compared to all baselines. Since the F1-score is the harmonic mean of precision and recall, the excellent performance of the F1-score demonstrates the comprehensive ability of the method. SSL-AML is robust and successfully addresses the challenges of large-scale datasets and heterogeneity. It was compared with advanced heterogeneous graph anomaly detection algorithms such as BWGNN (hetero). SSL-AML achieved a significant performance improvement, with a 20.1% increase in accuracy and a 4.0% increase in the F1-score.

[0114] The experimental results show that the proposed method solves the problems of scarcity of labeled data, class imbalance, and heterogeneity in the field of cryptocurrency money laundering detection, and outperforms the existing state-of-the-art methods in four performance indicators of cryptocurrency money laundering detection, effectively improving the generalization and accuracy of the detection model, highlighting its practical value and broad application prospects in cryptocurrency financial security.

Claims

1. A cryptocurrency anti-money laundering method based on semi-supervised learning, characterized in that: The implementation process of this method is as follows: Firstly, the cryptocurrency transaction data is constructed into a homogeneous graph or a heterogeneous graph. Then, in the pre-training stage, the pseudo labels generated by the student model are screened with a dynamic threshold using a label generator and a label evaluator to obtain high-quality pseudo-labeled data. Finally, the labeled data and the high-quality pseudo-labeled data are used for consistency training to calculate the relevant loss value and optimize the node classification effect of the model.

2. A cryptocurrency anti-money laundering method based on semi-supervised learning according to claim 1, characterized in that: The implementation steps of the method are as follows: Step (1) Graph network construction and meta-path aggregation based on homomorphic distribution: Transaction or wallet addresses are used as nodes and directed edges are used as the direction of fund flow to construct a graph network. Heterogeneous transaction data also introduces the meta-path method to construct a heterogeneous transaction network, and then the meta-path information is aggregated based on the concept of homomorphic distribution; Step (2) Pre-training phase: pre-train the evaluator to provide a threshold for selecting high-quality pseudo labels; to improve the generalization and flexibility of pseudo label selection, a similarity-based evaluation mechanism is used to evaluate pseudo labels; The evaluation module consists of a label generator and a label evaluator, which evaluates the teacher model f t (·) Perform pre-training and select pseudo labels; Step (3) Consistency training phase: Use learnable data augmentation methods to improve classification in cryptocurrency transaction data.

3. A cryptocurrency anti-money laundering method based on semi-supervised learning according to claim 2, characterized in that: In step (1), 1) heterogeneous graph and homogeneous graph: Defined as Where V and E represent the set of nodes and the set of edges, respectively. For each node v∈V, there exists a mapping function τ(v):V→A that maps the node to its type. Each edge e∈E has an associated type with a mapping function ξ(e):E→R. Where A and R represent the predefined sets of node types and edge types, respectively. When |A|+|R|>2, it is a heterogeneous graph, otherwise it is a homogeneous graph; in the heterogeneous cryptocurrency transaction graph, the node set V is divided into two subsets: the node set V of all cryptocurrency transaction entities TX And the node set V of all cryptocurrency wallet addresses AD ; 2) Meta-path construction of heterogeneous graph: Meta-path is constructed to capture the characteristic information between different nodes, which is defined as follows: the meta-path on the cryptocurrency transaction graph is represented as a triple (τ(i), ξ(e), τ(j)), which represents the heterogeneous relationship from the starting node i to the target node j, and the connecting edge is e; τ(·) represents the type category of the corresponding node, and ξ(·) represents the type of edge e; meta-path is constructed, and three types of meta-paths are defined according to the type of wallet address node, meta-path A: T X →Add←T X 、Metapath B: T X ←Add→T X and metapath C:T X →Add→T X ; Metapath A corresponds to an input-only address node, metapath B corresponds to an output-only address node, and metapath C corresponds to an input address node and outputs from a direct node; 3) Meta-path aggregation based on homomorphic distribution: In class-imbalanced datasets, the homomorphic distribution patterns of normal nodes and abnormal nodes are very different; specifically, normal nodes tend to form connections with other normal neighbors, resulting in high homogeneity; on the contrary, abnormal nodes are often surrounded by a larger proportion of normal neighbors, resulting in reduced homogeneity; homogeneous distribution is considered to be an effective way to distinguish normal and abnormal nodes; meta-path aggregation adopts a homogeneity-based encoder to distinguish between legitimate and illegal nodes; given a meta-path instance P(v,u), it connects the target node v and its neighbors based on the meta-path The embedding representation of node v at layer l Calculated as: Where θ h is the parameter set of MLP(·); is the embedding representation of the neighbors of node v in the l-1 layer; the homomorphism between two nodes is calculated using MLP(·) to obtain the homomorphic representation of the edge; the homomorphism of all connected edges is aggregated to describe the overall contextual homomorphic distribution of node v, which solves the class imbalance problem without changing the original distribution of the dataset; h is used v To represent the embedding of node v below; using a simplified graph neural network GNN, where the prediction p of node v v ∈R 2 Defined as: p v =F(h v ;θ p )=softmax(W p h v +b p ) (2) in, is the parameter of the prediction function F(·; ·); the corresponding prediction label y v The calculation is as follows: y v [argmaxp v ]=1(3) Because p v and v There is a direct correlation between , and the prediction is called the label.

4. A cryptocurrency anti-money laundering method based on semi-supervised learning according to claim 2, characterized in that: In step (2), 1) Label generator: The generator G extracts the embedding vector h of the labeled node v v Extract its true value label The semantic information of ; calculated as: G(h v ;θ G )=MLP(EMB(h v )) (4) In the formula, θ G is the set of parameters of G; the output of the generator G is not the student model f S The pseudo-label used in (·) is called a false label to distinguish it from the false label p′. v With the true label p v The similarity between is defined as the scaled cosine similarity: 2) Label Evaluator: The evaluator E receives the embedding h v The nodes and their pseudo labels p′ v As input, from h v Extract p v Semantic information, calculate p v and p′ v The semantic correlation between them is used to evaluate the similarity: Where θ E is the parameter set of E; 3) Pre-training process: In the student model f S (·) Before classifying nodes using the selected pseudo labels, the teacher model f t (·) Perform pre-training to find the selection threshold; auxiliary loss L aux Calculated as: L aux =L G +L E (7) Among them B R is the size of the random sub-sampling batch, i is the node for l2 regularization loss calculation, is the node embedding of the i-th node; in order to prevent mutual interference during training, the parameter of E, i.e., θ E , is frozen when updating G, and vice versa; Formula (8) is the loss of G, indicating that the generator is trained so that Instead of false labels The semantic correlation between them is close to 1, indicating the maximum correlation; Similarly, formula (9) is the loss of E, indicating that the estimator is trained to estimate the similarity Approximately similar to actual Evaluate the pre-trained pseudo labels using the similarity estimated by E; 4) High-quality nodes: Given an unlabeled node set V u and the labeled node set V l , high quality node V hq It can be defined as: V hq =V l ∪{v|v∈V v ∧E(h v ,G(h v ))≥τ} (10) Where τ is the threshold, set as the training batch E(h v , G(h v )).

5. The cryptocurrency anti-money laundering method based on semi-supervised learning according to claim 2, characterized in that: In step (4), the specific steps are as follows: 1) Given a high-quality node v∈V hq , sharpening its embedding h by masking the dimensions v , expressed as: in, is the synthetic embedding after sharpening, θ lda is the parameter set for learning data enhancement; ATT(·) is the calculation of h v Attention mechanism for the importance of each dimension, such as linear transformation Wh v +b.MASK(·) is a masking operation, according to h v The attention calculation is to set the dimension with smaller value to zero and the dimension with larger value to 1. The parameter θ msk Control threshold; 2) To optimize the learnable data enhancement, the loss function is designed using diversity and consistency metrics; the diversity metric emphasizes the original h v With the enhanced The difference between the corresponding loss L d Denoted as: Where V hq is a set of high-quality nodes; on the other hand, the consistency metric emphasizes correctness, that is, enhancing the representation The predicted label With the original pseudo label y″ v The approximation degree of c Denoted as: In the formula, for The predicted value, y" v [argmaxpv] = 1; 3) The loss function allows the injection of noise data to obtain data diversity and ensures consistency with the original data, thereby improving the quality of data enhancement. In supervised and unsupervised loss calculations, given a set of labeled nodes V l , a set of high-quality nodes V hq , the total loss L total Defined as the supervised loss L s With the unsupervised loss L u The sum is recorded as: L total =L s +L u (14) i ranges from 0 to 1, indicating two different categories or states: normal nodes and abnormal nodes.