Ethereum transaction user classification method and system based on semantic motif differentiation, medium and program product

By building a transaction subgraph containing semantic information and using the HGN model, the problem that traditional methods are difficult to identify Ethereum phishing users is solved, effectively detecting and classifying phishing user nodes in the Ethereum trading network is realized, and the ability to identify abnormal transaction behaviors is improved.

CN119939310APending Publication Date: 2025-05-06HARBIN ENG UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510010305.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The traditional phishing user classification method is difficult to adapt to the Ethereum scenario and cannot effectively identify and classify phishing users in Ethereum.

Method used

Using a method based on semantic model differentiation, a transaction subgraph containing semantic information is constructed, and topological features and semantic features are extracted using the Heterogeneous Graph Network (HGN) model to perform accurate subgraph-level classification.

Benefits of technology

It realizes effective detection and classification of phishing user nodes in the Ethereum trading network, improves the ability to identify abnormal transaction behaviors, and protects the user's property security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939310A_ABST
    Figure CN119939310A_ABST
Patent Text Reader

Abstract

The invention discloses an Ethereum transaction user classification method and system based on semantic motif differentiation, a medium and a program product. The invention provides an Ethereum transaction user classification model EPSD-motif based on semantic motif differentiation. The model comprises a transaction data acquisition and processing device, a transaction network construction device, a semantic motif construction device, a subgraph sampling device based on motif features, and a feature extraction and classification device. According to the method, after public Ethereum transaction data and a label data set are acquired, second-order transaction data of a phishing user node, a common user node and an ICO wallet node are acquired, and transaction characteristics are quantitatively analyzed; constructing a motif by using the semantic features, and constructing a node role according to motif frequency distribution; carrying out subgraph sampling by adopting transaction semantic features, and obtaining structural features through a model; and in combination with the structural features and the transaction features, result classification is carried out by using an MLP. According to the method, the accuracy and the interpretability of the Ethereum transaction user classification method are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Ethereum transaction user classification, and specifically relates to an Ethereum transaction user classification method, system, medium and program product based on semantic motif differentiation. Background Art

[0002] As a distributed computing platform, Ethereum is widely used in smart contracts, decentralized finance and other fields. However, the anonymity of Ethereum also facilitates some people to use it for improper activities, making Ethereum a carrier of some special economic behaviors. Related activities on the platform are gradually increasing, such as phishing transactions and fund-raising. Phishing transactions, as a typical transaction behavior, are widely seen on the platform. In the statistics of phishing transactions, in the first week of the activity, users reported asset losses of 645,000 US dollars, and the relevant parties made profits of more than 3 million US dollars in one month. Phishing transactions have caused great economic losses and have become an important concern for Ethereum transaction security.

[0003] It is precisely because of the diversity of Ethereum phishing methods that traditional phishing user classification methods cannot adapt well to Ethereum scenarios. Traditional phishing activities rely on building fake websites or software to collect sensitive information of victims or receive remittances, so traditional methods focus on mining fake platform patterns, such as website URLs, source code, CSS styles, and page layouts. However, traditional detection methods based solely on feature extraction are difficult to migrate to Ethereum, and these methods are also ineffective. Therefore, designing methods to identify phishing user detection in Ethereum has become an important topic, attracting the attention of experts in the field of cryptocurrency security.

[0004] As a decentralized public ledger, the transaction records in Ethereum are public and accessible to anyone. This also provides a valuable data source for the research on phishing user classification in Ethereum transactions. Modeling and analyzing complex transaction content in the form of a graph structure, thereby converting massive Ethereum transaction information into a classification task on the graph, can help researchers find the target node. Using relevant methods in the graph field, such as graph representation learning and graph deep learning, mining various features (account features, transaction features, topological structure features, etc.) from the transaction graph has become a common method for classifying Ethereum transaction users.

[0005] However, in existing research, it is usually ignored that Ethereum is a multi-label, highly heterogeneous transaction network. Therefore, the present invention designs and implements a method that can fully extract transaction semantic features. The method can be effectively applied to the highly heterogeneous Ethereum transaction network and identify various types of transaction user nodes in a multi-label scenario. Summary of the invention

[0006] The purpose of the present invention is to provide an Ethereum transaction user classification method, system, medium and program product based on semantic motif differentiation, to assist in the compliance management of the Ethereum transaction network and to effectively combat the identification of abnormal transaction behaviors using Ethereum.

[0007] The purpose of the present invention is achieved through the following technical solutions:

[0008] A method for classifying Ethereum transaction users based on semantic motif differentiation specifically includes the following steps:

[0009] Step 1: Dataset and transaction network construction;

[0010] Collect all Ethereum transaction data and three types of addresses marked as phishing user nodes, ordinary user nodes, and ICO wallet nodes on the Etherscan and Xblock platforms, and build respective second-order subgraphs centered on these three types of nodes;

[0011] Step 2:

[0012] First, the motif is reconstructed according to the transaction semantics. Then, the motif features of the nodes in the second-order subgraph obtained in step 1 are constructed according to the predefined semantic motifs, focusing on the motif features with higher frequency. Then, a sampling strategy combining semantic enhancement and complementarity is adopted. Starting from the central node, the first-order neighbors and second-order neighbors are sampled layer by layer to construct a second-order subgraph that contains semantics and has obvious differentiation.

[0013] Step 3: Feature extraction and classification;

[0014] The transaction subgraph containing semantic information obtained in step 2 is input into the HGN model designed specifically for the strong heterogeneity of Ethereum; the model learns the adjacency matrix and feature matrix of the transaction graph respectively to generate feature representations of the subgraph, including topological structure features and semantic features; finally, these feature vectors are sent to the multi-layer perceptron MLP for accurate subgraph-level classification.

[0015] Furthermore, the step 1 is specifically as follows:

[0016] Step 1.1: Collect all transaction records of Ethereum; download all transaction data through the well-known blockchain academic platform Xblock, which provides complete Ethereum transaction data processing. All transaction data can be synchronized to the server device used in the present invention simply through the command line; for each transaction field, retain the four fields of transaction initiator from, transaction recipient to, transaction amount amount and transaction timestamp timestamp;

[0017] Step 1.2: Collect account labels; in order to build a multi-category node network close to the real-life scenario of Ethereum, download from the Label Cloud section through the API interface provided by the Xblock academic platform and the Etherscan website, and collect labels of phishing accounts, ordinary accounts with non-specific functions, and ICO wallet accounts;

[0018] Step 1.3. Construct a second-order transaction network; based on the obtained account address labels and transaction data on the Ethereum chain, retain all second-order transaction data related to the label address in a traversal form, forming a second-order subgraph centered on these three label addresses;

[0019] Step 1.4: Redundant address screening; In order to solve the problem of redundant transaction information on the second-order subgraph, secondary filtering is performed to reduce the interference of irrelevant information; (1) The present invention deletes transactions with a transfer amount of 0, because such transactions have no reference significance. (2) When obtaining the first-order neighbors of the target node, the exchange nodes with larger degree values ​​are targeted. Due to their large transaction volume, these exchange nodes often appear in the transaction sequence, thereby reducing the uniqueness of the phishing user representation; therefore, those nodes that are connected to the exchange node but not directly connected to the target node are deleted; then the exchange node is retained, because the connection of the phishing user node to the exchange itself is its characteristic.

[0020] Furthermore, the step 2 is specifically as follows:

[0021] Step 2.1: Construct a motif with semantics; add the centrality of nodes to the motif, expand the number of motifs from 13 to 21, and the reconstructed 21 motifs contain rich transaction semantics;

[0022] Step 2.2: Motif feature construction: For each node v in the network, analyze all subgraphs G consisting of three nodes containing node v v , and identify the subgraphs that are related to the predefined motifs M={M1,M2,...,M 21} matching part; for each node v, calculate the frequency t of each motif appearing around node v v,i :

[0023]

[0024] Among them, |g∈G v :g∈M i | indicates a specific motif M i The number of occurrences of |G v | represents the number of subgraphs consisting of three nodes of node v;

[0025] The motif feature vector of node v is represented by T v :

[0026] T v =(t v,1 ,t v,2 ,...,t v,M )

[0027] Among them, each element t v,i represents the frequency of a specific motif i in the first-order neighbor range of node v, and |M| is the 21 types of semantic motifs;

[0028] Step 2.3: Subgraph sampling based on motif features; After calculating the number distribution of motif features for each node, the sampling process is carried out based on these motif features.

[0029] Furthermore, step 2.3 focuses on the most common motifs of each central node, that is, the motifs corresponding to the larger values ​​in the motif feature vector. This reflects the importance of a certain motif in the interaction pattern and semantic information of the node. Then, this motif recognition method is further applied to the first-order and second-order neighbors of the central node, giving priority to those nodes that show similar or identical high-frequency motif features as the central node, that is, nodes with highly similar semantic information. Through this strategy, targeted sampling of the second-order subgraph is achieved, so that the transaction semantics of the entire subnet are as close to the semantics of the central node as possible, thereby achieving a semantic enhancement effect.

[0030] Furthermore, the HGN model in step 3 simultaneously learns to extract topological relationships from the adjacency matrix and semantic information from the node feature matrix, and strengthens the retention and synergy of the two features through a fusion mechanism of splicing, linear transformation and direct addition, so that the model can comprehensively capture the diverse information of complex networks; the core of HGN is composed of two modules, MLP and LINK: the MLP module focuses on extracting semantic information from node features, and the LINK module focuses on extracting graph structure information from the adjacency matrix. The model adaptively adjusts the weights of the two modules according to the characteristics of the input data, thereby adapting to a variety of heterogeneous graph data scenarios; HGN inputs the generated comprehensive features into the multi-layer perceptron MLP for refined classification, which significantly improves the classification accuracy of heterogeneous graph data.

[0031] Furthermore, the step 3 is specifically as follows:

[0032] Step 3.1: The first part of the model does not consider the graph structure, that is, the MLP model that only requires node feature information is used;

[0033] h A =MLP A (A)∈R d×n

[0034] h x =MLPx (X)∈R d×n

[0035] Among them, h A Indicates that through MLP A The hidden representation extracted from the adjacency matrix A; h X Indicates that through MLP X Hidden representations extracted from the node feature matrix X;

[0036] Step 3.2: A logistic regression model is trained using the LINK model of the graph topology, where the features of each node are taken from the adjacency matrix A∈{0,1} n×n A column of W∈R c×n is a learnable weight matrix; with this setting, the LINK model can calculate the probability of a node belonging to each category, the specific formula is:

[0037] Y=soft max(WA)

[0038] For a specific node u∈{1,...,n} and a specific class k∈{1,...,c}, the logarithmic probability that node u belongs to class k is calculated by the following formula:

[0039] (WA) ku =Σ v∈N(u) W kv

[0040] Among them, N (u) Contains the one-hop neighbors of u. The formula is based on the weights of the first-order neighbors of node u and W kv ; If a particular node v has many neighbors belonging to category k, then W kv It may be very large, which means that there is a high probability that any neighbor of v belongs to category k;

[0041] Combine these two partial models through linear transformation; let X∈R d×n represents the node feature matrix with input dimension d, and [h1;h2] represents the concatenation operation of the vector h1 generated by MLP and the vector h2 generated by LINK; then the prediction result is obtained through the following process:

[0042] y=MLP f (σ(W[h1;h2]+h1+h2)).

[0043] Furthermore, in the HGN model, the hidden representation h1 of the feature matrix and the hidden representation h2 of the adjacency matrix are first calculated; then, the model realizes the integration of these two hidden representations by concatenating h1 and h2 and performing linear transformation; in order to further strengthen the model's retention of topological structure and node features, the model also adopts the strategy of directly adding h1 and h2, and finally classifies the results through MLP; the HGN model effectively solves the challenges faced by traditional graph models in processing highly heterogeneous graph data through the representation fusion mechanism, and demonstrates the ability to deeply explore and effectively utilize complex network structures and rich node features.

[0044] Further, an Ethereum transaction user classification system based on semantic motif differentiation includes a transaction data collection and processing device, a transaction network construction device, a semantic motif construction device, a subgraph sampling device based on motif features, and a feature extraction and classification device;

[0045] The main task of the transaction data collection and processing device is to collect all transaction records of Ethereum as well as fishing user nodes, ordinary users and ICO wallet tags;

[0046] The main task of the transaction network construction device is to traverse the transaction records of Ethereum, construct a second-order transaction subgraph with the central node as the starting node, and perform secondary filtering to reduce the interference of irrelevant information;

[0047] The semantic motif construction device reconstructs a motif with semantic information based on the transaction purpose of the Ethereum transaction user;

[0048] The subgraph sampling device based on motif features constructs role types according to the motif distribution frequency in the network to guide the sampling of subgraphs;

[0049] The feature extraction and classification device inputs the transaction subgraph containing semantic information into the HGN model to generate feature representation of the subgraph; finally, the basic features are combined with these feature vectors and sent to the multi-layer perceptron MLP for accurate subgraph level classification.

[0050] A computer-readable storage medium stores a computer program / instruction thereon, which, when executed by a processor, implements the steps of a method for classifying Ethereum transaction users based on semantic motif differentiation.

[0051] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of a method for classifying Ethereum transaction users based on semantic motif differentiation.

[0052] The beneficial effects of the present invention are:

[0053] The present invention detects and classifies transaction users on Ethereum so as to timely identify and prevent abnormal behaviors in the transaction network and protect the property safety of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic diagram of the system of the present invention;

[0055] Figure 2 It is a system structure block diagram of the present invention;

[0056] Figure 3 Schematic diagram of 21 semantic motifs reconstructed in the system of the present invention;

[0057] Figure 4 A schematic diagram of finding motif features in the system of the present invention;

[0058] Figure 5 A schematic diagram of an enhanced sampling strategy in the system of the present invention;

[0059] Figure 6 A schematic diagram of a complementary sampling strategy in the system of the present invention;

[0060] Figure 7 It is a structural diagram of the HGN model in the system of the present invention;

[0061] Figure 8 This is a comparison between using the enhanced sampling method alone and using a mixture of the two sampling methods;

[0062] Fig. 9 Comparison of semantic identification capabilities of different methods. DETAILED DESCRIPTION

[0063] The present invention is further described below in conjunction with the accompanying drawings.

[0064] according to Figure 1As shown, the present invention first collects three types of addresses marked as fishing user nodes, ordinary user nodes and ICO wallets from Etherscan and Xblock platforms, and constructs respective second-order subgraphs with these three types of nodes as the center. In the subgraph sampling stage based on motif features, the present invention constructs motif features for nodes in the subgraph according to predefined semantic motifs, focusing on motif features with higher frequency of occurrence. Afterwards, a sampling strategy combining semantic enhancement and complementarity is adopted, starting from the central node, and sampling is performed layer by layer for first-order neighbors and second-order neighbors to construct a second-order subgraph containing semantics and having obvious differentiation. In the feature extraction and classification stage, the present invention inputs the transaction subgraph containing semantics obtained in the previous step into the HGN model designed specifically for the strong heterogeneity of Ethereum. The model learns the adjacency matrix and feature matrix on the transaction graph respectively, and then generates feature representations of the subgraph, including topological structure features and semantic features. Finally, these feature vectors are spliced ​​and then processed by a multi-layer perceptron (MLP) to achieve subgraph level classification.

[0065] according to Figure 2 As shown, it is a structural block diagram of the system of the present invention, including:

[0066] Transaction data collection and processing device: The main function of this device is to collect data sets. The device works in two steps. The first step: collect all transaction records of Ethereum. This device downloaded all transaction data through the well-known blockchain academic platform Xblock. For each transaction field, the present invention only retains four fields: transaction initiator (from), transaction recipient (to), transaction amount (amount) and transaction timestamp (timestamp). The second step: collect account tags. In order to build a multi-category node network close to the real scene of Ethereum, the present invention collects three types of tags: fishing user nodes, ordinary user nodes and ICO wallets.

[0067] Transaction network construction device: The main task of this device is to build respective second-order subgraphs with the three types of collected tag data as the central node. In order to solve the problem of redundant transaction information on the second-order subgraph, this device performs secondary screening, filters out transactions with a transfer amount of 0, and deletes all transaction records that are connected to the exchange node but not directly involved in the target node transaction, so as to effectively reduce the interference of irrelevant information.

[0068] Semantic model construction device: The main task of this device is to increase the centrality of nodes in the model based on the model that does not contain semantics, thereby expanding the number of models from 13 to 21, and the reconstructed 21 models contain rich transaction semantics.

[0069] Subgraph sampling device based on motif features: The main task of this device is to calculate the motif distribution frequency in each second-order subgraph, and construct the role difference of the node based on the motif distribution frequency. Specifically, the sampling strategy is mainly divided into two directions: one is sampling based on the semantic enhancement of the motif; the other is sampling based on the semantic complementarity of the motif. The subgraph formed according to the above sampling strategy has semantic features.

[0070] Feature extraction and classification device: The main function of this device is to receive the sampled sub-graph as input, combine the acquired feature representation with the transaction features for classification, and obtain the classification result.

[0071] according to Figure 3 As shown, the schematic diagram of 21 kinds of semantic motifs after reconstruction in the system of the present invention is as follows:

[0072] Based on the motifs without semantics, the present invention adds the centrality of nodes in the motifs, thereby expanding the number of motifs from 13 to 21, and the reconstructed 21 motifs contain rich transaction semantics. By adding central nodes, the transaction edges of the present invention have richer semantic information and can better express the transaction characteristics of real scenarios.

[0073] The present invention divides the reconstructed motifs into six groups, each group has three to four motifs, and the division relationship between the groups is the transaction relationship between the central node and its two adjacent first-order neighbors. The division of motifs within the same group is the transaction relationship between first-order neighbors, including unidirectional transaction relationships and bidirectional transaction relationships.

[0074] Here, the present invention uses the motif M 16 and motif M 17 Let’s take an example to explain the similarities and differences between the two models. The central node of the two models and its two first-order neighbors are the outgoing transaction relationship and the incoming transaction relationship respectively. 16 The transaction relationship between the two first-order neighbors of is the node that has a transaction relationship with the central node and points to the node that transfers out transactions to the central node. 17 The transaction relationship between the two first-order neighbors is from the node that transfers outgoing transactions to the central node to the node that has a transaction relationship with the central node. Although the two motifs are extremely similar, they are very different in the semantic environment of real transactions.

[0075] It is worth emphasizing that when considering motifs in the network, even for different types of motifs, such as open triangles and closed triangles, the semantics of motifs in the same group still have certain similarities. Among these motifs, the transaction relationship between the central node and its two first-order neighbors is the focus of the analysis. Although the transaction relationship between first-order neighbors is equally critical, its relative importance is slightly inferior to the direct connection with the central node. This semantic similarity is crucial when developing subsequent sampling strategy algorithms, and provides reasonable guidelines for the design of diversified sampling algorithms in the present invention. This not only helps the present invention to capture and understand network behavior more accurately, but also ensures that the priority and relevance of core transaction patterns are fully considered during the sampling process.

[0076] Taking a non-closed triangle as an example, the following analyzes the semantics of the motif in the Ethereum phishing transaction network in the Ethereum transaction scenario:

[0077] M1 motif semantics: The phishing user node adopts a decentralized transfer strategy to transfer the obtained ether to multiple nodes in order to avoid direct tracking and increase the difficulty of detection. This strategy shows that the phishing node attempts to hide its true capital path by splitting the capital flow.

[0078] M4 motif semantics: The phishing user node first trades with multiple transaction objects, obtains ether, and then transfers the funds through other intermediate nodes. This shows that the phishing user not only interacts directly with the victim, but also uses third-party nodes to cover up the flow of funds.

[0079] M8 motif semantics: Phishing user nodes tend to trade with different victims rather than repeatedly trade with the same object. This may reflect the strategy of phishing user nodes, which is to widely identify and exploit potential victims to maximize profits.

[0080] M 11 Model definition: The phishing user node not only conducts one-way luring transactions, but also has back-and-forth transactions with some ordinary nodes. This behavior may build trust and reduce the vigilance of specific nodes, making phishing activities more difficult to identify and attracting more victims to invest in Ethereum.

[0081] M 15 Model definition: This model adds two-way transactions with ordinary nodes based on M8. This strategy is used to further confuse transaction tracks and reduce the vigilance of victims and other network participants by establishing superficial trust relationships, thereby increasing the concealment and effectiveness of phishing behavior.

[0082] M 19Model definition: Frequent two-way transactions between phishing user nodes and ordinary user nodes are intended to simulate normal trading behavior and reduce the vigilance of other network participants to phishing activities. This interaction may be intended to construct false business activities and induce more victims to invest or trade.

[0083] according to Figure 4 As shown, the schematic diagram of finding motif features in the system of the present invention is as follows:

[0084] For each node v in the network, the present invention analyzes all subgraphs G consisting of three nodes containing the node v. v , and identify the subgraphs that are related to the predefined motifs M={M1,M2,...,M 21 For each node v, the present invention calculates the frequency t of each motif appearing around the node v. v,i :

[0085]

[0086] Among them, |g∈G v :g∈M i | indicates a specific motif M i The number of occurrences of |G v | represents the number of subgraphs consisting of three nodes of node v. The motif feature vector of node v can be expressed as T v :

[0087] T v =(t v,1 ,t v,2 ,...,t v,|M| )

[0088] Each element t v,i represents the frequency of a specific motif i in the first-order neighborhood of node v, and |M| is one of the 21 types of semantic motifs.

[0089] Figure 4 It shows how to determine the frequency of occurrence of a specific motif. On the far left is a simplified version of the transaction network diagram centered on node A. {A,C,D},{A,E,C},{A,E,D} are three instances of the motif M1 centered on A, and {A,B,C},{A,E,B} are three instances of the motif M2 centered on A. 13 , and {A,D,B} constitutes the motif M centered on A. 14 Then in the transaction subnetwork composed of {A, B, C, D, E}, the motif feature of the node is: T A =(1 / 2,...,1 / 3,1 / 6,....).

[0090] The present invention constructs an adjacency matrix to traverse the frequency of occurrence of each motif of each node in the graph, and thereby forms their motif features. Since the transaction network of the present invention is a directed multigraph (edges are directional and allow multiple edges between two nodes), the adjacency matrix A needs to be processed to facilitate subsequent matrix operations. Matrix B represents a bidirectional adjacency matrix, which is obtained by A and its transposed A. T The product of is calculated, in B, the non-zero element B ij In other words, if there is an edge from node i to node j, and there is an edge from node j to node i, then B ij is 1, which reflects a "bidirectional" or "reciprocal" relationship. The matrix U represents the one-way adjacency matrix, which is obtained by subtracting the bidirectional adjacency matrix B from the adjacency matrix A. In U, the non-zero elements U ij It means that there is a one-way edge from node i to node j, but there is no edge from j to i.

[0091] according to Figure 5 As shown, the schematic diagram of the enhanced sampling strategy in the system of the present invention is as follows:

[0092] In this sampling strategy, the present invention focuses on the most common motifs of each central node, that is, the motifs corresponding to the larger values ​​in the motif feature vector. This reflects the importance of a certain motif in the interaction pattern and semantic information of the node. Then, this motif recognition method is further applied to the first-order neighbors and second-order neighbors of the central node, giving priority to those nodes that exhibit similar or identical high-frequency motif features as the central node, that is, nodes with highly similar semantic information. Through this strategy, targeted sampling of the second-order subgraph is achieved, so that the transaction semantics of the entire subnet are as close to the semantics of the central node as possible, to achieve a semantic enhancement effect.

[0093] For example, when analyzing the motif features centered on the fishing node, if it is found that the motif M1 appears most frequently in the node, this means that M1 is particularly critical in representing the network behavior and semantic content of the fishing node. Therefore, when sampling the first-order neighbors of the fishing node, the nodes with the M1 motif (including the group of motifs M1, M2, and M3) as the most significant feature will be selected.

[0094] according to Figure 6 As shown, the schematic diagram of the complementary sampling strategy in the system of the present invention is as follows:

[0095] In this sampling strategy, the present invention still focuses on the motifs with higher frequency of occurrence in each central node, and then determines the sampling range of first-order neighbors and second-order neighbors through the complementarity of nodes. In the complementary sampling process, the present invention will focus on the motifs with more important semantic correspondence in the transaction pattern, which may not simply be similar to the semantics of the central node. It tends to select those motifs that can reveal phishing user behavior and possible collusion relationships. This strategy focuses on key semantic features and samples the motifs corresponding to important semantics as the next hop.

[0096] Figure 6 The model M 16 The complementary sampling process is taken as an example. 16 The revealed semantics is the transaction relationship between the phishing node A and its neighbor B, which not only includes B transferring money to A, but also covers A luring B through some form of transaction. At the same time, the transaction between nodes B and C may indicate that they belong to different accounts of the same entity, or that they know each other in the real world, so there are transactions between them. Such a transaction relationship may lead to node C participating in the transaction. The phishing user node may adopt a variety of strategies to make other users trade with it. For example, the phishing user node first transfers a small amount of cryptocurrency to the address of an ordinary user. This method is called a "baiting attack" or "dusting attack". This strategy means that when the phishing user node transfers money to the ordinary user node, the accompanying data field may contain information that induces the user to visit a counterfeit website. Once the user visits these websites and enters the private key or other sensitive information, the phishing user can steal the assets.

[0097] Therefore, in such a semantic environment, for M 16 The sampling strategy of will be extended to the relevant motifs in the next step. Specifically, the motif M 16 The next hop sampled motifs include motif M1 represented by node B, which indicates that there is both a fraud experience and a friend transfer relationship. Another sampled motif is motif M8 represented by node C, which indicates that there is an experience of receiving friend transfers and a fraud experience.

[0098] According to the above 16 Taking the next hop sampling range of an example, the present invention designs a semantic-based complementary sampling strategy for all motifs. Complementarity will focus on more important motifs in transaction semantics, and use important motifs that meet transaction semantics as the next hop sampling range.

[0099] Furthermore, relying solely on enhanced sampling may lead to over-enhancement of features. The first-order and second-order neighbors retained are those with similar semantics to the central node, which makes the semantics of the entire subnet close to the central node and cannot fully reflect the diversity of semantic information in the transaction network. Therefore, in the actual sampling process, it is necessary to mix enhanced sampling and complementary sampling. For the motif sampling range of the next hop of the central node, not only can the enhanced sampling range be retained, but also the complementary sampling range can be added, thereby increasing the diversity of sampling. This sampling strategy can retain the semantic information of the network to the maximum extent, avoid the problem of over-enhancement of features, and ensure the richness and comprehensiveness of network features.

[0100] according to Figure 7 As shown, the HGN model structure diagram of the system of the present invention is as follows:

[0101] Compared with the traditional method of classifying a single node on a large-scale transaction network, the research of this invention classifies different types of node groups at the subgraph level based on the construction of a second-order subgraph. In the highly heterogeneous Ethereum network, the use of semantic-based motifs to construct subgraphs not only highlights the differences between nodes, but also amplifies the impact of these differences on the overall network. Therefore, the invention designs a model that conforms to the heterogeneity of the second-order subgraph to improve the detection performance.

[0102] HGN (Heterogeneous Graph Network) for heterogeneous networks not only reduces the complexity of the model, but also achieves good results on highly heterogeneous network graphs. The working principle of HGN is to use the adjacency matrix A and the feature matrix X separately, and then combine them with a multi-layer perceptron (MLP) and a linear transformation. This method is easy to train and evaluate in a small batch manner, and circumvents the problem of homogeneous GNN. Specifically, the model consists of two parts.

[0103] The first part focuses on exploring a model that does not consider the graph structure and only uses the feature information of the nodes, that is, using MLP. MLP classification based on node features ignores the topological structure of the graph. The second part focuses on exploring the LINK model that uses the graph topology, which is a method that only relies on graph structure information in node classification tasks. LINK trains a logistic regression model in which the features of each node are taken from the adjacency matrix A∈{0,1} n×n A column of W∈R c×n is a learnable weight matrix. With this setting, the LINK model can calculate the probability of a node belonging to each category. The specific formula is:

[0104] Y = softmax(WA)

[0105] For a specific node u∈{1,...,n} and a specific class k∈{1,...,c}, the logarithmic probability that node u belongs to class k can be calculated by the following formula:

[0106] (WA) ku =∑ v∈N(u) W kv

[0107] Where N( u (Including u's one-hop neighbors, the formula is based on the weights of node u's first-order neighbors and W kv If a particular node v has many neighbors belonging to category k, then W kv may be large, which means that there is a high probability that any of v's neighbors belongs to category k. The LINK model has shown good performance on node classification tasks including gender prediction in social networks. The key to its success lies in its role as a 2-hop (second-order) method. For example, in the gender prediction task, a person's direct friends (first-order neighbors) may not provide enough information to accurately predict the person's gender because the circle of friends may include people of various genders. However, if the friends of the person's friends (second-order neighbors) are considered, some gender-specific social patterns or preferences may be found, such as male users may be more connected to other male users, which is called the "single-sex preference" principle. This means that by observing a person's second-order social connections, more clues about their gender can be obtained. The LINK model is based on this approach to capture and utilize this information by considering the second-order neighbors of the node. This is highly consistent with the second-order subgraph formed by the Ethereum phishing user transaction data. By applying the LINK model to the semantically sampled subgraph of the present invention, the potential semantic relationships of strongly heterogeneous graph data can be fully mined and utilized while retaining the key information of the graph structure.

[0108] The two partial models are combined by linear transformation and. Let X∈R d×n represents the node feature matrix with input dimension d, and [h1;h2] represents the concatenation operation of vectors h1 and h2. Then the prediction result is obtained through the following process:

[0109] h A =MLP A (A)∈R d×n

[0110] h x =MLP x (X)∈R d×n

[0111] y=MLP f (σ(W[h A ;hX ]+h A +h X ))

[0112] where d is the hidden layer dimension, W∈R d×2d is the weight matrix, σ is the RELU activation function, and y is the predicted label result.

[0113] In the HGN model, we first calculate the hidden representation h of the adjacency matrix A and the hidden representation h of the feature matrix X . The model then transforms h A and h X The model also concatenates h and performs linear transformation to achieve the integration of these two hidden representations. In order to further strengthen the model's retention of topological structure and node features, the model also adopts h A and h X The strategy of direct addition is to finally classify the results through MLP.

[0114] The experimental scenario of the present invention is described in detail below, and the implementation results are analyzed in combination with the advantages of the present invention.

[0115] In the study of Ethereum transaction network, since nodes do not have any features in the initial state, in order to deeply analyze and understand the behavior of nodes in the network and their interactions, a series of statistical features must be artificially introduced in the early stage of the study. This paper proposes two key feature sets: transaction features and semantic motif features.

[0116] Transaction features are obtained through quantitative analysis based on the transaction activity data of the account. This includes but is not limited to statistical information such as the account's transaction frequency, transaction volume, capital inflow and outflow, and number of counterparties. These features capture the patterns and rules of the account at the transaction behavior level, and provide a basic quantitative basis for identifying and distinguishing different types of Ethereum accounts. The specific transaction features used in this invention are shown in Table 1.

[0117] Table 1 Transaction statistics characteristics of accounts

[0118]

[0119]

[0120] In addition to transaction statistics, semantic motif features focus on the structure and semantic information of the network, and also guide the subsequent sampling process. The present invention defines the frequency of 21 motifs appearing around nodes as motif features.

[0121] In the present invention, networkx is required to build the transaction network, and the implementation of the graph deep learning model is based on the pytorch framework. The experiment also involves various components. The environment and components required for the experiment are shown in Table 2.

[0122] Table 2 Experimental environment

[0123]

[0124] In terms of experimental settings, for the three comparative experiments and the EPSD-motif method proposed in the present invention, the present invention selects a ratio of 7:3 to divide the test set and the training set. The number of epochs for all models is 100. For the three groups of comparative experiments, the present invention needs to modify them from node-level classification to subgraph-level classification, and all select maximum pooling as the final representation of the subgraph. The parameters of the three groups of comparative experiments selected by the present invention are set as follows: in E-GCN, the input dimension d is set to 8, and the output node embedding dimension is 16; the number of leaves and the learning rate of the LightGBM model are fixed to 50 and 0.03, respectively. In Trans2Vec, the present invention sets the number of node walks r to 20, the walk length l to 5, the context size k to 10, p to 0.25, q to 0.75, and the search bias parameter to 0.5. In TTAGN, the present invention sets the attention hidden size to 2, the learning rate to 0.01, and sets two layers with a learning rate of 0.001 in the GCN module of TTAGN.

[0125] The experimental parameters of the EPSD-motif method proposed in the present invention are set as follows: in the process of sampling based on motif features, special attention is paid to the first five motifs with significant features. In the HGN model for heterogeneous networks, the initial dimension of each node is set to 32, including the statistical features and semantic motif features of the transaction, the learning rate is set to 0.001, the hidden layer dimension is set to 128, and the batch_size of the model is 64.

[0126] The evaluation indicators used in the present invention include accuracy, precision, recall and F1-score, and the value range is [0,1]. The closer to 1, the better the classification effect.

[0127] In order to evaluate the effectiveness of the proposed method, the present invention conducts account address recognition experiments on three data sets, and selects equal-proportion (1:1) balanced sample states and unbalanced sample states. The balanced sample state is to evaluate the effectiveness of the proposed method under ideal conditions, and the unbalanced sample state and multi-classification are to simulate the data imbalance in real scenarios.

[0128] The three comparative experiments selected by the present invention are as follows:

[0129] (1) E-GCN: A detection method based on graph convolutional networks and autoencoders. The transaction information of each node is counted as the node feature, and then the structural features are extracted using GCN and autoencoder technology.

[0130] (2) Trans2Vec: A biased sampling method based on transaction amount and timestamp, and using skip-gram to represent the features of the sampled nodes. This method assumes that the larger the transaction amount and the later the transaction time, the closer and more important the relationship between the two related nodes.

[0131] (3) TTAGN: A method that combines transaction, structural and statistical features to enhance Ethereum phishing detection performance. This method generates the final transaction features by modeling historical transaction time information and using an attention mechanism to capture similar transaction behaviors.

[0132] Table 3 Detection performance of phishing user nodes, ordinary account nodes, and ICO account wallet nodes under balanced data sets

[0133]

[0134] In order to fairly and accurately evaluate the effect of the method proposed in the present invention, in the balanced data set, the present invention chooses to reduce the fishing user nodes to achieve a 1:1 ratio with other label data (the results are shown in Table 3). The present invention first classifies the fishing user nodes and ordinary account nodes, and finds that due to the obvious differences in semantics and transaction patterns between fishing user nodes and ordinary user nodes, EPSD-motif can effectively extract the semantic information of the transaction, and achieves a good improvement effect on the four indicators. Then, a classification experiment is conducted on the data set composed of fishing user nodes and ICO wallet nodes to solve the situation in which fishing nodes are disguised as ICO nodes and are misclassified as illegal nodes with other similar transaction patterns in the real scene of Ethereum. The results show that EPSD-motif exceeds other comparative experiments in all indicators. The improvement of the effect achieved for different types of nodes further proves the generalization ability of the model considering semantics in the environment of processing multiple labels of the Ethereum network. Especially in terms of recall rate, it is shown that the model of the present invention can effectively identify most of the real fishing user nodes and reduce the situation of missed detection, which is very helpful for preventing fishing behavior while reducing the misclassification of fishing user nodes and identifying disguises.

[0135] Under the condition of an unbalanced data set, the high accuracy index no longer has a significant reference value, because even if the model only predicts most samples correctly, it may still obtain a high accuracy. Therefore, in this case, the present invention focuses on evaluating the recall rate and F1 value, and adopts the macro parameter (the results are shown in Table 4).

[0136] Table 4 Detection performance of phishing and Ponzi, ICO, gambling and multi-classification under unbalanced datasets

[0137]

[0138] Despite the imbalanced dataset, EPSD-motif still shows significant improvements in recall and F1-score. This shows that even in an unbalanced environment where there are more phishing user nodes and a relatively small number of other category nodes, EPSD-motif can still effectively identify and recall phishing user nodes. However, in this unbalanced data environment, the excess of phishing nodes may lead to incomplete expression of the model's semantic features, which affects the overall performance of the model. But this phenomenon also shows that EPSD-motif is effective in extracting semantic features, but when faced with incomplete semantic expressions, the model's recognition ability needs to be improved. This further verifies the ability of EPSD-motif in understanding the intrinsic semantics of transaction data.

[0139] Furthermore, the present invention also conducts a comparative analysis on the results of different sampling strategies.

[0140] Only using the enhanced sampling method will form subnetworks with highly similar semantics, which will lead to excessive semantic feature enhancement and cause overfitting problems, so the improvement effect will not be obvious. Figure 8 The comparison between using only enhanced sampling method and using a mixture of two sampling methods is shown. It can be found that the first-order and second-order neighbors retained by the enhanced sampling method are highly similar to the central node. At the same time, the sampling range of enhanced sampling is greatly reduced, and the extraction of effective information is relatively reduced. The mixed sampling method can retain rich semantic information, so the detection effect is the best.

[0141] Furthermore, the present invention also analyzes the semantic recognition capability of the model.

[0142] Fig. 9The comparison of semantic recognition capabilities of different methods is shown. EPSD-motif has a good detection effect on ordinary account and ICO wallet datasets. This is because phishing nodes show significant differences in transaction patterns and semantics from ordinary accounts, and ICO wallet addresses show unique transaction patterns and semantic information due to their financing functions, and are also different from illegal abnormal nodes such as phishing nodes. Therefore, the overall detection performance is good. This result is consistent with the expectation of the model's ability to extract semantic information: that is, the greater the semantic difference between nodes, the better the detection effect; when the semantic difference is small, the model's detection ability is relatively limited, although it is still better than other methods.

[0143] In real Ethereum transaction scenarios, phishing user nodes often disguise themselves as normal nodes with different semantics to evade monitoring by relevant departments, rather than disguise themselves as other abnormal user nodes. Therefore, the detection method for identifying semantic information proposed in the present invention can detect various nodes in Ethereum transactions to the greatest extent.

[0144] In contrast, E-GCN, Trans2Vec, and TTAGN do not show differences in detection performance when dealing with different types of nodes with obvious semantic differences. This indicates that these methods do not fully learn from the semantic level, thus affecting the overall performance.

[0145] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for classifying Ethereum transaction users based on semantic motif differentiation, characterized by: The specific steps include: Step 1: Dataset and transaction network construction; Collect all Ethereum transaction data and three types of addresses marked as phishing user nodes, ordinary user nodes, and ICO wallet nodes on the Etherscan and Xblock platforms, and build respective second-order subgraphs centered on these three types of nodes; Step 2: semantic motif construction and subgraph sampling based on motif features; First, the motif is reconstructed according to the transaction semantics. Then, the motif features of the nodes in the second-order subgraph obtained in step 1 are constructed according to the predefined semantic motifs, focusing on the motif features with higher frequency. Then, a sampling strategy combining semantic enhancement and complementarity is adopted. Starting from the central node, the first-order neighbors and second-order neighbors are sampled layer by layer to construct a second-order subgraph that contains semantics and has obvious differentiation. Step 3: Feature extraction and classification; The transaction subgraph containing semantic information obtained in step 2 is input into the HGN model designed specifically for the strong heterogeneity of Ethereum; the model learns the adjacency matrix and feature matrix of the transaction graph respectively to generate feature representations of the subgraph, including topological structure features and semantic features; finally, these feature vectors are sent to the multi-layer perceptron MLP for accurate subgraph-level classification.

2. The Ethereum transaction user classification method based on semantic motif differentiation according to claim 1 is characterized by: The step 1 is specifically as follows: Step 1.1: Collect all transaction records of Ethereum; use the blockchain academic platform Xblock to provide fully processed Ethereum transaction data, download all transaction data, and synchronize all transaction data to the server device through commands; for each transaction field, retain the transaction initiator from, transaction recipient to, transaction amount amount and transaction timestamp timestamp; Step 1.2: Collect account labels; download from the Label Cloud section through the API interface provided by the Xblock Academic Platform and the Etherscan website, and collect labels of phishing accounts, ordinary accounts with non-specific functions, and ICO wallet accounts; Step 1.

3. Construct a second-order transaction network; based on the obtained account address labels and transaction data on the Ethereum chain, retain all second-order transaction data related to the label address in a traversal form, forming a second-order subgraph centered on these three label addresses; Step 1.4: Redundant address screening; delete transactions with a transfer amount of 0; when obtaining the first-order neighbors of the target node, for exchange nodes with larger degree values, delete those nodes that are connected to the exchange node but not directly connected to the target node; then retain the exchange node.

3. The Ethereum transaction user classification method based on semantic motif differentiation according to claim 1 is characterized in that: The step 2 is specifically as follows: Step 2.1: Construct a motif with semantics; add the centrality of nodes to the motif, expand the number of motifs from 13 to 21, and the reconstructed 21 motifs contain rich transaction semantics; Step 2.2: Motif feature construction: For each node v in the network, analyze all subgraphs G consisting of three nodes containing node v v , and identify the subgraphs that match the predefined motif M= { M1,M2,...,M 21} Matching part; for each node v, calculate the frequency t of each motif appearing around node v v,i : in, | g∈G v :g∈M i| Represents a specific motif M i The number of occurrences of | G v| Represents the number of subgraphs consisting of three nodes of node v; The motif feature vector of node v is represented by T v : T v =(t v,1 ,t v,2 ,...,t v,M ) Among them, each element t v,i represents the frequency of a specific motif i in the first-order neighbor range of node v, | M | There are 21 types of semantic motifs; Step 2.3: Subgraph sampling based on motif features; After calculating the number distribution of motif features for each node, the sampling process is carried out based on these motif features.

4. The Ethereum transaction user classification method based on semantic motif differentiation according to claim 3 is characterized by: The motif corresponding to the larger value in the feature vector of the motif in step 2.3 is the most common motif of each central node, which reflects the importance of a certain motif in the interaction mode and semantic information of the node; this motif recognition method gives priority to those nodes that exhibit similar or identical high-frequency motif features as the central node, that is, nodes with highly similar semantic information; it realizes targeted sampling of the second-order subgraph, so that the transaction semantics of the entire subnet are as close to the semantics of the central node as possible, thereby achieving a semantic enhancement effect.

5. The Ethereum transaction user classification method based on semantic motif differentiation according to claim 1 is characterized in that: The HGN model in step 3 simultaneously learns to extract topological relationships from the adjacency matrix and semantic information from the node feature matrix, and strengthens the retention and synergy of the two features through a fusion mechanism of splicing, linear transformation and direct addition, so that the model can comprehensively capture the diversity information of complex networks; the core of HGN is composed of two modules, MLP and LINK: the MLP module focuses on extracting semantic information from node features, and the LINK module focuses on extracting graph structure information from the adjacency matrix. The model adaptively adjusts the weights of the two modules according to the characteristics of the input data, so as to adapt to a variety of heterogeneous graph data scenarios; HGN inputs the generated comprehensive features into the multi-layer perceptron MLP for refined classification, which significantly improves the classification accuracy of heterogeneous graph data.

6. The Ethereum transaction user classification method based on semantic motif differentiation according to claim 1 is characterized by: The step 3 is specifically as follows: Step 3.1: The first part of the model does not consider the graph structure, that is, the MLP model that only requires node feature information is used; h A =MLP A (A)∈R d×n h x =MLP x (X)∈R d×n Among them, h A Indicates that through MLP A The hidden representation extracted from the adjacency matrix A; h X Indicates that through MLP X Hidden representations extracted from the node feature matrix X; Step 3.2: A logistic regression model is trained using the LINK model of the graph topology, where the features of each node are taken from the adjacency matrix A∈{0,1} n×n A column of W∈R c×n is a learnable weight matrix; with this setting, the LINK model can calculate the probability of a node belonging to each category, the specific formula is: Y=soft max(WA) For a specific node u∈{1,...,n} and a specific class k∈{1,...,c}, the logarithmic probability that node u belongs to class k is calculated by the following formula: (WA) ku = ∑ v∈N(u) W kv Among them, N (u) Contains the one-hop neighbors of u. The formula is based on the weights of the first-order neighbors of node u and W kv ; If a particular node v has many neighbors belonging to category k, then W kv It may be very large, which means that there is a high probability that any neighbor of v belongs to category k; Combine these two partial models through linear transformation; let X∈R d×n represents the node feature matrix with input dimension d, and [h1;h2] represents the concatenation operation of the vector h1 generated by MLP and the vector h2 generated by LINK; then the prediction result is obtained through the following process: y=MLP f (σ(W[h1;h2]+h1+h2))。 7. The Ethereum transaction user classification method based on semantic motif differentiation according to claim 6 is characterized by: In the HGN model, the hidden representation h1 of the feature matrix and the hidden representation h2 of the adjacency matrix are first calculated; then, the model combines h1 and h2 by concatenating them and performing a linear transformation; in order to further enhance the model's retention of topological structure and node features, the model also adopts the strategy of directly adding h1 and h2, and finally classifies the results through MLP.

8. An Ethereum transaction user classification system based on semantic motif differentiation according to any one of claims 1 to 7, characterized in that: It includes a transaction data collection and processing device, a transaction network construction device, a semantic motif construction device, a subgraph sampling device based on motif features, and a feature extraction and classification device; The main task of the transaction data collection and processing device is to collect all transaction records of Ethereum as well as fishing user nodes, ordinary users and ICO wallet tags; The main task of the transaction network construction device is to traverse the transaction records of Ethereum, construct a second-order transaction subgraph with the central node as the starting node, and perform secondary filtering to reduce the interference of irrelevant information; The semantic motif construction device reconstructs a motif with semantic information based on the transaction purpose of the Ethereum transaction user; The subgraph sampling device based on motif features constructs role types according to the motif distribution frequency in the network to guide the sampling of subgraphs; The feature extraction and classification device inputs the transaction subgraph containing semantic information into the HGN model to generate feature representation of the subgraph; finally, the basic features are combined with these feature vectors and sent to the multi-layer perceptron MLP for accurate subgraph level classification.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Ethereum malicious sample detection method and system for homogeneous enhanced compression modeling

    CN120546989A

  • A method and system for detecting malicious Ethereum samples based on homogeneous enhanced compression modeling

    CN120546989B

  • Phishing account detection model training method, detection method and device

    CN121365975A

  • A phishing account detection model training method, detection method and device

    CN121365975B

  • Block chain transaction address role identification method and device

    CN121544388A