Pump and Dump transaction detection method, equipment and product based on time series behavior graph

By constructing a transaction graph for the ERC-20 token market and utilizing a time-series behavior graph detection method based on smart contracts and graph neural networks, we can identify and predict pump and dump transactions, addressing the issue of market manipulation and improving market transparency and stability.

CN119205113BActive Publication Date: 2025-09-12WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411081632.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2025-09-12
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

The lack of effective regulation and transparency in the ERC-20 token market has led to frequent "pump and dump" phenomena, allowing manipulators to quickly profit and investors to face huge risks, which has affected market confidence and stability.

Method used

A time-series behavior graph detection method based on smart contracts and graph neural networks is adopted. By constructing a transaction graph, account and transaction features are extracted, and transaction behaviors are learned using a memory mechanism and graph neural network to identify pump and dump transactions.

Benefits of technology

Effectively detecting pump and dump transactions improves the credibility of the ERC-20 token market, enabling timely identification of abnormal transactions, avoiding market manipulation, and improving market transparency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119205113B_ABST
    Figure CN119205113B_ABST
Patent Text Reader

Abstract

The present invention discloses a Pump and Dump detection method, device and product based on a time-series behavior graph. First, the original transaction data is constructed into a time-series behavior graph G(N,E); wherein N is the set of accounts in the graph G; E is the set of transaction information in the graph G. Then, the statistical features of each account and each transaction are extracted. Then, by aggregating the account's own transaction information and the features of multi-order neighbors, the real-time account embedding is calculated, and the account embeddings, account statistical features and transaction statistical features of both parties to the transaction are spliced ​​to obtain the real-time embedding of each transaction. Subsequently, two sets of historical normal transactions and Pump and Dump transactions are constructed respectively, and the distinguishability of the two sets is enhanced through comparative learning. Finally, the enhanced real-time embedding is used to predict whether each transaction is abnormal. The present invention can detect whether each transaction is abnormal and, to a certain extent, solve the "pseudo-loop" problem caused by one user controlling multiple accounts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of Ethereum technology and deep learning technology, and relates to a pump and dump transaction detection method, device and product based on a time series behavior graph, and specifically to a pump and dump transaction detection method, device and product based on a smart contract and a graph neural network. Background Art

[0002] ERC-20 tokens are standardized tokens implemented on the Ethereum blockchain through smart contracts. These smart contracts, running on the Ethereum Virtual Machine (EVM), define and implement the ERC-20 standard interface, including functions such as total supply, account balance query, token transfer, and authorization. By adhering to the ERC-20 standard, these smart contracts ensure consistent behavior across all tokens, enabling seamless interoperability across various decentralized applications (DApps), wallets, and exchanges. The Ethereum blockchain provides a decentralized, secure, and transparent infrastructure to support the creation and execution of these smart contracts.

[0003] Due to the lack of effective supervision and transparency, the ERC-20 token market has experienced a "pump and dump" phenomenon. Manipulators quickly profited by manipulating token prices, causing investors to face huge risks and losses in a short period of time, seriously affecting investor confidence and market stability, while also damaging the credibility and transparency of the overall market. Summary of the Invention

[0004] In order to detect the existence of "Pump and Dump" in the ERC-20 token market, the present invention uses smart contract technology and graph neural network technology to provide a pump and dump transaction detection method, device and product based on time-series behavior graph.

[0005] The technical solution adopted by the method of the present invention is: a pump and dump transaction detection method based on a time sequence behavior diagram, comprising the following steps:

[0006] Step 1: Construct the raw transaction data into a time-series behavior graph G(N,E); where N is the set of accounts in graph G, used to store the accounts in the transaction; E is the set of transaction information in graph G, used to store the transaction relationship between accounts and transaction information at the time of the transaction, including token type, transaction amount, transaction time, and token precision;

[0007] Step 2: Extract the statistical features α of each account i from the time series behavior graph Gi , the statistical characteristics of each transaction eβ e , and behavioral characteristics of each transaction e , and store these features in the transaction graph G;

[0008] Step 3: By aggregating the transaction behavior characteristics of each account and the characteristics of its multi-order neighbors, a real-time account embedding emb is calculated for both parties of each transaction e∈E. n (n∈N), and embed emb by splicing the source account src 、Target account embedded emb dst , Source account statistical characteristics α src , Target account statistical characteristics α dst and transaction statistics β e Get real-time embeds for each transaction e ;

[0009] Step 4: Construct historical normal transactions and pump and dump transactions into two sets respectively. By calculating the Euclidean distance of the real-time embeddings within the two sets and the Euclidean distance between the two sets, the Euclidean distance within the two sets is minimized and the Euclidean distance between the sets is maximized. This maximizes the similarity within the sets and the difference between the sets, and enhances the distinguishability of the real-time embeddings generated by the time series behavior graph G.

[0010] Step 5: The real-time embedding obtained after the enhancement will be used to predict whether each transaction is abnormal or not

[0011] As a preference, in step 2, the statistical feature α of each account i ,include:

[0012] closeness_centrality, which measures the inverse of the average distance between an account and other accounts, reflecting the proximity of the account in the network;

[0013] betweenness_centrality, which measures the importance of an account as a bridge in the network, i.e., how often the account appears in the shortest path;

[0014] indegree_centrality, which measures the degree to which the target account receives the number of connections in the directed network;

[0015] outdegree_centrality, which measures the degree to which the target account sends the number of connections in the directed network;

[0016] degree_centrality, which is the degree centrality of the account, i.e. the number of connections of the account;

[0017] eigenvector_centrality, which is the centrality of an account related to the centrality of its neighboring accounts, emphasizes the connection between an account and highly central accounts;

[0018] PageRank is used to measure the importance of a target account by calculating its stable state probability;

[0019] tx_per_account, which is the number of target account's trading accounts divided by the number of target account's transactions, is used to measure the degree of diversity of the target account's trading accounts;

[0020] In step 2, the statistical feature β of each transaction e ,include:

[0021] counts is the number of historical transactions between the two parties within the detection window up to the time of the transaction;

[0022] Mean is the historical transaction mean of both parties within the detection window up to the transaction;

[0023] var is the historical transaction variance between the two parties within the detection window up to the time of the transaction;

[0024] In step 2, the behavioral characteristics of each transaction ξ e ,include:

[0025] time, the time when the transaction occurred;

[0026] value, the amount of the transaction;

[0027] token_decimals, which is the precision of the tokens involved in the transaction;

[0028] token is the token contract address involved in this transaction.

[0029] As a preference, in step 3, a layer of real-time embedding z for account i at time t is generated. i The specific implementation of (t) includes the following sub-steps:

[0030] Step 3.1: Account feature transformation;

[0031] X′=XW n ;

[0032] in, is the account memory matrix, N is the number of accounts, F n is the dimension of account memory, is the linear transformation weight matrix of the account features, F′ is the feature dimension output by each attention head, and H is the number of attention heads;

[0033] Step 3.2: Transaction feature splicing;

[0034] The behavioral characteristics of all transactions are constructed as a matrix ξ and compared with the time difference φ(tt - ) together to obtain the behavioral characteristic matrix E of all transactions:

[0035] E=concat(ξ,φ(tt - ));

[0036] Among them, t is the timestamp of each transaction, t - The time when the memory of the source account of each transaction was last updated;

[0037] Step 3.3: Transaction feature transformation;

[0038] E′=EW e ;

[0039] in, is the behavioral feature matrix of transactions, M is the number of transactions, F e is the behavioral characteristic dimension of the transaction, is the linear transformation weight matrix of transaction features;

[0040] Step 3.4: Calculate the attention score e based on the transformed account features and transaction features;

[0041] e=LeakyReLU(X′a l +X′a r +E′a e );

[0042] Among them, a l ,a r ,a e ∈R F′×H×1 is the attention parameter matrix;

[0043] Step 3.5: Use the softmax function to normalize the attention scores to obtain the attention weight a of each account and its neighbor accounts;

[0044]

[0045] in, represents the set of neighbor accounts of account i, e k represents the attention score of account i’s neighbor account k;

[0046] Step 3.6: Use the calculated attention weights to perform weighted summation on the features of neighbor accounts to achieve message passing and feature aggregation, and obtain the real-time embedding of the account of one of the attention heads.

[0047]

[0048] Among them, h∈(1,…,H) is the serial number of the attention head, X j ′ represents the account feature transformation matrix of the neighboring accounts of account i;

[0049] Step 3.7: Concatenate the output features of all heads to get the real-time embedding of account i Here, H is the number of attention heads.

[0050] Preferably, in step 4, account embedding and transaction embedding are generated by adopting neighbor account sampling, time encoder, message mechanism, and memory mechanism, and the historical transaction behavior of the source account src and the target account dst is learned to enhance the detection capability of pump and dump transaction behavior;

[0051] The time encoder encodes time t using a finite Fourier series:

[0052] φ(t)=[cos(w1t+ψ1),cos(w2t+ψ2),…,cos(w η t+ψ η )];

[0053] Where η is the dimension of the Fourier series after encoding time t; weight ω j and bias ψ j is a learnable parameter, j∈{1,…,η};

[0054] The message mechanism, at any time t, for an interaction event e(t) involving the source account src and the target account dst, will generate a message msg(t) to update the memory of the source account src and the target account dst:

[0055] msg(t)=concat(m src (t - ),m dst (t - ),e(t),φ(t));

[0056] Among them, m src (t - ) is the account memory of the source account src before the update at time t, m dst (t - ) is the account memory of the target account dst before it is updated at time t, and e(t) is the behavioral characteristic of the transaction ξ E ; Update the memory of the source account src and the destination account dst using the same message msg(t);

[0057] For t1,…,t b ≤t, account i uses the aggregation mechanism to aggregate message msg i (t1),…,msg i (t b ):

[0058]

[0059] The memory mechanism mentioned above updates the memory of any account i after an event involving account i occurs. For an interaction event e(t) between a source account src and a target account dst, the memory of the source account src and the target account dst is updated as follows:

[0060]

[0061] Among them, mem is a learnable memory update function that can maintain the order information of the sequence while capturing the temporal dependencies in the sequence data, thereby identifying and understanding the dynamic change patterns of the input data over time.

[0062] Preferably, the neighbor account sampling first generates a generated subgraph g′ based on the sampling account set seed_nodes in the given graph to be sampled g, and then filters out transactions that do not meet the conditions based on the timestamp filter ts; then, based on the setting of the number of neighbors to be sampled for each account in each layer of the graph neural network, fanouts, the size of the sampling neighborhood is determined, and the corresponding frontier subgraph list frontier is returned;

[0063] The specific implementation includes the following sub-steps:

[0064] (1) The number of neighbors to be sampled for each account in each layer of the graph neural network, fanouts, is assigned to the sampling frontier list frontiers, indicating the number of neighbors to be sampled;

[0065] (2) The sampled frontier list frontiers is set to empty;

[0066] (3) Traverse fanouts and determine whether the timestamp of the transaction is less than the timestamp filter ts;

[0067] If so, continue to check whether the frontiers are full. If the frontiers are not full, add the transaction to the frontiers; if the frontiers are full, keep the latest frontiers;

[0068] If not, discard the transaction.

[0069] As a preference, for the interaction event e(t) between the source account src and the target account dst, after obtaining the real-time embedding z of the source account src and the target account dst src (t) and z dst (t), the real-time embedding h of the interactive event e(t) is obtained as follows e (t):

[0070] h e (t) = concat(z src (t),z dst (t),β e ,α src ,α dst );

[0071] Among them, β e is the statistical feature of the transaction between the source account src and the target account dst, α src and α dst It is the statistical characteristics of the source account src and the destination account dst at the account level.

[0072] Preferably, in step 4, the contrast loss function used in contrastive learning is:

[0073]

[0074] in, Calculate the positive sample set S P Internal sample h i 、h j The average of the sum of squared Euclidean distances between them; Calculate the negative sample set S N Internal sample h i 、h j The average of the sum of squared Euclidean distances between them; Calculate the distance difference between positive samples and negative samples, and set a threshold margin. If the distance between positive samples and negative samples is less than this threshold, a positive loss is generated to distinguish the two types of samples.

[0075] The technical solution adopted by the device of the present invention is: a Pump and Dump transaction detection device based on a timing behavior diagram, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the Pump and Dump transaction detection method based on the timing behavior diagram is implemented.

[0076] The technical solution adopted by the product of the present invention is: a Pump and Dump transaction detection product based on a timing behavior diagram, including a computer program, which implements the Pump and Dump transaction detection method based on the timing behavior diagram when executed by a processor.

[0077] Compared with the prior art, the beneficial effects of the present invention include:

[0078] (1) This paper proposes a transaction-based temporal behavior graph to detect pump and dump behavior in the ERC-20 token market, extending the detection of pump and dump behavior in traditional centralized exchanges off-chain to the blockchain.

[0079] (2) The present invention uses the token type as one of the parameters of the model, constructs the transactions of different tokens into graphs respectively, and selects the native token of each token as its value scale, which effectively avoids the interference of inconsistent precision and unstable exchange rates between different tokens on the model.

[0080] (3) This invention introduces a memory mechanism that can learn and record the historical transaction behavior of each ERC-20 account, and can keenly capture suspicious transactions of a certain account. It can further refine the detection granularity of pump and dump events from "one event occurred" to "one of the transactions", which is of great significance for maintaining the credit of the ERC-20 token market.

[0081] (4) This paper proposes a pump and dump detection model based on a transaction time-series behavior graph. This model combines graph neural networks with transaction time information to effectively capture and predict the evolution of edges in the graph. This model can promptly detect pump and dump behaviors that are both sequential and sudden in time. Furthermore, it can, to a certain extent, address the "pseudo-loop" problem caused by a single user controlling multiple accounts. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The technical solution of the present invention is further illustrated below using embodiments and specific implementation methods. In addition, some drawings are also used in the process of illustrating the technical solution. For those skilled in the art, other drawings and the intention of the present invention can be obtained based on these drawings without making any creative efforts.

[0083] Figure 1 A schematic diagram of a method according to an embodiment of the present invention;

[0084] Figure 2 This is a schematic diagram of a neighbor account sampling mechanism according to an embodiment of the present invention;

[0085] Figure 3 Schematic diagram of a message mechanism according to an embodiment of the present invention;

[0086] Figure 4 This is a schematic diagram of an account memory mechanism according to an embodiment of the present invention;

[0087] Figure 5 A schematic diagram of generating account embedding according to an embodiment of the present invention. DETAILED DESCRIPTION

[0088] In order to facilitate ordinary technicians in this field to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0089] This embodiment detects pump and dump behavior in the ERC-20 token market based on the time series behavior graph of transactions. First, the ERC-20 token transactions are constructed into a graph, and the statistical characteristics of the trading accounts and their transactions are collected on this basis. Then, the trading behavior of the account will be used to generate the embedding of each transaction. For a transaction e at time t, the n-order frontier subgraph of the source account src and the target account dst before time t is filtered out respectively, and the messages are aggregated layer by layer from the n-order neighbors towards the source account src and the target account dst, and the memory is updated to generate the embedding emb of the source account src and the target account dst. src and emb dst , and splice the source account statistical features α src , Target account statistical characteristics α dst and transaction statistics β e Get real-time embeds for each transaction e Subsequently, by comparing Pump and Dump transactions with normal transactions, the distinction between the two transactions is expanded and the embedding representation of both is enhanced. Finally, the LightGBM model is used to distinguish Pump and Dump transactions from normal transactions.

[0090] Please see Figure 1 This embodiment provides a method for detecting pump and dump transactions based on a time series behavior diagram, including the following steps:

[0091] Step 1: Construct the raw transaction data into a time-series behavior graph G(N,E); where N is the set of accounts in graph G, used to store the accounts in the transaction; E is the set of transaction information in graph G, used to store the transaction relationship between accounts and transaction information at the time of the transaction, including token type, transaction amount, transaction time, and token precision;

[0092] In one embodiment, the DGL library is used to construct raw transaction data into a graph. Considering that new tokens are constantly appearing in the ERC-20 token market every day, and different tokens have different trading characteristics in the market (for example, some tokens have higher precision and daily trading volume), transactions for each token are constructed into a separate graph. In different transaction graphs, the same ERC-20 account is marked as a different account, thus completely isolating the transaction graphs composed of different tokens from each other.

[0093] Step 2: Extract the statistical features α of each account i from the time series behavior graph G i , the statistical characteristics of each transaction eβ e , and behavioral characteristics of each transaction e , and store these features in the transaction graph G;

[0094] In one embodiment,

[0095] By collecting the statistical characteristics and transaction behaviors of accounts and transactions in the transaction graph, the embedding capability of each transaction is enhanced.

[0096] Among them, the statistical feature α of each account i include:

[0097] closeness_centrality: measures the inverse of the average distance between an account and other accounts, reflecting the proximity of the account in the network.

[0098] betweenness_centrality: measures the importance of an account as a bridge in the network, that is, the frequency with which the account appears in the shortest path.

[0099] indegree_centrality: A measure of how much the target account receives connections in a directed network.

[0100] outdegree_centrality: A measure of how far the target account outsources the number of connections in the directed network.

[0101] degree_centrality: The degree centrality of the account, that is, the number of connections of the account.

[0102] eigenvector_centrality: The centrality of an account is related to the centrality of its neighboring accounts, emphasizing the connection between an account and highly central accounts.

[0103] PageRank: measures the importance of a target account by calculating its stable state probability.

[0104] tx_per_account: The number of target account's trading accounts / the number of target account's transactions, used to measure the degree of diversity of the target account's trading accounts.

[0105] The statistical characteristics of each transaction β e ,include:

[0106] counts: The number of historical transactions between the two parties within the detection window up to the time of the transaction.

[0107] mean: The average of the historical transactions between the two parties within the detection window up to the transaction.

[0108] var: The historical transaction variance between the two parties within the detection window up to the transaction time.

[0109] The behavioral characteristics of each transaction e ,include:

[0110] time: The time when the transaction occurred.

[0111] value: The amount of the transaction.

[0112] token_decimals: The precision of the tokens involved in the transaction.

[0113] token: The token contract address involved in the transaction.

[0114] Among them, the behavioral features at the transaction level will be used for embedded learning based on the transaction-based temporal behavior graph.

[0115] Step 3: By aggregating the transaction behavior characteristics of each account and the characteristics of its multi-order neighbors, a real-time account embedding emb is calculated for both parties of each transaction e∈E. n (n∈N), and embed emb by splicing the source account src 、Target account embedded emb dst , Source account statistical characteristics α src , Target account statistical characteristics α dst and transaction statistics β e Get real-time embeds for each transaction e ;

[0116] In one embodiment, by adopting neighbor account sampling, time encoder, message mechanism, and memory mechanism, account embedding and transaction embedding are generated, and the historical transaction behavior of source account src and target account dst is learned to enhance the detection ability of pump and dump transaction behavior;

[0117] In one embodiment, the neighbor account sampling, graph neural network faces the problem of computational complexity and memory consumption when processing large-scale graph data. In order to cope with these challenges, sampling is usually required. Figure 2 As shown in the figure, it first generates a generated subgraph g′ based on the given sampling account set seed_nodes in the graph to be sampled g, and then filters out transactions that do not meet the conditions based on the timestamp filter ts; then, based on the setting of the number of neighbors to be sampled for each account in each layer of the graph neural network, fanouts, it determines the size of the sampling neighborhood and returns the corresponding frontier subgraph list frontier;

[0118] The following is a formal definition of neighbor account sampling:

[0119] The graph to be sampled g

[0120] The set of sampled accounts in the graph to be sampled, seed_nodes

[0121] The number of neighbors to sample for each account in each layer of the graph neural network, fanouts

[0122] Timestamp filter ts for sampling

[0123] List of sampled frontiers frontiers

[0124] The specific algorithm is shown in Algorithm 1:

[0125]

[0126] In one embodiment, the time encoder encodes time t using a finite Fourier series:

[0127] φ(t)=[cos(w1t+ψ1),cos(w2t+ψ2),…,cos(w η t+ψ η )];

[0128] Where η is the dimension of the Fourier series after encoding time t; weight ω j and bias ψ j is a learnable parameter, j∈{1,…,η};

[0129] In one embodiment, the messaging mechanism is a method for transmitting information within a dynamic graph, used to process the temporal evolution of accounts and transactions. This mechanism allows the model to efficiently transmit and aggregate information while taking into account the structure and temporal changes of the graph, thereby improving the modeling capabilities and predictive performance of dynamic graph data.

[0130] At any time t, for an interaction event e(t) involving the source account src and the target account dst, a message msg(t) will be generated to update the memory of the source account src and the target account dst:

[0131] msg(t)=concat(m src (t - ),m dst (t - ),e(t),φ(t));

[0132] Among them, m src (t - ) is the account memory of the source account src before the update at time t, m dst (t - ) is the account memory of the target account dst before it is updated at time t, and e(t) is the behavioral characteristic of the transaction ξ E Use the same message msg(t) to update the memory for both the source account src and the destination account dst. This approach can record the direction of the interaction event and reduce the extra time overhead of separately calculating the source account src and the destination account dst messages.

[0133] like Figure 3 As shown, for efficiency reasons, batch processing may result in multiple events involving the same account i in the same batch. Since the model generates a message for each event, for t1,…,t b ≤tUse aggregation mechanism to aggregate message msg i (t1),…,msg i (t b ):

[0134]

[0135] Among them, the max method is called "latest retention", which means retaining the latest message of a given account i.

[0136] In one embodiment, Figure 4 As shown, the memory mechanism records the account memory m of all accounts i encountered up to time t i (t). The memory of any account i is updated after an event involving account i occurs (such as an interaction with another account or a change to the account itself). This module memorizes the long-term behavioral habits of each account in the graph and can be viewed as a compressed representation of the account's historical behavior. When encountering a new account, its memory is initialized to a zero vector and updated for each event involving that account during training and testing.

[0137] The memory of any account i will be updated after an event involving account i occurs; for an interaction event e(t) between a source account src and a target account dst, the memory of the source account src and the target account dst will be updated as follows after the event occurs:

[0138]

[0139] Among them, mem is a learnable memory update function that can maintain the order information of the sequence while capturing the temporal dependencies in the sequence data, thereby identifying and understanding the dynamic change patterns of the input data over time.

[0140] In one embodiment, the generating of account embedding is used to generate a real-time embedding z of account i at any time t. i (t). The embedding module can avoid the so-called “memory aging” problem, that is, since the memory of account i is only updated when the account is involved in an interaction event, the memory of account i will become stale when there are no events for a long time.

[0141] The account embedding module introduces an attention mechanism to more effectively integrate the features of accounts and transactions. Figure 5 As shown in Figure 2, it receives a graph's account memory matrix and transaction feature matrix as input, aggregates neighboring account features, and performs feature transformation to generate a new feature representation for each account. The output is the convolutional and transformed account feature matrix, which preserves the graph structure and enhances the expressiveness of account features.

[0142] Generate a real-time embedding z for account i at time t using the following method: i (t):

[0143] Step 3.1: Account feature transformation;

[0144] X′=XW n ;

[0145] in, is the account memory matrix, N is the number of accounts, F n is the dimension of account memory, is the linear transformation weight matrix of the account features, F′ is the feature dimension output by each attention head, and H is the number of attention heads;

[0146] Step 3.2: Transaction feature splicing;

[0147] The behavioral characteristics of all transactions are constructed as a matrix ξ and compared with the time difference φ(tt - ) together to obtain the behavioral characteristic matrix E of all transactions:

[0148] e=concat(ξ,φ(tt - ));

[0149] Among them, t is the timestamp of each transaction, t - The time when the memory of the source account of each transaction was last updated;

[0150] Step 3.3: Transaction feature transformation;

[0151] E′=EW e ;

[0152] in, is the behavioral feature matrix of transactions, M is the number of transactions, F e is the behavioral characteristic dimension of the transaction, is the linear transformation weight matrix of transaction features;

[0153] Step 3.4: Calculate the attention score e based on the transformed account features and transaction features;

[0154] e=LeakyReLU(X′a l +X′a r +E′a e );

[0155] Among them, a l ,a r ,a e ∈R F′×H×1 is the attention parameter matrix;

[0156] Step 3.5: Use the softmax function to normalize the attention scores to obtain the attention weight a of each account and its neighbor accounts;

[0157]

[0158] in, represents the set of neighbor accounts of account i, e k represents the attention score of account i’s neighbor account k;

[0159] Step 3.6: Use the calculated attention weights to perform weighted summation on the features of neighbor accounts to achieve message passing and feature aggregation, and obtain the real-time embedding of the account of one of the attention heads.

[0160]

[0161] Among them, h∈(1,…,H) is the serial number of the attention head, X j ′ represents the account feature transformation matrix of the neighboring accounts of account i;

[0162] Step 3.7: Multi-head aggregation

[0163]

[0164] Where H is the number of attention heads. The output features of all heads are concatenated to obtain the real-time embedding of account i.

[0165] In one embodiment, the transaction embedding is generated, for the interaction event e(t) between the source account src and the target account dst, after obtaining the real-time embedding z of the source account src and the target account dst. src (t) and z dst (t), the real-time embedding h of the interactive event e(t) is obtained as follows e (t):

[0166] h e (t) = concat(z src (t),z dst (t),β e ,α src ,α dst );

[0167] Among them, β e is the statistical feature of the transaction between the source account src and the target account dst, α src and α dst It is the statistical characteristics of the source account src and the destination account dst at the account level.

[0168] Step 4: Construct historical normal transactions and pump and dump transactions into two sets respectively. By calculating the Euclidean distance of the real-time embeddings within the two sets and the Euclidean distance between the two sets, the Euclidean distance within the two sets is minimized and the Euclidean distance between the sets is maximized. This maximizes the similarity within the sets and the difference between the sets, and enhances the distinguishability of the real-time embeddings generated by the time series behavior graph G.

[0169] In one embodiment, the contrast loss function used in the contrastive learning is:

[0170]

[0171] in, Calculate the positive sample set S P Internal sample h i 、h j The average of the sum of squared Euclidean distances between them; Calculate the negative sample set S N Internal sample h i 、h j The average of the sum of squared Euclidean distances between them; Calculate the distance difference between positive samples and negative samples, and set a threshold margin. If the distance between positive samples and negative samples is less than this threshold, a positive loss is generated to distinguish the two types of samples.

[0172] Step 5: The real-time embedding obtained after the enhancement will be used to predict whether each transaction is abnormal or not.

[0173] In one implementation, a LightGBM model is used as a downstream classifier. LightGBM is an efficient decision tree-based gradient boosting framework designed for fast training and low memory consumption. It was developed by Microsoft and excels at handling large-scale datasets and high-dimensional data.

[0174] This embodiment also provides a Pump and Dump transaction detection device based on a timing behavior graph, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the Pump and Dump transaction detection method based on the timing behavior graph is implemented.

[0175] This embodiment further provides a Pump and Dump transaction detection product based on a timing behavior graph, including a computer program. When the computer program is executed by a processor, it implements the Pump and Dump transaction detection method based on a timing behavior graph.

[0176] The present invention is further described below through specific experiments.

[0177] The dataset used in this example covers the period from December 1, 2022, to December 31, 2022, and includes 858 tokens and 924,508 transactions, of which 799,518 are normal transactions, accounting for 86.48%; and 124,990 are pump and dump transactions, accounting for 13.52%.

[0178] The parameters used in this embodiment are given below:

[0179] Training set ratio: 70%;

[0180] Validation set ratio: 15%;

[0181] Test set ratio: 15%;

[0182] Training epochs: 30;

[0183] Batch size: 512;

[0184] Optimizer: Adam;

[0185] Learning rate: 0.0001;

[0186] Account embedding dimension: 128;

[0187] Account memory dimension; 128;

[0188] Time encoding dimension: 128;

[0189] Transaction embedding dimension: 275;

[0190] Account memory update method: GRU;

[0191] Message aggregation method: latest update;

[0192] Neighbor sampling set size: 10;

[0193] Neighbor sampling method: latest sampling;

[0194] Number of attention heads: 4;

[0195] Multi-stage neighbor sampling hop count: 2;

[0196] Contrastive learning margin: 0.5;

[0197] LightGBM leaf count: 127;

[0198] LightGBM learning rate: 0.1;

[0199] The present invention's detection of Pump and Dump transactions includes the following modules:

[0200] Module 1 constructs the raw transaction data into a time-series behavior graph G(N,E); where N is the set of accounts in graph G, used to store the accounts in the transaction; E is the set of transaction information in graph G, used to store the transaction relationship between accounts and transaction information at the time of the transaction, including token type, transaction amount, transaction time, and token precision;

[0201] Module 2: Extract the statistical features of each account from the time series behavior graph G N Statistical features of each transaction E , and store these features in the transaction graph G;

[0202] Module 3: In each round of training, for a transaction e at time t, with the source account src and the target account dst as the source points, the transactions before time t are filtered out and the second-order frontier subgraphs of the source account src and the target account dst are generated. Then, starting from the second-order neighbors, messages are aggregated layer by layer towards the source account src and the target account dst, the account memory is updated, and the account embedding module is used to generate the embedding emb of the source account src and the target account dst.src and emb dst , and splice the source account statistical features α src , Target account statistical characteristics α dst and transaction statistics β e Get real-time embeds for each transaction e .

[0203] Module 4 constructs historical normal transactions and pump and dump transactions into two sets respectively. By calculating the Euclidean distance of the real-time embeddings within the two sets and the Euclidean distance between the two sets, the Euclidean distance within the two sets is minimized and the Euclidean distance between the sets is maximized. This maximizes the similarity within the sets and the difference between the sets, strengthens the distinguishability of the real-time embeddings generated by the time series behavior graph G, and records the loss value generated by the comparative learning module in this round.

[0204] Module 5 compares the loss value of this training with the loss value of the previous training. If the loss value of this training is reduced, the parameters of each module of this training are saved; otherwise, the parameters of each module of the previous training are still saved. If the training is not completed after 30 rounds, continue to execute Module 3 and transfer the gradient back according to the loss value to optimize the parameters of each module; otherwise, end the training and execute Module 6;

[0205] Module 6 uses the enhanced real-time embedding to train the LightGBM model, implements the "early stopping" training strategy, and retains the optimal training results of the LightGBM model.

[0206] Module 7 executes modules 1-6 on the test set to generate prediction results for each transaction, and uses FPR, FNR, and BAC values ​​to measure the detection effect of the present invention on Pump and Dump transactions.

[0207] The prediction of the data set in this embodiment achieved FPR=0.077, FNR=0.077, and BAC=0.927.

[0208] It should be understood that the embodiments described above are only some of the embodiments of the present invention, rather than all of the embodiments. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention may be arbitrarily combined with each other to form a feasible technical solution. Such combination is not restricted by the order of steps and / or structural composition mode, but must be based on the ability of ordinary technicians in this field to implement it. When the combination of technical solutions is mutually inconsistent or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0209] It should be understood that the above description of the preferred embodiment is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.

Claims

1. A Pump and Dump transaction detection method based on a time series behavior graph, characterized in that: The following steps are involved: Step 1: Construct the raw transaction data into a time-series behavior graph G(N,E); where N is the set of accounts in graph G, used to store the accounts in the transaction; E is the set of transaction information in graph G, used to store the transaction relationship between accounts and transaction information at the time of the transaction, including token type, transaction amount, transaction time, and token precision; Step 2: Extract the statistical features α of each account i from the time series behavior graph G i , the statistical characteristics of each transaction eβ e , and behavioral characteristics of each transaction e , and store these features in the transaction graph G; Step 3: By aggregating the transaction behavior characteristics of each account and the characteristics of its multi-order neighbors, a real-time account embedding emb is calculated for both parties of each transaction e∈E. n , n∈N, and embed emb by concatenating the source accounts src 、Target account embedded emb dst , Source account statistical characteristics α src , Target account statistical characteristics α dst and transaction statistics β e Get real-time embeds for each transaction e ; Step 4: Construct historical normal transactions and pump and dump transactions into two sets respectively. By calculating the Euclidean distance of the real-time embeddings within the two sets and the Euclidean distance between the two sets, the Euclidean distance within the two sets is minimized and the Euclidean distance between the sets is maximized. This maximizes the similarity within the sets and the difference between the sets, and enhances the distinguishability of the real-time embeddings generated by the time series behavior graph G. Step 5: The real-time embedding obtained after the enhancement will be used to predict whether each transaction is abnormal or not.

2. The method for detecting pump and dump transactions based on a time series behavior graph according to claim 1, characterized in that: In step 2, the statistical feature α of each account i ,include: closeness_centrality, which measures the inverse of the average distance between an account and other accounts, reflecting the proximity of the account in the network; betweenness_centrality, which measures the importance of an account as a bridge in the network, i.e., how often the account appears in the shortest path; indegree_centrality, which measures the degree to which the target account receives the number of connections in the directed network; outdegree_centrality, which measures the degree to which the target account sends the number of connections in the directed network; degree_centrality, which is the degree centrality of the account, i.e. the number of connections of the account; eigenvector_centrality, which is the centrality of an account related to the centrality of its neighboring accounts, emphasizes the connection between an account and highly central accounts; PageRank is used to measure the importance of a target account by calculating its stable state probability; tx_per_account, which is the number of target account's trading accounts divided by the number of target account's transactions, is used to measure the degree of diversity of the target account's trading accounts; The statistical characteristics of each transaction β e ,include: counts is the number of historical transactions between the two parties within the detection window up to the time of the transaction; Mean is the historical transaction mean of both parties within the detection window up to the transaction; var is the historical transaction amount between the two parties within the detection window up to the time of the transaction; The behavioral characteristics of each transaction e ,include: time, the time when the transaction occurred; value, the amount of the transaction; token_decimals, which is the precision of the tokens involved in the transaction; token is the token contract address involved in this transaction.

3. The method for detecting pump and dump transactions based on a time series behavior graph according to claim 1, wherein: In step 3, a real-time embedding z of account i at time t is generated i The specific implementation of (t) includes the following sub-steps: Step 3.1: Account feature transformation; X′=XW n ; in, is the account memory matrix, F n is the dimension of account memory, is the linear transformation weight matrix of the account features, F′ is the feature dimension output by each attention head, and H is the number of attention heads; Step 3.2: Transaction feature splicing; The behavioral characteristics of all transactions are constructed as a matrix ξ and compared with the time difference φ(tt - ) together to obtain the behavioral characteristic matrix E of all transactions: E=concat(ξ,φ(t-t - )); Among them, t is the timestamp of each transaction, t - The time when the memory of the source account of each transaction was last updated; Step 3.3: Transaction feature transformation; E′=EW e ; in, is the behavioral feature matrix of transactions, M is the number of transactions, F e is the behavioral characteristic dimension of the transaction, is the linear transformation weight matrix of transaction features; Step 3.4: Calculate the attention score e based on the transformed account features and transaction features; e=LeakyReLU(X′a l +X′a r +E′a e ); Among them, a l ,a r ,a e ∈R F′×H×1 is the attention parameter matrix; Step 3.5: Use the softmax function to normalize the attention scores to obtain the attention weight a of each account and its neighbor accounts; in, represents the set of neighbor accounts of account i, e k represents the attention score of account i’s neighbor account k; Step 3.6: Use the calculated attention weights to perform weighted summation on the features of neighbor accounts to achieve message passing and feature aggregation, and obtain the real-time embedding of the account of one of the attention heads. Among them, h∈(1,…,H) is the serial number of the attention head, X j ′ represents the account feature transformation matrix of the neighboring accounts of account i; Step 3.7: Concatenate the output features of all heads to get the real-time embedding of account i Here, H is the number of attention heads.

4. The method for detecting pump and dump transactions based on a time series behavior graph according to claim 1, wherein: In step 4, by adopting neighbor account sampling, time encoder, message mechanism, and memory mechanism, account embedding and transaction embedding are generated, and the historical transaction behavior of source account src and target account dst is learned to enhance the detection ability of pump and dump transaction behavior; The time encoder encodes time t using a finite Fourier series: φ(t)=[cos(w1t+ψ1),cos(w2t+ψ2),…,cos(w η t+ψ η )]; Where η is the dimension of the Fourier series after encoding time t; weight ω j and bias ψ j is a learnable parameter, j∈{1,…,η}; The message mechanism, at any time t, for an interaction event e(t) involving the source account src and the target account dst, will generate a message msg(t) to update the memory of the source account src and the target account dst: msg(t)=concat(m src (t - ),m dst (t - ),e(t),φ(t)); Among them, m src (t - ) is the account memory of the source account src before the update at time t, m dst (t - ) is the account memory of the target account dst before it is updated at time t, and e(t) is the behavioral characteristic of the transaction ξ E ; Update the memory of the source account src and the destination account dst using the same message msg(t); For t1,…,t b ≤t, account i uses the aggregation mechanism to aggregate message msg i (t1),…,msg i (t b ): The memory mechanism mentioned above updates the memory of any account i after an event involving account i occurs. For an interaction event e(t) between a source account src and a target account dst, the memory of the source account src and the target account dst is updated as follows: Among them, mem is a learnable memory update function that can maintain the order information of the sequence while capturing the temporal dependencies in the sequence data, thereby identifying and understanding the dynamic change patterns of the input data over time.

5. The Pump and Dump transaction detection method based on a time series behavior graph according to claim 4 is characterized in that: The neighbor account sampling method first generates a subgraph g′ based on the initial account set seed_nodes in the given graph to be sampled g, and then filters out transactions that do not meet the conditions based on the timestamp filter ts. Then, the size of the sampling neighborhood is determined based on the number of neighbors to be sampled for each account in each layer of the graph neural network, fanouts, and the corresponding frontier subgraph list frontiers is returned. The specific implementation includes the following sub-steps: (1) The number of neighbors to be sampled for each account in each layer of the graph neural network, fanouts, is assigned to the sampling frontier subgraph list frontiers, indicating the number of neighbors to be sampled; (2) The sampled frontier subgraph list frontiers is set to empty; (3) Traverse fanouts and determine whether the timestamp of the transaction is less than the timestamp filter ts; If so, continue to check whether the frontiers are full. If the frontiers are not full, add the transaction to the frontiers. If the frontiers are full, keep the latest frontiers. If not, discard the transaction.

6. The method for detecting pump and dump transactions based on a time series behavior graph according to claim 4, characterized in that: For the interaction event e(t) between the source account src and the target account dst, after obtaining the real-time embedding z of the source account src and the target account dst src (t) and z dst (t), the real-time embedding h of the interactive event e(t) is obtained as follows e (t): h e (t)=concat(z src (t),z dst (t),β e ,α src ,α dst ); Among them, β e is the statistical feature of the transaction between the source account src and the target account dst, α src and α dst It is the statistical characteristics of the source account src and the destination account dst at the account level.

7. The method for detecting pump and dump transactions based on a time series behavior graph according to any one of claims 1 to 6, characterized in that: In step 4, the contrast loss function used in contrastive learning is: in, Calculate the positive sample set S P Internal sample h i 、h j The average of the sum of squared Euclidean distances between them; Calculate the negative sample set S N Internal sample h i 、h j The average of the sum of squared Euclidean distances between them; Calculate the distance difference between positive samples and negative samples, and set a threshold margin. If the distance between positive samples and negative samples is less than this threshold, a positive loss is generated to distinguish the two types of samples.

8. A Pump and Dump transaction detection device based on a time sequence behavior diagram, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the Pump and Dump transaction detection method based on the timing behavior diagram as described in any one of claims 1 to 7 is implemented.

9. A pump and dump transaction detection product based on a time series behavior graph, comprising a computer program, characterized in that: When the computer program is executed by a processor, the pump and dump transaction detection method based on the timing behavior diagram as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method for detecting abnormal entity in digital currency transaction and storage medium

    CN113506179A

  • Locating suspect transaction patterns in financial networks

    US20230214842A1