Block chain intelligent contract detection method
By combining reinforcement learning and graph neural networks, the semantic segmentation and feature fusion of blockchain smart contracts are dynamically optimized, solving the problems of insufficient semantic segmentation granularity and weak dynamic behavior correlation in existing technologies, and realizing efficient and real-time detection of blockchain smart contracts.
Patent Information
- Application Number
- CN202511055347.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-18
AI Technical Summary
Existing blockchain smart contract detection technologies suffer from insufficient robustness in dynamic address tracking when facing adversarial contract attacks, a disconnect between static and dynamic semantics, and coarse semantic segmentation granularity, resulting in a significant decrease in detection efficiency in complex attack scenarios.
We employ a reinforcement learning-based dynamic semantic segmentation algorithm, combined with graph neural networks and cross-modal Transformers, to construct a 'feature-behavior-intent' association analysis system. This enables fine-grained bytecode control flow segmentation and cross-modal feature fusion, dynamically optimizes contract logic structure, and detects new adversarial attacks in real time.
It achieves second-level real-time detection and dynamic address tracking, improving the accuracy and real-time performance of detection in complex attack scenarios, and meeting the efficiency and generalization requirements of blockchain proactive defense mechanisms.
Smart Images

Figure CN120979703A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of blockchains, and particularly relates to a blockchain smart contract detection method. BACKGROUND
[0002] The security of blockchain smart contracts has become a core challenge to guarantee the ecology of decentralized applications (DApps). In recent years, adversarial contracts (Adversarial Contracts) have been widely used by hackers as a new type of attack carrier. By disguising as a legitimate contract or nesting calling other contracts on the chain, it realizes the automatic exploitation of the victim contract, such as re-entrant attack, flash loan attack. To cope with such threats, active defense mechanisms need to complete threat identification and blocking in the pre-attack and pre-emergency stages. Research shows that the average time window from deploying malicious contracts to launching attacks is 17 hours and 24 minutes, and the median time window is only 4 minutes and 51 seconds, which puts very high requirements on the real-time and accuracy of detection technology.
[0003] Current adversarial contract detection technology is mainly based on static code feature analysis and local behavior pattern modeling:
[0004] Static feature extraction technology: such as the prior art generates text features by disassembling EVM operation codes, extracts n-gram operation code combinations using TF-IDF, such as high-frequency malicious operation code PUSH20, and trains a logistic regression classifier. Although this method can capture the static code patterns of malicious contracts, it cannot track dynamic address interactions, such as multi-layer nested calls between attacker contracts and victim contracts, resulting in the loss of semantic context in complex attack scenarios.
[0005] Dynamic behavior analysis technology: such as the prior art proposes a pruning semantic control flow tokenization (PSCFT) method, which trains a machine learning model based on function calls and control flow features, but its dynamic behavior modeling is limited to single-contract internal logic and cannot associate global semantics across contract attack links, such as dynamic jumps in the flow of funds in flash loan attacks.
[0006] Special vulnerability detection technology: such as the prior art detects re-entrant vulnerabilities through cross-contract static data flow analysis, but only targets specific vulnerability types and lacks the generalization ability to new attack patterns such as combined attacks and multi-stage attacks.
[0007] The existing methods have the following key bottleneck problems:
[0008] Dynamic address tracking lacks robustness: Attackers often use dynamically generated addresses, such as CREATE2 instructions or nested calls, or multi-layered contract jumps in flash loan attacks, to hide the attack chain. Existing technologies rely on static address tagging or local control flow analysis, which leads to breaks in the semantic context of cross-contract interactions and makes it impossible to accurately capture the attacker's global intent.
[0009] Separation of static features and dynamic behavior: Traditional methods separate the opcode sequence, which is a static code feature, from the control flow path, which is a dynamic behavior. They lack the fusion analysis of the heterogeneous semantics of the two, making it difficult to establish a "feature-behavior-intent" association mapping, which in turn leads to high false positive and false negative rates.
[0010] Insufficient semantic awareness granularity: Existing semantic segmentation technologies use a fixed-granularity control flow segmentation strategy, which cannot adapt to the complexity of contract logic, such as recursive calls and multi-branch jumps, resulting in the incorrect merging or splitting of key semantic segments in complex attack patterns.
[0011] In summary, existing adversarial contract detection technologies suffer from insufficient dynamic address tracking capabilities, separation of static and dynamic semantics, and coarse semantic segmentation granularity, resulting in a significant decrease in detection efficiency in complex attack scenarios such as nested flash loan calls and multi-contract collaborative attacks. There is an urgent need for a new detection framework that can adapt to the evolution of contract logic and integrate multimodal semantic perception. To this end, this invention proposes a blockchain smart contract detection method. Summary of the Invention
[0012] The purpose of this invention is to provide a blockchain smart contract detection method that identifies adversarial contracts based on adaptive semantic perception, aiming to solve the core problems of insufficient semantic segmentation granularity and weak correlation with dynamic behavior in existing technologies. This method firstly designs a fine-grained bytecode control flow segmentation algorithm based on reinforcement learning to achieve dynamic semantic deconstruction of contract logic. Secondly, it innovatively integrates graph neural networks and cross-modal Transformers to heterogeneously align static features with dynamic semantic fragments, constructing a "feature-behavior-intent" correlation analysis system. Finally, it forms a dynamically evolving detection system that can perceive new adversarial attacks in real time and complete the detection of blockchain smart contracts.
[0013] The specific technical solution adopted by this invention is as follows:
[0014] The method for detecting blockchain smart contracts includes the following steps:
[0015] Step 1: Dynamic semantic segmentation optimization based on reinforcement learning: To address the multi-stage path concealment and cross-contract semantic association complexity of attack contracts in price manipulation vulnerabilities, a dynamic semantic segmentation optimization framework based on reinforcement learning is designed.
[0016] Preferably, in step 1, a targeted state space feature s is first defined. t The targeted state-space feature s t It includes: basic features, dynamic context features, and runtime features; the basic features are: opcode distribution; cross-contract dependency vector; the dynamic context features are: flash loan call depth, reflecting arbitrage complexity; oracle dependency chain length; token transfer imbalance index; the runtime features are: function call depth; historical segmentation behavior; a graph attention network is used to model cross-contract dependencies, combined with bi-LSTM to capture temporal dynamic features, generating a unified multimodal feature vector.
[0017] Preferably, in step 1, the objective of the action space is to dynamically adjust three segmentation parameters: segment length threshold increment l, number of iterations r, and merging probability threshold θ. A reward function is designed based on a multi-objective weighted approach, as follows:
[0018] R(s t ,a t )=λ1·Cov PM +λ2·Integ PM -λ3·Frag Key (1)
[0019] Where λ1 represents the importance of attack path coverage in the reward function, λ2 represents the weight of semantic integrity, and λ3 represents the weight of the critical path breakage penalty term.
[0020] Among them, Cov PM For coverage calculations, it represents the proportion of detected nested flash loan paths to the known attack pattern library; Integ PM For completeness assessment, represents the mean semantic similarity between adjacent segments based on CodeBERT; Frag Key The breakage penalty term represents the proportion of broken oracle dependency chains to the total number of chains. For the PPO algorithm implementation, the policy network (Actor) adopts a three-layer fully connected network with a 256-dimensional input layer, a 128-dimensional hidden layer, and a 3-dimensional output layer, which outputs the action parameter distribution. The value network (Critic) has a homogeneous network structure and outputs a state value scalar. The training mechanism adopts a course learning strategy to optimize the attack path coverage and semantic integrity goals in stages.
[0021] Preferably, in step 1, the dynamic semantic segmentation algorithm is implemented through the following steps:
[0022] In the initialization phase, the bytecode is first segmented into initial segments according to the JUMPDEST instruction, and a control flow graph is constructed. Then, the work queue is initialized as the leaf node set of the CFG. In the reinforcement learning decision loop, the algorithm pops a node n to be processed from the work queue and establishes associations by traversing all parent nodes p of node n. At this time, the algorithm extracts the state space features of p in real time and constructs the state vector s. t and s t Input to the policy network to generate action parameters a t = (Δl,r,θ); if the semantic similarity sim(p,n)>θ between p and n is true, then the output is merged into n′, otherwise the output is split; in addition, segment length constraints are also introduced; by dynamically adjusting l←l+Δl, if the merged segment length exceeds the preset threshold l, then a forced split is performed; in the post-processing optimization stage, the algorithm first performs critical path verification, detects the breakage of flash loan loops or oracle dependency chains, and corrects these paths; then, by sorting by the program counter, the final semantic segments are sorted according to the execution order of the bytecode.
[0023] Step 2: "Feature-Behavior-Intent" Heterogeneous Feature Fusion: The semantic segments output by the dynamic semantic segmentation optimization based on reinforcement learning are fused with "feature-behavior-intent" heterogeneous features to provide highly discriminative fused feature inputs for adversarial contract recognition based on attack intent reasoning and adaptive evolution.
[0024] Preferably, step 2 specifically includes feature fusion at three levels: feature layer, behavior layer, and intent layer;
[0025] The feature layer focuses on extracting static suspiciousness indicators of adversarial contracts, providing a basis for locating sensitive areas in the behavior layer. In addition to basic static features such as opcode distribution, function call graph complexity, and state variable dependency depth, it also extends adversarial features, including the code obfuscation index C. obfusc The complexity based on the Abstract Syntax Tree (AST) structure is calculated as follows:
[0026]
[0027] Among them, C loop Represents the number of nested loops, C jump Indicates the number of redundant jump instructions, C lines Represents the total number of lines of code; similarity S between dummy functions disguise That is, the cosine distance is used to measure the degree of deviation between the public function name and the normal contract, and it is calculated as follows:
[0028]
[0029] Among them, v funcv represents the feature vector of the common functions in the contract to be detected. normal This represents the average eigenvector of similar functions in a normal contract library;
[0030] The behavioral layer aims to capture anomalous logical jumps and cross-contract attack chains during adversarial contract execution. First, it receives a set of semantic segments {SS, SS, ..., SS} optimized through reinforcement learning and performs time-sensitive behavior modeling; this includes control flow anomaly detection. A cf This indicates the frequency of jump instructions deviating from the normal pattern within a statistical semantic segment, where Jump... i This represents the i-th jump instruction within the semantic segment. This represents the set of legal jump targets in safe mode, ∏() represents the indicator function; and it involves extracting cross-contract attack chains, constructing a cross-segment call graph, and identifying high-frequency callback paths; finally, bi-LSTM is used to encode the semantic segment execution trajectory h. traj =BiLSTM(X) opcode ,X gas This allows for the capture of temporal dependencies, providing more accurate dynamic information for attack pattern analysis.
[0031] Bi-LSTM() refers to a bidirectional long short-term memory network used to encode temporal data while capturing both forward and backward dependencies; X opcode X represents the input opcode sequence. gas h represents the input gas consumption sequence. traj This represents the trajectory encoding vector generated by the behavior layer;
[0032] The intent layer: By inferring the attacker's potential targets, a causal chain of multi-stage attack logic is constructed; combining reentrancy attack intent and gas exhaustion attack intent, an attack pattern rule base is established and analyzed through a causal reasoning network; in this network, nodes define different stages of the attack, focusing on key links such as "induced invocation," "state tampering," and "fund transfer," and performing causal weight calculations. This allows us to deduce the attacker's potential behavioral intentions; MLP stands for Multilayer Perceptron, used for nonlinear feature fusion; This represents the static characteristics of the i-th attack phase; This represents the behavioral trajectory characteristics of the i-th attack phase.
[0033] Preferably, in step 2, a cross-modal bidirectional association mechanism is employed, guiding attention to sensitive regions from features to behavior; the cross-modal attention localization is calculated using the following formula:
[0034] Q = W q h static ,K,V=W kh traj W v h traj (4)
[0035] Among them, W q W k W v It refers to a trainable parameter matrix, which is used to project input features into the query, key, and value spaces, respectively.
[0036] Among them, the static feature vector h static As a query, dynamic trajectory is embedded in h traj As keys and values, the output attention weight matrix focuses on highly suspicious semantic segments; from behavior to intent, it models through contextual dependencies. The dynamic behavior sequence is input into the Transformer encoder to capture the temporal dependencies of cross-segment attack chains, and the attack type P(y|E) is predicted based on the behavioral context. behavior = softmax(W) c ·E behavior +b c );
[0037] Among them, E behavior W refers to the behavior context embedding vector output by the Transformer encoder. c and b c This refers to the weight matrix and bias term of the classification layer, P(y|E behavior) This refers to the probability distribution of attack types calculated using Softmax.
[0038] From intent to features, under the adaptive evolutionary feedback mechanism, after detecting novel price manipulation attack intent, adversarial examples are generated and key features are extracted, updating the static feature library F←F∪f. new Where F: static feature library, f new Key features extracted from novel attack intentions
[0039] Meanwhile, inverse reinforcement learning was used to adjust the weights of the reward function, thereby increasing the coverage R of novel attack paths. new =R+λ·Cov;
[0040] Where R is the original reward function; λ is the weight coefficient of the novel attack path; and Cov is the coverage of the novel attack path.
[0041] Step 3: Adaptive Evolution-Based Adversarial Contract Identification: To address the challenges of dynamic evolution and intent concealment in adversarial contract price manipulation attacks, the system rapidly adapts to new price manipulation attacks and continuously evolves by providing real-time feedback on attack intent and dynamically optimizing the feature library and segmentation strategy.
[0042] Preferably, in step 3, during the mapping process from intent to feature, the detected attack intent is inverted into feature enhancement rules:
[0043] f new =α·f flash_depth +β·f oracle_chain (5)
[0044] f flash_depth It is a deep feature call for flash loans; f oracle_chain This is an oracle-dependent chain length feature;
[0045] Where α and β are dynamic weights; when a new type of price manipulation attack is detected, the system generates corresponding feature enhancement rules and automatically adjusts these rules based on the attack path coverage; the feature library updater is F←F∪{f new}, PCA dimensionality reduction is used to retain key information in the intent layer, while high-contribution features are selected through Top-k filtering, i.e.:
[0046] f new =PCA(Top-k(h) intent (6)
[0047] h intent This represents the feature vector generated by the intent layer.
[0048] Preferably, in step 3, the reward weights are optimized based on the coverage of the novel attack paths through reverse reinforcement learning tuning.
[0049]
[0050] Where, N new_attack_path This represents the number of new attack paths detected; N total_attack_path The total number of attack paths is the sum of historical and new attacks; λ is a balancing coefficient that controls the optimization intensity of new attacks.
[0051] This updates the policy network; the PPO algorithm is used to optimize the segmentation policy network parameters online.
[0052] Where θ is the policy network parameter; η is the learning rate;
[0053] Incremental learning anti-forgetting mechanisms constrain the update magnitude of model parameters through elastic weight consolidation, thus protecting existing knowledge. The formula is as follows:
[0054] L total =L new +γ·∑ i F i (θ i -θ old,i ) 2 (8)
[0055] Where F is the Fisher information matrix;
[0056] L new It is the loss function for novel attack missions; F i These are the diagonal elements of the Fisher information matrix, with the metric parameter θ. i The importance of old tasks γ: is a regularization coefficient that controls the strength of old knowledge retention; in addition, the experience replay buffer is used to store historical attack samples and replay them periodically to consolidate old knowledge; ultimately, it can continuously evolve.
[0057] The technical effects achieved by this invention are as follows:
[0058] This invention is based on adaptive semantic perception to identify adversarial contracts, aiming to solve the core problems of insufficient semantic segmentation granularity and weak correlation with dynamic behavior in existing technologies. The method first designs a fine-grained bytecode control flow segmentation algorithm based on reinforcement learning to realize the dynamic semantic deconstruction of contract logic. Secondly, it innovatively integrates graph neural networks and cross-modal Transformers to heterogeneously align static features with dynamic semantic fragments, constructing a "feature-behavior-intent" correlation analysis system. Finally, it forms a dynamically evolving detection system that can detect new adversarial attacks in real time and complete the detection of blockchain smart contracts.
[0059] This invention utilizes a reinforcement learning-driven fine-grained semantic control flow segmentation algorithm to dynamically optimize the logical structure of smart contracts, overcoming the limitations of insufficient granularity in traditional semantic segmentation. By combining graph neural networks and cross-modal Transformer fusion technology, it achieves heterogeneous semantic alignment between static features and dynamic behaviors, constructing a three-layer association analysis system of "feature-behavior-intent," effectively solving the problems of semantic fragmentation and context loss in complex attack scenarios. The system adaptively evolves based on a reinforcement learning framework, supporting second-level real-time detection and dynamic address tracking, meeting the core requirements of blockchain proactive defense mechanisms for efficiency, generalization, and real-time performance. Attached Figure Description
[0060] Fig. 1 This is a system block diagram of the blockchain smart contract detection method of the present invention;
[0061] Fig. 2 This is a system block diagram of step 1 of the present invention, which is based on reinforcement learning for dynamic semantic segmentation optimization.
[0062] Fig. 3 This is a system block diagram of the heterogeneous feature fusion in step 2 of this invention, which involves "feature-behavior-intent".
[0063] Fig. 4 This is a system block diagram of step 3 of the present invention, which is based on adaptive evolution for adversarial contract identification. Detailed Implementation
[0064] To make the objectives and advantages of this invention clearer, the invention will be specifically described below with reference to embodiments. It should be understood that the following text is merely used to describe one or more specific embodiments of the invention and does not strictly limit the scope of protection specifically claimed by the invention.
[0065] like Figs. 1-4 As shown, the blockchain smart contract detection method includes the following steps:
[0066] Step 1: Dynamic Semantic Segmentation Optimization Based on Reinforcement Learning: To address the multi-stage path concealment and cross-contract semantic association complexity in price manipulation vulnerabilities (e.g., nested flash loan calls, liquidity pool state tampering); and cross-contract semantic association complexity (e.g., oracle price-dependent chains), a dynamic semantic segmentation optimization framework based on reinforcement learning is designed. By guiding the segmentation strategy through price manipulation sensitive features, the semantic integrity capture capability of the attack path is improved, and the false negative rate is reduced.
[0067] like Figs. 1-4 As shown, in step 1, the targeted state space feature s is first defined. t ; Targeted state-space features s t This includes: basic features, dynamic context features, and runtime features; basic features: opcode distribution, i.e., the proportion of total computation, storage, and jump instructions; cross-contract dependency vector, i.e., the number and type of external contract calls, such as oracles and liquidity pools; dynamic context features: flash loan call depth, i.e., the nested call level, reflecting arbitrage complexity; oracle dependency chain length, i.e., the number of consecutive oracle queries; token transfer imbalance index, i.e., the variance of the token transfer-in / transfer-out ratio, which detects liquidity pool manipulation; runtime features: function call depth, i.e., the nesting level of the current code segment; historical segmentation behavior, i.e., the mean and variance of the segment length of the most recent merge operation; a graph attention network is used to model cross-contract dependencies, combined with bi-LSTM to capture temporal dynamic features, generating a unified multimodal feature vector.
[0068] The goal of the action space is to dynamically adjust three segmentation parameters: segment length threshold increment l, number of iterations r, and merging probability threshold θ. A reward function is designed based on a multi-objective weighted approach, as follows:
[0069] R(st a t )=λ1·Cov PM +λ2·Integ PM -λ3·Frag Key (1)
[0070] Where λ1 represents the importance of attack path coverage in the reward function, λ2 represents the weight of semantic integrity, and λ3 represents the weight of the critical path breakage penalty term.
[0071] Among them, Cov PM For coverage calculations, it represents the proportion of detected nested flash loan paths to the known attack pattern library; Integ PM For completeness assessment, represents the mean semantic similarity between adjacent segments based on CodeBERT; Frag Key The breakage penalty term represents the proportion of broken oracle dependency chains to the total number of chains. For the PPO algorithm implementation, the policy network (Actor) adopts a three-layer fully connected network with a 256-dimensional input layer, a 128-dimensional hidden layer, and a 3-dimensional output layer, which outputs the action parameter distribution. The value network (Critic) has a homogeneous network structure and outputs a state value scalar. The training mechanism adopts a course learning strategy to optimize the attack path coverage and semantic integrity goals in stages.
[0072] Based on the above definition, the dynamic semantic segmentation algorithm is implemented through the following steps:
[0073] In the initialization phase, the bytecode is first segmented into initial segments according to the JUMPDEST instruction, and a control flow graph (CFG) is constructed. Through these segments, the algorithm can clearly identify the control flow structure in the bytecode, ensuring that each basic block is correctly distinguished. Subsequently, the work queue is initialized as the leaf node set of the CFG, providing the basis for processing nodes in subsequent reinforcement learning decisions. In the reinforcement learning decision loop, the algorithm pops a node n to be processed from the work queue and establishes associations by traversing all parent nodes p of node n. At this time, the algorithm extracts the state space features of p in real time and constructs the state vector s. t and will s The input t is fed into the policy network to generate action parameters a. t= (Δl,r,θ); if the semantic similarity sim(p,n)>θ between p and n is true, then the output is merged as n′, otherwise the output is split; in addition, segment length constraints are also introduced to ensure that the segmented segments are not too long; by dynamically adjusting l←l+Δl, if the length of the merged segment exceeds the preset threshold l, then a forced split is performed; in the post-processing optimization stage, the algorithm first performs critical path verification, detects the breakage of flash loan loops or oracle dependency chains, and corrects these paths to avoid potential vulnerabilities caused by erroneous merging; then, by sorting by the program counter, the final semantic segments are sorted according to the execution order of the bytecode, thereby ensuring that the segmented segments are logically coherent and conform to the actual execution order of the program.
[0074] Step 2: "Feature-Behavior-Intent" Heterogeneous Feature Fusion: The semantic segments output by the dynamic semantic segmentation optimization based on reinforcement learning are fused with "feature-behavior-intent" heterogeneous features to provide highly discriminative fused feature inputs for adversarial contract recognition based on attack intent reasoning and adaptive evolution.
[0075] like Figs. 1-4 As shown, step 2 specifically includes feature fusion at three levels: feature layer, behavior layer, and intent layer.
[0076] Feature layer: Focuses on extracting static suspiciousness indicators of adversarial contracts, providing a basis for locating sensitive areas in the behavior layer; in addition to basic static features such as opcode distribution, function call graph complexity, and state variable dependency depth, it also extends adversarial features, including the code obfuscation index C. obfusc The complexity based on the Abstract Syntax Tree (AST) structure is calculated as follows:
[0077]
[0078] Among them, C loop Represents the number of nested loops, C jump Indicates the number of redundant jump instructions, C lines Represents the total number of lines of code; similarity S between dummy functions disguise That is, the cosine distance is used to measure the degree of deviation between the public function name and the normal contract, and it is calculated as follows:
[0079]
[0080] Among them, v func v represents the feature vector of the common functions in the contract to be detected. normal This represents the average eigenvector of similar functions in a normal contract library;
[0081] Behavioral Layer: The goal is to capture anomalous logical jumps and cross-contract attack chains during adversarial contract execution. First, it receives a set of semantic segments {SS,SS,…,SS} optimized by reinforcement learning and performs time-sensitive behavior modeling; this includes control flow anomaly detection. A cf This indicates the frequency of jump instructions deviating from the normal pattern within a statistical semantic segment, where Jump... i This represents the i-th jump instruction within the semantic segment. This represents the set of legal jump targets in safe mode, ∏() represents the indicator function; and it involves extracting cross-contract attack chains, constructing a cross-segment call graph, and identifying high-frequency callback paths; finally, bi-LSTM is used to encode the semantic segment execution trajectory h. traj =BiLSTM(X) opcode ,X gas This allows for the capture of temporal dependencies, providing more accurate dynamic information for attack pattern analysis.
[0082] Bi-LSTM() refers to a bidirectional long short-term memory network used to encode temporal data while capturing both forward and backward dependencies; X opcode X represents the input opcode sequence. gas h represents the input gas consumption sequence. traj This represents the trajectory encoding vector generated by the behavior layer.
[0083] Intent Layer: By inferring the attacker's potential targets, a causal chain of multi-stage attack logic is constructed; combining reentrancy attack intent and gas exhaustion attack intent, an attack pattern rule base is established and analyzed through a causal inference network; in this network, nodes define different stages of the attack, focusing on key links such as "induced invocation," "state tampering," and "fund transfer," and performing causal weight calculations. This allows us to deduce the attacker's potential behavioral intentions; MLP stands for Multilayer Perceptron, used for nonlinear feature fusion; This represents the static characteristics of the i-th attack phase; This represents the behavioral trajectory characteristics of the i-th attack phase.
[0084] In step 2, a cross-modal bidirectional association mechanism is adopted, from features to behavior, with sensitive regions receiving attention through guidance; cross-modal attention localization is achieved, and the calculation formula is as follows:
[0085] Q = W q h static K,V = W k h traj W v h traj (4)
[0086] Among them, Wq W k W v It refers to a trainable parameter matrix, which is used to project input features into the query, key, and value spaces, respectively.
[0087] Among them, the static feature vector h static As a query, dynamic trajectory is embedded in h traj As keys and values, the output attention weight matrix focuses on highly suspicious semantic segments; from behavior to intent, it models through contextual dependencies. The dynamic behavior sequence is input into the Transformer encoder to capture the temporal dependencies of cross-segment attack chains and predict the attack type based on the behavioral context.
[0088] P(y|E behavior = softmax(W) c ·E behavior +b c );
[0089] Among them, E behavior W refers to the behavior context embedding vector output by the Transformer encoder. c and b c This refers to the weight matrix and bias term of the classification layer, P(y|E behavior () refers to the probability distribution of attack types calculated using Softmax;
[0090] From intent to features, under the adaptive evolutionary feedback mechanism, after detecting novel price manipulation attack intent, adversarial examples are generated and key features are extracted, updating the static feature library F←F∪f. new Where F: static feature library, f new Key features were extracted from novel attack intentions; simultaneously, inverse reinforcement learning was used to adjust the weights of the reward function, enhancing the coverage R of novel attack paths. new =R + λ·Cov, where R is the original reward function; λ is the weight coefficient of the novel attack path; Cov is the coverage of the novel attack path; improve the model's ability to detect novel price manipulation attack patterns.
[0091] Step 3: Adaptive Evolution-Based Adversarial Contract Recognition: To address the challenges of dynamic evolution and intent concealment in adversarial contract price manipulation attacks, the system can quickly adapt to new price manipulation attacks and continuously evolve by providing real-time feedback on attack intent and dynamically optimizing the feature library and segmentation strategy, thereby improving the robustness of adversarial contract recognition.
[0092] like Figs. 1-4As shown, in step 3, during the mapping process from intent to feature, the detected attack intent is inverted into feature enhancement rules:
[0093] f new =α·f flash_depth +β·f oracle_chain (5)
[0094] Where α and β are dynamic weights; for example, when a new type of price manipulation attack is detected, the system generates corresponding feature enhancement rules and automatically adjusts these rules based on the attack path coverage; the feature library updater is F←F∪{f new}, PCA dimensionality reduction is used to retain key information in the intent layer, while high-contribution features are selected through Top-k filtering, i.e.:
[0095] f new =PCA(Top-k(h) intent (6)
[0096] h intent This represents the feature vectors generated by the intent layer; ensuring the accuracy and effectiveness of the feature library.
[0097] By optimizing through reverse reinforcement learning, the reward weights are adjusted based on the coverage of novel attack paths:
[0098]
[0099] Where, N new_attack_path This represents the number of new attack paths detected; N total_attack_path The total number of attack paths is the sum of historical and new attacks; λ is a balancing coefficient that controls the optimization intensity of new attacks.
[0100] This updates the policy network; the PPO algorithm is used to optimize the segmentation policy network parameters online. This allows the model to continuously adapt to new attack methods. To avoid catastrophic forgetting, the incremental learning anti-forgetting mechanism uses elastic weight consolidation to constrain the update magnitude of model parameters and protect existing knowledge. The formula is as follows:
[0101] L totatl =L new +γ·Σ i (θ i -θ old,i ) 2 (8)
[0102] Where F is the Fisher information matrix, which measures the importance of parameters; in addition, the experience replay buffer is used to store historical attack samples and replay them periodically to consolidate old knowledge; ultimately, it can continuously evolve and effectively improve the ability to identify adversarial contracts in dynamically changing attack environments.
[0103] This invention is based on adaptive semantic perception to identify adversarial contracts, aiming to solve the core problems of insufficient semantic segmentation granularity and weak correlation with dynamic behavior in existing technologies. The method first designs a fine-grained bytecode control flow segmentation algorithm based on reinforcement learning to realize the dynamic semantic deconstruction of contract logic. Secondly, it innovatively integrates graph neural networks and cross-modal Transformers to heterogeneously align static features with dynamic semantic fragments, constructing a "feature-behavior-intent" correlation analysis system. Finally, it forms a dynamically evolving detection system that can detect new adversarial attacks in real time and complete the detection of blockchain smart contracts.
[0104] This invention utilizes a reinforcement learning-driven fine-grained semantic control flow segmentation algorithm to dynamically optimize the logical structure of smart contracts, overcoming the limitations of insufficient granularity in traditional semantic segmentation. By combining graph neural networks and cross-modal Transformer fusion technology, it achieves heterogeneous semantic alignment between static features and dynamic behaviors, constructing a three-layer association analysis system of "feature-behavior-intent," effectively solving the problems of semantic fragmentation and context loss in complex attack scenarios. The system adaptively evolves based on a reinforcement learning framework, supporting second-level real-time detection and dynamic address tracking, meeting the core requirements of blockchain proactive defense mechanisms for efficiency, generalization, and real-time performance.
[0105] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.
Claims
1. A method for detecting blockchain smart contracts, characterized in that: Includes the following steps: Step 1: Dynamic semantic segmentation optimization based on reinforcement learning: To address the multi-stage path concealment and cross-contract semantic association complexity of attack contracts in price manipulation vulnerabilities, a dynamic semantic segmentation optimization framework based on reinforcement learning is designed. Step 2: "Feature-Behavior-Intent" Heterogeneous Feature Fusion: The semantic segments output by the dynamic semantic segmentation optimization based on reinforcement learning are fused with "feature-behavior-intent" heterogeneous features to provide highly discriminative fused feature inputs for adversarial contract recognition based on attack intent reasoning and adaptive evolution. Step 3: Adaptive Evolution-Based Adversarial Contract Identification: To address the challenges of dynamic evolution and intent concealment in adversarial contract price manipulation attacks, the system rapidly adapts to new price manipulation attacks and continuously evolves by providing real-time feedback on attack intent and dynamically optimizing the feature library and segmentation strategy.
2. The blockchain smart contract detection method according to claim 1, characterized in that: In step 1, the specific state space feature s is first defined. t The targeted state-space feature s t It includes: basic features, dynamic context features, and runtime features; the basic features are: opcode distribution; cross-contract dependency vector; the dynamic context features are: flash loan call depth, reflecting arbitrage complexity; oracle dependency chain length; token transfer imbalance index; the runtime features are: function call depth; historical segmentation behavior; a graph attention network is used to model cross-contract dependencies, combined with bi-LSTM to capture temporal dynamic features, generating a unified multimodal feature vector.
3. The blockchain smart contract detection method according to claim 2, characterized in that: In step 1, the objective of the action space is to dynamically adjust three segmentation parameters: segment length threshold increment l, number of iterations r, and merging probability threshold θ. A reward function is designed based on a multi-objective weighted approach, as follows: R(s t ,a t )=λ1·Cov PM +λ2·Integ PM -λ3·Frag Key (1) Where λ1 represents the importance of attack path coverage in the reward function, λ2 represents the weight of semantic integrity, and λ3 represents the weight of the critical path breakage penalty term. Among them, Cov PM For coverage calculations, it represents the proportion of detected nested flash loan paths to the known attack pattern library; Integ PM For completeness assessment, represents the mean semantic similarity between adjacent segments based on CodeBERT; Frag Key The breakage penalty term represents the proportion of broken oracle dependency chains to the total number of chains. For the PPO algorithm implementation, the policy network (Actor) adopts a three-layer fully connected network with a 256-dimensional input layer, a 128-dimensional hidden layer, and a 3-dimensional output layer, which outputs the action parameter distribution. The value network (Critic) has a homogeneous network structure and outputs a state value scalar. The training mechanism adopts a course learning strategy to optimize the attack path coverage and semantic integrity goals in stages.
4. The blockchain smart contract detection method according to claim 3, characterized in that: In step 1, the dynamic semantic segmentation algorithm is implemented through the following steps: In the initialization phase, the bytecode is first segmented into initial segments according to the JUMPDEST instruction, and a control flow graph is constructed. Then, the work queue is initialized as the leaf node set of the CFG. In the reinforcement learning decision loop, the algorithm pops a node n to be processed from the work queue and establishes associations by traversing all parent nodes p of node n. At this time, the algorithm extracts the state space features of p in real time and constructs the state vector s. t and s t Input to the policy network to generate action parameters a t = (Δl,r,θ); if the semantic similarity between p and n, sim(p,n)>θ, then the output is merged into n′, otherwise the output is split; in addition, segment length constraints are also introduced; By dynamically adjusting l←l+Δl, if the length of the merged segment exceeds the preset threshold l, a forced split is performed. In the post-processing optimization stage, the algorithm first performs critical path verification, detects the breakage of flash loan cycles or oracle dependency chains, and corrects these paths. Then, it sorts the final semantic segments according to the execution order of the bytecode by sorting using the program counter.
5. The blockchain smart contract detection method according to claim 4, characterized in that: Step 2 specifically includes feature fusion at three levels: feature layer, behavior layer, and intent layer. The feature layer focuses on extracting static suspiciousness indicators of adversarial contracts, providing a basis for locating sensitive areas in the behavior layer. In addition to basic static features such as opcode distribution, function call graph complexity, and state variable dependency depth, it also extends adversarial features, including the code obfuscation index C. obfusc The complexity based on the Abstract Syntax Tree (AST) structure is calculated as follows: Among them, C loop Represents the number of nested loops, C jump Indicates the number of redundant jump instructions, C lines Represents the total number of lines of code; similarity S between dummy functions disguise That is, the cosine distance is used to measure the degree of deviation between the public function name and the normal contract, and it is calculated as follows: Among them, v func v represents the feature vector of the common functions in the contract to be detected. normal This represents the average eigenvector of similar functions in a normal contract library; The behavioral layer aims to capture anomalous logical jumps and cross-contract attack chains during adversarial contract execution. First, it receives a set of semantic segments {SS, SS, ..., SS} optimized through reinforcement learning and performs time-sensitive behavior modeling; this includes control flow anomaly detection. A cf This indicates the frequency of jump instructions deviating from the normal pattern within a statistical semantic segment, where Jump... i This represents the i-th jump instruction within the semantic segment. This represents the set of legal jump targets in safe mode, ∏() represents the indicator function; and it involves extracting cross-contract attack chains, constructing a cross-segment call graph, and identifying high-frequency callback paths; finally, bi-LSTM is used to encode the semantic segment execution trajectory h. traj =BiLSTM(X) opcode ,X gas This allows for the capture of temporal dependencies, providing more accurate dynamic information for attack pattern analysis. Bi-LSTM() refers to a bidirectional long short-term memory network used to encode temporal data while capturing both forward and backward dependencies; X opcode X represents the input opcode sequence. gas h represents the input gas consumption sequence. traj This represents the trajectory encoding vector generated by the behavior layer; The intent layer: By inferring the attacker's potential targets, a causal chain of multi-stage attack logic is constructed; combining reentrancy attack intent and gas exhaustion attack intent, an attack pattern rule base is established and analyzed through a causal reasoning network; in this network, nodes define different stages of the attack, focusing on key links such as "induced invocation," "state tampering," and "fund transfer," and performing causal weight calculations. This allows us to deduce the attacker's potential behavioral intentions; MLP stands for Multilayer Perceptron, used for nonlinear feature fusion; This represents the static characteristics of the i-th attack phase; This represents the behavioral trajectory characteristics of the i-th attack phase.
6. The blockchain smart contract detection method according to claim 5, characterized in that: In step 2, a cross-modal bidirectional association mechanism is adopted, from features to behavior, and sensitive regions are guided to focus; the cross-modal attention localization is calculated using the following formula: Q=W q h static ,K,V=W k h traj ,W v h traj (4) Among them, W q W k W v It refers to a trainable parameter matrix, which is used to project input features into the query, key, and value spaces, respectively. Among them, the static feature vector h static As a query, dynamic trajectory is embedded in h traj As keys and values, the output attention weight matrix focuses on highly suspicious semantic segments; from behavior to intent, it models through contextual dependencies. The dynamic behavior sequence is input into the Transformer encoder to capture the temporal dependencies of cross-segment attack chains, and the attack type P(y|E) is predicted based on the behavioral context. behavior = softmax(W) c ·E behavior +b c ); Among them, E behavior W refers to the behavior context embedding vector output by the Transformer encoder. c and b c This refers to the weight matrix and bias term of the classification layer, P(y|E behavior () refers to the probability distribution of attack types calculated using Softmax; From intent to features, under the adaptive evolutionary feedback mechanism, after detecting novel price manipulation attack intent, adversarial examples are generated and key features are extracted, updating the static feature library F←F∪f. new Where F: static feature library, f new Key features extracted from novel attack intentions Meanwhile, inverse reinforcement learning was used to adjust the weights of the reward function, thereby increasing the coverage R of novel attack paths. new =R+λ·Cov; Where R is the original reward function; λ is the weight coefficient of the novel attack path; and Cov is the coverage of the novel attack path.
7. The blockchain smart contract detection method according to claim 6, characterized in that: In step 3, during the mapping process from intent to features, the detected attack intent is inverted into feature enhancement rules: f new =α·f flash_depth +β·f oracle_chain (5) f flash_depth It is a deep feature call for flash loans; f oracle_chain This is an oracle-dependent chain length feature; Where α and β are dynamic weights; when a new type of price manipulation attack is detected, the system generates corresponding feature enhancement rules and automatically adjusts these rules based on the attack path coverage; the feature library updater is F←F∪{f new }, PCA dimensionality reduction is used to retain key information in the intent layer, while high-contribution features are selected through Top-k filtering, i.e.: f new =PCA(Top-kh intent )) (6) h intent This represents the feature vector generated by the intent layer.
8. The blockchain smart contract detection method according to claim 7, characterized in that: In step 3, the reward weights are optimized based on the coverage of novel attack paths through reverse reinforcement learning tuning. Where, N new_attack_path This represents the number of new attack paths detected; N total_attack_path The total number of attack paths is the sum of historical and new attacks; λ is a balancing coefficient that controls the optimization intensity of new attacks. This updates the policy network; the PPO algorithm is used to optimize the segmentation policy network parameters online. Where θ is the policy network parameter; η is the learning rate; Incremental learning anti-forgetting mechanisms constrain the update magnitude of model parameters through elastic weight consolidation, thus protecting existing knowledge. The formula is as follows: L total =L new +γ·∑ i F i (i i -θ old,i ) 2 (8) Where F is the Fisher information matrix; L new It is the loss function for novel attack missions; F i These are the diagonal elements of the Fisher information matrix, with the metric parameter θ. i The importance of old tasks γ: is a regularization coefficient that controls the strength of old knowledge retention; in addition, the experience replay buffer is used to store historical attack samples and replay them periodically to consolidate old knowledge; ultimately, it can continuously evolve.