A blockchain transaction security detection dynamic incentive method based on reinforcement learning

By using a dynamic incentive method based on reinforcement learning to adjust the blockchain transaction detection reward in real time, the problems of static and fixed rewards and low budget utilization efficiency in existing technologies are solved, and the system is made fast, safe and stable, improving the effectiveness of incentives and operational stability.

CN122335294APending Publication Date: 2026-07-03BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing blockchain transaction security testing incentive mechanisms suffer from static and fixed rewards, delayed adjustments, low budget utilization efficiency, and difficulty in balancing convergence speed and system stability, making it impossible to achieve rapid, secure, and stable system operation under limited budget constraints.

Method used

A dynamic incentive method based on reinforcement learning is adopted. By sensing the system state in real time, the detection reward is dynamically adjusted, and an adaptive incentive mechanism with dual constraints of incentive compatibility and budget is established. The incentive strategy is optimized by reinforcement learning algorithm to form a closed-loop feedback mechanism, which quickly guides the system into a safe and stable region.

Benefits of technology

It enables dynamic adjustment of reward levels under incentive compatibility and budget constraints, improving incentive effectiveness and sustainability, increasing system convergence speed and operational stability, reducing incentive costs, and possessing engineering practicality and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122335294A_ABST
    Figure CN122335294A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning's blockchain transaction security detection dynamic incentive method, it is related to blockchain security governance technical field.The method first constructs incentive environment and sets group behavior parameter, forms system state vector based on multidimensional index;Then according to incentive compatibility and budget constraint, the dynamic reward feasible region is calculated, the original action of reinforcement learning is projected as actual reward;Then the proportion of block builder detection and user normal transaction is updated by evolutionary game coupling, combined with convergence progress, reward cost, oscillation suppression and other factors to calculate immediate return, and the incentive strategy is optimized using reinforcement learning algorithm;Finally, when the system continuously meets stable threshold, it enters safe and stable state and switches to low maintenance incentive.The application can adaptively adjust reward, improve budget utilization, speed up system convergence and suppress oscillation, without modifying blockchain bottom layer, and compatible with existing security detection system, suitable for various blockchain transaction security detection incentive scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of blockchain security governance technology, and in particular to a dynamic incentive method for blockchain transaction security detection based on reinforcement learning. Background Technology

[0002] In the blockchain transaction security governance system, the initiative of block builders to perform security checks on candidate transactions is a crucial link in intercepting abnormal transactions and ensuring the trustworthiness of on-chain data and the stability of the ecosystem. To enhance builders' willingness to perform checks, existing technologies generally set up checks as rewards to compensate for their checks, opportunity costs, and potential ecosystem risk losses, thereby guiding builders to actively perform transaction screening. This type of incentive mechanism has become an important supporting means for blockchain security governance.

[0003] Current mainstream incentive schemes mostly adopt static rules such as fixed rewards, constant strong incentives, or preset linear decay rewards. While these schemes have the advantages of simple design and ease of theoretical analysis and engineering implementation, they have significant drawbacks in complex and dynamic blockchain operation scenarios: First, the reward level cannot adaptively adjust with the real-time state of the system, which can easily lead to insufficient incentives and low detection coverage in the early stages, and excessive incentives and wasted budget resources in the later stages. Second, the reward rules rely on only a single environmental variable, making it difficult to simultaneously take into account multiple constraints such as the proportion of builder detection, the proportion of normal user transactions, the stability of the system, and the remaining incentive budget. Third, the incentive design and the evolution of the system's collective behavior are disconnected, making it impossible to form a closed-loop control.

[0004] Due to the aforementioned limitations, existing static incentive methods struggle to simultaneously achieve the three objectives of rapid system convergence to a safe and stable state, maintaining operational stability, and improving cost efficiency under limited budget constraints. To address this issue, there is an urgent need for an adaptive incentive method that can perceive the system state in real time, dynamically adjust reward intensity, and satisfy both incentive compatibility and budget constraints. This method would enable the blockchain system to efficiently enter a safe equilibrium region where "a majority of builders actively detect activity and a majority of users initiate normal transactions." Summary of the Invention

[0005] The purpose of this invention is to propose a dynamic incentive method for blockchain transaction security detection based on reinforcement learning. This method solves the problems of static and fixed rewards, delayed adjustment, low budget utilization efficiency, and difficulty in balancing convergence speed and system stability in existing blockchain transaction security detection incentive mechanisms. It achieves dynamic adjustment of detection rewards based on system state under the dual constraints of incentive compatibility and budget, quickly guides the blockchain system into a safe and stable equilibrium region, and reduces incentive costs and system oscillations.

[0006] To achieve the above objectives, this invention proposes a dynamic incentive method for blockchain transaction security detection based on reinforcement learning, the specific steps of which are as follows: Step S1: Establish a secure detection incentive environment for blockchain transactions by setting the block builder detection ratio and the normal user transaction ratio, and configuring environmental parameters. Step S2: Construct the current system state based on the builder detection ratio, the user normal transaction ratio, the ratio change, the system deviation, the oscillation intensity, the remaining budget ratio, and the reward value of the previous moment; Step S3: Calculate the dynamic reward feasible region based on the incentive compatibility constraint and budget constraint, and always project the original action within the dynamic reward feasible region to obtain the actual reward value; Step S4: Inject the actual reward value into the evolutionary game environment, couple the update builder detection ratio with the user normal transaction ratio, and obtain the system state at the next moment; Step S5: Calculate the immediate reward by combining the system convergence progress, reward expenditure, oscillation suppression and positive reward after entering the stable region, and update the incentive strategy through reinforcement learning algorithm; Step S6: When the system meets the stability threshold for multiple consecutive steps, it is determined to have entered a safe and stable state and training is terminated or switched to a low-maintenance stimulus mode.

[0007] Preferably, in step S1, the environmental parameters include detection cost, detection rate, false positive rate, handling fee difference, ecological penalty, budget limit and target stable area threshold.

[0008] Preferably, in step S2, the current system state is represented as follows: ; in, for The system state at any given moment. for The block builder detection rate at any given time. for The proportion of normal transactions by users at any given time. for The change in the proportion of block builder detection at any given time. for The change in the proportion of normal transactions by users at any given time. for The degree of system deviation at any given time for The intensity of system oscillation at time t, for The percentage of the remaining budget at any given time. for The actual reward value at any given moment.

[0009] Preferably, the degree of system deviation The calculation formula is as follows: ; in, , This is the deviation weighting coefficient; System oscillation intensity The calculation formula is as follows: ; in, The difference in the detection ratio of block builders at adjacent time points. The difference between the proportion of normal transactions by users at adjacent time points. for The block builder detection rate at any given time. for The proportion of normal transactions by users at any given time; the initial time. When =0, ; Remaining budget percentage The calculation formula is as follows: ; in, As of Cumulative reward expenditure at any time This is the budget ceiling.

[0010] Preferably, in step S3, the minimum reward lower bound is... and dynamic reward upper limit This constitutes the feasible domain of dynamic rewards. ; Calculate the minimum reward lower bound based on incentive compatibility constraints. The formula is as follows: ; in, To reduce testing costs, This represents the difference in average transaction fees between abnormal and normal transactions. This represents the probability that an abnormal transaction will be identified after it has been recorded on the blockchain. The intensity of the equivalent ecological loss to the block builder after an abnormal transaction is identified. To account for the anticipated ecological penalties that may be incurred if testing is not carried out. The probability share of a builder's block being adopted. It is a numerically stable term; The dynamic upper limit of the reward is determined based on the reward cap and the proportion of the remaining budget. The formula is as follows: ; in, This is the maximum reward.

[0011] Preferably, the actual reward is projected onto the dynamic reward feasible region through a linear mapping, as shown in the following formula: ; in, for The actual reward value at any given moment. This is the original action.

[0012] Preferably, in step S4, the builder detection ratio is coupled with the normal user transaction ratio, and the specific steps are as follows: Step S41: Read the current system status and environmental parameter set; The current system status includes: Block builder detection ratio at any given time , Normal transaction ratio of users at any time , Changes in the proportion of block builder detections at any given time , Changes in the proportion of normal transactions by users at any given time , Remaining budget percentage at any given time and actual reward value ; The specific set of environmental parameters is as follows: ; in, A set of environmental parameters For detection rate, For false positive rate, This is the average transaction fee for normal transactions. The average transaction fee for abnormal transactions. For waiting cost coefficient, This is the normal transaction confirmation time. This is the confirmation time for abnormal transactions. This represents the maximum time cost for users to wait for confirmation after a transaction has been filtered. The utility and benefits for users to complete normal transactions. These are the illegal profits obtained after the successful execution of an abnormal transaction; Step S42: Calculate the expected benefit and benefit difference between the builder performing detection and not performing detection; When the builder performs the detection, the expected benefit is calculated as follows: ; in, Expected benefits when performing detection for the builder; When the builder does not perform the detection, the expected revenue is calculated as follows: ; in, The expected benefit when the builder does not perform detection; The expected revenue difference on the builder side is calculated using the following formula: ; in, The difference in expected revenue between performing detection and not performing detection for the builder; when When the value is >0, it indicates that under the current state and reward level, performing detection has higher fitness than not performing detection, which will drive up the proportion of builder detections; when When the value is less than 0, it will drive down the proportion of builders detected. Step S43: Calculate the expected profit and profit difference between normal and abnormal transactions initiated by the user; The expected return when a user initiates a normal transaction is calculated as follows: ; in, The expected return for a user when initiating a normal transaction; The expected return when a user initiates an abnormal transaction is calculated as follows: ; in, The expected return when a user initiates an abnormal transaction; The formula for calculating the expected revenue difference on the user side is as follows: ; in, The difference between the expected returns of normal and abnormal transactions on the user's side; when A value greater than 0 indicates that, under the current detection environment, users are more inclined to initiate normal transactions, thereby driving... Rise; when If the value is less than 0, it indicates that abnormal transactions are more attractive at the current stage, thus leading to... decline; Step S44: Use discrete replication dynamic equations to couple and update the builder detection ratio and the user normal transaction ratio, and apply interval projection constraints; The formula for updating the builder detection ratio is as follows: ; in, Before interval projection Candidate values ​​for the proportion of block builder detection at any given time. For the evolution rate coefficient on the builder side; The formula for updating the user's normal transaction ratio is as follows: ; in, Before interval projection Candidate values ​​for the proportion of normal transactions by users at any given time. The evolution rate coefficient on the user side; To prevent numerical updates from causing the scale to exceed the limit, an interval projection is performed on the candidate update values, as shown in the following formula: ; ; in, To crop the input variable to the [0,1] interval, for The block builder detection rate at time +1 for The percentage of users making normal transactions at time +1.

[0013] Preferably, in step S5, the immediate return is calculated as follows: ; in, For immediate returns, - These are non-negative weighting coefficients. The nearest neighbor threshold, for The degree of system deviation at time +1 The system deviation threshold, The system oscillation threshold, For indicator functions; The reinforcement learning algorithm employs the proximal policy optimization (PPO) algorithm, updating the incentive policy by pruning the policy probability ratio, as shown in the following formula: ; in, Let PPO be the objective function for the pruning strategy. For policy network parameters, The probability ratio between the old and new strategies. This is the estimated value of the dominance function. This is the cutting factor. For the clipping function, To calculate the expectation over the sampling time step or sample trajectory; The formula for the probability ratio between the old and new strategies is as follows: ; in, The ratio of the probabilities of the new and old strategies. For the current policy in state Down Output Action The probability, For the old strategy in state Down Output Action The probability of.

[0014] Preferably, in step S6, it is determined whether the system has entered the stable region based on the system deviation threshold and the system oscillation threshold: The evolutionary game environment is considered to have entered the target stable region if and only if the system simultaneously satisfies the constraints of the system deviation threshold and the system oscillation threshold for K consecutive time steps. ; in, The deviation threshold, This is the oscillation threshold.

[0015] This invention also provides a reinforcement learning-based dynamic incentive system for blockchain transaction security detection, used to implement the above-mentioned reinforcement learning-based dynamic incentive method for blockchain transaction security detection, comprising: The status acquisition module is used to collect the block builder detection ratio, the user normal transaction ratio, and the ratio change between adjacent time points in real time. It calculates the system deviation, oscillation intensity, and remaining budget ratio, and combines them with the actual reward value of the previous time point to construct a multi-dimensional system state vector. The policy reasoning module is used to input the system state vector output by the state acquisition module into the trained reinforcement learning policy network, perform forward reasoning, and output the original continuous actions. The reward interval calculation module is used to calculate the minimum reward lower bound based on incentive compatibility constraints, and determine the dynamic reward upper bound based on the ratio of the budget upper limit to the remaining budget, forming a reward feasible region that changes in real time with the system status. The reward projection module is used to project the original continuous actions output by the strategy reasoning module onto the dynamic reward feasible region through a linear mapping method, generate the actual reward value that meets the budget and incentive constraints, and perform reward truncation when the budget is insufficient. The environment evolution module is used to inject actual rewards into the evolutionary game environment, calculate the expected revenue and revenue difference for block builder detection or non-detection, and normal or abnormal user transactions, respectively, and use discrete replication dynamic equations to couple and update the builder detection ratio and the user normal transaction ratio, and perform interval constraint projection. The reward calculation module is used to calculate the comprehensive real-time reward based on the system's convergence progress toward the stable region, the reward expenditure cost, the oscillation intensity after convergence, and whether it has entered the target stable region. The policy update module is used to iteratively update the policy network based on the state vector, actual reward, immediate reward and state transition samples, using a reinforcement learning algorithm to obtain the optimal dynamic incentive policy. The online deployment module is used to load the trained policy network, replace the training process to perform off-chain real-time incentive control, and switch to a low-intensity sustain incentive mode after the system enters the stable region.

[0016] Therefore, this invention proposes a dynamic incentive method for blockchain transaction security detection based on reinforcement learning, the beneficial effects of which are as follows: (1) Based on the block builder detection ratio, the user normal transaction ratio, the system deviation, the oscillation intensity and the remaining budget, the present invention adaptively adjusts the reward level, which fundamentally overcomes the defects of traditional fixed rewards and preset decay rewards such as insufficient early incentives, excessive late incentives and budget waste, so that the reward distribution is highly matched with the system demand, and improves the effectiveness and sustainability of incentives.

[0017] (2) This invention explicitly integrates incentive compatibility constraints and budget constraints into the reinforcement learning decision-making process, dynamically calculates the feasible region of rewards, optimizes the reinforcement learning strategy within the constraint space, avoids meaningless exploration, and significantly improves training efficiency; at the same time, the reward output has a clear theoretical basis and boundary, greatly enhances the interpretability, stability and security of the online control process, and avoids system fluctuations caused by abnormal reward values.

[0018] (3) The present invention constructs a two-layer architecture of upper-layer reinforcement learning controller and lower-layer evolutionary game environment. Through the closed-loop feedback of reward distribution, builder detection behavior, user transaction behavior and system stability, it rapidly promotes the system to converge to a safe and stable region. When it is close to equilibrium, it automatically reduces the reward and suppresses system oscillation. Under limited budget, it achieves high convergence speed, strong operation stability and good cost efficiency at the same time.

[0019] (4) The present invention adopts a multi-dimensional state vector that includes group proportion, change trend, deviation, oscillation, budget margin and historical reward to comprehensively depict the system operation status, so that the incentive strategy can cope with complex environmental changes, user behavior fluctuations and budget constraints, and remain stable and reliable under different initial states and operating conditions.

[0020] (5) This invention is implemented as an upper-layer component of the incentive layer / governance layer. It does not modify the underlying consensus, node logic and transaction process of the blockchain. It can be directly connected to the existing transaction security detection system, and is compatible with various blockchain architectures such as public chain and consortium chain. It is easy to deploy, requires little modification, and has strong engineering practicality and scenario adaptability.

[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0022] Figure 1 This is a flowchart of a dynamic incentive method for blockchain transaction security detection based on reinforcement learning, according to the present invention. Figure 2 This is a schematic diagram of the overall framework of a dynamic incentive method for blockchain transaction security detection based on reinforcement learning, as proposed in this invention. Figure 3 This is a flowchart illustrating the coupling of the update builder detection ratio and the normal user transaction ratio in this invention. Detailed Implementation

[0023] To make the technical solutions, advantages, and objectives of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below. The described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the protection scope of the present invention.

[0024] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0025] Example 1 In this embodiment, the system operates in an application scenario where block builders perform security checks on candidate transaction sets. During initialization, the detection cost is configured first. Detection performance parameters, budget limit Reward cap Threshold deviation threshold in stable region With oscillation threshold And state transition parameters used to characterize environmental evolution. Then, the block builder detection ratio is randomly initialized. Ratio of normal transactions with users And based on this, an initial state vector is formed.

[0026] like Figures 1-2 As shown, this invention provides a dynamic incentive method for blockchain transaction security detection based on reinforcement learning, and the specific steps are as follows: Step S1: By setting the block builder detection ratio and the normal user transaction ratio, and configuring environmental parameters, including detection cost, detection rate, false positive rate, transaction fee difference, ecological penalty, budget limit and target stable area threshold, a secure detection incentive environment for blockchain transactions is established. Step S2: Construct the current system state based on the builder detection ratio, the user normal transaction ratio, the ratio change, the system deviation, the oscillation intensity, the remaining budget ratio, and the reward value of the previous moment; The current system status is as follows: ; in, for The system state at any given moment. for The block builder detection rate at any given time. for The proportion of normal transactions by users at any given time. for The change in the proportion of block builder detection at any given time. for The change in the proportion of normal transactions by users at any given time. for The degree of system deviation at any given time for The intensity of system oscillation at time t, for The percentage of the remaining budget at any given time. for The actual reward value at any given moment.

[0027] System deviation The calculation formula is as follows: ; in, , This is the deviation weighting coefficient; System oscillation intensity The calculation formula is as follows: ; in, The difference in the detection ratio of block builders at adjacent time points. The difference between the proportion of normal transactions by users at adjacent time points. for The block builder detection rate at any given time. for The proportion of normal transactions by users at any given time; the initial time. When =0, ; Remaining budget percentage The calculation formula is as follows: ; in, As of Cumulative reward expenditure at any time This is the budget ceiling.

[0028] Step S3: Calculate the dynamic reward feasible region based on the incentive compatibility constraint and budget constraint, and always project the original action within the dynamic reward feasible region to obtain the actual reward value; Minimum reward lower bound and dynamic reward upper limit This constitutes the feasible domain of dynamic rewards. ; Calculate the minimum reward lower bound based on incentive compatibility constraints. The formula is as follows: ; in, To reduce testing costs, This represents the difference in average transaction fees between abnormal and normal transactions. This represents the probability that an abnormal transaction will be identified after it has been recorded on the blockchain. The intensity of the equivalent ecological loss to the block builder after an abnormal transaction is identified. To account for the anticipated ecological penalties that may be incurred if testing is not carried out. The probability share of a builder's block being adopted. It is a numerically stable term; The dynamic upper limit of the reward is determined based on the reward cap and the proportion of the remaining budget. The formula is as follows: ; in, This is the maximum reward amount. The actual reward is achieved by projecting the original action onto the dynamic reward feasible region through a linear mapping, as shown in the following formula: ; in, for The actual reward value at any given moment. This is the original action.

[0029] Step S4: Inject the actual reward value into the evolutionary game environment, couple the update builder detection ratio with the user normal transaction ratio, and obtain the system state at the next moment; like Figure 3 As shown, the coupling update builder detection ratio and the user normal transaction ratio are implemented through the following steps: Step S41: Read the current system status and environmental parameter set; The current system status includes: Block builder detection ratio at any given time , Normal transaction ratio of users at any time , Changes in the proportion of block builder detections at any given time , Changes in the proportion of normal transactions by users at any given time , Remaining budget percentage at any given time and actual reward value ; The specific set of environmental parameters is as follows: ; in, A set of environmental parameters For detection rate, For false positive rate, This is the average transaction fee for normal transactions. The average transaction fee for abnormal transactions. For waiting cost coefficient, This is the normal transaction confirmation time. This is the confirmation time for abnormal transactions. This represents the maximum time cost for users to wait for confirmation after a transaction has been filtered. The utility and benefits for users to complete normal transactions. These are the illegal profits obtained after the successful execution of an abnormal transaction; Step S42: Calculate the expected benefit and benefit difference between the builder performing detection and not performing detection; When the builder performs the detection, the expected benefit is calculated as follows: ; in, Expected benefits when performing detection for the builder; When the builder does not perform the detection, the expected revenue is calculated as follows: ; in, The expected benefit when the builder does not perform detection; The expected revenue difference on the builder side is calculated using the following formula: ; in, The difference in expected revenue between performing detection and not performing detection for the builder; when When the value is >0, it indicates that under the current state and reward level, performing detection has higher fitness than not performing detection, which will drive up the proportion of builder detections; when When the value is less than 0, it will drive down the proportion of builders detected. Step S43: Calculate the expected profit and profit difference between normal and abnormal transactions initiated by the user; The expected return when a user initiates a normal transaction is calculated as follows: ; in, The expected return for a user when initiating a normal transaction. The utility and benefits for users to complete normal transactions; The expected return when a user initiates an abnormal transaction is calculated as follows: ; in, The expected return when a user initiates an abnormal transaction. These are the illegal profits obtained after the successful execution of an abnormal transaction; The formula for calculating the expected revenue difference on the user side is as follows: ; in, The difference between the expected returns of normal and abnormal transactions on the user's side; when A value greater than 0 indicates that, under the current detection environment, users are more inclined to initiate normal transactions, thereby driving... Rise; when If the value is less than 0, it indicates that abnormal transactions are more attractive at the current stage, thus leading to... decline; Step S44: Use discrete replication dynamic equations to couple and update the builder detection ratio and the user normal transaction ratio, and apply interval projection constraints; The formula for updating the builder detection ratio is as follows: ; in, Before interval projection Candidate values ​​for the proportion of block builder detection at any given time. For the evolution rate coefficient on the builder side; The formula for updating the user's normal transaction ratio is as follows: ; in, Before interval projection Candidate values ​​for the proportion of normal transactions by users at any given time. The evolution rate coefficient on the user side; To prevent numerical updates from causing the scale to exceed the limit, an interval projection is performed on the candidate update values, as shown in the following formula: ; ; in, To crop the input variable to the [0,1] interval, for The block builder detection rate at time +1 for The percentage of users making normal transactions at time +1.

[0030] Step S5: Calculate the immediate reward by combining the system convergence progress, reward expenditure, oscillation suppression and positive reward after entering the stable region, and update the incentive strategy through reinforcement learning algorithm; The calculation of immediate returns is as follows: ; in, For immediate returns, - These are non-negative weighting coefficients. The nearest neighbor threshold, for The degree of system deviation at time +1 The system deviation threshold, The system oscillation threshold, For indicator functions; The reinforcement learning algorithm employs the proximal policy optimization (PPO) algorithm, updating the incentive policy by pruning the policy probability ratio, as shown in the following formula: ; in, Let PPO be the objective function for the pruning strategy. For policy network parameters, The probability ratio between the old and new strategies. This is the estimated value of the dominance function. This is the cutting factor. For the clipping function, To calculate the expectation over the sampling time step or sample trajectory; The formula for the probability ratio between the old and new strategies is as follows: ; in, The ratio of the probabilities of the new and old strategies. For the current policy in state Down Output Action The probability, For the old strategy in state Down Output Action The probability of.

[0031] Step S6: When the system meets the stability threshold for multiple consecutive steps, it is determined to have entered a safe and stable state and training is terminated or switched to a low-maintenance stimulus mode.

[0032] Determine whether the system has entered the stable region based on the system deviation threshold and the system oscillation threshold: The evolutionary game environment is considered to have entered the target stable region if and only if the system simultaneously satisfies the constraints of the system deviation threshold and the system oscillation threshold for K consecutive time steps. ; in, The deviation threshold, This is the oscillation threshold.

[0033] The method of the present invention will be further illustrated below through specific implementation examples.

[0034] Example 2 This embodiment conducts experiments around three objectives: First, to verify whether the method of the present invention can drive the system to a safe and stable state faster with lower incentive costs under budget constraints; second, to verify whether the method of the present invention can achieve higher stability and better cost efficiency compared with fixed reward methods and rule-based reward methods; and third, to verify whether the components of the method of the present invention, such as state awareness, reward constraints, and temporal modeling, are necessary for the overall performance improvement.

[0035] This embodiment uses comparison objects including a fixed low-reward method, a fixed high-reward method, and three typical rule-based reward methods. The fixed low-reward method consistently provides a low constant reward to the block builder performing the detection throughout the entire operation; the fixed high-reward method consistently provides a high constant reward; and the rule-based reward methods adjust the reward through linear decay, deviation correction, and rule decay, respectively. The above methods are compared with the method of this invention under the same test scenario and evaluation criteria.

[0036] 1. High-complexity dynamic disturbance test scenario.

[0037] This embodiment constructs a high-complexity dynamic perturbation test scenario. In the process of the evolution of the block builder detection ratio, the normal user transaction ratio and related environmental parameters, random parameter fluctuations and sudden single-point variable mutations are introduced to simulate common state changes, profit fluctuations and sudden risk shocks in the operation of the blockchain.

[0038] 2. Running parameter settings.

[0039] Each method was repeated 10 times in independent experiments, with each experiment running 500 decision steps; training and inference were both performed on a single-machine server. The Proximal Policy Optimization (PPO) algorithm was preferred for policy updates.

[0040] 3. Comparison of experimental results.

[0041] In the same high-complexity dynamic perturbation test scenario, the method of the present invention is compared with the fixed low reward method and the fixed high reward method, and the results are shown in Table 1.

[0042] Table 1. Comparison of the proposed method and the fixed reward method under high-complexity dynamic perturbation test scenarios.

[0043] As shown in Table 1, although the fixed low-reward method has the lowest incentive expenditure, its safe and stable region occupancy rate is only 0.21, and the time required to reach a stable state is as high as 148.9 steps, indicating that it is difficult to maintain system stability continuously. The fixed high-reward method can shorten the stabilization time to 59.1 steps, but the cumulative additional reward consumption increases to 31.51 × 10 4 The previous method relied heavily on high incentives to achieve faster convergence. In contrast, the method of this invention, while maintaining a convergence speed close to that of the fixed high-reward method, reduces the cumulative additional reward cost to 23.92 × 10⁻⁶. 4 Furthermore, the occupancy rate of the safe and stable region was increased to 0.79, indicating that the method of the present invention can achieve a more balanced optimization between cost, convergence speed and stability.

[0044] To further verify the advantages of the method of the present invention compared with existing interpretable rule strategies, the method of the present invention was compared with three typical rule-based reward methods, and the results are shown in Table 2.

[0045] Table 2. Comparison of the method of this invention with three types of rule-based reward methods under high-complexity dynamic perturbation test scenarios.

[0046] As shown in Table 2, the cumulative additional reward cost for the three types of rule-based reward methods ranges from 26.90 × 10 4 Up to 28.27×10 4 Between these values, the occupancy rate of the safe and stable region ranges from 0.51 to 0.56, and the time required to reach a stable state ranges from 84.2 to 93.6 steps; in contrast, the cumulative additional reward cost of the method of this invention is 23.92 × 10⁻⁶. 4 Compared to linear decay, distance correction, and regular decay methods, the time required to reach a stable state is reduced by 2.98, 4.35, and 3.64, respectively; the occupancy rate of the safe and stable region is increased to 0.79, an improvement of 0.28, 0.23, and 0.27, respectively; and the time required to reach a stable state is shortened to 61.4 steps, a reduction of 29.6, 22.8, and 32.2 steps, respectively. Therefore, the method of this invention does not merely improve a single indicator, but demonstrates superior overall performance across all three core indicators.

[0047] Example 3 To demonstrate that the method of the present invention has good adaptability in dynamic environments ranging from low to high complexity, this embodiment further constructs four types of test scenarios with increasing complexity and statistically analyzes the main mean results of the method of the present invention in different scenarios, as shown in Table 3.

[0048] Table 3. Adaptability verification results of the method of the present invention under different complexity test scenarios.

[0049] As shown in Table 3, with the increase in the complexity of the test scenario, the cumulative additional reward consumption of the method of this invention increases from 21.6 × 10⁻⁶. 4 Increased to 23.92×10 4 The increase was approximately 10.7%; the occupancy of the safe and stable region decreased from 0.95 to 0.79; and the time required to reach a stable state increased from 35.6 steps to 61.4 steps. These changes indicate that under more complex dynamic disturbance conditions, the system requires higher excitation costs and undergoes a longer stabilization process. However, the overall performance degradation is gradual rather than abrupt. The method of this invention still maintains a high occupancy of the safe and stable region and an acceptable convergence speed, demonstrating good environmental adaptability and robustness.

[0050] Example 4 This invention also provides a dynamic incentive system for blockchain transaction security detection based on reinforcement learning, comprising: The status acquisition module is used to collect the block builder detection ratio, the user normal transaction ratio, and the ratio change between adjacent time points in real time. It calculates the system deviation, oscillation intensity, and remaining budget ratio, and combines them with the actual reward value of the previous time point to construct a multi-dimensional system state vector. The policy reasoning module is used to input the system state vector output by the state acquisition module into the trained reinforcement learning policy network, perform forward reasoning, and output the original continuous actions. The reward interval calculation module is used to calculate the minimum reward lower bound based on incentive compatibility constraints, and determine the dynamic reward upper bound based on the ratio of the budget upper limit to the remaining budget, forming a reward feasible region that changes in real time with the system status. The reward projection module is used to project the original continuous actions output by the strategy reasoning module onto the dynamic reward feasible region through a linear mapping method, generate the actual reward value that meets the budget and incentive constraints, and perform reward truncation when the budget is insufficient. The environment evolution module is used to inject actual rewards into the evolutionary game environment, calculate the expected revenue and revenue difference for block builder detection or non-detection, and normal or abnormal user transactions, respectively, and use discrete replication dynamic equations to couple and update the builder detection ratio and the user normal transaction ratio, and perform interval constraint projection. The reward calculation module is used to calculate the comprehensive real-time reward based on the system's convergence progress toward the stable region, the reward expenditure cost, the oscillation intensity after convergence, and whether it has entered the target stable region. The policy update module is used to iteratively update the policy network based on the state vector, actual reward, immediate reward and state transition samples, using a reinforcement learning algorithm to obtain the optimal dynamic incentive policy. The online deployment module is used to load the trained policy network, replace the training process to perform off-chain real-time incentive control, and switch to a low-intensity sustain incentive mode after the system enters the stable region.

[0051] In this embodiment, the system operates in an application scenario where block builders perform security checks on candidate transaction sets. During initialization, the detection cost is configured first. Detection performance parameters, budget limit Reward cap Threshold deviation threshold in stable region With oscillation threshold And state transition parameters used to characterize environmental evolution. Then, the block builder detection ratio is randomly initialized. Ratio of normal transactions with users And based on this, an initial state vector is formed.

[0052] In this embodiment, the specific workflow of the system is as follows: 1. The status acquisition module obtains the current time. , And historical changes, and calculate the deviation. Oscillation Remaining budget percentage Rewards from the previous moment .

[0053] 2. The strategy reasoning module outputs the original continuous actions based on the current state vector. .

[0054] 3. The reward interval calculation module calculates the current minimum reward lower bound based on parameters such as the current user status, detection cost, detection rate, false positive rate, and ecosystem penalties. Then calculate the upper limit of the reward based on the reward cap and the remaining budget percentage. .

[0055] 4. The reward projection module will project the original action. Mapped to actual rewards This places the actual reward in Within the budget; if the current budget is insufficient to cover the theoretical lower bound, the maximum executable reward within the budget limit is output using a truncation method.

[0056] 5. The environmental evolution module will provide actual rewards. Input the evolutionary game environment and update the block builder detection ratio for the next time step. Ratio of normal transactions with users In this process, rewards first influence whether builders are willing to perform checks, and then indirectly affect user behavior through changes in builder strategies.

[0057] 6. The reward calculation module generates real-time rewards based on factors such as the reduction in the system's distance from the target stable region, reward expenditure, oscillation magnitude, and whether the system has entered the stable region. .

[0058] 7. In training mode, the policy update module uses the PPO algorithm to iteratively update the policy network and value network based on the collected state transition samples until a stable dynamic incentive policy is obtained.

[0059] 8. In online deployment mode, the training process of the policy update module is replaced by a trained policy network. Only the closed loop of state awareness, reward interval calculation, reward projection and environment update is retained to output dynamic reward values ​​based on the real-time observed state on or off chain.

[0060] 9. When the system meets the system deviation threshold and system oscillation threshold for K consecutive time steps, it is determined that it has entered the target safe and stable region. The current round can be ended, or the system can be switched to a low-intensity reward maintenance state to continue to control budget expenditure.

[0061] It is worth noting that all contents not described in detail in this invention are existing technologies and are well known to those skilled in the art.

[0062] Therefore, this invention provides a dynamic incentive method for blockchain transaction security detection based on reinforcement learning. Through a two-layer closed-loop architecture combining reinforcement learning and evolutionary game theory, it achieves adaptive dynamic adjustment of rewards, quickly guides the system to a safe and stable state under budget and incentive compatibility constraints, improves incentive efficiency and governance effect, and does not require modification of the underlying blockchain and is easy to implement in engineering.

[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A dynamic incentive method for blockchain transaction security detection based on reinforcement learning, characterized in that, The steps are as follows: Step S1: Establish a secure detection incentive environment for blockchain transactions by setting the block builder detection ratio and the normal user transaction ratio, and configuring environmental parameters. Step S2: Construct the current system state based on the builder detection ratio, the user normal transaction ratio, the ratio change, the system deviation, the oscillation intensity, the remaining budget ratio, and the reward value of the previous moment; Step S3: Calculate the dynamic reward feasible region based on the incentive compatibility constraint and budget constraint, and always project the original action within the dynamic reward feasible region to obtain the actual reward value; Step S4: Inject the actual reward value into the evolutionary game environment, couple the update builder detection ratio with the user normal transaction ratio, and obtain the system state at the next moment; Step S5: Calculate the immediate reward by combining the system convergence progress, reward expenditure, oscillation suppression and positive reward after entering the stable region, and update the incentive strategy through reinforcement learning algorithm; Step S6: When the system meets the stability threshold for multiple consecutive steps, it is determined to have entered a safe and stable state and training is terminated or switched to a low-maintenance stimulus mode.

2. The dynamic incentive method for blockchain transaction security detection based on reinforcement learning according to claim 1, characterized in that, In step S1, the environmental parameters include detection cost, detection rate, false positive rate, handling fee difference, ecological penalty, budget limit and target stable area threshold.

3. The dynamic incentive method for blockchain transaction security detection based on reinforcement learning according to claim 2, characterized in that, In step S2, the current system state is represented as follows: ; in, for The system state at any given moment. for The block builder detection rate at any given time. for The proportion of normal transactions by users at any given time. for The change in the proportion of block builder detection at any given time. for The change in the proportion of normal transactions by users at any given time. for The degree of system deviation at any given time for The intensity of system oscillation at time t, for The percentage of the remaining budget at any given time. for The actual reward value at any given moment.

4. The dynamic incentive method for blockchain transaction security detection based on reinforcement learning according to claim 3, characterized in that, System deviation The calculation formula is as follows: ; in, , This is the deviation weighting coefficient; System oscillation intensity The calculation formula is as follows: ; in, The difference in the detection ratio of block builders at adjacent time points. The difference between the proportion of normal transactions by users at adjacent time points. for The block builder detection rate at any given time. for The proportion of normal transactions by users at any given time; the initial time. When =0, ; Remaining budget percentage The calculation formula is as follows: ; in, As of Cumulative reward expenditure at any time This is the budget ceiling.

5. The dynamic incentive method for blockchain transaction security detection based on reinforcement learning according to claim 4, characterized in that, In step S3, the minimum reward lower bound and dynamic reward upper limit This constitutes the feasible domain of dynamic rewards. ; Calculate the minimum reward lower bound based on incentive compatibility constraints. The formula is as follows: ; in, To reduce testing costs, This represents the difference in average transaction fees between abnormal and normal transactions. This represents the probability that an abnormal transaction will be identified after it has been recorded on the blockchain. The intensity of the equivalent ecological loss to the block builder after an abnormal transaction is identified. To account for the anticipated ecological penalties that may be incurred if testing is not carried out. The probability share of a builder's block being adopted. It is a numerically stable term; The dynamic upper limit of the reward is determined based on the reward cap and the proportion of the remaining budget. The formula is as follows: ; in, This is the maximum reward.

6. The dynamic incentive method for blockchain transaction security detection based on reinforcement learning according to claim 5, characterized in that, The actual reward is achieved by projecting the original action onto the dynamic reward feasible region through a linear mapping, as shown in the following formula: ; in, for The actual reward value at any given moment. This is the original action.

7. The dynamic incentive method for blockchain transaction security detection based on reinforcement learning according to claim 6, characterized in that, In step S4, the ratio of builder detection is coupled with the ratio of normal user transactions. The specific steps are as follows: Step S41: Read the current system status and environmental parameter set; The current system status includes: Block builder detection ratio at any given time , Normal transaction ratio of users at any time , Changes in the proportion of block builder detections at any given time , Changes in the proportion of normal transactions by users at any given time , Remaining budget percentage at any given time and actual reward value ; The specific set of environmental parameters is as follows: ; in, A set of environmental parameters For detection rate, For false positive rate, This is the average transaction fee for normal transactions. The average transaction fee for abnormal transactions. For waiting cost coefficient, This is the normal transaction confirmation time. This is the confirmation time for abnormal transactions. This represents the maximum time cost for users to wait for confirmation after a transaction has been filtered. The utility and benefits for users to complete normal transactions. These are the illegal profits obtained after the successful execution of an abnormal transaction; Step S42: Calculate the expected benefit and benefit difference between the builder performing detection and not performing detection; When the builder performs the detection, the expected benefit is calculated as follows: ; in, Expected benefits when performing detection for the builder; When the builder does not perform the detection, the expected revenue is calculated as follows: ; in, The expected benefit when the builder does not perform detection; The expected revenue difference on the builder side is calculated using the following formula: ; in, The difference in expected revenue between performing detection and not performing detection for the builder; when When the value is >0, it indicates that under the current state and reward level, performing detection has higher fitness than not performing detection, which will drive up the proportion of builder detections; when When the value is less than 0, it will drive down the proportion of builders detected. Step S43: Calculate the expected profit and profit difference between normal and abnormal transactions initiated by the user; The expected return when a user initiates a normal transaction is calculated as follows: ; in, The expected return for a user when initiating a normal transaction; The expected return when a user initiates an abnormal transaction is calculated as follows: ; in, The expected return when a user initiates an abnormal transaction; The formula for calculating the expected revenue difference on the user side is as follows: ; in, The difference between the expected returns of normal and abnormal transactions on the user's side; when A value greater than 0 indicates that, under the current detection environment, users are more inclined to initiate normal transactions, thereby driving... Rise; when If the value is less than 0, it indicates that abnormal transactions are more attractive at the current stage, thus leading to... decline; Step S44: Use discrete replication dynamic equations to couple and update the builder detection ratio and the user normal transaction ratio, and apply interval projection constraints; The formula for updating the builder detection ratio is as follows: ; in, Before interval projection Candidate values ​​for the proportion of block builder detection at any given time. For the evolution rate coefficient on the builder side; The formula for updating the user's normal transaction ratio is as follows: ; in, Before interval projection Candidate values ​​for the proportion of normal transactions by users at any given time. The evolution rate coefficient on the user side; To prevent numerical updates from causing the scale to exceed the limit, an interval projection is performed on the candidate update values, as shown in the following formula: ; ; in, To crop the input variable to the [0,1] interval, for The block builder detection rate at time +1 for The percentage of users making normal transactions at time +1.

8. The dynamic incentive method for blockchain transaction security detection based on reinforcement learning according to claim 7, characterized in that, In step S5, the immediate return is calculated as follows: ; in, For immediate returns, - These are non-negative weighting coefficients. The nearest neighbor threshold, for The degree of system deviation at time +1 The system deviation threshold, The system oscillation threshold, For indicator functions; The reinforcement learning algorithm employs the proximal policy optimization (PPO) algorithm, updating the incentive policy by pruning the policy probability ratio, as shown in the following formula: ; in, Let PPO be the objective function for the pruning strategy. For policy network parameters, The probability ratio between the old and new strategies. This is the estimated value of the dominance function. This is the cutting factor. For the clipping function, To calculate the expectation over the sampling time step or sample trajectory; The formula for the probability ratio between the old and new strategies is as follows: ; in, The ratio of the probabilities of the new and old strategies. For the current policy in state Down Output Action The probability, For the old strategy in state Down Output Action The probability of.

9. The dynamic incentive method for blockchain transaction security detection based on reinforcement learning according to claim 8, characterized in that, In step S6, the system is determined to have entered the stable region based on the system deviation threshold and the system oscillation threshold. The evolutionary game environment is considered to have entered the target stable region if and only if the system simultaneously satisfies the constraints of the system deviation threshold and the system oscillation threshold for K consecutive time steps. ; in, The deviation threshold, This is the oscillation threshold.