Anti-noise federated smart contract vulnerability detection method and system with probability label estimation

Through the combination of probability label estimation and meta-learning training algorithm, the problems of label noise and heterogeneous noise in smart contract vulnerability detection are solved, the accuracy and adaptability of the model are improved, and efficient detection in complex scenarios is ensured.

CN120337214AActive Publication Date: 2025-07-18GUANGDONG UNIVERSITY OF FOREIGN STUDIES
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510311531.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-18
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

In the existing smart contract vulnerability detection technology, there are problems such as label noise affecting model performance, limited ability to deal with heterogeneous label noise, lack of effective label calibration methods, and poor adaptability in actual industrial scenarios. Especially in federated learning, noise propagation and data privacy protection limit the effectiveness and accuracy of the model.

Method used

The anti-noise federal smart contract vulnerability detection method using probability label estimation is used to construct a probability label model and a meta-learning training algorithm, and the real label probability is estimated using Bayes theorem and Markov random field model, and label calibration is performed based on global and local label information, and the model performance is improved through iterative optimization.

Benefits of technology

Effectively handle label noise, improve the accuracy and reliability of the model's vulnerability detection in smart contracts, adapt to complex and diverse data environments, enhance the adaptability and performance of the model in actual industrial scenarios, and improve the overall effect of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337214A_ABST
    Figure CN120337214A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent contract vulnerability detection, and discloses an anti-noise federated intelligent contract vulnerability detection method and system with probability tag estimation, and the method comprises the steps: taking a pseudo tag and a local tag as observation noise tags through a probability tag estimation module, constructing a probability tag model, and carrying out the recognition of the probability tag model; calculating a real label probability by comprehensively considering information of a global model pseudo label and a local label; a meta-learning training algorithm is introduced to optimize the probability label estimation network, and parameters of the probability label estimation network are continuously adjusted by utilizing feedback of a local model through iterative internal and external training steps, so that the accuracy of label calibration is improved. According to the method, the problems in the prior art can be effectively solved, wide applicability is achieved, a reliable and efficient solution can be provided for intelligent contract vulnerability detection in different scenes, the safety of the intelligent contract in a block chain system is powerfully guaranteed, and the safety application of the intelligent contract technology in wider fields is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to, but is not limited to, the technical field of intelligent contract vulnerability detection, and particularly relates to a noise-resistant federated intelligent contract vulnerability detection method and system with probabilistic label estimation. Background Art

[0002] With the development of blockchain technology, the application of intelligent contracts has become increasingly widespread. However, the problem of intelligent contract vulnerabilities seriously threatens the security of blockchain systems and may lead to financial losses and system attacks. To ensure the stable operation of blockchain systems, intelligent contract vulnerability detection has become a key technical field. In this field, due to its powerful end-to-end feature learning ability, deep learning technology has been widely concerned and applied to intelligent contract vulnerability detection. According to the representation learning methods used, existing methods can be roughly divided into three categories: sequence learning-based methods analyze the behavior of programs by extracting execution order features, graph learning-based methods represent intelligent contract codes as graphs to capture operation logic, and hybrid methods combine expert-defined vulnerability patterns to enrich vulnerability features. There is an obvious gap between existing deep learning-based solutions and actual industrial applications. Although deep learning methods have improved detection performance, they rely on supervised learning and often ignore the problem of label noise in datasets.

[0003] Currently, in the field of intelligent contract vulnerability detection, most existing deep learning methods assume that the labels of the training dataset are accurate. However, in reality, it is difficult to ensure the correctness of labels. On the one hand, as the logic of intelligent contracts becomes more complex, accurate annotation becomes increasingly difficult; on the other hand, although expert teams can create clean datasets, the high cost forces developers to rely on automatic annotation tools. However, these automatic annotation tools have limitations in dealing with complex intelligent contract logic and have a relatively high annotation error rate.

[0004] At the same time, federated learning technology has also been applied to this field to solve the data silo problem by training local models on the data owner side and aggregating them into a global model, combining global knowledge. Recent research uses pseudo-labels from the global model to improve label accuracy. However, different label capabilities among data owners will introduce different levels and types of label noise, resulting in noise propagation when aggregating models.

[0005] In view of the above analysis, the technical problems urgently to be solved in the prior art are as follows:

[0006] (1) Label noise affects model performance: In the existing technologies for intelligent contract vulnerability detection, due to the difficulty in ensuring the accuracy of the training dataset labels, the problem of label noise is widespread. For example, automated annotation tools can lead to a high mislabeling rate and a high false negative rate. When a deep detection model is trained on a noisy dataset, it will learn incorrect patterns, resulting in inaccurate predictions during the inference phase. For instance, in actual detection, an intelligent contract with vulnerabilities may be misjudged as a normal contract, or vice versa, greatly reducing the accuracy and reliability of the detection.

[0007] (2) Limited ability to handle heterogeneous label noise: In actual application scenarios, there are differences in the annotation capabilities of different data owners in federated learning, which makes the label noise in the dataset exhibit different levels and types. Existing federated learning methods are prone to being affected by noise propagation when aggregating local models and cannot effectively handle this heterogeneous label noise. This limits the effectiveness of the model in actual industrial scenarios, unable to fully utilize its detection capabilities in complex and diverse data environments, and difficult to meet actual requirements.

[0008] (3) Lack of effective label calibration methods: Although some federated learning methods attempt to use the pseudo-labels of the global model for label calibration, the accuracy of these pseudo-labels highly depends on the reliability of the global task model. In the case of diverse levels and types of label noise among data owners, it is difficult to ensure the reliability of the global task model, resulting in low accuracy of label calibration. For example, in the absence of comprehensive knowledge, relying solely on local data for label inference cannot accurately calibrate noisy labels, affecting the overall performance of the model.

[0009] (4) Poor adaptability to actual industrial scenarios: In actual industrial scenarios, the requirements for privacy protection limit the sharing and integration of data, and at the same time, the number of clean samples is often limited. Existing technologies are difficult to effectively utilize limited resources for model training and optimization in this situation and cannot fully adapt to the data characteristics and constraints in industrial scenarios. For example, under the premise of data privacy protection, it is difficult to obtain enough clean samples to accurately train the model, resulting in limited performance of the model in actual applications and unable to provide reliable security for intelligent contracts. Summary of the Invention

[0010] In view of the problems existing in the prior art, the present invention provides an anti-noise federated intelligent contract vulnerability detection method and system with probability label estimation.

[0011] The present invention is implemented as follows. An anti-noise federated intelligent contract vulnerability detection method with probability label estimation, characterized in that the anti-noise federated intelligent contract vulnerability detection method with probability label estimation specifically includes:

[0012] S1: Create a clean smart contract vulnerability dataset containing Reentrancy and Timestamp vulnerabilities;

[0013] S2: Construct a noisy dataset with two types of label noise: symmetric label noise and false negative label noise;

[0014] S3: The dataset is divided into a training set, a validation set, and a test set in proportion;

[0015] S4: Evenly distribute the training set to multiple data owners. Each data owner randomly selects smart contracts from its local training set as the clean validation set for training the probability label model, and the remaining smart contracts in the local training set are used to generate a noisy training set;

[0016] S5: Set the maximum noise levels of the two types of noise. Each smart contract in the training set is randomly flipped with the probability of different symmetric label noise levels to generate a dataset of symmetric label noise. For false negative label noise, the contracts marked as vulnerable are randomly flipped with the probability of different false negative label noise levels;

[0017] S6: Establish a deep learning model. The data center initializes the global deep learning model parameters and broadcasts them to all participating data owners in the federated learning;

[0018] S7: Each data owner in the federated learning uses the probability label estimation model to calibrate the local noisy dataset;

[0019] S8: Each data owner in the federated learning updates the local model using the calibrated dataset and the parameters of the received global model;

[0020] S9: Each data owner in the federated learning uploads the updated local model parameters to the data center. The data center weights and aggregates the parameters of all local models to update the global model parameters;

[0021] S10: Broadcast the globally updated model parameters again to the data owners, repeat the training process, and repeat S6 - S9;

[0022] S11: Use accuracy and F1-score as evaluation metrics to evaluate the performance of the present invention.

[0023] Further, in S2, the false negative label noise is an asymmetric label noise without false positive label noise.

[0024] Further, in S4, the training set is evenly distributed to multiple data owners. For multiple data owners, it means that in the federated learning, it consists of a data center and multiple data owners, and each data owner has a local dataset.

[0025] Further, in S7, the local noise dataset is calibrated using a probability label estimation module. The probability label estimation model includes: probability label modeling, a probability label estimation network, and optimization.

[0026] (1) Probability label modeling: By modeling the relationship between local labels, pseudo-labels, and true labels, the true label probability of each sample is estimated. The local label and the pseudo-label are regarded as the observed noise label y nl , and according to Bayes' theorem and the Markov random field model, the probability label model is derived. The probability label model is expressed as follows:

[0027]

[0028] where y nl is the observed noise label, y* is the true label of the smart contract, x is the smart contract feature, P(y* = 1) represents the prior probability that the true label of the smart contract is 1 (i.e., contains vulnerabilities), P(y* = 1) is a constant estimated from the local clean validation dataset of each data owner. The probability label model estimates the probability label of a given smart contract by weighted aggregation of the observed noise labels according to its accuracy score;

[0029] (2) Probability label estimation network: A deep neural network is used as the probability label estimation network to parameterize the probability label model for each data owner k. In the network, the accuracy score function s1(y nl , x) is parameterized by the scoring network . For any labeled smart contract (x, y) ∈ D k , first generate the pseudo-label and the one-hot representation . Remove the output layer from the global model to obtain the feature embedding h(x) of the smart contract. Estimate the class distribution P(y* = 1) of the local vulnerability dataset D on the local clean validation dataset k . Then the scoring network takes and h(x) as inputs, predicts the accuracy score vector of the observed noise label. The last layer is a softmax transformation layer for normalizing the accuracy score; Finally, the aggregator combines the local class distribution P(y* = 1), the one-hot representation of the observed noise label, and the accuracy score s x to generate the probability label

[0030] (3) Optimization. The joint optimization of the probability label estimation network and the local depth detection model ensures that the probability labels provided by the label estimation network can improve the detection performance of the local depth detection model, which in turn improves the label accuracy. The optimization objective is to minimize the loss of the detection model on the local clean validation dataset:

[0031]

[0032] where is the local clean validation dataset, is the probability label estimation network, h(x) is the feature of the smart contract, and y nl is the observed noise label, P(y* = 1) is the prior probability that the true label of the smart contract is 1 (i.e., contains vulnerabilities), is the prediction result of the local depth detection model of the k-th data owner for the smart contract x.

[0033] Furthermore, in S8, to adapt to the real industrial scenario with limited clean samples, a meta-learning-based training algorithm is used to jointly optimize the probability label estimation network and the local depth detection model with the validation set. This algorithm works through iterative inner and outer training steps:

[0034] (1) In the inner training step, the local depth detection model is optimized using the calibration labels generated by the probability label estimation network. The probability label estimation network generates a probability label y~ for each smart contract x. At the same time, the parameters of the local depth detection model are initialized using the parameters of the global model to generate a prediction model The inner training loss is calculated and the local model parameters are updated. The local depth detection model is updated as follows:

[0035]

[0036] where η1 is the inner training learning rate, is the inner training loss on the local training set D k .

[0037] (2) In the outer training step, the probability label estimation network is updated based on the validation loss of the updated local depth detection model on the validation dataset as follows:

[0038]

[0039] The calculation method is as follows:

[0040]

[0041] Among them, η2 is the external training learning rate, which represents the expected validation loss on the validation dataset.

[0042] Furthermore, in S9, the global model parameters are updated by aggregating the parameters from all data owners in the data center, and the process is as follows:

[0043]

[0044] where θ k is the local model parameter, θ is the global model parameter, n k is the number of samples owned by each data owner, and K represents the total number of data owners.

[0045] Another object of the present invention is to provide an anti-noise federated smart contract vulnerability detection system with probability label estimation, and the system specifically includes:

[0046] A dataset construction module for constructing a dataset and dividing it into a training set, a validation set, and a test set;

[0047] A probability label estimation module for estimating the true label probability of each sample by modeling the relationship between the local label, the pseudo label, and the true label;

[0048] A meta-learning training module for jointly optimizing the probability label estimation network and the local deep detection model with the validation set based on the training algorithm of meta-learning.

[0049] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0050] First, effectively handle label noise and improve model performance: Aiming at the problem that label noise affects the model performance in the prior art, the present invention regards the pseudo label and the local label as observed noise labels through the probability label estimation module, and constructs a probability label model to calculate the true label probability. This model can effectively capture the relationship between the observed noise label and the true label, thereby more accurately evaluating the reliability of the label, performing effective label calibration, improving the accuracy and reliability of the model in smart contract vulnerability detection, and further improving the overall performance. As shown in Table 6, the performance of GBGRU and CGE under different noise levels and Figure 4 、 Figure 5As a result, the performance of the depth detection model remains relatively stable until the training set contains 20% symmetric label noise or 15% to 25% false negative label noise. Specifically, in the case of symmetric label noise, the average performance metric of CBGRU does not show a significant decline until the noise level reaches 20%.

[0051] Adapting to heterogeneous label noise and enhancing model adaptability: For the problem of limited ability of existing technologies to handle heterogeneous label noise, the probability label model of the present invention calculates the true label probability by comprehensively considering the information of the global model pseudo-label and the local label, and no longer simply relies on the reliability of the global model. When facing different levels and types of label noise from different data owners, it can effectively integrate multi-source knowledge for label calibration. In the heterogeneous label noise scenario, the present invention can adapt to complex actual industrial scenarios, effectively handle heterogeneous label noise among different data owners, and enhance the adaptability and effectiveness of the model in diverse data environments. As shown in Tables 3 and 4, the performance of Reentrancy of CBGRU and CGE under symmetric label noise and false negative label noise at the heterogeneous level. The present invention achieves the highest F1 scores of CBGRU and CGE on the Reentrancy dataset, which are 82.34% and 83.69% respectively, and 90.75% and 84.79% on the Timestamp dataset respectively. Under symmetric label noise and false negative label noise with heterogeneous levels, the present invention is always superior to the second-best result. As shown in Table 5, the performance comparison of Reentrancy and Timestamp under heterogeneous type label noise, the present invention obtains the highest F1 scores, recording 84.79% for CBGRU and 87.22% for CGE on the Reentrancy dataset, and 85.87% for CBGRU and 85.89% for CGE on the Timestamp dataset. These results demonstrate the effectiveness of the present invention under heterogeneous level and heterogeneous type of noisy labels.

[0052] Improve label calibration accuracy and optimize model performance: In view of the lack of effective label calibration methods in the existing technology, the present invention introduces a meta-learning training algorithm to optimize the probability label estimation network. In an actual industrial scenario where clean samples are limited, this algorithm can continuously adjust the parameters of the probability label estimation network by using the feedback of the local model through iterative inner and outer training steps, thereby improving the accuracy of label calibration. The meta-learning training algorithm enables the present invention to better utilize the validation set information to optimize label calibration under the condition of limited clean samples, and further improve the overall detection performance of the model and optimize the effect of the model in intelligent contract vulnerability detection. By introducing a probability label model and a meta-learning-based training algorithm, the present invention can calibrate labels more accurately. In the ablation study, the effect of using the probability label model alone is limited, but after combining with the meta-learning-based training algorithm, under symmetric label noise, the F1 score of the Reentrancy dataset increases from 57.7% to 63.08%, and the Timestamp dataset increases from 72.78% to 77.62%; under false negative label noise, the Reentrancy dataset increases from 60.43% to 80.34%, and the Timestamp dataset increases from 72.93% to 85.24%, indicating that this algorithm can effectively utilize limited clean samples, improve label calibration accuracy, and further improve model performance.

[0053] Comprehensive advantages bring overall performance improvement and wide applicability: Through the above improvements in label noise processing, heterogeneous noise adaptation, and label calibration optimization, the present invention demonstrates a significant improvement in overall performance in different vulnerability detection tasks of smart contracts. Different from traditional federated learning-based label calibration methods, the present invention generates real label probabilities by using pseudo-labels from the global model and local labels provided by data owners. The present invention achieves this through a probabilistic label model, which captures the conditional dependence relationship between observed noisy labels and real labels. This model is parameterized by a dedicated probabilistic label estimation network. In addition, the present invention introduces a meta-learning-based training algorithm, enhancing the adaptability to real industrial scenarios where clean samples are usually limited. When dealing with heterogeneous noise, whether it is the difference in noise level or type, the present invention performs excellently. For example, in different types of label noise experiments, compared with other advanced methods, the present invention achieved F1 scores of 84.79% and 87.22% for CBGRU and CGE respectively on the Reentrancy dataset, and F1 scores of 85.87% and 85.89% respectively on the Timestamp dataset, both exceeding other methods. These results strongly prove that the present invention can effectively handle various complex noise situations and demonstrates an improvement in overall performance in different vulnerability detection tasks of smart contracts. The present invention promotes the development of high-performance global depth detection models with noisy and distributed datasets for the method of anti-noise federated smart contract vulnerability detection with probabilistic label estimation. Its wide applicability can provide reliable and efficient solutions for smart contract vulnerability detection in different scenarios, effectively guarantee the security of smart contracts in the blockchain system, greatly promote the secure application of smart contract technology in a wider range of fields, and bring new breakthroughs to the development of smart contract vulnerability detection technology.

[0054] Second, as an auxiliary evidence for the inventiveness of the claims of the present invention, it is also reflected in the following important aspects: (1) The technical solution of the present invention solves the technical problems that people have been eager to solve but have never succeeded in: due to the huge losses caused by hacker attacks, the detection of smart contract vulnerabilities is developing rapidly together with blockchain technology. Due to the powerful end-to-end feature learning ability of deep learning, it has shown significant advantages in this field. However, the existence of label noise in the smart contract vulnerability dataset will weaken the performance of the deep detection model. Existing research attempts to solve this technical problem by using the comprehensive task knowledge from multiple data owners to correct the noisy labels. These methods usually adopt federated learning to train the global task model and use the pseudo-labels from this model as the ground truth for further training. Despite these efforts, the current methods are still insufficient to handle the labeled noise with heterogeneous levels and types generated due to different labeling capabilities among data owners, which limits their effectiveness in actual industrial scenarios. On this basis, the present invention proposes an anti-noise federated smart contract vulnerability detection method with probability label estimation. This method regards the pseudo-labels and local labels as observed noisy labels and introduces a new probability label model to calculate the true label probability by capturing the conditional dependence relationship between the observed noisy labels and the true labels. This method effectively evaluates the reliability of the global and local vulnerability knowledge in label calibration, thereby improving the accuracy of label calibration. In addition, the present invention also implements a meta-learning-based training algorithm for the probability label model, enabling it to adapt to real industrial scenarios with limited clean samples.

[0055] Third, by introducing a probability label estimation mechanism, the present invention effectively solves the problem of the impact of noisy labels on the performance of the federated learning model in the prior art. In traditional federated learning, the existence of noisy labels will significantly reduce the accuracy and generalization ability of the model. However, the present invention calibrates the local noisy dataset by combining the global and local vulnerability knowledge to generate a label probability closer to the truth, thereby improving the robustness and detection performance of the model.

[0056] The prior art mostly focuses on dealing with a single type of label noise (such as symmetric noise or false negative noise), while ignoring the coexistence problem of multiple label noises in complex scenarios. The present invention innovatively deals with both symmetric label noise and false negative label noise at the same time. Through a flexible noise modeling method, the applicability of the federated learning system in diverse noise scenarios is significantly improved, overcoming the limitations of the prior art.

[0057] In traditional federated learning, the existence of noisy data not only degrades model performance but also may increase computational and communication overheads. Through the introduction of a probabilistic label estimation module, the present invention achieves effective calibration of the local dataset, reducing unnecessary computational and transmission burdens. The weight aggregation strategy further optimizes the efficiency of global model updates, significantly enhancing the performance of federated learning in large-scale distributed scenarios.

[0058] Aiming at the characteristics of limited clean samples in industrial scenarios, the present invention adopts a joint optimization algorithm based on meta-learning to ensure the co-optimization of the local deep detection model and the probabilistic label estimation network while reducing the direct dependence on the original data. This solution not only addresses the problem of the difficulty in balancing privacy and performance in traditional federated learning but also significantly enhances the adaptability and practicality of federated learning to noise and privacy protection in industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flowchart of an anti-noise federated smart contract vulnerability detection method with probabilistic label estimation provided by an embodiment of the present invention;

[0060] Figure 2 is a flowchart of the operation of the probabilistic label estimation module provided by an embodiment of the present invention;

[0061] Figure 3 is a training algorithm diagram based on meta-learning provided by an embodiment of the present invention;

[0062] Figure 4 are the accuracies and F1 scores of CBGRU and CGE with symmetric label noise (0%, 5%, 10%, 15%, 20%, 20%, 25% and 30%) on Reentrancy and Timestamp provided by an embodiment of the present invention;

[0063] Figure 5 are the accuracies and F1 scores of CBGRU and CGE with false negative label noise (0%, 5%, 10%, 15%, 20%, 20%, 25% and 30%) on Reentrancy and Timestamp provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0065] As Figure 1 shown, an embodiment of the present invention provides an anti-noise federated smart contract vulnerability detection method with probabilistic label estimation, which specifically includes:

[0066] S1: Create a clean vulnerability dataset containing Reentrancy and Timestamp vulnerabilities;

[0067] S2: Construct a noisy dataset with two types of label noise: symmetric label noise and false negative label noise (asymmetric label noise has no false positive label noise);

[0068] S3: The dataset is divided into a training set, a validation set, and a test set in the ratio of 7:1:2;

[0069] S4: Evenly distribute the training set to multiple data owners. Each data owner randomly selects 60 smart contracts from its local training set as a clean validation set for training the probability label model. The remaining smart contracts in the local training set are used to generate a noisy training set;

[0070] S5: Set the maximum noise level of the two types of noise to 30%. For a given symmetric label noise level (5%, 10%, 15%, 20%, 25%, 30%), each smart contract in the training set is randomly flipped with a probability of (5%, 10%, 15%, 20%, 25%, 30%) to generate a dataset with (5%, 10%, 15%, 20%, 25%, 30%) symmetric label noise. For different false negative label noise levels (5%, 10%, 15%, 20%, 25%, 30%), the contracts labeled as vulnerable are randomly flipped with a probability of (5%, 10%, 15%, 20%, 25%, 30%);

[0071] S6: Establish a deep learning model to evaluate label noise. Use the model CBGRU based on sequence learning and the model CGE that integrates graph learning and expert-defined features. The data center of federated learning initializes the global deep learning model parameters and broadcasts them to all participating data owners in the federated learning;

[0072] S7: Each data owner in the federated learning uses the probability label estimation model to calibrate the local noisy dataset. Use a simple multi-layer perceptron (MLP) as the scoring network. The network includes a hidden layer with a dimension of 64, batch normalization, and a ReLU activation function. Use the Adam optimizer to optimize the probability label estimation network;

[0073] S8: Each data owner in the federated learning updates the local model using the calibrated dataset and the parameters of the received global model;

[0074] S9: Each data owner in the federated learning uploads the updated local model parameters to the data center, and the data center weights and aggregates the parameters of all local models to update the global model parameters;

[0075] S10: Broadcast the updated global model parameters to the data owners again, repeat the training process, and repeat S6 - S9;

[0076] S11: Use accuracy and F1 - score as evaluation metrics to measure the performance of smart contract vulnerability detection.

[0077] In this method, first, a clean smart contract vulnerability dataset is created, and a noise dataset with symmetric label noise and false - negative label noise is constructed for possible label errors in the actual scenario. Specifically, symmetric label noise means randomly assigning a certain proportion of wrong labels to all categories, while false - negative label noise specifically simulates the "missed detection" situation in the vulnerability detection results. By artificially adding these two types of noise to the clean dataset, a noise dataset closer to the real distribution is obtained, providing a basic sample for subsequent noise - resistant federated learning detection.

[0078] After the construction of the noise dataset, the entire dataset is divided into a training set, a validation set, and a test set, and the training set is evenly distributed to multiple data owners according to the principle of roughly equal data volume. The data center initializes the parameters of the global deep - learning model at this stage and broadcasts the initial model parameters to all data owners so that each party can conduct training locally. This process ensures that in a distributed environment, each owner has the same starting model and can perform personalized training iterations on their local data.

[0079] In the local training stage, each data owner calibrates the labels of the local noise dataset based on the probability label estimation model to correct the wrong label information as much as possible. This estimation model will re - evaluate the noise samples in the dataset or assign more credible label probabilities according to the characteristics of the known symmetric label noise and false - negative label noise. After calibration, each data owner uses the calibrated local dataset and the received global model parameters to train the local deep - learning detection model, so as to maintain a high model accuracy and robustness in a noisy environment.

[0080] After the local model training is completed, each data owner uploads the trained model parameters to the data center. The data center performs weighted aggregation on the model parameters from different owners to update the global model parameters and broadcasts the updated global model to each data owner again. This process is repeated according to the preset number of training rounds or performance metric thresholds until the global model meets the preset requirements on the validation set or the test set. Through this cycle, the federated learning system can still obtain a robust and high - precision detection model in a noisy smart contract vulnerability data environment.

[0081] In S4, the training set is evenly distributed to multiple data owners. For multiple data owners, it means that in federated learning, it consists of a data center and multiple data owners, and each data owner has a local dataset.

[0082] In S6, it is broadcast to all data owners in federated learning, aiming to transfer the parameters of the global depth detection model from the data center to all participating data owners to ensure that the local training starting models of all data owners are consistent.

[0083] In S7, the local noisy dataset is calibrated using a probability label estimation module. For the probability label estimation model, it includes: probability label modeling, probability label estimation network, and optimization.

[0084] (1) Probability label modeling. Probability label modeling is a key technical step in the present invention, aiming to infer the probability of the true label from the noisy label by modeling the conditional dependence relationship between labels. Its purpose is to estimate the true label probability of each sample by modeling the relationship between the local label, pseudo-label, and true label. The local label and pseudo-label are regarded as the observed noisy label y nl . According to Bayes' theorem and the Markov random field model, the probability label model is derived. The probability label model is expressed as follows:

[0085]

[0086] where y nl is the observed noisy label, y* is the true label of the smart contract, x is the smart contract feature, and P(y* = 1) represents the prior probability that the true label of the smart contract is 1 (i.e., contains vulnerabilities).

[0087] P(y* = 1) is a constant estimated from the local clean validation dataset of each data owner. The probability label model estimates the probability label of a given smart contract by weighted aggregation of the observed noisy labels according to its accuracy score. This weighted aggregation mechanism enables the model to assign greater importance to more reliable labels, thereby more accurately estimating the true label.

[0088] (2) Probability label estimation network. To optimize the accuracy of the probability label, the accuracy score of the observed noisy label in each smart contract is estimated and continuously updated through an appropriate accuracy score function. As Figure 2 shown, a deep neural network is used as the probability label estimation network to parameterize the probability label model for each data owner k. In the network, the accuracy score function s1(y nl , x) is parameterized by the scoring network . For any labeled smart contract (x, y) ∈ D k , first generate the pseudo-label and one-hot representation Remove the output layer from the global model to obtain the feature embedding h(x) of the smart contract. Further, on the local clean validation dataset estimate the class distribution P(y* = 1) of the local vulnerability dataset D k Then the scoring network takes and h(x) as inputs and predicts the accuracy scoring vector of the observed noisy labels The last layer is a softmax transformation layer for normalizing the accuracy scores. Finally, the aggregator v combines the local class distribution P(y* = 1), the one-hot representation of the observed noisy labels and the accuracy score s x to generate the probability label:

[0089] (3) Optimization. To improve the label calibration accuracy, joint optimization of the probability label estimation network and the local deep detection model. The optimization objective is to minimize the loss of the detection model on the local clean validation dataset.

[0090] It can be expressed as:

[0091]

[0092] where is the local clean validation dataset, is the probability label estimation network, h(x) is the feature of the smart contract, y nl is the observed noisy label, P(y* = 1) is the prior probability that the true label of the smart contract is 1 (i.e., contains vulnerabilities), is the prediction result of the local deep detection model of the k-th data owner for the smart contract x.

[0093] To implement the development of the probability label estimation network, a meta-learning-based training algorithm is used to jointly optimize the probability label estimation network and the local deep detection model with the validation set. As Figure 3 shown, this algorithm works through iterative inner and outer training steps.

[0094] In the inner training step, the local deep detection model is optimized using the calibrated labels generated by the probability label estimation network. The probability label estimation network generates the probability label for each smart contract x initializes the parameters of the local deep detection model using the parameters of the global model, generates the prediction calculates the inner training loss and updates the local model parameters. The local deep detection model is updated as follows:

[0095]

[0096] Among them, η1 is the internal training learning rate, is the internal training loss on the local training set D k above.

[0097] In the external training step, based on the updated local depth detection model on the validation data set update the probability label estimation network is expressed as follows:

[0098]

[0099] The calculation method is as follows:

[0100]

[0101] Among them, η2 is the external training learning rate, is the expected validation loss on the validation data set above.

[0102] Furthermore, in S9, the global model parameters are updated by aggregating the parameters from all data owners in the data center, and the process is expressed as follows:

[0103]

[0104] Among them, θ k is the local model parameter, θ is the global model parameter, n k is the number of samples owned by each data owner, and K represents the total number of data owners.

[0105] An anti-noise federated smart contract vulnerability detection system with probability label estimation provided by an embodiment of the present invention specifically includes:

[0106] A data set construction module for constructing a data set and dividing it into a training set, a validation set, and a test set;

[0107] A probability label estimation module for estimating the true label probability of each sample by modeling the relationship between the local label, the pseudo label, and the true label;

[0108] A meta-learning training module for jointly optimizing the probability label estimation network and the local depth detection model with the validation set based on the training algorithm of meta-learning.

[0109] The present invention is mainly applied to the field of intelligent contract vulnerability detection. In practical applications, it can be applied to the security detection of intelligent contracts on various blockchain platforms. For example, in blockchain applications in the financial field, intelligent contracts involving a large number of fund transactions are of crucial security importance. The present invention can perform noise-resistant federated vulnerability detection on intelligent contracts from different platforms. In blockchain applications in the supply chain field, intelligent contracts are used to manage links such as the tracking and transaction of goods. The present invention can perform noise-resistant federated vulnerability detection on the intelligent contracts of the supply chain, ensuring the safe and stable operation of the supply chain process and preventing problems such as the loss of goods and transaction errors caused by contract vulnerabilities.

[0110] In terms of experimental implementation, the researchers constructed a dataset of intelligent contracts containing different types and levels of label noise and divided it into a training set, a validation set, and a test set. Then, accuracy and F1-score evaluation metrics were used to measure its performance in the intelligent contract vulnerability detection task. Under different noise settings (such as homogeneous noise, heterogeneous noise levels and types), its effectiveness and technical value in the field of intelligent contract vulnerability detection were demonstrated, and it can effectively solve the problems of existing technologies in dealing with label noise and improving detection accuracy, providing strong technical support for the security guarantee of intelligent contracts.

[0111] The present invention was compared with the state-of-the-art noise label learning methods. The baseline was implemented based on FedAvg, and it was also used as an additional baseline experimental result. The comparison under homogeneous noise levels is shown in Tables 1 and 2. The comparison under heterogeneous noise levels is shown in Tables 3 and 4.

[0112] On two vulnerability datasets corrupted by symmetric or false negative label noise, CBGRU and CGE are consistently superior to all baseline methods in terms of F1-score. Specifically, under false negative label noise, compared with the second-best results, the present invention increased the F1-scores of CBGRU and CGE by 8.08% and 2.68% on the Reentrancy dataset and by 12.31% and 2.9% on the Timestamp dataset. Similarly, under symmetric label noise, the present invention increased the F1-score of CBGRU by 4.43% on Reentrancy and by 3.47% on Timestamp. For CGE, the F1-score was increased by 16.13% on Reentrancy and by 2.43% on Timestamp. This proves the effectiveness of the present invention in solving homogeneous label noise in intelligent contract vulnerability detection.

[0113] Table 1 Performance of CBGRU and CGE on Reentrancy under 30% symmetric label noise and 30% false negative label noise

[0114]

[0115] Performance of Timestamp of CBGRU and CGE under 30% symmetric label noise and 30% false negative label noise in Table 2

[0116]

[0117] To evaluate the performance of the present invention under heterogeneous label noise levels, the label noise of four data owners was set at the 20% level, while the label noise of the remaining data owners was set at 30%. The present invention achieved the highest F1 scores of CBGRU and CGE on the Reentrancy dataset, which were 82.34% and 83.69% respectively, and on the Timestamp dataset were 90.75% and 84.79% respectively. Under symmetric label noise and false negative label noise with heterogeneous levels, the present invention was always superior to the second-best results. These results confirm the effectiveness of the present invention in practical scenarios involving heterogeneous label noise levels.

[0118] Performance of Reentrancy of CBGRU and CGE under symmetric label noise and false negative label noise at heterogeneous levels in Table 3

[0119]

[0120]

[0121] Performance of Timestamp of CBGRU and CGE under symmetric label noise and false negative label noise at heterogeneous levels in Table 4

[0122]

[0123] To evaluate the effectiveness in solving label noise of heterogeneous types, the label noise of four data owners was set as 30% symmetric label noise, while the remaining data owners were set as 30% false negative label noise. The present invention obtained the highest F1 scores. For CBGRU it was 84.79% and for CGE it was 87.22% on the Reentrancy dataset, and for CBGRU it was 85.87% and for CGE it was 85.89% on the Timestamp dataset. These results show that the proposed present invention effectively integrates reliable global and local vulnerability knowledge from data owners, can accurately calibrate labels and mitigate the impact of heterogeneous types of label noise. The results are shown in Table 5.

[0124] Comparison of Reentrancy and Timestamp performance under heterogeneous type label noise. "-" indicates no result

[0125]

[0126] The present invention evaluates the performance of the present invention under different noise levels, as shown in Table 6. Regardless of the type of noise, as the label noise level increases, the average F1 scores of CBGRU and CGE both decrease consistently. Specifically, at 30% label noise, the average F1 score drops by 13.76% to 24.098%. This performance degradation is mainly attributed to the tendency of the model to overfit the noisy samples during training. Table 6 shows that the performance of the deep detection model remains relatively stable until the training set contains 20% symmetric label noise or 15% to 25% false negative label noise. Specifically, in the case of symmetric label noise, the average performance metrics of CBGRU do not show a significant decline until the noise level reaches 20%. A similar trend is observed for CGE as the symmetric label noise level increases. The trend graph in Table 6 can be seen Figure 4 and Figure 5 .

[0127] Table 6 Average performance of CBGRU and CGE at different levels of Reentrancy and Timestamp

[0128]

[0129] To evaluate the contributions of different components of the present invention (i.e., the probabilistic label estimation network and the meta-learning based training algorithm), ablation experiments are conducted under symmetric label noise and false negative label noise, as shown in Table 7. The results of this ablation study verify that the combined components of the present invention work together to improve the performance of the deep detection model in the presence of label noise. Using the probabilistic label model alone results in a global deep detection model with limited utility. This deficiency is largely due to the lack of clean samples, which hinders the effective training of the probabilistic label model. Conversely, when the meta-learning based training algorithm is combined with the probabilistic label model, significant improvements in the F1 score are achieved: under symmetric label noise, the F1 score of Reentrancy reaches 63.08% and the F1 score of Timestamp reaches 77.62%. For the global deep detection model under false negative label noise, the F1 scores are 80.34% and 85.24% respectively. These results show that compared with the baseline method FedAvg, in the symmetric label noise environment, the F1 scores of Reentrancy and Timestamp are increased by 5.38% and 4.84% respectively, and in the false negative label noise environment, they are increased by 19.91% and 12.31% respectively. These findings highlight the important role of the meta-learning based training algorithm in improving the calibration accuracy of the probabilistic label model. By leveraging the meta-learning paradigm, the algorithm effectively improves the model's ability to correct noisy labels even when only a limited number of clean samples are available.

[0130] Ablation experiments under 30% symmetric label noise and 30% false negative label noise are shown in Table 7. √ indicates the use of the corresponding part. "_" indicates no result.

[0131]

[0132] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and their modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.

[0133] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. An anti-noise federated smart contract vulnerability detection method with probability label estimation, characterized in that The method includes the following steps: Create a clean smart contract vulnerability dataset and construct a noisy dataset with symmetric label noise and false negative label noise; Divide the dataset into a training set, a validation set, and a test set, and evenly distribute the training set to multiple data owners; The data center initializes the global deep learning model parameters and broadcasts them to all participating data owners in the federated learning; Use the probability label estimation model to calibrate the local noisy dataset; Based on the calibrated local noisy dataset and the received global model parameters, train the local deep detection model; Upload the local model parameters to the data center, and the data center weights and aggregates all local model parameters to update the global model parameters; The data center broadcasts the updated global model parameters to the data owners, and repeats the training process until completion.

2. The anti-noise federated smart contract vulnerability detection method with probability label estimation according to claim 1, wherein In the step of constructing the noisy dataset, the symmetric label noise is generated by randomly flipping the smart contract labels with different probabilities, and the false negative label noise is generated by randomly flipping the labels of the smart contracts marked as vulnerable with different probabilities.

3. The anti-noise federated smart contract vulnerability detection method with probability label estimation according to claim 1, wherein The probability label estimation model includes probability label modeling, a probability label estimation network, and an optimization step, where: Probability label modeling models the relationship between local labels, pseudo-labels, and true labels, and derives the probability label model using Bayes' theorem and Markov random fields; The probability label estimation network is implemented based on a deep neural network, normalizes the accuracy score of the observed noisy labels, and generates probability labels; The probability label estimation network and the local deep detection model are jointly optimized to minimize the loss on the local validation dataset.

4. The anti-noise federated smart contract vulnerability detection method with probability label estimation according to claim 1, characterized in that, The training algorithm based on meta-learning includes an inner training step and an outer training step, where: In the inner training step, use the calibrated labels generated by the probability label estimation network to optimize the local deep detection model; In the outer training step, update the parameters of the probability label estimation network through the validation loss.

5. The anti-noise federated smart contract vulnerability detection method with probability label estimation according to claim 1, wherein The data center aggregates the local model parameters according to the weights, and the weights are determined by the number of samples of each data owner. The weight aggregation process is based on the following formula: Among them, W j is the local model parameter, n j is the number of samples of data owner j, and N is the total number of data owners.

6. The anti-noise federated smart contract vulnerability detection method with probability label estimation as described in claim 1, wherein During the federated learning process, the steps of global model parameter update and broadcast are carried out in fixed rounds until the performance metrics of the global model on the test set meet the preset requirements or the number of training rounds reaches the upper limit; among them, the global training rounds set for the CBGRU model is 50, and the global training rounds set for the CGE model is 40.

7. An anti-noise federated smart contract vulnerability detection system with probability label estimation, characterized in that, The system includes: A data center for initializing, aggregating, and broadcasting the global model; Multiple data owners, each data owner including a local data storage module, a model training module, a probability label estimation module, and a communication module; A communication network for connecting the data center and multiple data owners, supporting encrypted data transmission and model parameter exchange.

8. The system according to claim 7, wherein The data center includes: A parameter initialization unit for generating the initial parameters of the global deep learning model and distributing them to multiple data owners; An aggregation unit for receiving the local model parameters uploaded by multiple data owners, aggregating the model parameters based on the weights, and updating the global model parameters; A parameter broadcast unit for distributing the updated global model parameters to multiple data owners.

9. The system according to claim 7, wherein The data owner includes: A data storage module for storing a local smart contract data set, including a training set, a validation set, and a test set; A model training module for training a model according to the local data set and global model parameters, and generating updated local model parameters; A probability label estimation module for calibrating the local noisy data set, generating probability labels, and optimizing the local depth detection model; A communication module for communicating model parameters with the data center.

10. The system according to claim 7, wherein The communication network supports data transmission based on a security protocol to ensure the integrity and security of data during the upload of local model parameters and the broadcast of global model parameters.

Citation Information

Patent Citations

  • Training method of machine learning model, prediction method and device of machine learning model, and electronic equipment

    CN114239863A

  • Health data federal learning method and system based on block chain

    CN118471479A

  • Fine-grained vulnerability detection method, system and device for smart contract and storage medium

    CN119598475A

  • Smart contract code vulnerability detection method and apparatus, computer device and storage medium

    WO2021037196A1

  • Method and apparatus for service allocation based on reinforcement learning

    WO2021208720A1