Self-supervision anomaly detection method based on mask self-coding

By calculating the feature importance score in the mask autoencoder and dynamically adjusting the mask probability, the problem of ignoring the feature importance in the prior art is solved, and the accuracy and robustness of abnormal detection are improved.

CN120146850APending Publication Date: 2025-06-13YUNNAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510162055.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing mask autoencoder ignores the differences in importance of different features to model learning in exception detection, which leads to the model being unable to fully learn the correlation between key features, affecting the accuracy of abnormal detection.

Method used

By calculating the importance score of the feature, dynamically adjust the mask probability to improve the accuracy of abnormal detection. Specific steps include data preprocessing, mask autoencoder design, dynamic masking strategy, self-supervised training, exception detection, evaluation and optimization, and deployment and application.

Benefits of technology

By dynamically adjusting the mask probability, the model pays more attention to key features, improving the accuracy of feature extraction and model learning efficiency, and improving the accuracy and robustness of abnormal detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146850A_ABST
    Figure CN120146850A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervision anomaly detection method based on mask self-encoding, and relates to the technical field of anomaly detection, and the method comprises the steps of S1, data preprocessing, S2, mask self-encoder design, S3, mask strategy, S4, self-supervision training, S5, anomaly detection, S6, evaluation and optimization, and S7, deployment and application. By introducing a dynamic feature importance mask strategy, the model can dynamically adjust the mask probability according to the importance score of the feature, so as to pay more attention to the incidence relation between key features, specifically, the importance score of each feature is calculated through a mutual information method or a gradient saliency method, and then the mask probability is dynamically generated; according to the method, the model can preferentially learn features which contribute more to anomaly detection in the training process, excessive attention to unimportant features is avoided, the model can capture internal rules of data more efficiently, and the accuracy of feature extraction and the learning efficiency of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of anomaly detection, and particularly relates to a self-supervised anomaly detection method based on masked autoencoders. Background Art

[0002] In recent years, self-supervised learning techniques have received extensive attention in the field of anomaly detection. Self-supervised learning enables the model to learn useful feature representations from unlabeled data by designing pre-training tasks, thereby reducing the dependence on labeled data. As a typical self-supervised learning method, the masked autoencoder can effectively capture the internal structure and feature correlations of the data by randomly masking some features of the input data and training the model to recover the masked parts.

[0003] The masking strategy usually adopts random masking, which ignores the importance differences of different features for model learning. This random masking method may cause the model to fail to fully learn the correlation relationships between key features, thereby affecting the accuracy of anomaly detection. To address the above problems, the following solutions are proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide a self-supervised anomaly detection method based on masked autoencoders, which improves the accuracy of anomaly detection by calculating the importance scores of features and dynamically adjusting the masking probability, and solves the problem that the existing random masking method ignores the importance differences of different features for model learning.

[0005] To solve the above technical problems, the present invention is implemented through the following technical solutions:

[0006] The present invention is a self-supervised anomaly detection method based on masked autoencoders, including:

[0007] Step S1, data preprocessing: cleaning, denoising blockchain transaction data, performing standardization and normalization processing, converting the format, and calculating the feature importance scores;

[0008] Step S2, masked autoencoder design: constructing an encoder and a decoder, where the encoder uses a deep neural network structure to extract latent features, and the decoder restores the complete features from the masked data based on the attention mechanism;

[0009] Step S3, masking strategy: calculating the masking probability, masking features according to the probability, controlling the total masking ratio, and dynamically adjusting the temperature parameter;

[0010] Step S4, self-supervised training: training the model with the goal of minimizing the reconstruction error, and adjusting the learning rate and the masking strategy;

[0011] Step S5, anomaly detection: calculating the reconstruction error of the test data, setting a threshold, and determining whether the transaction is abnormal;

[0012] Step S6, Evaluation and Optimization: Compare the dynamic and random masking metrics, evaluate the model using cross-validation, and adjust the hyperparameters;

[0013] Step S7, Deployment and Application: Integrate the optimized model into the blockchain platform and monitor abnormal transactions in real time;

[0014] The specific steps of the said Step S3, Masking Strategy are as follows:

[0015] Step S31: Calculate the masking probability of each feature;

[0016] Step S32: Randomly mask the input features according to the calculated probability distribution, while ensuring that the total masking ratio is controlled within the set range;

[0017] The calculation formula for the masking probability in the said Step S31 is:

[0018]

[0019] where p i is the masking probability of the i-th feature, s i is the importance score of the i-th feature, τ is the temperature parameter used to control the steepness of the probability distribution, initially set to 1 and dynamically adjusted according to the training loss, n is the total number of features, and s j is the importance score of the j-th feature.

[0020] Furthermore, the specific steps of the said Step S1, Data Preprocessing are as follows:

[0021] Step S11: Clean and denoise the input blockchain transaction data, and remove invalid and error data records;

[0022] Step S12: Perform standardization and normalization processing;

[0023] Step S13: According to the temporal and graph structure characteristics of the blockchain transaction data, convert the data into a format suitable for input to the deep learning model;

[0024] Step S14: Train the reverse propagation gradient of the model, statistically calculate the average absolute value of the gradients of the feature dimensions, and calculate the importance scores of each feature.

[0025] Furthermore, the calculation formula for the importance score in the said Step S14 is:

[0026]

[0027] where s i is the importance score of the i-th feature, K is the number of training samples, k is the sample index, and H is the loss function. It is the partial derivative of the loss function H with respect to the i-th feature of the k-th sample.

[0028] Furthermore, the specific steps of the self-supervised training in step S4 are as follows:

[0029] Step S41: Take minimizing the reconstruction error as the objective function;

[0030] Step S42: Use the unlabeled blockchain transaction data for training. When training, adopt the Adam optimizer to dynamically adjust the learning rate and the masking strategy.

[0031] Furthermore, the formula for the reconstruction error is:

[0032]

[0033] In the formula, ι is the loss function, representing the reconstruction error of the model, N is the total number of samples, b is the sample index, used to traverse all samples, x b is the original data of the b-th sample, is the reconstructed data of the b-th sample.

[0034] Furthermore, the specific steps of the anomaly detection in step S5 are as follows:

[0035] Step S51: For the test data, use the trained model to calculate the reconstruction error ι b ;

[0036] Step S52: Set a dynamic threshold. When ι b is greater than the threshold, mark the transaction as abnormal.

[0037] Furthermore, the formula for setting the threshold in step S52 is:

[0038] Threshold = μ + α·σ;

[0039] In the formula, Threshold is the final anomaly detection threshold, μ is the mean of the reconstruction errors of all samples on the validation set, α is an adjustable parameter used to control the strictness of the threshold, and σ is the standard deviation of the reconstruction errors of all samples on the validation set.

[0040] Furthermore, the specific steps of the evaluation and optimization in step S6 are as follows:

[0041] Step S61: Compare the dynamic masking and random masking metrics to verify the effectiveness of the dynamic masking strategy;

[0042] Step S62: Adopt the cross-validation method and use multiple training and validation sets to evaluate the stability and generalization ability of the model;

[0043] Step S63: Adjust the hyperparameters according to the evaluation results.

[0044] The present invention has the following beneficial effects:

[0045] 1. By introducing a dynamic feature importance masking strategy, the model of the present invention can dynamically adjust the masking probability according to the importance scores of features, thereby paying more attention to the correlation relationships between key features. Specifically, the importance scores of each feature are calculated by the mutual information method or the gradient saliency method, and then the masking probability is dynamically generated, enabling the model to preferentially learn the features that contribute more to anomaly detection during the training process, avoiding excessive attention to unimportant features, and thus enabling the model to more efficiently capture the internal laws of the data, improving the accuracy of feature extraction and the learning efficiency of the model.

[0046] 2. By introducing a temperature parameter, the present invention dynamically controls the steepness of the masking probability distribution to ensure that the model can balance exploration and exploitation at different training stages. Specifically, at the initial stage of training, the temperature parameter is set relatively high, and the masking probability distribution is relatively smooth, enabling the model to widely explore the relationships between different features; as training progresses, it is gradually decreased, making the masking probability distribution more concentrated, and the model gradually focuses on key features. This adaptive adjustment mechanism effectively avoids the model falling into local optima during the training process, enhances the stability and generalization ability of the model, and enables it to better adapt to complex blockchain transaction data.

[0047] 3. The present invention calculates the reconstruction error by the mean square error and combines a dynamic threshold setting strategy to improve the robustness of anomaly detection. Specifically, the model optimizes the parameters by minimizing the reconstruction error during the training process to ensure that the reconstructed data is as close as possible to the original data. During the detection stage, anomalies are judged by a dynamic threshold, and thus can adapt to the distribution characteristics of the data, effectively balancing the false alarm rate and the miss rate, and improving the accuracy and reliability of anomaly detection.

[0048] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0050] Figure 1 It is a schematic flow chart of a self-supervised anomaly detection method based on masked auto-encoding of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0052] Please refer to Figure 1 As shown, the present invention is a self-supervised anomaly detection method based on masked autoencoders, including:

[0053] Step S1, data preprocessing: cleaning, denoising blockchain transaction data, performing standardization and normalization processing, converting the format, and calculating feature importance scores;

[0054] Step S2, masked autoencoder design: constructing an encoder and a decoder, where the encoder uses a deep neural network structure to extract latent features, and the decoder restores complete features from the masked data based on the attention mechanism;

[0055] Step S3, masking strategy: calculating the masking probability, masking features according to the probability, controlling the total masking ratio, and dynamically adjusting the temperature parameter;

[0056] Step S4, self-supervised training: training the model with the goal of minimizing the reconstruction error, and adjusting the learning rate and masking strategy;

[0057] Step S5, anomaly detection: calculating the reconstruction error of the test data, setting a threshold, and determining whether the transaction is abnormal;

[0058] Step S6, evaluation and optimization: comparing dynamic and random masking metrics, evaluating the model using cross-validation, and adjusting hyperparameters;

[0059] Step S7, deployment and application: integrating the optimized model into the blockchain platform to monitor abnormal transactions in real time;

[0060] The specific steps of Step S3, the masking strategy, are as follows:

[0061] Step S31: Calculate the masking probability of each feature;

[0062] Step S32: Randomly mask the input features according to the calculated probability distribution, while ensuring that the total masking ratio is controlled within the set range;

[0063] The calculation formula for the masking probability in Step S31 is:

[0064]

[0065] In the formula, p i is the masking probability of the i-th feature, s iis the importance score of the i-th feature, τ is the temperature parameter used to control the steepness of the probability distribution, initially set to 1 and dynamically adjusted according to the training loss, n is the total number of features, s j is the importance score of the j-th feature.

[0066] Step S1, the specific steps of data preprocessing are as follows:

[0067] Step S11: Clean and denoise the input blockchain transaction data to remove invalid and error data records;

[0068] Step S12: Perform standardization and normalization processing;

[0069] Step S13: According to the temporal and graph structure characteristics of the blockchain transaction data, convert the data into a format suitable for input to the deep learning model;

[0070] Step S14: Train the model to backpropagate the gradient, calculate the average absolute value of the gradient of the feature dimension, and calculate the importance score of each feature.

[0071] The formula for calculating the importance score in Step S14 is:

[0072]

[0073] In the formula, s i is the importance score of the i-th feature, K is the number of training samples, k is the sample index, H is the loss function, is the partial derivative of the loss function H with respect to the i-th feature of the k-th sample.

[0074] Step S4, the specific steps of self-supervised training are as follows:

[0075] Step S41: Use minimizing the reconstruction error as the objective function;

[0076] Step S42: Use unlabeled blockchain transaction data for training. During training, use the Adam optimizer to dynamically adjust the learning rate and masking strategy.

[0077] The formula for the reconstruction error is:

[0078]

[0079] In the formula, ι is the loss function representing the reconstruction error of the model, N is the total number of samples, b is the sample index used to iterate through all samples, x b is the original data of the b-th sample, is the reconstructed data of the b-th sample.

[0080] Step S5, the specific steps of anomaly detection are as follows:

[0081] Step S51: Calculate the reconstruction error ι for the test data using the trained model b ;

[0082] Step S52: Set a dynamic threshold. When ι b is greater than the threshold, mark the transaction as abnormal.

[0083] The formula for setting the threshold in Step S52 is:

[0084] Threshold = μ + α·σ;

[0085] In the formula, Threshold is the final anomaly detection threshold, μ is the mean of the reconstruction errors of all samples on the validation set, α is an adjustable parameter used to control the strictness of the threshold, and σ is the standard deviation of the reconstruction errors of all samples on the validation set.

[0086] Steps S6, the specific steps of evaluation and optimization are as follows:

[0087] Step S61: Compare the dynamic mask and random mask metrics to verify the effectiveness of the dynamic mask strategy;

[0088] Step S62: Adopt the cross-validation method and use multiple training and validation sets to evaluate the stability and generalization ability of the model;

[0089] Step S63: Adjust the hyperparameters according to the evaluation results.

[0090] A specific application of this embodiment is:

[0091] Step S1, Data preprocessing:

[0092] Step S11: Clean and denoise the input blockchain transaction data, remove invalid or incorrect data records, and improve the data quality;

[0093] Step S12: Perform standardization and normalization processing to ensure that each feature dimension has the same scale and make different features equally important in model training;

[0094] Step S13: According to the temporal and graph structure characteristics of the blockchain transaction data, convert it into a format suitable for input to a deep learning model, such as a temporal data matrix or graph structure data;

[0095] Step S14: Calculate the importance score s of each feature i , and through the pre-trained model, backpropagate the gradient and statistically calculate the mean of the absolute values of the gradients of the feature dimensions, that is:

[0096]

[0097] In the formula, s iis the importance score of the i-th feature, K is the number of training samples, k is the sample index, and H is the loss function. is the partial derivative of the loss function H with respect to the i-th feature of the k-th sample;

[0098] Step S2, Mask Autoencoder Design:

[0099] Encoder: Adopt deep neural network structures such as multi-layer Transformer or Graph Convolutional Network (GCN) to encode the input data, learn and extract the latent feature representation of the data;

[0100] Decoder: Constructed based on the attention mechanism. After inputting the data in the masked part, it recovers the complete transaction record from the partially missing data and reconstructs the masked information by learning the associations of other part features;

[0101] Step S3, Masking Strategy:

[0102] Step S31: According to the formula calculate the masking probability of each feature, where p i is the masking probability of the i-th feature, s i is the importance score of the i-th feature, τ is the temperature parameter used to control the steepness of the probability distribution, initially set to 1 and dynamically adjusted according to the training loss, n is the total number of features, and s j is the importance score of the j-th feature;

[0103] Step S32: Randomly mask the input features according to the calculated probability distribution {p i}, while ensuring that the total masking ratio is controlled within a set range (such as 20%); the temperature parameter τ is initially set to 1 and dynamically adjusted during training. For example, when the loss decreases slowly, τ is reduced to strengthen the importance difference;

[0104] Step S4, Self-Supervised Training:

[0105] Step S41: Take minimizing the reconstruction error as the objective function, and the formula is:

[0106]

[0107] In the formula, ι is the loss function representing the reconstruction error of the model, N is the total number of samples, b is the sample index used to traverse all samples, and x b is the original data of the b-th sample, is the reconstructed data of the b-th sample;

[0108] Step S42: Use a large amount of unlabeled blockchain transaction data for training. Adopt the Adam optimizer, dynamically adjust the learning rate and masking strategy during the training process, and continuously update the model parameters to enable the model to accurately recover the masked feature information;

[0109] Step S5, Anomaly Detection:

[0110] Step S51: For the test data, use the trained model to calculate the reconstruction error ι b ;

[0111] Step S52: Set a dynamic threshold, and the formula is:

[0112] Threshold = μ + α·σ;

[0113] In the formula, Threshold is the final anomaly detection threshold, μ is the mean of the reconstruction errors of all samples on the validation set, α is an adjustable parameter used to control the strictness of the threshold, and σ is the standard deviation of the reconstruction errors of all samples on the validation set;

[0114] If ι b > Threshold, then mark this transaction as an anomaly;

[0115] Step S6, Evaluation and Optimization:

[0116] Step S61: Compare metrics such as the F1 score and AUC value of the dynamic mask and the random mask to verify the effectiveness of the dynamic masking strategy;

[0117] Step S62: Adopt the cross-validation method and use multiple training and validation sets to evaluate the stability and generalization ability of the model;

[0118] Step S63: According to the evaluation results, adjust hyperparameters such as τ, masking ratio, network depth, and learning rate to improve the model performance;

[0119] Step S7, Deployment and Application: Integrate the optimized model into the blockchain platform, monitor the transaction flow in real time, mark the transactions with high reconstruction errors, and trigger the alarm mechanism. This method can also be extended and applied to the identification of abnormal behaviors in other fields such as finance and social networks.

[0120] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0121] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A self-supervised anomaly detection method based on masked self-encoding, characterized in that: The detection method comprises the following steps: Step S1, data preprocessing: cleaning, denoising blockchain transaction data, standardizing and normalizing, converting formats, and calculating feature importance scores; Step S2, masked autoencoder design: construct an encoder and a decoder, the encoder uses a deep neural network structure to extract latent features, and the decoder recovers complete features from masked data based on an attention mechanism; Step S3, mask strategy: calculate the mask probability, mask the features according to the probability, control the total mask ratio, and dynamically adjust the temperature parameters; Step S4, self-supervised training: train the model with the goal of minimizing the reconstruction error, and adjust the learning rate and mask strategy; Step S5, anomaly detection: calculate the test data reconstruction error, set the threshold, and determine whether the transaction is abnormal; Step S6, evaluation and optimization: compare dynamic and random mask indicators, evaluate the model using cross-validation, and adjust hyperparameters; Step S7, deployment and application: integrating the optimization model into the blockchain platform to monitor abnormal transactions in real time; The specific steps of step S3, mask strategy are as follows: Step S31: Calculate the mask probability of each feature; Step S32: randomly masking the input features according to the calculated probability distribution, while ensuring that the total mask ratio is controlled within a set range; The calculation formula of the mask probability in step S31 is: In the formula, p i is the mask probability of the i-th feature, s i is the importance score of the i-th feature, τ is the temperature parameter used to control the steepness of the probability distribution, initially set to 1 and dynamically adjusted according to the training loss, n is the total number of features, and s j is the importance score of the jth feature.

2. The self-supervised anomaly detection method based on masked self-encoding according to claim 1, characterized in that: The specific steps of step S1, data preprocessing are as follows: Step S11: Clean and denoise the input blockchain transaction data to remove invalid and erroneous data records; Step S12: performing standardization and normalization processing; Step S13: According to the temporal and graph structure characteristics of blockchain transaction data, the data is converted into a format for deep learning model input; Step S14: Train the model to back-propagate the gradient, calculate the absolute mean of the gradient in the feature dimension, and calculate the importance score of each feature.

3. The self-supervised anomaly detection method based on masked self-encoding according to claim 2, characterized in that: The importance score calculation formula in step S14 is: In the formula, s i is the importance score of the i-th feature, K is the number of training samples, k is the sample index, H is the loss function, is the partial derivative of the loss function H with respect to the i-th feature of the k-th sample.

4. The self-supervised anomaly detection method based on masked self-encoding according to claim 1, characterized in that: The specific steps of step S4, self-supervised training, are as follows: Step S41: minimizing the reconstruction error is taken as the objective function; Step S42: Use unlabeled blockchain transaction data for training, use the Adam optimizer during training, and dynamically adjust the learning rate and mask strategy.

5. The self-supervised anomaly detection method based on masked self-encoding according to claim 4, characterized in that: The formula for the reconstruction error is: In the formula, ι is the loss function, which represents the reconstruction error of the model, N is the total number of samples, b is the sample index, which is used to traverse all samples, and x b is the original data of the bth sample, is the reconstructed data of the bth sample.

6. The self-supervised anomaly detection method based on masked self-encoding according to claim 1, characterized in that: The specific steps of step S5, abnormality detection are as follows: Step S51: For the test data, the reconstruction error ι is calculated using the trained model b ; Step S52: Setting a dynamic threshold value. b When it is greater than the threshold, the transaction is marked as abnormal.

7. The self-supervised anomaly detection method based on masked self-encoding according to claim 6, characterized in that: The formula for setting the threshold in step S52 is: Threshold=μ+α·σ; Where Threshold is the final anomaly detection threshold, μ is the mean of the reconstruction errors of all samples in the validation set, α is an adjustable parameter used to control the strictness of the threshold, and σ is the standard deviation of the reconstruction errors of all samples in the validation set.

8. The self-supervised anomaly detection method based on masked self-encoding according to claim 1, characterized in that: The specific steps of step S6, evaluation and optimization are as follows: Step S61: Compare the dynamic mask and random mask indicators to verify the effectiveness of the dynamic mask strategy; Step S62: using a cross-validation method to evaluate the stability and generalization ability of the model using multiple training and validation sets; Step S63: Adjust the hyperparameters according to the evaluation results.

Citation Information

Cited By

  • Abnormality detection method and device for time series data, electronic equipment and storage medium

    CN120804898A

  • Methods, devices, electronic equipment, and storage media for anomaly detection of time-series data

    CN120804898B

  • Soybean high-temperature-resistant grading method based on vegetation index prior and self-supervised learning

    CN121982552A