Malicious deletion traceability method of distributed file system based on block chain

By combining CNN and LSTM network models for real-time risk assessment, and using detection anchor points and detection stack mechanisms, the problems of inaccurate identification of malicious deletion operations and low traceability efficiency are solved, and efficient malicious deletion traceability is achieved.

CN120386768AActive Publication Date: 2025-07-29CHINA ACADEMY OF INFORMATION & COMM
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510873436.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In a distributed file system based on blockchain, the identification of malicious deletion operations is inaccurate and the traceability efficiency is low, resulting in system performance degradation and waste of resources.

Method used

A neural network model combining CNN network and LSTM network is used for real-time risk assessment, and through detection of anchor points and detection stack mechanisms, malicious deletion operations are accurately identified and traced.

Benefits of technology

It improves the identification accuracy and traceability speed of malicious deletion operations, reduces computing overhead, and improves the data security and response efficiency of distributed file systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386768A_ABST
    Figure CN120386768A_ABST
Patent Text Reader

Abstract

The invention relates to a malicious deletion traceability method for a distributed file system based on a block chain, belongs to the technical field of data security, and solves the problems of inaccurate positioning and low traceability efficiency of the existing malicious deletion operation. Comprising the steps that when deletion operation is written, behavior characteristics of the current deletion operation are extracted, and a risk level is obtained through a deletion behavior recognition model; the deletion behavior recognition model is a neural network model which combines a CNN network and an LSTM network and introduces a self-attention mechanism; creating or updating a detection anchor point according to the risk levels of the current deletion operation and the previous deletion operation, and setting the detection anchor point to which the current deletion operation belongs; writing each detection anchor point into a detection stack with a corresponding priority; and when the malicious deletion operation occurs, according to the malicious deletion file identifier and the priority of the detection stack, verifying the deletion operation contained in each detection anchor point in the detection stack until the malicious deletion operation is found. And rapid source tracing of the malicious deletion operation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security, and particularly to a method for tracing malicious deletion in a blockchain-based distributed file system. Background Art

[0002] In a blockchain-based distributed file system, malicious deletion operations may pose a serious threat to the integrity of files and the security of data. Therefore, how to accurately identify malicious deletion behaviors and effectively trace their starting positions and influence scopes has become a key issue to be solved urgently.

[0003] The existing technologies lack a real-time risk assessment mechanism for operation behaviors, resulting in the difficulty of dynamically assessing the risk levels of deletion operations. Some high-risk malicious operations are difficult to be identified and processed in a timely manner, while low-risk regular operations may waste tracing resources and slow down the tracing process. In addition, some static detection rules cannot adapt to the dynamic changes of the operation environment, leading to the lag and inaccuracy of detection results.

[0004] In addition, when tracing malicious deletion behaviors traditionally, the entire operation records are often scanned and compared item by item, resulting in low tracing efficiency. Especially in a distributed system storing massive data, there are a large number of operation records. Such a processing method not only has a large computational overhead, but may also cause a serious decline in system performance. Summary of the Invention

[0005] In view of the above analysis, embodiments of the present invention aim to provide a method for tracing malicious deletion in a blockchain-based distributed file system to solve the problems of inaccurate positioning and low tracing efficiency of existing malicious deletion operations.

[0006] Embodiments of the present invention provide a method for tracing malicious deletion in a blockchain-based distributed file system, including the following steps: When a deletion operation is written into the file operation log of the blockchain, extract the behavior characteristics of the current deletion operation, and obtain the risk level of the current deletion operation by using a deletion behavior recognition model; the deletion behavior recognition model is a neural network model that combines a CNN network and an LSTM network and introduces a self-attention mechanism; According to the risk levels of the current deletion operation and the previous deletion operation, create or update a detection anchor point, and set the detection anchor point to which the current deletion operation belongs; according to the risk levels of each detection anchor point, write each detection anchor point into a detection stack with a corresponding priority; When a malicious deletion operation occurs, according to the maliciously deleted file identifier and the priority of the detection stack, verify the deletion operations included in each detection anchor point in the detection stack until the malicious deletion operation is found.

[0007] Based on further improvements to the above method, the behavioral characteristics of the current deletion operation include: deletion mode, number of deleted files, size of deleted files, highest category of deleted files, deletion interval time, and node reputation; the node reputation is the reputation of the node initiating the deletion operation in the distributed file system, which is dynamically updated based on the initial reputation value and according to the node's stability score, contribution score, and compliance score after each round of consensus cycle.

[0008] Based on further improvements to the above method, the risk level of the current deletion operation is obtained using a deletion behavior recognition model, including: Receiving the behavioral characteristics of the current deletion operation through a CNN network and outputting the first feature vector of the current deletion operation; Taking multiple first feature vectors of historical deletion operations from the feature queue and combining them with the first feature vector of the current deletion operation to form a first feature vector sequence, which is passed into an LSTM network to output the hidden state corresponding to each first feature vector and passed into the attention layer; The attention layer uses the self-attention mechanism to calculate the weights of each hidden state and perform weighted summation to obtain a second feature vector, which is passed into the fully connected layer; The fully connected layer maps the second feature vector to the probabilities of multiple risk levels, and takes the risk level corresponding to the maximum probability as the risk level of the current deletion operation.

[0009] Based on further improvements to the above method, the risk levels include high risk, medium risk, and low risk; the detection anchor includes one or more deletion operations with a risk level of medium risk or high risk.

[0010] Based on further improvements to the above method, according to the risk levels of the current deletion operation and the previous deletion operation, a detection anchor is created or updated, and the detection anchor to which the current deletion operation belongs is set, including: If the risk level of the current deletion operation is high risk or medium risk, and the risk level of the previous deletion operation is low risk, then a new detection anchor is created, and the current deletion operation is set as the starting operation of the new detection anchor; If the risk level of the current deletion operation is low risk, and the risk level of the previous deletion operation is high risk or medium risk, then the previous deletion operation is set as the ending operation of the last created detection anchor; If the risk level of the current deletion operation is high risk or medium risk, and the risk level of the previous deletion operation is high risk or medium risk, then the current deletion operation is set as one of the deletion operations in the last created detection anchor; If the risk levels of both the current deletion operation and the previous deletion operation are low risk, then it does not belong to any detection anchor.

[0011] Based on the further improvement of the above method, each detection anchor point includes all deletion operations within the time range of its start operation and its end operation. The highest risk level among all the deletion operations it contains is taken as the risk level of the detection anchor point, and the time of the end operation is taken as the time of the detection anchor point.

[0012] Based on the further improvement of the above method, according to the malicious file deletion identifier and the priority of the detection stack, the deletion operations included in each detection anchor point in the detection stack are verified until a malicious deletion operation is found, including: Obtain the deletion operation containing the file deletion identifier according to the file deletion identifier, and then determine the detection anchor point to which it belongs as the first candidate anchor point; Obtain other detection anchor points within the preset time period before the time of the first candidate anchor point from the detection stacks with high priority and low priority respectively as the second candidate anchor point and the third candidate anchor point; Regard the deletion operations with a low risk level within the preset time before the time of the anchor point to be verified as the fourth candidate anchor point; Verify the deletion operations in the first candidate anchor point, the second candidate anchor point, the third candidate anchor point, and the fourth candidate anchor point in sequence. If a malicious deletion operation is found, the verification ends.

[0013] Based on the further improvement of the above method, the malicious deletion operation is verified successively from three dimensions: file content integrity, file permission matching, and associated log operation compliance. If there is an abnormality in any one dimension, it is determined as a malicious deletion operation.

[0014] Based on the further improvement of the above method, file content integrity means that the content hash value of the deleted file is the same as that of its backup file; file permission matching includes: the user who deletes the file has the file deletion permission, and the file permission has not changed before deletion; associated log operation compliance includes: the associated log timestamps are continuous within a short time period after the file is deleted, and there are no batch deletion and creation operations.

[0015] Based on the further improvement of the above method, during the training process of the deletion behavior recognition model, the loss function value is calculated by the following formula to update the model parameters: , where represents the loss function value of each batch in the training sample set, represents the number of samples in a batch, represents the th sample category weight, represents the proportion of the number of samples of category in the training sample set to the total number, represents the proportion of the number of the th category in the training sample set to the total number, represents the total number of categories, represents the probability that the th sample belongs to the true category, represents the weight for adjusting easy and difficult samples, represents the basic weight of easy and difficult samples, represents the adjustment coefficient.

[0016] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects: 1. By combining deep learning, detection anchor points, and detection stack mechanisms, the deficiencies in the recognition and traceability efficiency of malicious deletion operations in the prior art are solved, and the data security and traceability response speed of the distributed file system are improved; 2. Using a neural network model combining CNN and LSTM, fully analyze the multi-dimensional behavioral characteristics of deletion operations, conduct real-time risk assessment on deletion operations, accurately mark the risk level, significantly improve the accuracy of malicious deletion operation recognition, and reduce the occurrence of missed detections and false alarms; 3. Classify deletion operations with different risk levels into detection anchor points with different risk levels, which is convenient for quickly locating the area where malicious operations are located during malicious deletion traceability, and realizing the detection order through the priority of the detection stack, avoiding global scanning of all operation records, reducing unnecessary calculations and verifications while paying attention to the context-related operations of malicious deletions, so as to achieve efficient and fast traceability.

[0017] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can be made obvious from the description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the content specifically pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs represent the same components; Figure 1 is a flowchart of a method for tracing malicious deletions in a blockchain-based distributed file system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings, where the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.

[0020] A specific embodiment of the present invention discloses a method for tracing malicious deletion in a blockchain-based distributed file system, as Figure 1 shown, which includes the following steps: S1. When a deletion operation is written into the file operation log of the blockchain, extract the behavioral characteristics of the current deletion operation, and use the deletion behavior recognition model to obtain the risk level of the current deletion operation; the deletion behavior recognition model is a neural network model that combines a CNN network and an LSTM network and introduces a self-attention mechanism.

[0021] In this embodiment, the file operation log of the blockchain is used to record the operations of each node on the file in the distributed file system. In order to accurately analyze the starting position and influence range of malicious deletion operations, real-time analysis and monitoring of deletion operations are carried out.

[0022] It should be noted that the original information recorded for the deletion operation in the file operation log includes: deletion mode, number of deleted files, size of deleted files, names of each deleted file, categories of each deleted file, content hash values of each deleted file, request node identifier, node reputation, operator user identifier, and deletion time. Among them, the deletion mode includes: single-file deletion and batch deletion; the size of the deleted files is the total size of the files deleted at one time; the categories of the deleted files are set by the distributed file system, such as personal files, temporary files, backup files, and system files, etc., and each category corresponds to a corresponding level; the request node identifier is the node in the distributed file system that initiates the deletion operation; the node reputation is the reputation of the node that initiates the deletion operation in the distributed file system, which is dynamically updated based on the initial reputation value and according to the node's stability score, contribution score, and compliance score after each round of consensus cycle.

[0023] It should be noted that a round of consensus cycle refers to the process from the submission of the transaction corresponding to the file operation to the blockchain network until the transaction is packaged into a block and the block is finally confirmed by the network through the consensus algorithm.

[0024] The stability score of the node is calculated according to the packet loss rate and the network delay variance as follows: .

[0025] The contribution score of the node is obtained by weighting the storage contribution rate and response rate of the node. Among them, the storage contribution rate of the node is the ratio of the node's storage capacity to the maximum storage capacity, and the response rate is the ratio of the number of responses of the node within a preset time period to the total number of responses.

[0026] The compliance score of a node is the score rewarded by calculating the number of consecutive non-malicious deletion behaviors, or the score punished for malicious deletion behaviors.

[0027] When performing dynamic updates, each part of the score is compared with the corresponding threshold, and the initial reputation value is increased or deducted according to the value exceeding or falling below.

[0028] According to the original information of the deletion operation record, the extracted behavioral characteristics include: deletion pattern, number of deleted files, size of deleted files, highest category of deleted files, deletion interval time, and node reputation.

[0029] Furthermore, in this embodiment, a deletion behavior recognition model is used to identify the risk level of the deletion operation, where the risk level includes high risk, medium risk, and low risk.

[0030] It should be noted that the deletion behavior recognition model successively includes: a CNN network, an LSTM network, an attention layer, and a fully connected layer; among them, the CNN network is used to extract the spatial features of the deletion operation, the LSTM network is used to extract the temporal features of the deletion operation, the weight of key features is increased through the attention mechanism layer, and finally a fully connected layer is connected. The features output by the self-attention mechanism are mapped to the classification space through the softmax activation function, and the prediction probabilities of different risk levels are output.

[0031] When training the deletion behavior recognition model, a sliding window (window size is T, step size is S) is used to obtain the behavioral characteristics of the deletion operation from the historical file operation log to generate each sample sequence, that is, the behavioral characteristics of every T deletion operations are taken out by sliding according to the fixed length and step size as a sample, and the true risk level of the last deletion operation in each sample is taken as the label of the entire sample sequence, and N samples are used as a batch.

[0032] Furthermore, the CNN network processes the D-dimensional behavioral characteristics of each deletion operation in each sample in turn, performs convolution on the D-dimensional behavioral characteristics, and converts them into a high-dimensional first feature vector, such as D', that is, N×T D'-dimensional first feature vectors are output. The LSTM network processes the T D'-dimensional feature vectors in each sample in sequence order, and calculates an H-dimensional hidden state respectively. The attention layer processes the T H-dimensional hidden states output by the LSTM network for each sample respectively, calculates the weights of each hidden state and performs weighted summation to obtain an H-dimensional second feature vector representing each sample. The fully connected layer maps the H-dimensional second feature vector of each sample to the probabilities of 3 risk levels, and finally outputs N 3-dimensional vectors for a batch. Each 3-dimensional vector corresponds to the probabilities of 3 risk levels of one sample, and the risk level corresponding to the maximum probability is taken as the predicted risk level.

[0033] Considering that there are relatively few samples corresponding to malicious file deletion operations in the actual scenario, in this embodiment, an improved Focal Loss is used as the loss function to dynamically adjust the weights of various types of samples, reduce the contribution of easily classified samples, and guide the learning focus of the model to difficult-to-classify malicious samples, thereby improving the detection accuracy of malicious file deletion behaviors.

[0034] Specifically, during the training process of the deletion behavior recognition model, the loss function value is calculated through the following formula to update the model parameters: , where, represents the loss function value of each batch in the training sample set, represents the number of samples in a batch, represents the th sample's class 's class weight, represents the proportion of the number of samples of class in the training sample set to the total number, represents the proportion of the number of samples of the th class in the training sample set to the total number, represents the total number of classes, represents the probability that the th sample belongs to the true class, represents the weight for adjusting easy and difficult samples, represents the basic weight for easy and difficult samples, , represents the adjustment coefficient, .

[0035] As can be seen from the above formula, the class weight is calculated based on the quantity ratio of each class. The fewer the number of classes, the greater the weight. The weight for adjusting easy and difficult samples is dynamically calculated based on the average prediction confidence of the current batch of samples. If the average prediction confidence is low, it indicates that the overall difficulty of the current batch of samples is high, and a larger is needed to focus on difficult samples, otherwise reduce to avoid over - focusing.

[0036] After the deletion behavior recognition model is trained, during actual application, in order to improve the processing efficiency, the CNN network only processes the behavior features of the currently newly collected deletion operations. Since the input of the LSTM is a sequence, in this embodiment, a feature queue with a length of T - 1 is maintained to store the first feature vectors output by the CNN network for the behavior features of the last T - 1 deletion operations. When the deletion behavior recognition model predicts the risk level of the current deletion operation, the first feature vector output by the CNN network for the current deletion operation is added to the feature queue, and the oldest first feature vector in the feature queue is removed to keep the length of the feature queue at T - 1.

[0037] Specifically, the deletion behavior recognition model obtains the risk level of the current deletion operation through the following steps: Receive the behavior features of the current deletion operation through the CNN network and output the first feature vector of the current deletion operation; Take out the first feature vectors of multiple historical deletion operations from the feature queue and form a first feature vector sequence with the first feature vector of the current deletion operation, and input it into the LSTM network to output the hidden state corresponding to each first feature vector and input it into the attention layer; The attention layer uses the self-attention mechanism to calculate the weights of each hidden state and perform weighted summation to obtain the second feature vector, and input it into the fully connected layer; The fully connected layer maps the second feature vector to the probabilities of multiple risk levels, and takes the risk level corresponding to the maximum probability as the risk level of the current deletion operation.

[0038] This step makes full use of the characteristics that the file operation log has both spatial behavior features and time series features, uses the CNN network and the LSTM network, and introduces the attention mechanism to focus on the most suspicious operations, significantly improving the ability to identify the risk level of deletion operations.

[0039] S2. If the risk levels of the current deletion operation and the previous deletion operation are different, create or update the detection anchor point and set the detection anchor point to which the current deletion operation belongs; according to the risk levels of each detection anchor point, write each detection anchor point into the detection stack with the corresponding priority.

[0040] It should be noted that creating or updating the detection anchor point and setting the detection anchor point to which the current deletion operation belongs according to the risk levels of the current deletion operation and the previous deletion operation includes: If the risk level of the current deletion operation is high risk or medium risk and the risk level of the previous deletion operation is low risk, create a new detection anchor point and set the current deletion operation as the starting operation of the new detection anchor point; If the risk level of the current deletion operation is low risk and the risk level of the previous deletion operation is high risk or medium risk, then set the previous deletion operation as the end operation of the last detected anchor point that has been created; If the risk level of the current deletion operation is high risk or medium risk and the risk level of the previous deletion operation is high risk or medium risk, then set the current deletion operation as one of the deletion operations in the last detected anchor point that has been created; If the risk levels of both the current deletion operation and the previous deletion operation are low risk, then it does not belong to any detected anchor point.

[0041] Based on the above, the detected anchor points in this embodiment include one or more deletion operations with a risk level of medium risk or high risk. When there are multiple ones, the detected anchor point is equivalent to a detection interval, with a clear start operation and end operation. The deletion operations within the time range of the start operation and the end operation all belong to this detected anchor point. Take the highest risk level of all the deletion operations it contains as the risk level of the detected anchor point, and take the time of the end operation as the time of the detected anchor point. That is to say, a high-risk detected anchor point contains at least one deletion operation with a risk level of high risk, and the deletion operations contained in a medium-risk detected anchor point are all with a risk level of medium risk.

[0042] Organize the relevant information (start operation, end operation, all the deletion operations it contains, risk level, and time) of each detected anchor point in JSON format and store it in the NoSQL database to obtain the identifier of each detected anchor point.

[0043] From the perspective of deleting files, record the relevant information of the deletion operations with a risk level of medium risk and high risk in the deletion index record. The deletion index record includes: deletion file identifier, risk level, last deletion time, and the identifier of the detected anchor point to which it belongs; Store the deletion index record in a relational database or a NoSQL database, which is convenient for subsequent rapid retrieval according to the deletion file identifier without having to traverse the file operation logs of each node in the blockchain, greatly improving the retrieval efficiency.

[0044] Considering that some files perform logical deletions and may be deleted again after being restored later, therefore, each time a record is added to the deletion index record, first judge whether there is already a deletion index record according to the deletion file identifier. If it already exists, update its information. If it does not exist, add a new record.

[0045] Furthermore, the priorities of the detection stacks include high priority and low priority; Store the identifiers of the high-risk detected anchor points in the database into the high-priority detection stack, and store the identifiers of the medium-risk detected anchor points in the database into the low-priority detection stack.

[0046] S3. When a malicious deletion operation occurs, according to the malicious deletion file identifier and the priority of the detection stack, verify the deletion operations included in each detection anchor in the detection stack until the malicious deletion operation is found.

[0047] In this embodiment, it is considered that when a malicious deletion operation occurs, if only based on the malicious deletion file identifier, the source of the malicious deletion operation will be ignored, and the vulnerability existing in the distributed system cannot be accurately located, thus effectively preventing the occurrence of similar events. Therefore, using the malicious deletion file identifier to quickly locate the corresponding detection anchor, verify the high-risk and medium-risk deletion operations within the corresponding range, reduce unnecessary calculations and verifications, and achieve efficient and rapid traceability.

[0048] Specifically, according to the malicious deletion file identifier and the priority of the detection stack, verify the deletion operations included in each detection anchor in the detection stack until the malicious deletion operation is found, including: Obtain the deletion operation containing the deletion file identifier according to the deletion file identifier, and then determine the detection anchor to which it belongs as the first candidate anchor; Obtain other detection anchors within a preset time period before the time of the first candidate anchor from the detection stacks with high priority and low priority respectively as the second candidate anchor and the third candidate anchor; Regard the deletion operation with a low risk level within a preset time before the time of the anchor to be verified as the fourth candidate anchor; Verify the deletion operations in the first candidate anchor, the second candidate anchor, the third candidate anchor, and the fourth candidate anchor in sequence. If the malicious deletion operation is found, end the verification.

[0049] It should be noted that according to the deletion file identifier, quickly obtain the deletion operation containing the deletion file identifier from the deletion index record, and get the first candidate anchor to which it belongs. When verifying the deletion operations of each candidate anchor, the verification is carried out sequentially starting from the initial operation.

[0050] Verify the malicious deletion operation from three dimensions: file content integrity, file permission matching, and associated log operation compliance in sequence. If there is an abnormality in any dimension, it is determined as a malicious deletion operation and submitted to manual review.

[0051] Among them, file content integrity means that the content hash value of the deleted file is the same as that of its backup file, indicating that the content of the deleted file has not been tampered with or lost. If there are multiple versions of the backup file, it is only necessary to be the same as the content hash value of one of the versions of the backup file.

[0052] File permission matching includes: the user who deletes the file has the file deletion permission, and the file permission has not changed before deletion; among them, the user who deletes the file having the file deletion permission includes a specific user group authorized to delete the file by the user, or the user is directly granted the deletion permission for the file; the file permission not changing before deletion is to avoid the permission being suddenly relaxed shortly before the deletion operation occurs.

[0053] The compliance of associated log operations includes: within a short period after deleting the file, the associated log timestamps are continuous, and there are no batch deletion, creation, or modification operations. Among them, the associated logs include key logs such as system logs, audit logs, and operation tracking logs; the timestamps being continuous and there being no batch deletion and creation operations are to avoid missing records caused by deletion, forging records by creation, or covering up the suspicion of malicious behavior by modification.

[0054] Through the verification of the above three dimensions, malicious deletion operations are identified comprehensively and accurately, providing solid and reliable support and basis for subsequent in-depth security incident investigation, liability tracing, and system recovery work.

[0055] Compared with the prior art, a method for tracing malicious deletion in a blockchain-based distributed file system provided in this embodiment solves the deficiencies in the existing malicious operation identification and tracing efficiency through the combination of deep learning, detection anchor points, and detection stack mechanisms, improving the data security and tracing response speed of the distributed file system. By using a neural network model combining CNN and LSTM, the multi-dimensional behavioral characteristics of the deletion operation are fully analyzed, the deletion operation is subjected to real-time risk assessment, the risk level is accurately marked, significantly improving the accuracy of malicious deletion operation identification and reducing the occurrence of missed detections and false alarms. Classifying deletion operations with different risk levels into detection anchor points with different risk levels facilitates quickly locating the area where the malicious operation is located during malicious deletion tracing, and realizing the detection order through the priority of the detection stack, avoiding global scanning of all operation records, reducing unnecessary calculations and verifications while paying attention to the context-related operations of malicious deletion, thereby achieving efficient and fast tracing.

[0056] Those skilled in the art can understand that all or part of the processes of implementing the method of the above embodiment can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.

[0057] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for tracing malicious deletion in a blockchain-based distributed file system, characterized in that, The following steps are involved: When a deletion operation is recorded in the blockchain's file operation log, the behavioral characteristics of the current deletion operation are extracted, and the risk level of the current deletion operation is determined using a deletion behavior recognition model. The deletion behavior recognition model is a neural network model that combines a CNN network and an LSTM network and introduces a self-attention mechanism. Create or update a detection anchor point based on the risk level of the current deletion operation and the previous deletion operation, and set the detection anchor point to which the current deletion operation belongs; According to the risk level of each detection anchor point, each detection anchor point is written into the detection stack of the corresponding priority; When a malicious deletion operation occurs, the deletion operations contained in each detection anchor point in the detection stack are verified according to the malicious deletion file identifier and the priority of the detection stack until the malicious deletion operation is found.

2. The malicious deletion traceability method of the blockchain-based distributed file system according to claim 1, wherein, The behavioral characteristics of the current deletion operation include: deletion mode, number of deleted files, size of deleted files, highest category of deleted files, deletion interval time and node reputation; the node reputation is the reputation of the node that initiates the deletion operation in the distributed file system, which is based on the initial reputation value and is dynamically updated according to the node's stability score, contribution score and compliance score after each consensus cycle is completed.

3. The method for tracing malicious deletions in a distributed file system based on blockchain according to claim 1, characterized in that: The method of using the deletion behavior recognition model to obtain the risk level of the current deletion operation includes: Receive the behavioral features of the current deletion operation through the CNN network and output the first feature vector of the current deletion operation; Take the first feature vectors of multiple historical deletion operations from the feature queue and combine them with the first feature vector of the current deletion operation to form a first feature vector sequence and pass it into the LSTM network. Output the hidden state corresponding to each first feature vector and pass it to the attention layer. The attention layer uses the self-attention mechanism to calculate the weight of each hidden state and the weighted sum is used to obtain the second eigenvector, which is then passed to the fully connected layer. The fully connected layer maps the second eigenvector to the probabilities of multiple risk levels, and takes the risk level corresponding to the maximum probability as the risk level of the current deletion operation.

4. The method for tracing malicious deletions in a distributed file system based on blockchain according to claim 1, characterized in that: The risk levels include high risk, medium risk and low risk; the detection anchor points include one or more deletion operations with a risk level of medium risk or high risk.

5. The malicious deletion tracing method of the blockchain-based distributed file system according to claim 4, wherein The step of creating or updating a detection anchor point based on the risk level of the current deletion operation and the previous deletion operation, and setting the detection anchor point to which the current deletion operation belongs, includes: If the risk level of the current deletion operation is high or medium, and the risk level of the previous deletion operation is low, a new detection anchor point is created, and the current deletion operation is set as the starting operation of the new detection anchor point; If the risk level of the current deletion operation is low risk and the risk level of the previous deletion operation is high risk or medium risk, then the previous deletion operation is set as the end operation of the last detection anchor point created; If the risk level of the current deletion operation is high risk or medium risk, and the risk level of the previous deletion operation is also high risk or medium risk, then the current deletion operation is set to be a deletion operation in the last detection anchor point created; If the risk level of the current deletion operation and the previous deletion operation are both low risk, they do not belong to any detection anchor point.

6. The malicious deletion traceability method of the blockchain-based distributed file system according to claim 5, characterized in that, Each detection anchor point contains all deletion operations within the time range of its start operation and its end operation. The highest risk level of all the deletion operations it contains is taken as the risk level of the detection anchor point, and the time of the end operation is taken as the time of the detection anchor point.

7. The method for tracing malicious deletions in a distributed file system based on blockchain according to claim 5, characterized in that: The method of verifying the deletion operations contained in each detection anchor point in the detection stack according to the malicious deletion file identifier and the priority of the detection stack until the malicious deletion operation is found includes: Obtaining a deletion operation including the deleted file identifier according to the deleted file identifier, and then determining the detection anchor point to which it belongs as the first candidate anchor point; Respectively obtaining other detection anchor points within a preset period before the time of the first candidate anchor point from the high-priority and low-priority detection stacks as the second candidate anchor point and the third candidate anchor point; Deletion operations with a low risk level within a preset time period before the time of the anchor point to be verified are taken as the fourth candidate anchor point; The deletion operations in the first candidate anchor point, the second candidate anchor point, the third candidate anchor point and the fourth candidate anchor point are verified in sequence. If a malicious deletion operation is found, the verification is terminated.

8. The malicious deletion traceability method of the blockchain-based distributed file system according to claim 1 or 7, characterized in that The malicious deletion operation is verified in three dimensions: file content integrity, file permission matching, and compliance of associated log operations. If any dimension has an abnormality, it is determined to be a malicious deletion operation.

9. The method for tracing malicious deletions in a distributed file system based on blockchain according to claim 8, characterized in that: The file content integrity is that the content hash value of the deleted file is consistent with the content hash value of its backup file; The file permission matching includes: the user who deletes the file has the file deletion permission, and the file permission has not been changed before deletion; the associated log operation compliance includes: the associated log timestamps are continuous within a short period after the file is deleted, and there are no batch deletion and creation operations.

10. The malicious deletion traceability method of the blockchain-based distributed file system according to claim 1, wherein, The deletion behavior recognition model calculates the loss function value and updates the model parameters using the following formula during the training process: , in, Represents the loss function value of each batch in the training sample set, represents the number of samples in a batch, Indicates the The category of samples The category weights, Represents the category in the training sample set The proportion of the total number of Indicates the first The proportion of the number of categories to the total number, Indicates the total number of categories, Indicates the The probability that a sample belongs to the true category, represents the adjustment weight of difficult and easy samples, represents the basic weight of difficult and easy samples, Represents the adjustment coefficient.

Citation Information

Patent Citations

  • Intrusion detection method and system based on traceability clustering and graph serialization

    CN114238958A

  • File processing monitoring method and device, equipment and medium

    CN115687244A

  • Block chain malicious node detection method based on graph neural network

    CN116662992A

  • Traceability graph intrusion detection-oriented malicious behavior subgraph hierarchical extraction method and system

    CN116881903A

  • Advanced threat detection method and device based on graph neural network

    CN117332413A