Data security auditing method and device based on block chain, equipment and medium
Through blockchain-based timing analysis and multi-dimensional feature analysis of autoencoder, the problem of poor accuracy of data security audit in the prior art is solved, and efficient and accurate data security audit is achieved.
Patent Information
- Application Number
- CN202510652927.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-15
AI Technical Summary
Existing data security audit methods rely on pre-set static rules, making it difficult to deal with complex and changeable data, resulting in poor accuracy.
The blockchain-based data security audit method is adopted, and the timing analysis model, autoencoder and multi-dimensional feature analysis is carried out, combined with the immutable characteristics of the blockchain, to realize the evidence storage and security audit of the original audit data.
It improves the accuracy and efficiency of data security audits, can adapt to dynamic changes in data, and reduces manual operation and maintenance costs.
Smart Images

Figure CN120493315A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of blockchain technology, and in particular relates to a blockchain-based data security audit method, device, equipment and medium. Background Art
[0002] With the growing demand for data security audits, intelligent auditing technology is demonstrating its significant value in a variety of fields, including enterprise operations, information technology, and e-commerce. Whether through real-time monitoring of data access behavior or detecting anomalous activity in complex network environments, accurate anomaly detection is key to effective data security audits. Intelligent auditing technology can help enterprises promptly identify potential security risks, thereby ensuring the security of their data assets and business continuity.
[0003] Currently, data security audits typically rely on pre-set static rules and thresholds to analyze collected system logs, database logs, application logs, and other data to identify abnormal behavior and trigger alerts or take action. This approach of pre-setting static rules is difficult to handle with complex and changing data, resulting in poor data security audit accuracy. Summary of the Invention
[0004] The embodiments of the present application provide a blockchain-based data security audit method, apparatus, device, and medium, which can improve the accuracy of data security audits through time series analysis models, autoencoders, and multi-dimensional feature analysis.
[0005] In a first aspect, an embodiment of the present application provides a data security audit method based on blockchain, which is applied to an audit server. The method includes:
[0006] Obtaining original audit data collected from each data center and uploading the original audit data to the blockchain;
[0007] Sending an acquisition request to the blockchain, and receiving the target audit data stored after consensus verification returned by the blockchain;
[0008] Determining a time series sequence, a feature vector, and original attributes of the target audit data, wherein the time series sequence represents changes in user operation behaviors in the target audit data over a time dimension, the feature vector is a vector representation of the user operation behaviors in the target audit data, and the original attributes include at least one of an access subject and object, an access time, and an access location of the user operation behaviors in the target audit data;
[0009] According to the time series, analyzing the difference between the time series pattern of the target audit data and the time series pattern of the historical audit data by using a time series analysis model to obtain a time series anomaly degree;
[0010] Reconstructing the feature vector using an autoencoder according to the feature vector, and determining a reconstruction error between the reconstructed feature vector and the feature vector;
[0011] Extracting target behavior features from the original attributes, the target behavior features including at least one of target access frequency, target access location, target access time, and target access data type;
[0012] determining a characteristic abnormality degree based on a difference between the target behavior characteristic and the historical behavior characteristic of the historical audit data;
[0013] A security audit result is determined according to the time series abnormality, the reconstruction error, and the feature abnormality.
[0014] In one practicable implementation, obtaining original audit data collected from each data center and uploading the original audit data to a blockchain node specifically includes:
[0015] Assign the collection tasks of each data center to each collection node;
[0016] Receive original audit data collected by each collection node;
[0017] The raw audit data is preprocessed and uploaded to the blockchain, wherein the preprocessing includes at least one of data filtering and data compression.
[0018] In one feasible implementation, allocating the collection tasks of each data center to each collection node specifically includes:
[0019] Sort the collection nodes according to their load ratio;
[0020] According to the data importance level of each data center, for each data center, select the target node of the data center from the sorted collection nodes;
[0021] Determining a target sampling frequency for the data center based on the load ratio of the target node and the initial sampling frequency corresponding to the data importance level of the data center;
[0022] According to the target sampling frequency, the collection task of the data center is determined, and the collection task is allocated to the target node.
[0023] In one possible implementation, the method further includes:
[0024] Determine the vector clock of each acquisition node;
[0025] updating the vector clock when an event occurs, the event including at least the collection of the original audit data, the transmission of the original audit data, and the modification of the original audit data;
[0026] The causal relationship between events is determined based on the vector clock when each event occurs.
[0027] In one possible implementation, the method further includes:
[0028] For each piece of audit information in the original audit data, determining a consistency weight of the audit information based on at least one of a data importance level corresponding to the audit information, a data age of the audit information, and a privacy level corresponding to the audit information;
[0029] Identify conflicting audit information based on the timestamp, location, and content of each audit information;
[0030] Merge conflicting audit information based on the consistency weight.
[0031] In one feasible implementation, determining the characteristic abnormality based on the difference between the target behavior characteristic and the historical behavior characteristic of the historical audit data specifically includes:
[0032] Determine historical access frequency, historical access location, historical access time, and historical access data type through feature extraction from original attributes of the historical audit data;
[0033] determining an access frequency deviation based on a difference between the target access frequency and the historical access frequency;
[0034] determining an access location deviation based on a difference between the target access location and the historical access location;
[0035] determining an access time deviation based on a difference between the target access time and the historical access time;
[0036] determining an access type deviation according to a difference between the target access type and the historical access type;
[0037] The characteristic abnormality degree is determined by weighted summation based on the access frequency deviation, the access position deviation, the access time deviation, and the access type deviation.
[0038] In one feasible implementation, determining a security audit result based on the time series abnormality, the reconstruction error, and the feature abnormality specifically includes:
[0039] Determining a comprehensive anomaly score according to the time series anomaly, the reconstruction error, and the feature anomaly;
[0040] Determine the target score range for the comprehensive anomaly score based on the score ranges corresponding to the preset risk levels;
[0041] The risk level to which the target score range belongs is used as the target risk level;
[0042] Determine security audit results that include the target risk level.
[0043] In one possible implementation, the method further includes:
[0044] Obtaining accuracy feedback data on the security audit results;
[0045] Based on the accuracy feedback data, updating the model parameters of the time series analysis model and the model parameters of the autoencoder through an online learning mechanism;
[0046] Adjusting the hyperparameters of the updated timing analysis model according to the performance indicators of the updated timing analysis model;
[0047] According to the performance indicators of the updated autoencoder, the hyperparameters of the updated autoencoder are adjusted.
[0048] In a second aspect, an embodiment of the present application provides a data security audit method based on blockchain, which is applied to blockchain and includes:
[0049] Receive original audit data uploaded by the audit server;
[0050] Based on the original audit data, multiple audit data blocks are obtained through hash calculation and block packaging;
[0051] For each audit data block, consensus verification is performed on the audit data block through a consensus verification mechanism, and after the consensus verification passes, the audit data block is stored;
[0052] In response to the acquisition request sent by the audit server, the target audit data is determined according to the preset smart contract, and the target audit data is returned to the audit server. The audit server is used to determine the security audit result of the target audit data through a time series analysis model, an autoencoder and multi-dimensional feature comparison.
[0053] In one feasible implementation, multiple audit data blocks are obtained based on the original audit data through hash calculation and block packaging, specifically including:
[0054] For each piece of audit information in the original audit data, determine the blockchain shard to which the audit information belongs based on the time dimension identifier and the business dimension identifier of the audit information;
[0055] Assign the audit information to the blockchain shard to which it belongs;
[0056] For each blockchain shard, a leaf node is generated through hash calculation based on the audit information assigned to the blockchain shard.
[0057] Constructing a weighted hash tree based on each leaf node, and determining a hash value of a root node of the weighted hash tree;
[0058] According to the root node hash value, each audit information in the blockchain shard is packaged into blocks to obtain an audit data block;
[0059] The consensus verification mechanism is used to perform consensus verification on the audit data block, and after the consensus verification is passed, the audit data block is stored, specifically including:
[0060] The audit data block is verified through intra-shard consensus and inter-shard consensus, and after verification, the audit data block is stored on the blockchain.
[0061] In one achievable implementation, for each audit data block, the audit data block is verified through intra-shard consensus and inter-shard consensus, and after passing the verification, the audit data block is stored on the blockchain, specifically including:
[0062] For each audit data block, if the audit data block passes the Byzantine fault-tolerant consensus within the shard, the audit data block is broadcast to each node on the blockchain;
[0063] According to the reference relationship of each audit data block, the audit data blocks are globally sorted and the consistency is verified through a directed acyclic graph structure;
[0064] Based on the node weight of each node on the blockchain, vote on each path in the directed acyclic graph structure and determine the backbone path. The node weight represents the processing capacity and reliability of the node;
[0065] The audit data blocks on the backbone path are written into the blockchain.
[0066] In a third aspect, an embodiment of the present application provides a data security audit device based on blockchain, the device comprising:
[0067] The collection module is used to obtain the original audit data collected from each data center and upload the original audit data to the blockchain;
[0068] A request module, configured to send an acquisition request to the blockchain and receive target audit data stored after consensus verification returned by the blockchain;
[0069] a determination module, configured to determine a time series sequence, a feature vector, and original attributes of the target audit data, wherein the time series sequence represents changes in the time dimension of user operation behaviors in the target audit data, the feature vector is a vector representation of the user operation behaviors in the target audit data, and the original attributes include at least one of an access subject and object, an access time, and an access location of the user operation behaviors in the target audit data;
[0070] A time series anomaly module is used to analyze the difference between the time series pattern of the target audit data and the time series pattern of the historical audit data through a time series analysis model according to the time series sequence to obtain a time series anomaly degree;
[0071] a reconstruction error module, configured to reconstruct the feature vector using an autoencoder according to the feature vector, and determine a reconstruction error between the reconstructed feature vector and the feature vector;
[0072] A behavior feature module, configured to extract target behavior features from the original attributes, wherein the target behavior features include at least one of target access frequency, target access location, target access time, and target access data type;
[0073] a feature anomaly module, configured to determine a feature anomaly degree based on a difference between the target behavior feature and the historical behavior feature of the historical audit data;
[0074] An audit result module is used to determine a security audit result based on the time series abnormality, the reconstruction error and the feature abnormality.
[0075] In a fourth aspect, an embodiment of the present application provides a data security audit device based on blockchain, the device comprising:
[0076] Receiving module, used to receive the original audit data uploaded by the audit server;
[0077] A block module is used to obtain multiple audit data blocks through hash calculation and block packaging based on the original audit data;
[0078] A verification module is used to perform consensus verification on each audit data block through a Byzantine fault-tolerant consensus mechanism, and store the audit data block after the consensus verification passes;
[0079] The response module is used to respond to the acquisition request sent by the audit server, determine the target audit data according to the preset smart contract, and return the target audit data to the audit server. The audit server is used to determine the security audit result of the target audit data through a time series analysis model, an autoencoder and multi-dimensional feature comparison.
[0080] In a fifth aspect, an embodiment of the present application provides a blockchain-based data security audit device, the device comprising: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement any one of the above-mentioned blockchain-based data security audit methods.
[0081] In a sixth aspect, an embodiment of the present application provides a computer storage medium having computer program instructions stored thereon, which, when executed by a processor, implement any one of the above-mentioned blockchain-based data security audit methods.
[0082] In the seventh aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by the processor of an electronic device, the electronic device executes any one of the above-mentioned blockchain-based data security audit methods.
[0083] The data security audit method, device, equipment and computer storage medium based on blockchain of the embodiment of the present application can upload the original audit data of each data center obtained to the blockchain, and use the tamper-proof characteristics of the blockchain to realize the evidence of the original audit data, providing a reliable chain of evidence for subsequent security audits. Send an acquisition request to the blockchain, and receive the target audit data returned by the blockchain after consensus verification. Based on the time series sequence of the target audit data, the difference between the time series pattern of the target audit data and the time series pattern of the historical audit data is analyzed by the time series analysis model to determine the time series anomaly. Based on the feature vector of the target audit data, the feature vector is reconstructed by the autoencoder to determine the reconstruction error. Based on the original attributes of the target audit data, the target behavior characteristics are extracted to determine the difference between the target behavior characteristics and the historical behavior characteristics to obtain the feature anomaly. The security audit result is determined based on the time series anomaly, reconstruction error and feature anomaly. The autoencoder and time series analysis model can automatically learn new audit data through machine learning technology, adapt to the dynamic changes of data, and analyze the target audit data through multiple dimensions to improve the audit accuracy.
[0084] In addition, distributed collection is adopted, and parallel processing is performed through multiple collection nodes, which is highly efficient. Based on the load ratio of each collection node and the data importance level of each data center, tasks are assigned and the sampling frequency is dynamically adjusted. That is, by balancing resource consumption and sampling frequency, the overall audit efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0086] Figure 1 This is a flowchart of a data security audit method based on blockchain provided by an embodiment of the present application;
[0087] Figure 2 This is a schematic diagram of the principle of determining the security audit results provided by the embodiment of the present application;
[0088] Figure 3 This is a schematic diagram of a layered architecture for data security auditing provided by an embodiment of the present application;
[0089] Figure 4 This is a schematic diagram of the structure of a blockchain-based data security audit device provided in an embodiment of the present application;
[0090] Figure 5 This is a schematic diagram of the structure of a blockchain-based data security audit device provided in an embodiment of the present application;
[0091] Figure 6 This is a structural diagram of a blockchain-based data security audit device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0092] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0093] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0094] With the growing demand for data security audits, intelligent auditing technology is demonstrating its significant value in a variety of fields, including enterprise operations, information technology, and e-commerce. Whether through real-time monitoring of data access behavior or detecting anomalous activity in complex network environments, accurate anomaly detection is key to effective data security audits. Intelligent auditing technology can help enterprises promptly identify potential security risks, thereby ensuring the security of their data assets and business continuity.
[0095] Currently, data security audits typically rely on pre-set static rules and thresholds to analyze collected system logs, database logs, application logs, and other data to identify abnormal behavior and trigger alerts or take action. This approach of pre-setting static rules is difficult to handle with complex and changing data, resulting in poor data security audit accuracy.
[0096] In order to solve the problems of the existing technology, the embodiments of the present application provide a data security audit method, device, equipment and computer storage medium based on blockchain.
[0097] The following first describes a blockchain-based data security audit method provided in an embodiment of the present application.
[0098] Figure 1 The following is a flow chart of a data security audit method based on blockchain provided by an embodiment of the present application. Figure 1 As shown, the method includes the following steps:
[0099] S100: The audit server obtains the original audit data collected from each data center and uploads the original audit data to the blockchain.
[0100] In one or more embodiments of the present application, in order to store the audit data through the blockchain in subsequent steps, in this step, the audit server needs to obtain the original audit data collected from each data center and upload the original audit data to the blockchain.
[0101] Specifically, the audit server obtains the original audit data collected from each data center and uploads the original audit data to the blockchain.
[0102] It should be noted that the number and distinction of data centers are not limited in this application and can be set according to actual needs, such as databases, application systems, file systems or other data sources. Taking the operator's user access data security audit as an example, data centers can be distinguished based on the type of data and business stored, and divided into data centers corresponding to multiple data types such as call records, location information, billing information, business logs, etc. The specific method of obtaining original audit data is not limited in this application and can be set according to actual needs, such as directly extracting original audit data from the data center, or using log files to collect original audit data. Of course, in order to improve the efficiency and stability of collecting original audit data, in one or more embodiments of this application, the audit server can assign the collection task of original audit data to multiple collection nodes, so that the original audit data in each data center can be collected in parallel by each collection node, wherein the collection node and the data center can be a one-to-one or one-to-many relationship, which is not limited in this application. At the same time, in order to reduce resource consumption and improve scalability, a lightweight collection agent can be deployed to realize the collection of original audit data of each data center through the lightweight collection agent.
[0103] In addition, before uploading the original audit data to the blockchain, the audit server may pre-process the original audit data to improve the quality of the original audit data and the efficiency of transmission to the audit server. In one or more embodiments of the present application, the audit server may receive the original audit data collected by each collection node, and pre-process the original audit data, and then upload the pre-processed original audit data to the blockchain. The specific method of pre-processing is not limited in this application. For example, the pre-processing may be data filtering (removing redundant data or duplicate data in the original audit data), data compression (reducing the amount of transmitted data), format unification, etc., which can be set according to actual needs.
[0104] S101: The blockchain receives the original audit data uploaded by the audit server.
[0105] S102: The blockchain obtains multiple audit data blocks through hash calculation and block packaging based on the original audit data.
[0106] In one or more embodiments of the present application, in order to verify the original audit data through the consensus verification mechanism in subsequent steps to achieve the preservation of the original audit data, in this step, the blockchain needs to obtain multiple audit data blocks through hash calculation and block packaging.
[0107] Specifically, the blockchain can obtain multiple audit data blocks through hash calculation and block packaging based on the original audit data received in step S101.
[0108] It should be noted that the specific method of determining the audit data block is not limited in this application and can be set according to actual needs, such as hash chain.
[0109] S103: The blockchain performs consensus verification on each audit data block through a consensus verification mechanism, and stores the audit data block after the consensus verification passes.
[0110] In one or more embodiments of the present application, the blockchain may perform consensus verification on each audit data block determined in step S102, and store the audit data block after the consensus verification is passed, thereby realizing the storage of the original audit data.
[0111] Specifically, the blockchain performs consensus verification on each audit data block through a consensus verification mechanism, and stores the audit data block after the consensus verification is passed.
[0112] It should be noted that in order to achieve efficient and secure audit data storage and consensus verification, a large amount of original audit data can be dispersed into different blockchain shards, that is, each shard only contains part of the original audit data, thereby improving data processing efficiency. Of course, in order to facilitate subsequent data query and analysis, the original audit data can be sharded according to the time dimension and business dimension to improve the orderliness of storage. In one or more embodiments of the present application, the block packaging and consensus verification of the original audit data can be achieved by constructing a weighted hash tree (Merkle) and combining a two-layer consensus mechanism of intra-shard consensus and inter-shard consensus, as follows:
[0113] First, for each piece of audit information in each piece of raw audit data, the blockchain determines the blockchain shard to which the audit information belongs based on its time dimension identifier and business dimension identifier, and then assigns the audit information to the blockchain shard to which it belongs. The business dimension identifier represents the business logic characteristics of the raw audit information.
[0114] It should be noted that the specific type of business dimension identifier is not limited in this application and can be determined according to business needs and the characteristics of the original audit data. For example, the business dimension identifier can be location information, user group identifier, business type, etc. Continuing with the example of user access data security audit in step S100, the original audit data is segmented according to the time dimension identifier and the location information identifier. For details, refer to the following formula:
[0115] TimeSlice={startTime, endTime, interval} (1)
[0116] N time =(endTime-startTime) / interval (2)
[0117] GeoSlice={province, city, district} (3)
[0118] ShardID=Hash(TimeSlice||GeoSlice)%TotalShards (4)
[0119] Among them, in formula (1), TimeSlice is the time dimension identifier; startTime is the start timestamp of the original audit data; endTime is the end timestamp of the original audit data; interval is the slice interval time in the time dimension. In formula (2), N time The number of time slices divided into which the original audit data is divided. In formula (3), GeoSlice is the location information identifier; province is the province of the audit information; city is the city of the audit information; and district is the district of the audit information. In formula (4), ShardID is the shard identifier; TotalShards is the total number of shards of the original data. Formula (4) represents the concatenation of the time dimension identifier TimeSlice and the location information identifier GeoSlice, then calculating the hash value using the hash function Hash(*), and finally taking the modulus of the total number of shards TotalShards to obtain the shard identifier ShardID.
[0120] Secondly, for each blockchain shard, the blockchain generates leaf nodes through hash calculation based on the audit information allocated to the blockchain shard, and constructs a weighted hash tree layer by layer based on each leaf node, and then determines the hash value of the root node of the weighted hash tree.
[0121] H(x)=SHA256(x) (5)
[0122] Leaf i =H(Data_i‖Timestampi ‖Features_i) (6)
[0123] Parent=H(w1Child1+w2Child2) (7)
[0124] MerkleRoot=H(∑(wi*NodeHash_i)) (8)
[0125] Wherein, in formula (5), H(x) is the hash value of data x; SHA256(*) is a hash function that outputs a 256-bit hash value. In formula (6), Leaf_i represents the hash value of the i-th leaf node; Data_i represents the i-th audit information, which is the data content to be recorded in the blockchain, such as transaction records, log information, user behavior data, etc.; Timestamp i is the timestamp, which indicates the time when the audit information is generated or recorded; Features_i is the data feature value, such as data type, data center, and importance level. In formula (7), Parent is the hash value of the parent node; child1 and Child2 are the hash values of the child nodes; w1 and w2 are weight coefficients. Formula (7) is the calculation formula for the parent node. In formula (8), MerkleRoot is the hash value of the root node of the Merkle tree, NodeHash_i is the hash value of the i-th node; wi is the weight coefficient of the i-th node, which is determined by the importance of the node data and is used to adjust the proportion of each node in the calculation of the root node hash value. The method for determining the weight coefficient is not limited in this application and can be set according to actual needs.
[0126] Finally, the blockchain packages the audit information within the blockchain shard into blocks based on the root node hash value, obtaining audit data blocks. Of course, the number and specific method of block packaging are not limited in this application and can be set according to actual needs. For example, the hash values of all leaf nodes and related metadata (timestamps, shard identifiers, etc.) can be organized into a block structure, and the root node hash value can be used as the summary of the block to ensure the integrity and consistency of the data within the block.
[0127] In addition, after shard storage, the validity of the audit data block can be ensured through intra-shard consensus and inter-shard consensus. After the consensus verification is passed, the audit data block is stored on the blockchain to complete the storage of the original audit data.
[0128] S104: The audit server sends an acquisition request to the blockchain, and receives the target audit data stored after consensus verification returned by the blockchain.
[0129] S105: The blockchain responds to the acquisition request sent by the audit server, determines the target audit data according to the preset smart contract, and returns the target audit data to the audit server.
[0130] It should be noted that a request includes at least a request identifier, a target data description (including characteristics of the audit data being queried, such as time range, data type, and business dimensions), and a smart contract address. A smart contract is an automated program deployed on the blockchain that parses the request, retrieves the target audit data based on the target data description, and returns it to the audit server.
[0131] S106: The audit server determines the time series sequence, feature vector and original attributes of the target audit data, wherein the time series sequence represents the changes in the user operation behavior in the target audit data in the time dimension, the feature vector is the vector representation of the user operation behavior in the target audit data, and the original attributes include at least one of the access subject and object, access time and access location of the user operation behavior in the target audit data.
[0132] In one or more embodiments of the present application, in order to determine the security audit results in subsequent steps through the time series analysis model, autoencoder, and multi-dimensional behavioral feature comparison, in this step, the audit server needs to determine the time series sequence, feature vector, and original attributes of the target audit data.
[0133] Specifically, the audit server determines the time series, feature vectors, and original attributes of the target audit data. The time series represents the temporal changes in the user operation behavior in the target audit data, the feature vector is a vector representation of the user operation behavior in the target audit data, and the original attributes include at least one of the access subject and object, access time, and access location of the user operation behavior in the target audit data.
[0134] It should be noted that the specific method of determining the time series, feature vectors and original attributes is not limited in this application and can be set according to actual needs. The audit server can extract the timestamp from each audit information of the target audit data, and determine the time series based on the timestamp and each audit information. For example, based on the timestamp of the audit information (user access record) and the audit information (such as user operation behavior), the time series can be determined as {(2025-04-3012:00:00, access page A), (2025-04-30 12:05:00, access page B), (2025-04-30 12:10:00, access page C)}. The feature vector can be a numerical vector extracted from the audit information to describe the data characteristics. The original attribute refers to the information directly recorded in the audit information, such as the access subject and object, access time, access location, etc., that is, "who", "what to access", "when", and "where" in the audit information.
[0135] S107: The audit server analyzes the difference between the time series pattern of the target audit data and the time series pattern of the historical audit data through a time series analysis model according to the time series sequence to obtain a time series anomaly degree.
[0136] It should be noted that historical audit data refers to audit data without anomalies after a data security audit has been conducted. The time series analysis (LSTM) model is based on time series sequences, capturing the time series patterns in the target audit data and comparing them with the time series patterns of the historical audit data to obtain the time series anomaly degree, which reflects the deviation in the time series patterns between the target audit data and the historical audit data.
[0137] S108: The audit server reconstructs the feature vector through an autoencoder based on the feature vector, and determines a reconstruction error between the reconstructed feature vector and the feature vector.
[0138] It should be noted that the autoencoder performs an encoding-decoding reconstruction process on the feature vector of the target audit data, and determines the reconstruction error of the feature vector before and after reconstruction. Because the autoencoder learns the characteristic representation of normal user operation behavior based on the feature vectors of historical audit data, the more abnormal user operation behaviors in the target audit data, the greater the reconstruction error. Conversely, the fewer abnormal user operation behaviors, the smaller the reconstruction error.
[0139] S109: The audit server extracts target behavior features from the original attributes, where the target behavior features include at least one of target access frequency, target access location, target access time, and target access data type.
[0140] S110: The audit server determines a characteristic abnormality degree according to a difference between the target behavior characteristic and the historical behavior characteristic of the historical audit data.
[0141] Specifically, the audit server determines the characteristic abnormality by comparing the difference between the target behavior characteristics and the historical behavior characteristics of the historical audit data.
[0142] It is important to note that target behavior features are divided into operational behavior features, contextual features, and relationship features. Operational behavior features characterize user behavior patterns and operating habits, such as access frequency and access mode; contextual features characterize the background information of user behavior, such as access location, access time, and access device information; and relationship features characterize the interactive relationships between users, such as user interaction information and user collaborative operations. Therefore, the specific content of the target behavior features is not limited in this application and can be set according to actual needs.
[0143] In one or more embodiments of the present application, the audit server may first determine the historical access frequency, historical access location, historical access time, and historical access data type from the original attributes of the historical audit data through feature extraction.
[0144] Secondly, the audit server determines the access frequency deviation based on the difference between the target access frequency and the historical access frequency; determines the access location deviation based on the difference between the target access location and the historical access location; determines the access time deviation based on the difference between the target access time and the historical access time; and determines the access type deviation based on the difference between the target access data type and the historical access data type.
[0145] Finally, the audit server can determine the feature abnormality by weighted summation based on the access frequency deviation, access location deviation, access time deviation, and access type deviation. For details, please refer to the following formula:
[0146] AnomalyScore = Σ(wi * Fi) (9)
[0147] Wherein, in formula (9), wi is the weight of the access frequency deviation, access location deviation, access time deviation, and access type deviation, and Fi is the access frequency deviation, access location deviation, access time deviation, and access type deviation. The access frequency deviation can be the ratio of the absolute difference in access frequency to the standard deviation of the historical access frequency; the access location deviation can be the abnormal probability of the target access location calculated based on the historical access location distribution; the access time deviation can be the probability density value of the target access time in the historical access time distribution; the second similarity between the target access data type and the historical access data type is determined, and the access type deviation can be the difference between the first similarity of the expected data type (which can be set to 1, for example) and the second similarity.
[0148] The anomaly detection process from steps S107 to S110 can detect abnormal behavior in target audit data by implementing an anomaly detection engine that combines multiple machine learning techniques. Specifically, a deep learning model is used to analyze the target audit data stored on the blockchain and identify abnormal behavior. An LSTM network is used to capture the temporal patterns of the target audit data and compare them with those of normal audit data to determine the degree of temporal anomaly. The autoencoder learns the characteristic representation of normal behavior, resulting in a larger reconstruction error when targeting abnormal behavior. Multi-dimensional feature extraction is used to analyze the differences between the target behavior characteristics and historical behavior characteristics to determine the characteristic anomaly degree.
[0149] S111: The audit server determines a security audit result based on the timing anomaly, the reconstruction error, and the feature anomaly.
[0150] Specifically, the audit server may determine the security audit result based on the timing anomaly, reconstruction error, and feature anomaly.
[0151] It should be noted that the specific method of determining the security audit results is not limited in this application and can be set according to actual needs. For example, when the timing anomaly is less than the preset timing anomaly threshold, the reconstruction error is less than the preset reconstruction error threshold, and the feature anomaly is less than the preset feature anomaly threshold, a normal security audit result is determined, and in other cases, an abnormal security audit result is determined. Or, based on the timing anomaly, the reconstruction error, and the feature anomaly, a comprehensive anomaly score is determined by weighted summation. When the comprehensive anomaly score is not less than the preset warning threshold, an abnormal security audit result is determined, and when the comprehensive anomaly score is less than the preset warning threshold, a normal security audit result is determined, wherein the weight of the weighted summation can be set according to actual needs. At the same time, hierarchical processing can be used to implement corresponding protective measures for different risk levels. In one or more embodiments of the present application, the audit server can determine the target score range to which the comprehensive anomaly score belongs based on the score range corresponding to each preset risk level, and then use the risk level to which the target score range belongs as the target risk level to determine the security audit results including the target risk level. The audit server can also determine the risk control measures corresponding to the target risk level, such as warning reminders, setting access permissions, account locking, etc. The correspondence between risk levels and risk control measures is not restricted in this application and can be set according to actual needs.
[0152] like Figure 2As shown, this is a schematic diagram of the principle of determining the security audit results provided by an embodiment of the present application. Taking the security audit of the operator's user location data in step S100 as an example, the diagram shows the process of determining the security audit results through distributed collection, sharded storage, blockchain evidence storage, time series analysis model, autoencoder and multi-dimensional feature analysis.
[0153] In the above method, the audit server uploads the raw audit data obtained from each data center to the blockchain, leveraging the blockchain's immutable nature to store the raw audit data and provide a reliable chain of evidence for subsequent security audits. A request is sent to the blockchain, and the consensus-verified, stored target audit data is returned. Based on the time series sequence of the target audit data, a time series analysis model analyzes the differences between the target audit data's time series patterns and those of historical audit data to determine the time series anomaly. An autoencoder reconstructs the feature vectors of the target audit data, determining the reconstruction error. Based on the original attributes of the target audit data, target behavioral features are extracted, and the differences between the target behavioral features and historical behavioral features are determined to obtain the feature anomaly. The security audit results are determined based on the time series anomaly, reconstruction error, and feature anomaly. The autoencoder and time series analysis model automatically learn from new audit data through machine learning techniques, adapting to dynamic data changes. This significantly reduces manual operation and maintenance costs and improves audit efficiency. Furthermore, by analyzing the target audit data from multiple dimensions, audit accuracy is enhanced.
[0154] In step S100, multiple collection nodes are set up to achieve parallel collection of raw audit data from each data center. However, the audit data stored in different data centers has different data types, and the impact of different data types on security audits varies. Therefore, in one or more embodiments of the present application, the audit server can dynamically adjust the sampling frequency and adaptively allocate collection tasks based on the data importance level of each data center and the load ratio of each collection node. In other words, by balancing the resource consumption of each collection node, the overall sampling efficiency is improved, as follows:
[0155] First, the audit server can sort the collection nodes according to the load ratio.
[0156] Secondly, the audit server can select the target node of each data center from the sorted collection nodes according to the data importance level of each data center. Then, the audit server can determine the target sampling frequency of the data center based on the load share of the target node and the initial sampling frequency corresponding to the data importance of the data center.
[0157] Finally, the audit server may determine the collection task of the data center according to the target sampling frequency, and assign the collection task to the target node of the data center.
[0158] It should be noted that the present application does not limit the setting of the data importance level of each data center. It can be set according to actual needs and can be set based on the degree of influence / importance of the data type of the data in each data center in the security audit. Using the example in step S100, since the billing information involves the user's account data, in the security audit, the billing information is a key audit object, so the data importance level of the billing information can be set to the highest level; since the location information involves the user's whereabouts and can be used to analyze the user's behavioral characteristics, the data importance level of the location information can be set to the second highest level; since the call log contains communication information between users, the data importance level of the call log can be set to a medium level; since the business log records the user's operation behavior, the data importance level of the business log can be set to the second lowest level.
[0159] In addition, during the distributed collection process, the audit server can monitor fault conditions to ensure timely detection and resolution of faults, thereby ensuring the continuity and stability of data collection. Fault conditions include, but are not limited to, offline collection nodes caused by hardware failures or network connection failures, interrupted data transmission between collection nodes and data centers, and collection task failures caused by collection permissions. Of course, this application does not limit the specific methods of fault detection, such as heartbeat detection, network connectivity detection based on network tools, etc., and can be set according to actual needs.
[0160] In step S103, in order to improve the verification efficiency of the audit data block and enhance the fault tolerance capability, this application provides a two-layer consensus mechanism that combines the intra-shard Byzantine Fault Tolerance (PBFT) consensus with the inter-shard Directed Acyclic Graph (DAG) consensus to achieve consensus verification of the audit data block, as follows:
[0161] First, the blockchain performs a Byzantine fault-tolerant consensus on each audit data block within the blockchain shard to which the audit data block belongs. Once the audit data block passes the consensus within that shard, the audit data block is broadcast to all nodes on the blockchain. Second, the blockchain uses a directed acyclic graph (DAG) structure to globally sort and verify the consistency of each audit data block based on its reference relationships. Reference relationships can represent the order or dependency relationships of the audit data blocks. Finally, based on the node weights of each node on the blockchain, a vote is performed on each path in the DAG structure to determine the backbone path. The audit data blocks on this backbone path are then written to the blockchain. The node weights represent the processing power and reliability of the nodes.
[0162] It should be noted that a DAG is a graph structure in which nodes represent audit data blocks and edges represent references between them. Global sorting determines the order or dependency relationships of audit data blocks within the entire blockchain, ensuring the order of audit data block records. Consistency verification verifies the hash value, digital signature, timestamp, and other information of each audit data block, ensuring the authenticity and integrity of its content and conforming to the rules and protocols specified by the blockchain. Within this directed acyclic graph, there are multiple possible paths, each representing a specific order or dependency relationship among audit data blocks. Based on the node weights of each node on the blockchain, voting is performed on each path to determine a primary path. This primary path represents the final order or dependency relationship among audit data blocks that is recognized by a majority of nodes on the blockchain. This means that the audit data blocks on this primary path are deemed valid and reliable, and are therefore written into the blockchain.
[0163] The above-mentioned intra-shard PBFT consensus and inter-shard DAG consensus process is as follows:
[0164] PrePrepare=Sign_leader(v‖n‖d) PrepareCount≥2f+1 (10)
[0165] Reputation_new=α*Reputation_old+(1-α)*Performance (11)
[0166] Weight_i=(Capacity_i*Reliability_i) / ∑Capacity_j*Reliability_i (12)
[0167] In formula (10), PrePrepare represents the initial proposal message sent by the leader to other nodes; v is the view number, which identifies the current view where the leader is located; n is the sequence number, which is used to sort messages and ensure that messages are processed in the correct order; d is the data digest, which is the digest value obtained after processing the message content (such as hash calculation) and is used to verify the integrity and consistency of the message; Sign_leader(v‖n‖d) means that the leader digitally signs v, n, and d to ensure the reliability of the message and that it has not been tampered with; PrepareCount≥2f+1 means that during the PBFT consensus process, the number of nodes that confirm the message is at least 2f+1, ensuring that even if f nodes fail or act maliciously, consensus can still be reached.
[0168] In formula (11), Reputation_new represents the reputation value of the node after the update, that is, the new reputation value; Reputation_old represents the reputation value of the node after the last update, that is, the old reputation value; Performance refers to the performance of the node in the current time period, such as the node's processing speed, correctness, response time, etc.; α is a smoothing factor, which is a constant between 0 and 1. It is used to balance the weights of the new and old reputation values during the reputation value update process to avoid drastic changes in the reputation value caused by large fluctuations in a single performance.
[0169] In formula (12), Weight_i represents the node weight of node i in the inter-shard DAG consensus, which characterizes the influence of the node in the consensus process. The node weight is positively correlated with the processing power and the reliability. Capacity_i represents the processing power of node i in the inter-shard DAG consensus, which is determined by the hardware performance and software performance such as CPU, memory, and network bandwidth. Reliability_i represents the reliability of node i, that is, the probability that the node runs and executes tasks correctly during the consensus process.
[0170] In step S103, the blockchain shards the original audit data based on the time dimension and the business dimension. To further ensure load balancing, in one or more embodiments of the present application, the blockchain may re-shard each audit information when the load imbalance metric of each shard is greater than a preset trigger threshold, as follows:
[0171] ShardID=Hash(DataKey)mod N (13)
[0172]
[0173] Among them, in formula (13), ShardID is the shard identifier, which is used to determine which shard the data is stored in; DataKey is the data feature key value, which is the unique identifier of the data; N is the total number of shards; Hash is a hash function, which is used to map the data feature key value to a value within a fixed range. In formula (14), LoadImbalance is the load imbalance metric, which is used to determine whether load balancing needs to be triggered; Load i is the load of the i-th shard, which refers to the number of transactions, data volume, or other resource usage currently processed by the shard; Avg(Load i ) refers to the average value of each shard's load; Max(Load i ) refers to the maximum load in each shard; Threshold is the trigger threshold, that is, when the load imbalance metric exceeds the trigger threshold, the load balancing operation is triggered.
[0174] In addition, the optimizer can also be used to update the model parameters of the timing analysis model and the autoencoder according to the target audit data, so as to improve the detection capabilities of the timing analysis model and the autoencoder through an online learning mechanism. In one or more embodiments of the present application, the audit server can obtain the accuracy feedback data of the security audit result in step S111, and then based on the accuracy feedback data, update the model parameters of the timing analysis model and the model parameters of the autoencoder through an online learning mechanism, and then adjust the hyperparameters of the updated timing analysis model according to the performance indicators of the updated autoencoder, and adjust the hyperparameters of the updated autoencoder according to the performance indicators of the updated autoencoder.
[0175] In the above, by continuously collecting accuracy feedback data on security audit results, the model parameters are adjusted using the newly collected accuracy feedback data, and then through model evaluation, based on performance indicators such as detection accuracy obtained by model evaluation, hyperparameters such as learning rate and regularization parameters are adjusted.
[0176] To ensure event order and consistency in a distributed system, the causal relationship between events can be determined by determining and updating the vector clock of the collection node. This ensures that the order and causal relationship of events in a distributed environment can be accurately tracked and understood, thereby improving reliability. In one or more embodiments of the present application, the audit server can determine the vector clock of each collection node and update the vector clock when an event occurs. Then, based on the vector clock at the time of each event, the causal relationship of each event can be determined. The event includes at least the collection of original audit data, the transmission of original audit data, and the modification of original audit data.
[0177] In the above, a vector clock is an array whose elements correspond to the logical clock value of a collection node. Initially, the vector clock of each collection node is zero. When an event occurs at a collection node, the vector clock of the corresponding position is incremented by 1. By comparing the vector clocks of two events, the causal relationship between the two events can be determined. For example, if the vector clock of event A is no greater than the vector clock of event B at every position, then event A occurred before event B.
[0178] Continuing with the example of the operator user access data security audit, to handle cross-data center access dependencies, the vector clock is designed as follows:
[0179] VC.initial[i]=0,foriin[1…N] (15)
[0180] VC.local[i]=VC.local[i]+1 (16)
[0181] VC.new[k]=max(VC.local[k],VC.received[k]),for kin[1…N] (17)
[0182] Where VC.initial[i] is the initial state of the vector clock, and the logical clock value of each data center is 0; N is the number of data centers; VC.local[i] is the local vector clock value of the data center; VC.received[k] is the vector clock value of the event received from other data centers; and VC.new[k] is the updated vector clock value.
[0183] For the causal relationship determination between event e1 and event e2, please refer to the following formula:
[0184] VC.e1[i]≤VC.e2[i],for all i∈[1…N] (18)
[0185]
[0186] In equation (18), VC.e1[i] is the vector clock of event e1, and VC.e2[i] is the vector clock of event e2. If equation (18) holds, that is, the vector clock of event e1 is less than or equal to the vector clock of event e2, equation (19) holds. This means that there exists at least one data center k such that the vector clock of event e1 is strictly less than the vector clock of event e2. This ensures that events e1 and e2 are causally related, rather than concurrently related.
[0187] Due to network delays, node failures, and other reasons, data on different nodes may be inconsistent. By comparing the timestamps, locations, and content of conflicting audit information, the audit information with high reliability is retained to ensure data consistency. To this end, in one or more embodiments of the present application, the audit server can determine the consistency weight of each audit information in the original audit data based on at least one of the data importance level corresponding to the audit information, the data age of the audit information, and the privacy level corresponding to the audit information. Based on the timestamps, locations, and content of each audit information, the conflicting audit information is identified and merged based on the size of the consistency weight.
[0188] Using the example of the operator user access data security audit mentioned above, the consistency weight of the audit information is determined based on at least one of the data importance level corresponding to the audit information, the data age of the audit information, and the privacy level corresponding to the audit information. Specifically, refer to the following formula:
[0189] W.base = {call log 0.4, location information 0.3, billing information 0.2, other information 0.1} (20)
[0190] W.time=e -λt (twenty one)
[0191] W.privacy=Privacy level coefficient (1-3) (22)
[0192] W.final=W.base*W.time*W.privacy (23)
[0193] Among them, in formula (20), W.base is a pre-set basic weight, which can be set according to the level of data importance. In formula (21), W.time is the timeliness weight, which is determined based on the age of the data; λ is the decay coefficient, which is a positive number used to control the degree of influence of data age on the weight, which can be set according to actual needs; t is the age of the data, which can be represented by the difference between the time when the timestamp or audit data was generated and the current moment. In formula (22), W.privacy is the privacy level, which is set to a value between 1 and 3. In formula (23), W.final is the consistency weight, which is the product of the basic weight, the timeliness weight, and the privacy weight.
[0194] Assume that the privacy level of a call record is 2, the data age is 1 hour, and the decay coefficient is 0.1. Then the basic weight of the call record is 0.4, the timeliness weight is 0.90, and the privacy weight is 2. Then the product of the basic weight, timeliness weight, and privacy weight, that is, the consistency weight is 0.72.
[0195] It should be noted that data age refers to the timeliness of audit information, that is, newer data has higher reliability, and the privacy level reflects the sensitivity of audit information. The specific method of identifying conflicting audit information is not limited in this application and can be set according to actual needs. For example, the consistency weight of audit information 1 "User A, timestamp 10:05, location B" is 0.85, and the consistency weight of audit information 2 "User A, timestamp 10:06, location C" is 0.92, that is, the difference in timestamps between audit information 1 and audit information 2 is less than the first threshold. Audit information 1 and audit information 2 are both records of the location information of user A at a certain point in time, but there is a difference in the location information of the two records of audit information 1 and audit information 2. Therefore, the consistency weights of audit information 1 and audit information 2 can be compared, and audit information 2 with a large consistency weight can be selected to retain.
[0196] Audit detection functions and conflict resolution strategies can be used to detect conflicting audit information and merge conflicting audit information based on consistency weights, as follows:
[0197] Conflict detection function:
[0198] ConflictDetect(record1, record2) = {timing conflict: |T1-T2| <MinInterval
[0199] Spatial conflict: Location1≠Location2 and Time1≈Time2
[0200] Value conflict: |Value1-Value2|>Threshold}
[0201] Conflict Resolution Strategies:
[0202]
[0203] To ensure the consistency of audit data in a distributed environment, the above mentioned vector clocks of each collection node are maintained to ensure consistency in the time dimension, and the causal relationship between events is ensured by the sequence of the vector clocks of each event. Based on this, this application can verify the consistency of data from three dimensions: time (the sequence of events), space (the location of events), and causality (the causal relationship between events), as follows:
[0204] LC_new=max(LC_local+1,LC_received+1) (24)
[0205] CausalCheck(e1,e2)=(VC_1[i]≤VC_2[i]for all i (25)
[0206] ConsistencyScore=1-(NumConflicts / TotalChecks)(26)
[0207] In equation (24), LC_new is the updated logical clock value; LC_local is the local logical clock, representing the current timestamp of the local node; and LC_received is the received message clock value, representing the timestamp of messages received from other nodes. Logical clocks are used to record the order of events. When an event occurs on the local node, the logical clock increments by 1. When a message is received from another node, the logical clock is updated to the maximum of the local logical clock and the received message clock, plus 1. This ensures that the order of events remains consistent in a distributed environment.
[0208] In formula (25), CausalCheck(e1, e2) is a causal consistency check function used to determine whether event e1 causally precedes event e2; VC_1 and VC_2 are the vector clocks of the events, which are vectors with each component representing the logical clock value of a node. Vector clocks are used to record the causal relationship between events. If the vector clock of event e1 is less than or equal to the vector clock of event e2 in each component, then e1 is considered to be causally prior to e2. That is, the causal order of events in a distributed environment is maintained by vector clocks.
[0209] In formula (26), ConsistencyScore is the data consistency metric, which indicates the data consistency level of the system; NumConflicts is the number of conflicting events found in the consistency check; TotalChecks is the total number of checks, that is, the total number of consistency checks performed.
[0210] In addition, the present application can adjust the warning threshold in step S111, specifically referring to the following formula:
[0211] Threshold.new=β*Threshold_old+(1-β)*RecentAnomalyAvg(27)
[0212] Among them, β is the smoothing factor, which can be set according to actual needs, such as 0.7 to 0.9; RecentAnomalyAvg is the average of the comprehensive anomaly score of historical audit data; Threshold_old is the unupdated warning threshold; Threshold.new is the updated warning threshold.
[0213] It should be noted that in the security audit of operator user access data, distributed collection can support the parallel processing of large amounts of audit data, ensure the causal consistency of audit data, effectively identify and handle abnormal access behaviors, and adapt to the characteristics of operators' multi-geographical distribution, achieving data processing delays of less than 100ms, consistency accuracy greater than 99.9%, anomaly detection accuracy greater than 95%, and a 40% reduction in storage overhead.
[0214] In addition, this application provides a layered architecture design that can monitor data access behavior in real time and store audit data through the immutability of blockchain technology, such as Figure 3 , which is a schematic diagram of the data security audit layered architecture provided in an embodiment of the present application.
[0215] Figure 4 This is a schematic diagram of the structure of a data security audit device based on blockchain provided by an embodiment of the present application. Figure 4 As shown, the apparatus may include a collection module 401 , a request module 402 , a determination module 403 , a timing anomaly module 404 , a reconstruction error module 405 , a behavior feature module 406 , a feature anomaly module 407 and an audit result module 408 .
[0216] The collection module 401 is used to obtain the original audit data collected from each data center and upload the original audit data to the blockchain;
[0217] The request module 402 is configured to send an acquisition request to the blockchain and receive the target audit data stored after consensus verification returned by the blockchain;
[0218] Determination module 403 is configured to determine a time series sequence, a feature vector, and original attributes of the target audit data, wherein the time series sequence represents changes in the time dimension of user operation behaviors in the target audit data, the feature vector is a vector representation of the user operation behaviors in the target audit data, and the original attributes include at least one of an access subject and object, an access time, and an access location of the user operation behaviors in the target audit data;
[0219] A time series anomaly module 404 is configured to analyze the difference between the time series pattern of the target audit data and the time series pattern of the historical audit data using a time series analysis model according to the time series sequence to obtain a time series anomaly degree;
[0220] A reconstruction error module 405 is configured to reconstruct the feature vector using an autoencoder according to the feature vector, and determine a reconstruction error between the reconstructed feature vector and the feature vector;
[0221] A behavior feature module 406 is configured to extract target behavior features from the original attributes, wherein the target behavior features include at least one of target access frequency, target access location, target access time, and target access data type;
[0222] A feature anomaly module 407 is configured to determine a feature anomaly degree based on a difference between the target behavior feature and the historical behavior feature of the historical audit data;
[0223] The audit result module 408 is used to determine a security audit result based on the time series abnormality, the reconstruction error and the feature abnormality.
[0224] In one feasible implementation, the collection module 401 is specifically configured to assign the collection tasks of each data center to each collection node; receive the original audit data collected by each collection node; pre-process the original audit data, and upload the pre-processed original audit data to the blockchain, wherein the pre-processing includes at least one of data filtering and data compression.
[0225] In one feasible implementation, the acquisition module 401 is specifically used to sort the acquisition nodes according to the load share; select the target node of each data center from the sorted acquisition nodes for each data center in turn according to the data importance level of each data center; determine the target sampling frequency of the data center according to the load share of the target node and the initial sampling frequency corresponding to the data importance level of the data center; determine the acquisition task of the data center according to the target sampling frequency, and assign the acquisition task to the target node.
[0226] In one feasible embodiment, the feature anomaly module 407 is specifically used to determine the historical access frequency, historical access location, historical access time and historical access data type from the original attributes of the historical audit data through feature extraction; determine the access frequency deviation based on the difference between the target access frequency and the historical access frequency; determine the access location deviation based on the difference between the target access location and the historical access location; determine the access time deviation based on the difference between the target access time and the historical access time; determine the access type deviation based on the difference between the target access data type and the historical access data type; determine the feature anomaly degree through weighted summation based on the access frequency deviation, the access location deviation, the access time deviation and the access type deviation.
[0227] In one feasible implementation, the audit result module 408 is specifically used to determine a comprehensive anomaly score based on the time series anomaly, the reconstruction error, and the feature anomaly; determine a target score range to which the comprehensive anomaly score belongs based on the score ranges corresponding to the preset risk levels; use the risk level to which the target score range belongs as the target risk level; and determine a security audit result including the target risk level.
[0228] Figure 5 This is a schematic diagram of the structure of a data security audit device based on blockchain provided by an embodiment of the present application. Figure 5 As shown, the apparatus may include a receiving module 501 , a block module 502 , a verification module 503 and a response module 504 .
[0229] Receiving module 501, used to receive the original audit data uploaded by the audit server;
[0230] The block module 502 is configured to obtain multiple audit data blocks through hash calculation and block packaging based on the original audit data;
[0231] Verification module 503, configured to perform consensus verification on each audit data block through a Byzantine fault-tolerant consensus mechanism, and store the audit data block after the consensus verification passes;
[0232] The response module 504 is used to respond to the acquisition request sent by the audit server, determine the target audit data according to the preset smart contract, and return the target audit data to the audit server. The audit server is used to determine the security audit result of the target audit data through a time series analysis model, an autoencoder and multi-dimensional feature comparison.
[0233] In one feasible implementation, the block module 502 is specifically used to determine, for each audit information in the original audit data, the blockchain shard to which the audit information belongs based on the time dimension identifier and business dimension identifier of the audit information; assign the audit information to the blockchain shard to which it belongs; for each blockchain shard, generate a leaf node through hash calculation based on the audit information assigned to the blockchain shard; based on each leaf node, construct a weighted hash tree and determine the root node hash value of the weighted hash tree; based on the root node hash value, block package the audit information in the blockchain shard to obtain an audit data block; the consensus verification mechanism is used to perform consensus verification on the audit data block, and the audit data block is stored after the consensus verification is passed, which specifically includes: verifying the audit data block through intra-shard consensus and inter-shard consensus, and storing the audit data block on the blockchain after the verification is passed.
[0234] In one feasible implementation, the verification module 503 is specifically configured to, for each audit data block, broadcast the audit data block to each node on the blockchain when the audit data block passes the Byzantine fault-tolerant consensus within the shard; perform global sorting and consistency verification on the audit data blocks through a directed acyclic graph structure based on the reference relationship between the audit data blocks; vote on each path in the directed acyclic graph structure based on the node weight of each node on the blockchain, and determine the backbone path, wherein the node weight represents the processing capability and reliability of the node; and write the audit data block on the backbone path into the blockchain.
[0235] Figure 6 A schematic diagram of the hardware structure of a blockchain-based data security audit device provided in an embodiment of the present application is shown.
[0236] A blockchain-based data security audit device may include a processor 601 and a memory 602 storing computer program instructions.
[0237] Specifically, the processor 601 may include a central processing unit (CPU) or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0238] The memory 602 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In one example, the memory 602 may include a removable or non-removable (or fixed) medium, or the memory 602 may be a non-volatile solid-state memory. The memory 602 may be inside or outside the integrated gateway disaster recovery device.
[0239] In one example, the memory 602 may be a read-only memory (ROM). In one example, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0240] The memory 602 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method for data security audit based on blockchain according to the present application.
[0241] The processor 601 reads and executes the computer program instructions stored in the memory 602 to implement Figure 1 A data security audit method based on blockchain in the illustrated embodiment.
[0242] In one example, a data security audit device based on blockchain may further include a communication interface 603 and a bus 604. Figure 6 As shown, the processor 601 , the memory 602 , and the communication interface 603 are connected via a bus 604 and communicate with each other.
[0243] The communication interface 603 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0244] Bus 604 includes hardware, software or both, and couples the components of online data flow metering equipment to each other. For example, and not limitation, bus may include Accelerated Graphics Port (AGP) or other graphics bus, Enhanced Industry Standard Architecture (EISA) bus, Front Side Bus (FSB), Hyper Transport (HT) interconnection, Industry Standard Architecture (ISA) bus, InfiniBand interconnection, Low Pin Count (LPC) bus, memory bus, Micro Channel Architecture (MCA) bus, Peripheral Component Interconnect (PCI) bus, PCI-Express (PCI-X) bus, Serial Advanced Technology Attachment (SATA) bus, Video Electronics Standards Association Local (VLB) bus or other suitable bus or a combination of two or more of these. Where appropriate, bus 604 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the present application considers any suitable bus or interconnection.
[0245] In addition, in conjunction with the blockchain-based data security audit method in the above embodiment, the present application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the blockchain-based data security audit methods in the above embodiment is implemented.
[0246] An embodiment of the present application also provides a computer program product, including a computer program, which, when processed and executed, implements any one of the blockchain-based data security audit methods in the above embodiments.
[0247] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0248] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The programs or code segments can be stored in a machine-readable medium, or transmitted on a transmission medium or communication link via a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memories (ROMs), flash memories, erasable read-only memories (EROMs), floppy disks, compact disc read-only memories (CD-ROMs), optical discs, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segments can be downloaded via computer networks such as the Internet and intranets.
[0249] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0250] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed via the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. This processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or the flowchart and the combination of the boxes in the block diagram and / or the flowchart can also be implemented by the dedicated hardware that performs the specified function or action, or can be implemented by the combination of dedicated hardware and computer instructions.
[0251] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.
Claims
1. A data security audit method based on blockchain, characterized in that: Applied to an audit server, the method includes: Obtaining original audit data collected from each data center and uploading the original audit data to the blockchain; Sending an acquisition request to the blockchain, and receiving the target audit data stored after consensus verification returned by the blockchain; Determining a time series sequence, a feature vector, and original attributes of the target audit data, wherein the time series sequence represents changes in user operation behaviors in the target audit data over a time dimension, the feature vector is a vector representation of the user operation behaviors in the target audit data, and the original attributes include at least one of an access subject and object, an access time, and an access location of the user operation behaviors in the target audit data; According to the time series, analyzing the difference between the time series pattern of the target audit data and the time series pattern of the historical audit data by using a time series analysis model to obtain a time series anomaly degree; Reconstructing the feature vector using an autoencoder according to the feature vector, and determining a reconstruction error between the reconstructed feature vector and the feature vector; Extracting target behavior features from the original attributes, the target behavior features including at least one of target access frequency, target access location, target access time, and target access data type; determining a characteristic abnormality degree based on a difference between the target behavior characteristic and the historical behavior characteristic of the historical audit data; A security audit result is determined according to the time series abnormality, the reconstruction error, and the feature abnormality.
2. The method according to claim 1, characterized in that Obtaining original audit data collected from each data center and uploading the original audit data to the blockchain node, specifically including: Assign the collection tasks of each data center to each collection node; Receive original audit data collected by each collection node; The raw audit data is preprocessed and uploaded to the blockchain, wherein the preprocessing includes at least one of data filtering and data compression.
3. The method according to claim 2, characterized in that Assign the collection tasks of each data center to each collection node, including: Sort the collection nodes according to their load ratio; According to the data importance level of each data center, for each data center, select the target node of the data center from the sorted collection nodes; Determining a target sampling frequency for the data center based on the load ratio of the target node and the initial sampling frequency corresponding to the data importance level of the data center; According to the target sampling frequency, the collection task of the data center is determined, and the collection task is allocated to the target node.
4. The method according to claim 2, characterized in that The method further comprises: Determine the vector clock of each acquisition node; updating the vector clock when an event occurs, the event including at least the collection of the original audit data, the transmission of the original audit data, and the modification of the original audit data; The causal relationship between events is determined based on the vector clock when each event occurs.
5. The method according to claim 3, characterized in that The method further comprises: For each piece of audit information in the original audit data, determining a consistency weight of the audit information based on at least one of a data importance level corresponding to the audit information, a data age of the audit information, and a privacy level corresponding to the audit information; Identify conflicting audit information based on the timestamp, location, and content of each audit information; Merge conflicting audit information based on the consistency weight.
6. The method according to claim 1, characterized in that Determine the characteristic abnormality based on the difference between the target behavior characteristic and the historical behavior characteristic of the historical audit data, specifically including: Determine historical access frequency, historical access location, historical access time, and historical access data type through feature extraction from original attributes of the historical audit data; determining an access frequency deviation based on a difference between the target access frequency and the historical access frequency; determining an access location deviation based on a difference between the target access location and the historical access location; determining an access time deviation based on a difference between the target access time and the historical access time; determining an access type deviation according to a difference between the target access type and the historical access type; The characteristic abnormality degree is determined by weighted summation based on the access frequency deviation, the access position deviation, the access time deviation, and the access type deviation.
7. The method according to claim 1, characterized in that Determining a security audit result based on the time series abnormality, the reconstruction error, and the feature abnormality specifically includes: Determining a comprehensive anomaly score according to the time series anomaly, the reconstruction error, and the feature anomaly; Determine the target score range for the comprehensive anomaly score based on the score ranges corresponding to the preset risk levels; The risk level to which the target score range belongs is used as the target risk level; Determine security audit results that include the target risk level.
8. The method according to claim 1, characterized in that The method further comprises: Obtaining accuracy feedback data on the security audit results; Based on the accuracy feedback data, updating the model parameters of the time series analysis model and the model parameters of the autoencoder through an online learning mechanism; Adjusting the hyperparameters of the updated timing analysis model according to the performance indicators of the updated timing analysis model; According to the performance indicators of the updated autoencoder, the hyperparameters of the updated autoencoder are adjusted.
9. A data security audit method based on blockchain, characterized in that: Applied to blockchain, the method includes: Receive original audit data uploaded by the audit server; Based on the original audit data, multiple audit data blocks are obtained through hash calculation and block packaging; For each audit data block, consensus verification is performed on the audit data block through a consensus verification mechanism, and after the consensus verification passes, the audit data block is stored; In response to the acquisition request sent by the audit server, the target audit data is determined according to the preset smart contract, and the target audit data is returned to the audit server. The audit server is used to determine the security audit result of the target audit data through a time series analysis model, an autoencoder and multi-dimensional feature comparison.
10. The method according to claim 1, characterized in that Based on the original audit data, multiple audit data blocks are obtained through hash calculation and block packaging, specifically including: For each piece of audit information in the original audit data, determine the blockchain shard to which the audit information belongs based on the time dimension identifier and the business dimension identifier of the audit information; Assign the audit information to the blockchain shard to which it belongs; For each blockchain shard, a leaf node is generated through hash calculation based on the audit information assigned to the blockchain shard. Constructing a weighted hash tree based on each leaf node, and determining a hash value of a root node of the weighted hash tree; According to the root node hash value, each audit information in the blockchain shard is packaged into blocks to obtain an audit data block; The consensus verification mechanism is used to perform consensus verification on the audit data block, and after the consensus verification is passed, the audit data block is stored, specifically including: The audit data block is verified through intra-shard consensus and inter-shard consensus, and after verification, the audit data block is stored on the blockchain.
11. The method according to claim 1, wherein For each audit data block, the audit data block is verified through intra-shard consensus and inter-shard consensus. After verification, the audit data block is stored on the blockchain, specifically including: For each audit data block, if the audit data block passes the Byzantine fault-tolerant consensus within the shard, the audit data block is broadcast to each node on the blockchain; According to the reference relationship of each audit data block, the audit data blocks are globally sorted and the consistency is verified through a directed acyclic graph structure; Based on the node weight of each node on the blockchain, each path in the directed acyclic graph structure is voted and the backbone path is determined. The node weight represents the processing capacity and reliability of the node; The audit data blocks on the backbone path are written into the blockchain.
12. A data security audit device based on blockchain, characterized in that: The device comprises: The collection module is used to obtain the original audit data collected from each data center and upload the original audit data to the blockchain; A request module, configured to send an acquisition request to the blockchain and receive target audit data stored after consensus verification returned by the blockchain; a determination module, configured to determine a time series sequence, a feature vector, and original attributes of the target audit data, wherein the time series sequence represents changes in the time dimension of user operation behaviors in the target audit data, the feature vector is a vector representation of the user operation behaviors in the target audit data, and the original attributes include at least one of an access subject and object, an access time, and an access location of the user operation behaviors in the target audit data; A time series anomaly module is used to analyze the difference between the time series pattern of the target audit data and the time series pattern of the historical audit data through a time series analysis model according to the time series sequence to obtain a time series anomaly degree; a reconstruction error module, configured to reconstruct the feature vector using an autoencoder according to the feature vector, and determine a reconstruction error between the reconstructed feature vector and the feature vector; A behavior feature module, configured to extract target behavior features from the original attributes, wherein the target behavior features include at least one of target access frequency, target access location, target access time, and target access data type; a feature anomaly module, configured to determine a feature anomaly degree based on a difference between the target behavior feature and the historical behavior feature of the historical audit data; An audit result module is used to determine a security audit result based on the time series abnormality, the reconstruction error and the feature abnormality.
13. A data security audit device based on blockchain, characterized in that: The device comprises: Receiving module, used to receive the original audit data uploaded by the audit server; A block module is used to obtain multiple audit data blocks through hash calculation and block packaging based on the original audit data; A verification module is used to perform consensus verification on each audit data block through a Byzantine fault-tolerant consensus mechanism, and store the audit data block after the consensus verification passes; The response module is used to respond to the acquisition request sent by the audit server, determine the target audit data according to the preset smart contract, and return the target audit data to the audit server. The audit server is used to determine the security audit result of the target audit data through a time series analysis model, an autoencoder and multi-dimensional feature comparison.
14. A data security audit device based on blockchain, characterized in that: The device includes: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the blockchain-based data security audit method as described in any one of claims 1 to 9.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the blockchain-based data security audit method according to any one of claims 1 to 9.