Audit data distributed storage method and related products based on multi-layer encryption strategy

By adopting multi-layer encryption strategies and pseudo-random sharding redundant coding technology in a distributed storage environment, the problems of encryption protection and data integrity of audit data with different sensitivity are solved in a distributed storage environment, and efficient and flexible data protection and stable system operation are achieved.

CN119441229BActive Publication Date: 2025-05-06CHENGDU BIG DATA GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510031012.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

In a distributed storage environment, how to provide flexible and efficient encryption protection for audit data of different sensitivity, ensuring data integrity and data recovery capabilities in the event of storage node failure.

Method used

A multi-layer encryption strategy is used to divide it into multiple levels according to the sensitivity of audit data, and symmetric encryption, asymmetric encryption and fully homomorphic encryption strategies are used for encryption according to different levels. Through pseudo-random sharding and redundant coding technology, the reliability of data storage is improved, and combined with an efficient integrity verification mechanism, the security and integrity of data in a distributed environment are guaranteed.

Benefits of technology

It realizes flexible encryption protection for data with different sensitivity, improves the security of highly sensitive data, and takes into account the encryption efficiency of ordinary data, improves the reliability and fault tolerance of distributed storage systems, reduces the resource overhead caused by storage redundancy, and ensures the stable operation of the system when node failure or load is too high.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441229B_ABST
    Figure CN119441229B_ABST
Patent Text Reader

Abstract

The technical field of audit data storage of the present invention specifically relates to a distributed storage method for audit data based on a multi-layer encryption strategy and related products. The method comprises: dividing the audit data into multiple levels, and adopting different encryption strategies for the audit data according to different levels to obtain encrypted audit data; dynamically sharding the encrypted audit data based on sharding rules to generate multiple data fragments; distributing and storing the data fragments in multiple distributed storage nodes; and verifying the integrity of the stored data through hash calculation, Merkle tree verification and verification information storage mechanism. The present invention combines the multi-layer encryption strategy to realize flexible encryption protection for data of different sensitivities, while taking into account the encryption efficiency of ordinary data. The fault tolerance capability is improved through the combination of sharding and redundant coding technology. The dynamic monitoring and migration mechanism of the distributed storage nodes ensures the stable operation of the system under the condition of node failure or excessive load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audit data storage, and in particular to a distributed storage method for audit data based on a multi-layer encryption strategy and related products. Background Art

[0002] With the rapid development of informatization and big data technology, the security and integrity of audit data, as an important type of business data, are particularly important. However, traditional audit data storage methods face multiple challenges, including data leakage risks, data loss caused by storage node failure, and insufficient data integrity verification. These challenges place higher demands on the protection of sensitive audit data.

[0003] At present, common data storage technologies usually use symmetric encryption or asymmetric encryption to protect data. However, these encryption methods often lack flexibility when facing data of different sensitivities, and it is difficult to achieve the expected security for the protection of highly sensitive data. In addition, the encryption efficiency may decrease due to the use of a single encryption strategy during the data encryption process, especially when it is necessary to support encrypted state operations, the limitations of traditional encryption methods are more prominent.

[0004] On the other hand, distributed storage has become a mainstream technology for large-scale data storage. By distributing and storing data in multiple nodes, distributed storage systems can effectively improve the reliability and fault tolerance of data storage. However, due to the dynamic and instability of storage nodes, distributed storage systems still face the risk of data loss and service interruption. Therefore, the introduction of redundant coding technology in a distributed environment, such as erasure coding (Reed-Solomon coding), has become an important means to improve data availability. Summary of the invention

[0005] The technical problem to be solved by the present invention is how to provide flexible and efficient encryption protection for audit data of different sensitivities in a distributed storage environment, while ensuring the integrity of the data and the ability to recover the data in the event of a storage node failure. The purpose is to provide a distributed storage method and related products for audit data based on a multi-layer encryption strategy, which realizes flexible selection of encryption strategies according to data sensitivity, improves the reliability of data storage through pseudo-random sharding and redundant coding, and combines an efficient integrity verification mechanism to ensure the security and integrity of data in a distributed environment.

[0006] The present invention is achieved through the following technical solutions:

[0007] A distributed storage method for audit data based on a multi-layer encryption strategy, comprising:

[0008] Data encryption processing: audit data is divided into multiple levels according to their sensitivity, and different encryption strategies are used for audit data according to different levels to obtain encrypted audit data;

[0009] Data sharding processing: Generate sharding rules through a pseudo-random algorithm, and dynamically shard the encrypted audit data based on the sharding rules to generate multiple data fragments;

[0010] Distributed storage, storing data fragments in multiple distributed storage nodes;

[0011] Data integrity verification, through hash calculation, Merkle tree verification and verification information storage mechanism, the integrity of the stored data is verified.

[0012] Specifically, the method for performing data encryption processing includes:

[0013] According to the sensitivity of audit data, audit data is divided into multiple levels, including ordinary data, sensitive data and highly sensitive data;

[0014] Symmetric encryption strategy is used to encrypt ordinary data;

[0015] Encrypt sensitive data by combining symmetric encryption and asymmetric encryption, wherein after the sensitive data is encrypted using symmetric encryption to obtain the initial ciphertext, the initial ciphertext is encrypted again using asymmetric encryption;

[0016] Use fully homomorphic encryption strategy to encrypt highly sensitive data;

[0017] Get encrypted audit data.

[0018] Optionally, methods for grading audit data include:

[0019] Determine the real-time sensitivity of audit data , ,in, For the A sensitivity index function, For the The weight of the sensitivity index is is the total number of sensitivity indicators;

[0020] Determining the sensitivity threshold and , , and classify audit data based on sensitivity thresholds;

[0021] like , then the corresponding audit data belongs to ordinary data and is encrypted using a symmetric encryption strategy;

[0022] like , then the corresponding audit data is sensitive data and is encrypted by combining symmetric encryption and asymmetric encryption;

[0023] like , the corresponding audit data is highly sensitive data and is encrypted using a fully homomorphic encryption strategy.

[0024] Methods for encrypting ordinary data include:

[0025] Initial encryption of common data. ,in, is the initial ciphertext, For ordinary data plaintext, is a symmetric key, is a random number, It is symmetric encryption;

[0026] Through a pseudo-random number generator Generate a perturbation matrix and perturb the initial ciphertext to generate ordinary data ciphertext ,in, For The perturbation matrices of equal length, It is a bitwise XOR operation;

[0027] Methods for encrypting sensitive data include:

[0028] Segment sensitive data and Divided into The size is of equal length , , is plain text data;

[0029] For each piece of data Two levels of encryption are performed; the first level of encryption is symmetric encryption , the second level of encryption is asymmetric encryption ,in, for The hash value of the segment, is a symmetric key, is the asymmetric encryption public key, For the The random number of the segment, For symmetric encryption, It is asymmetric encryption;

[0030] Combine to obtain sensitive data ciphertext ;

[0031] Methods for encrypting highly sensitive data include:

[0032] Get dynamic parameters ,in, is the complexity of encryption operation, is a linear mapping or a nonlinear function, For bitwise XOR operation, Classify keys for highly sensitive data; , is the number of encrypted state addition operations, is the number of encrypted state multiplication operations, is the multiplication complexity weight;

[0033] Encrypt highly sensitive data ,in, For highly sensitive data plaintext, It is fully homomorphic encryption;

[0034] Methods for obtaining the key include:

[0035] Obtaining the master key through quantum key distribution , and generate hierarchical keys based on the master key ,in, is the timestamp, is a hash function, It is a hierarchical identifier, including general data, sensitive data and highly sensitive data;

[0036] If the hierarchical identifier is ordinary data, a symmetric key is generated ; If the hierarchical identifier is sensitive data, generate a symmetric key , asymmetric encryption public key ; If the classification identifier is highly sensitive data, a highly sensitive data classification key is generated .

[0037] Specifically, the method for performing data sharding processing includes:

[0038] Generate a sharding rule through a pseudo-random algorithm, wherein the sharding rule dynamically determines the sharding boundary according to the size of the encrypted audit data and the preset number of shards;

[0039] Dividing the encrypted audit data into a plurality of fragments according to the fragmentation rule;

[0040] Add pseudo-random perturbation information to each fragment to obtain perturbed fragment data;

[0041] Redundant encoding is performed on the disturbed sharded data to obtain the final data fragments.

[0042] Optionally, the sharding rule is expressed as: using a pseudo-random sharding sequence generator Determine the sharding order and boundaries, ,in, is a pseudo-random seed, is the number of shards, For the The fragmentation boundary of the segment satisfies , is the fragment size, is the total size of the encrypted audit data, is the number of shards, The offset for pseudo-random number generation;

[0043] Encrypt audit data across shard boundaries Divided into Clips , and for each shard Adding random perturbations , get the perturbed shard data , Generate a key for the perturbation, is a pseudo-random number generator;

[0044] Through erasure coding Redundant encoding is performed to generate redundant shards, ,in, is the Reed-Solomon encoding function, , is the redundancy rate;

[0045] Get the final data fragment .

[0046] Specifically, the method for performing distributed storage includes:

[0047] Allocate data fragments to multiple distributed storage nodes and store them in different storage nodes;

[0048] Record the storage location information of each data fragment, including the fragment index, storage node address and verification information;

[0049] Monitor the health status of storage nodes in real time. If a node failure or excessive load is detected, readjust the storage allocation rules and migrate the affected data fragments to healthy nodes.

[0050] Optionally, the method for allocating data segments includes:

[0051] Determine the input parameters, including: data segment set , Storage Node , Node health status and node load ;in, is the total number of storage nodes, Representation Node Available, Representation Node Not available, For Node The current load;

[0052] Assign a target storage node to each data fragment ,in, For Sharding The target storage node, Assign values ​​to storage based on a pseudo-random generator, For the current node The load, is the load weight;

[0053] Storage location information including shard index , storage node address and verification information ;

[0054] Perform health check on all storage nodes to obtain node health status and node load; if , then mark the node as invalid; if If the maximum preset load is exceeded, the node load is marked as too high;

[0055] If a node failure or excessive load is detected, data migration is performed. The migration methods include:

[0056] Reselect the target node for each affected shard using a new pseudo-random assignment rule ;

[0057] Shards that will be affected Slave Node Migrate to , if the verification information after migration is the same as the verification information before migration, the migration is completed.

[0058] Specifically, the method for performing data integrity verification includes:

[0059] Generate a checksum hash value for each stored data fragment, and build a Merkle tree based on all the checksum hash values ​​to generate a root hash value that represents the integrity of the overall data;

[0060] The verification hash value of each data fragment is associated with the corresponding storage node location information and recorded in the verification information storage mechanism;

[0061] When storing or accessing data, the integrity of the data fragment is verified by comparing the stored check hash value with the check hash value calculated in real time;

[0062] When the overall data integrity needs to be verified, the root hash value of the Merkle tree is recalculated and compared with the stored root hash value;

[0063] If it is found that the integrity check of the data fragment fails, the damaged data fragment is located through the verification information storage mechanism, and restored in combination with the redundant data fragment to ensure the integrity of the stored data.

[0064] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the audit data distributed storage method based on a multi-layer encryption strategy as described above.

[0065] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the audit data distributed storage method based on a multi-layer encryption strategy as described above.

[0066] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0067] The present invention combines a multi-layer encryption strategy to achieve flexible encryption protection for data of different sensitivities, improves the security of highly sensitive data, and takes into account the encryption efficiency of ordinary data. Through the combination of sharding and redundant coding technology, it not only improves the reliability and fault tolerance of the distributed storage system, but also reduces the resource overhead caused by storage redundancy; the dynamic monitoring and migration mechanism of distributed storage nodes ensures the stable operation of the system in the event of node failure or excessive load. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The accompanying drawings illustrate exemplary embodiments of the present invention and, together with the description thereof, are used to explain the principles of the present invention. These drawings are included to provide a further understanding of the present invention, and the accompanying drawings are included in and constitute a part of this specification and do not constitute a limitation of the embodiments of the present invention.

[0069] Figure 1 It is a flow chart of a method for distributed storage of audit data based on a multi-layer encryption strategy according to the present invention.

[0070] Figure 2 It is a schematic diagram of the flow of data encryption processing according to the present invention. DETAILED DESCRIPTION

[0071] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and implementation methods. It is understood that the specific implementation methods described herein are only used to explain the relevant content, rather than to limit the present invention.

[0072] It should also be noted that, for the convenience of description, only the parts related to the present invention are shown in the drawings.

[0073] In the absence of conflict, the embodiments and features of the embodiments of the present invention may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0074] Embodiment 1

[0075] like Figure 1 As shown, a distributed storage method for audit data based on a multi-layer encryption strategy is provided, including:

[0076] Data encryption processing divides audit data into multiple levels according to their sensitivity, and adopts different encryption strategies for audit data according to different levels to obtain encrypted audit data.

[0077] The purpose of data encryption processing is to provide targeted encryption strategies for different levels of data according to the sensitivity of audit data. Through quantitative analysis of data sensitivity, data is divided into three levels: ordinary data, sensitive data, and highly sensitive data. Different encryption strategies are used for different levels of data to meet the balance between security and efficiency. Ordinary data is protected by efficient symmetric encryption methods, sensitive data combines symmetric encryption and asymmetric encryption strategies to enhance transmission and storage security, and highly sensitive data uses fully homomorphic encryption technology, allowing data to be operated in an encrypted state to ensure the security of its processing process.

[0078] Data sharding processing generates sharding rules through a pseudo-random algorithm, and dynamically shards the encrypted audit data based on the sharding rules to generate multiple data fragments.

[0079] After data encryption is completed, the system dynamically shards the encrypted data. The sharding rules are dynamically generated through a pseudo-random algorithm to ensure the randomness and security of the distribution of data fragments. The boundaries of each shard are determined by the rules, and the encrypted data is divided into several fragments. In order to enhance security, pseudo-random perturbation information is added to each fragment during the sharding process to further confuse the data content. This process not only improves the randomness of data sharding, but also provides security for subsequent distributed storage.

[0080] Distributed storage, storing data fragments in multiple distributed storage nodes.

[0081] After data sharding is completed, the fragments are stored in multiple distributed storage nodes. The health status and load of the storage nodes are considered during the allocation process to ensure that the data is evenly distributed among the storage nodes. The storage location information of each data fragment, including the shard index and storage node address, will be recorded to support fast retrieval. Distributed storage can effectively improve the system's fault tolerance and the reliability of data storage. Even if some nodes fail or the load is too high, the system can still reallocate data fragments through a dynamic migration mechanism to maintain system stability.

[0082] Data integrity verification, through hash calculation, Merkle tree verification and verification information storage mechanism, the integrity of the stored data is verified.

[0083] By calculating the verification information of each storage fragment, it is ensured that the data has not been tampered with during storage, migration or access. The check value is generated by the hash function, and the integrity of all data fragments is quickly verified by building a Merkle tree. The root hash value is used for overall verification, while the child node verification information can quickly locate and confirm the damaged fragment. In the event of damage or loss, the system can restore the original data through the verification mechanism combined with the stored redundant data fragments to ensure the integrity and consistency of data storage.

[0084] Embodiment 2

[0085] This embodiment specifically describes the data encryption process.

[0086] like Figure 2 As shown, the summary method of data encryption processing includes:

[0087] According to the sensitivity of audit data, audit data is divided into multiple levels, including ordinary data, sensitive data and highly sensitive data;

[0088] Symmetric encryption strategy is used to encrypt ordinary data;

[0089] Encrypt sensitive data by combining symmetric encryption and asymmetric encryption, wherein after the sensitive data is encrypted using symmetric encryption to obtain the initial ciphertext, the initial ciphertext is encrypted again using asymmetric encryption;

[0090] Use fully homomorphic encryption strategy to encrypt highly sensitive data;

[0091] Get encrypted audit data.

[0092] Specific methods include:

[0093] Determine the real-time sensitivity of audit data through quantitative assessment of multiple sensitivity indicators , ,in, For the A sensitivity index function is used to measure the sensitivity of audit data on a certain indicator, such as the privacy level of the data, access frequency, etc. For the The indicator weight of a sensitivity indicator indicates the importance of the indicator to the sensitivity calculation. is the total number of sensitivity indicators.

[0094] Determining the sensitivity threshold and , , and classify audit data based on sensitivity thresholds;

[0095] like , then the corresponding audit data belongs to ordinary data, indicating that the data risk is low, and symmetric encryption strategy is used for encryption;

[0096] like , then the corresponding audit data is sensitive data, indicating that the data has certain risks, and is encrypted by combining symmetric encryption and asymmetric encryption;

[0097] like , then the corresponding audit data is highly sensitive data, indicating that the data risk is extremely high, and it is encrypted using a fully homomorphic encryption strategy;

[0098] Methods for encrypting ordinary data include:

[0099] Symmetric encryption is used for ordinary data, and the data is initially encrypted through the AES-GCM mode. ,in, is the initial ciphertext, For ordinary data plaintext, is a symmetric key, is a random number, It is symmetric encryption;

[0100] Through a pseudo-random number generator Generate a perturbation matrix and perturb the initial ciphertext to generate ordinary data ciphertext ,in, For The perturbation matrices of equal length, It is a bitwise XOR operation; the perturbation processing further enhances the encryption strength of ordinary data and prevents direct analysis attacks on the ciphertext.

[0101] Methods for encrypting sensitive data include:

[0102] Segment sensitive data and Divided into The size is of equal length , , Segmentation is the process of breaking down a large data set into multiple small segments, which helps improve encryption efficiency and supports more flexible encryption strategies.

[0103] For each piece of data Two-level encryption is performed; the first level of encryption is symmetric encryption, using AES-GCM mode to generate intermediate ciphertext: The second level of encryption is asymmetric encryption, which performs secondary encryption on the intermediate ciphertext and its hash value. ,in, for The hash value of the segment, is a symmetric key, is the asymmetric encryption public key, For the The random number of the segment, For symmetric encryption, It is asymmetric encryption;

[0104] Combine to obtain sensitive data ciphertext ;

[0105] Methods for encrypting highly sensitive data include:

[0106] Get dynamic parameters ,in, is the complexity of encryption state operations, including the weights of addition and multiplication. is a linear mapping or a nonlinear function, For bitwise XOR operation, Classify keys for highly sensitive data; , is the number of encrypted state addition operations, is the number of encrypted state multiplication operations, is the multiplication complexity weight;

[0107] Encrypt highly sensitive data ,in, For highly sensitive data plaintext, Fully Homomorphic Encryption (FHE) is an encryption technology that allows operations (such as addition and multiplication) to be performed directly on ciphertext. The result is still ciphertext, which is consistent with the result of the plaintext after decryption. This technology supports computing while protecting data privacy, and is suitable for processing highly sensitive data.

[0108] Methods for obtaining the key include:

[0109] Obtaining the master key through quantum key distribution , and generate hierarchical keys based on the master key ,in, is the timestamp, is a hash function, It is a hierarchical identifier, including general data, sensitive data and highly sensitive data;

[0110] If the hierarchical identifier is ordinary data, a symmetric key is generated ; If the hierarchical identifier is sensitive data, generate a symmetric key , asymmetric encryption public key ; If the classification identifier is highly sensitive data, a highly sensitive data classification key is generated .

[0111] The hierarchical encryption method of this embodiment achieves efficient protection for different types of audit data by combining data sensitivity quantification, multi-level encryption strategies and dynamic key generation mechanisms. Symmetric encryption combined with perturbation matrices is used for ordinary data, taking into account both security and performance; sensitive data is combined with segmented encryption and multi-level encryption to enhance the security of data storage and transmission; highly sensitive data is encrypted through full homomorphic encryption and dynamic parameter adjustment to ensure data privacy in complex computing scenarios. The hierarchical key management mechanism further enhances the flexibility and security of the system, and provides strong technical support for the audit data storage needs of different scenarios.

[0112] Embodiment 3

[0113] Methods for data sharding include:

[0114] The sharding rules are generated by a pseudo-random algorithm, and the sharding rules dynamically determine the sharding boundaries according to the size of the encrypted audit data and the preset number of shards; the sharding rules are dynamically generated by a pseudo-random algorithm to determine the sharding boundaries of the encrypted audit data. The pseudo-random algorithm takes a random seed as input and generates a pseudo-random sharding sequence according to the size of the data and the number of shards for dividing the data. The pseudo-random algorithm is an algorithm that generates a random number sequence through a deterministic method. Although the generated sequence is random on the surface, it is actually reproducible, and only the same seed needs to be provided. The sharding rules are the rules for dividing the data according to the pseudo-random sequence, which determine the starting and ending points of each shard, ensuring the uncertainty and security of the sharding position.

[0115] According to the sharding rule, the encrypted audit data is divided into multiple fragments; based on the generated sharding rule, the encrypted audit data is divided into multiple fragments. The number of fragments and the size of each fragment are determined by the total size of the data and the preset number of fragments.

[0116] Pseudo-random perturbation information is added to each fragment to obtain the perturbed fragment data; pseudo-random perturbation information is added to the data content of each fragment to further enhance the security of the data. The pseudo-random perturbation generates a perturbation matrix based on the fragment index and key through a pseudo-random number generator.

[0117] Redundant encoding is performed on the disturbed shard data to obtain the final data fragments. Redundant encoding (such as Reed-Solomon encoding) is applied to all disturbed shard data to generate additional redundant fragments to improve data reliability and fault tolerance.

[0118] The data sharding processing method generates sharding rules through a pseudo-random algorithm, achieving randomness and dynamic adjustment capabilities in the data sharding process. Adding pseudo-random perturbation information enhances the security of the sharding content, and redundant coding provides high reliability and fault tolerance for sharding data.

[0119] By generating additional redundant fragments based on the original data, redundant coding allows the complete data to be restored through the remaining original fragments and redundant fragments when some data fragments are lost or damaged. For example, even if some storage nodes fail, the data can still be recovered. Redundant coding ensures the reliability of data in distributed storage systems, and data is still available even in poor network conditions or storage hardware failures. Compared with simple multi-copy storage methods, redundant coding generates a small amount of redundant data through mathematical methods, which significantly reduces storage costs while providing the same fault tolerance.

[0120] The sharding rule is expressed as: using a pseudo-random sharding sequence generator Determine the sharding order and boundaries, ,in, is a pseudo-random seed, is the number of shards, For the The fragmentation boundary of the segment satisfies , is the fragment size, is the total size of the encrypted audit data, is the number of shards, The offset generated by the pseudo-random number is used to introduce randomness and make the shard size slightly change. The size and number of shards can be flexibly adjusted through dynamic sharding rules.

[0121] Encrypting audit data across shard boundaries Divided into Clips , and for each shard Adding random perturbations , get the perturbed shard data , Generate a key for the perturbation to ensure the uniqueness of the perturbation sequence. is a pseudo-random number generator; Index the shards, ensuring that the perturbations are different for each shard. The range of each shard is determined by the shard boundaries, e.g. Indicates from arrive data.

[0122] Through erasure coding Redundant encoding is performed to generate redundant shards, ,in, is the Reed-Solomon encoding function, , Redundancy rate; Redundant coding provides strong fault tolerance. If part of the fragment is lost or damaged, the system can reconstruct the original data from the remaining fragments and the redundant fragments.

[0123] Get the final data fragment .

[0124] Embodiment 4

[0125] Methods for distributed storage include:

[0126] Data fragments are distributed to multiple distributed storage nodes and stored in different storage nodes. During the distribution process, a pseudo-random distribution algorithm is used to ensure the uniform distribution of data fragments between nodes, combined with the health status and load of the storage nodes. The system selects the target storage node based on the pseudo-random algorithm to ensure the randomness of the shard distribution, thereby improving storage security. At the same time, the health status and load balancing of the nodes are considered, and more data is allocated to healthy and less loaded nodes to optimize system performance.

[0127] Record the storage location information of each data fragment, including the fragment index, storage node address and verification information; the fragment index uniquely identifies the location of each fragment in the original data. The storage node address is the physical or logical address of the storage node where the fragment is located. The verification information is generated by a hash algorithm (such as SHA-256) and is used to verify the integrity of the data.

[0128] Monitor the health status of storage nodes in real time. If a node failure or excessive load is detected, readjust the storage allocation rules and migrate the affected data fragments to healthy nodes.

[0129] The health status is to determine whether the node is available. For example, the heartbeat mechanism is used to detect whether the node is online. If the node fails (such as downtime or inaccessible), the node is marked as "failed".

[0130] The load condition monitors the storage utilization or access load of the node to prevent the performance of some nodes from being degraded due to overload. If the node load exceeds the preset threshold, it is marked as "overloaded".

[0131] If a node is abnormal, data migration is required. The migration steps are as follows:

[0132] Reselect target nodes: A new storage node is selected for each affected fragment based on a pseudo-random algorithm. The new target node must be in a normal state and have a load below a preset threshold.

[0133] Migrate data fragments: Migrate fragments from failed nodes or high-load nodes to new target nodes. After the migration is completed, verify the consistency of the fragment data through verification information to ensure that no data errors occurred during the migration process.

[0134] Update metadata: Update the storage node address record of the migrated fragment to keep the data fragment consistent with the metadata.

[0135] Methods for allocating data segments include:

[0136] Before allocating data segments, you first need to determine the following input parameters

[0137] Data fragment collection ,include The original data fragments and Redundant data fragments.

[0138] Storage Node , which represents the storage system Distributed storage nodes.

[0139] Node health status , Representation Node Available, indicating a node Not available.

[0140] Node Load ; For Node The current load;

[0141] Assign a target storage node to each data fragment ,in, For Sharding The target storage node, Assign values ​​to storage based on a pseudo-random generator, For the current node The load, The load weight is the weight of the load. Storage nodes with good health and low load are selected first to ensure the security of data storage and the performance of the system.

[0142] Storage location information including shard index , identifies the location of the data fragment in the original data. Storage node address , record the target node address where the fragment is stored. Verification information , generated by a hash algorithm (such as SHA-256) and used for subsequent verification of data integrity.

[0143] Perform health check on all storage nodes to obtain node health status and node load; if , then mark the node as invalid; if If the maximum preset load is exceeded, the node load is marked as too high;

[0144] If a node failure or excessive load is detected, data migration is performed. The migration methods include:

[0145] Reselect the target node for each affected shard using a new pseudo-random assignment rule ;The target node must be healthy and not include the current failed node.

[0146] Shards that will be affected Slave Node Migrate to , if the verification information after migration is the same as the verification information before migration, the migration is completed.

[0147] Embodiment 5

[0148] Methods for data integrity verification include:

[0149] Generate a verification hash value for each stored data fragment, and build a Merkle tree based on all the verification hash values ​​to generate a root hash value that represents the integrity of the overall data.

[0150] In distributed storage, each stored data fragment generates a unique check hash value (such as calculated by the SHA-256 algorithm), which serves as an indicator of data integrity. Subsequently, the system constructs a Merkle tree based on the check hash values ​​of all fragments. A Merkle tree is a tree-like data structure whose leaf nodes store the check hash values ​​of data fragments, and non-leaf nodes store the combined results of the hash values ​​of their child nodes. The top node of the tree (root hash value) represents the integrity of the entire data.

[0151] The verification hash value of each data fragment is associated with the corresponding storage node location information and recorded in the verification information storage mechanism; the verification hash value of each data fragment is associated with its location information in the storage node and recorded in an independent verification information storage mechanism. This information includes: data fragment index, verification hash value, storage node address.

[0152] When storing or accessing data, the integrity of the data fragment is verified by comparing the stored verification hash value with the real-time calculated verification hash value; after reading the stored fragment, its verification hash value is recalculated. The stored verification hash value is compared with the real-time calculated value to see if it is consistent.

[0153] When the overall data integrity needs to be verified, the root hash value of the Merkle tree is recalculated and compared with the stored root hash value; the Merkle tree is reconstructed using the stored verification hash value. The newly calculated root hash value is compared with the stored root hash value to see if they are consistent.

[0154] If the integrity check of a data segment fails, the damaged data segment is located through the verification information storage mechanism, and restored in combination with the redundant data segment to ensure the integrity of the stored data. The storage node and specific location of the damaged segment are determined through the recorded verification information. The damaged segment is reconstructed using the redundant segments generated by the redundant coding (such as Reed-Solomon coding).

[0155] This embodiment realizes efficient integrity verification of data fragments and overall data by combining the verification hash value and the Merkle tree. The verification information storage mechanism supports the rapid location of damaged data, and combined with the recovery capability of redundant fragments, ensures the security and consistency of data in a distributed storage environment.

[0156] Embodiment 6

[0157] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the audit data distributed storage method based on a multi-layer encryption strategy as described above.

[0158] Without loss of generality, computer readable media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer readable instruction data structures, program modules or other data. Computer storage media include RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage technology, CD-ROM, DVD or other optical storage, cassettes, magnetic tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that computer storage media are not limited to the above. The above-mentioned system memory and mass storage devices can be collectively referred to as memory.

[0159] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the audit data distributed storage method based on a multi-layer encryption strategy as described above.

[0160] A computer program product includes a computer program or set of instructions for performing specific tasks or implementing specific functions. These programs or instructions are designed to be executed by a processor to implement a series of predefined steps or operations. The program product may be stored in various forms of computer storage media, such as memory, hard disk, solid-state drive, optical disk or other forms of digital storage devices. It may exist in the form of compiled binary code or in the form of scripts or bytecodes that can be executed by an interpreter. The program product uses carefully designed algorithms and logical instructions to enable the processor to process data in a specific order and manner to complete various functions such as data analysis, user interaction, device control, etc.

[0161] In the description of this specification, the description with reference to the terms "one embodiment / method", "some embodiments / methods", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment / method or example are included in at least one embodiment / method or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment / method or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments / methods or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments / methods or examples described in this specification and the features of the different embodiments / methods or examples, unless they are contradictory.

[0162] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0163] It should be understood by those skilled in the art that the above embodiments are only for the purpose of clearly illustrating the present invention, and are not intended to limit the scope of the present invention. For those skilled in the art, other changes or modifications may be made based on the above invention, and these changes or modifications are still within the scope of the present invention.

Claims

1. A distributed storage method for audit data based on a multi-layer encryption strategy, characterized in that: Including: Data encryption processing, classifying audit data into multiple levels according to the sensitivity of the audit data, including ordinary data, sensitive data, and highly sensitive data, and adopting different encryption strategies for the audit data according to different levels to obtain encrypted audit data; Data sharding processing, generating a sharding rule through a pseudo-random algorithm, and dynamically sharding the encrypted audit data based on the sharding rule to generate multiple data fragments; Distributed storage, storing the data fragments in multiple distributed storage nodes; Data integrity verification, verifying the integrity of the stored data through hash calculation, Merkle tree verification, and verification information storage mechanism; The encryption method for ordinary data includes: Initial encryption of common data, C1 = E AES-GCM (K sym1 , P low , nonce), where C1 is the initial ciphertext, P low is ordinary data plaintext, K sym1 is a symmetric key, nonce is a random number, E AES-GCM It is symmetric encryption; Generate a perturbation matrix through the pseudo-random number generator G, and perturb the initial ciphertext to generate the normal data ciphertext Where R = G(K syml , nonce) is a perturbation matrix of the same length as C1; The encryption method for sensitive data includes: The sensitive data is segmented and the sensitive data P mid Divide into N equal length segments P of size B mid = {P1, P2, ..., P N }, For each data segment P b Perform two-level encryption; the first level of encryption is symmetric encryption C b,1 =E AES-GCM (K sym2 , P b , nonce b ), the second level of encryption is asymmetric encryption Among them, H(P b ) is P b The hash value of the segment, K sym2 is the symmetric key, K pub For asymmetric encryption public key, nonce b is the random number of segment b, E AES-GCM For symmetric encryption, E RSA It is asymmetric encryption; Combine to obtain sensitive data ciphertext The encryption method for highly sensitive data includes: Get dynamic parameters Among them, O complexity is the complexity of encryption state operation, f(·) is a linear mapping or nonlinear function, is a bitwise XOR operation, K high Classify keys for highly sensitive data; complexity =n add +w O ·n mult , n add is the number of encrypted addition operations, n mult is the number of encrypted state multiplication operations, w o is the multiplication complexity weight; Encrypt highly sensitive data high =E FHE (P high ,λ), where P high For highly sensitive data plaintext, E FHE It is fully homomorphic encryption.

2. The audit data distributed storage method based on a multi-layer encryption strategy according to claim 1 is characterized in that: The method for grading audit data includes: Determine the real-time sensitivity S(D) of the audit data, Among them, f a (D) is the ath sensitivity index function, w a is the indicator weight of the ath sensitivity indicator, and A is the total number of sensitivity indicators; Determining sensitivity thresholds T1 and T2, where T1 < T2, and classifying the audit data based on the sensitivity thresholds; If S(D) < T1, the corresponding audit data belongs to ordinary data and is encrypted using a symmetric encryption strategy; If T1 ≤ S(D) ≤ T2, the corresponding audit data belongs to sensitive data and is encrypted by combining symmetric encryption and asymmetric encryption; If S(D) > T2, the corresponding audit data belongs to highly sensitive data and is encrypted using a fully homomorphic encryption strategy.

3. The audit data distributed storage method based on a multi-layer encryption strategy according to claim 1 is characterized in that: The method for obtaining keys includes: Obtain the master key K through quantum key distribution master , and generate hierarchical keys based on the master key Among them, T is the timestamp, H is the hash function, ID level It is a hierarchical identifier, including general data, sensitive data and highly sensitive data; If the hierarchical identifier is common data, generate a symmetric key K sym1 ; If the hierarchical identifier is sensitive data, generate a symmetric key K sym2 , asymmetric encryption public key K pub ; If the hierarchical identifier is highly sensitive data, then generate a highly sensitive data hierarchical key K high .

4. The audit data distributed storage method based on a multi-layer encryption strategy according to claim 1 is characterized in that: The method for performing data sharding processing includes: Generating a sharding rule through a pseudo-random algorithm, where the sharding rule dynamically determines the sharding boundary according to the size of the encrypted audit data and a preset number of shards; Dividing the encrypted audit data into multiple fragments according to the sharding rule; Adding pseudo-random perturbation information to each fragment to obtain perturbed sharded data; Performing redundant encoding on the perturbed sharded data to obtain the final data fragments.

5. The audit data distributed storage method based on a multi-layer encryption strategy according to claim 4 is characterized in that: The sharding rule is expressed as: using a pseudo-random sharding sequence generator G split (s, n) determines the fragmentation order and boundaries, {p1, p2, ..., p n =G split (s, n), where s is the pseudo-random seed, n is the number of shards, and p is the i is the shard boundary of the i-th segment, satisfying is the fragment size, |D enc | is the total size of the encrypted audit data, n is the number of shards, Δ i The offset for pseudo-random number generation; The encrypted audit data D is passed through the shard boundary enc Divided into n segments {D1, D2, ..., D n }, and for each shard D i Add random perturbation R i =G(k, i), get the disturbed shard data k is the perturbation generation key, G is the pseudo-random number generator; Through the erasure code pair {D′1, D′2, ..., D′ n } Perform redundant encoding to generate m redundant fragments, {D″1, D″2, ..., D″ n+m }=RS(D′1,D′2,...,D′ n ), where RS is the Reed-Solomon encoding function, R r is the redundancy rate; Get the final data segment {D″1, D″2, ..., D″ n+m }.

6. The audit data distributed storage method based on a multi-layer encryption strategy according to claim 1 is characterized in that: The method for performing distributed storage includes: Allocating the data fragments to multiple distributed storage nodes and storing them in different storage nodes; Recording the storage location information of each data fragment, including shard index, storage node address, and verification information; Real-time monitoring the health status of the storage nodes. If a node failure or high load is detected, readjust the storage allocation rule and migrate the affected data fragments to healthy nodes.

7. The audit data distributed storage method based on a multi-layer encryption strategy according to claim 6 is characterized in that: The method for allocating data fragments includes: Determine input parameters, including: data segment set {D″1, D″2, ..., D″ n+m }、Storage nodes {N1,N2,...,N k }, node health status {J(N1), J(N2), ..., J(N k )} and node loads {L(N1), L(N2), ..., L(N k )}; where k is the total number of storage nodes, J(N i )=1 indicates node N i Available, J(N i )=0 indicates node N i Not available, L(N i ) is node N i The current load; Assign a target storage node to each data fragment Among them, N(D″ i ) is the fragment D″ i The target storage node, G(D″ i , j) is the storage allocation value based on the pseudo-random generator, L(N j ) is the current node N j The load, w is the load weight; The storage location information includes the shard index i, the storage node address N(D″) i ) and verification information H(D″ i ); Perform health check on all storage nodes to obtain node health status and node load; if J(N i )=0, then mark the node as failed; if L(N j ) exceeds the maximum preset load, the node load is marked as too high; If a node failure or high load is detected, perform data migration. The migration method includes: Reselect the target node for each affected shard using a new pseudo-random assignment rule The affected shard D″ i From node N(D″) i ) migrates to N′(D″ i ), if the verification information after migration is the same as the verification information before migration, the migration is completed.

8. The audit data distributed storage method based on a multi-layer encryption strategy according to claim 1 is characterized in that: The method for performing data integrity verification includes: Generating a verification hash value for each stored data fragment, and constructing a Merkle tree based on all the verification hash values to generate a root hash value representing the overall data integrity; Associating the verification hash value of each data fragment with the corresponding storage node location information and recording it in the verification information storage mechanism; When storing or accessing data, verify the integrity of the data fragment by comparing the stored verification hash value with the real-time calculated verification hash value; When it is necessary to verify the overall data integrity, recalculate the root hash value of the Merkle tree and compare it with the stored root hash value; If it is found that the integrity check of the data fragment fails, the damaged data fragment is located through the verification information storage mechanism, and restored in combination with the redundant data fragment to ensure the integrity of the stored data.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the audit data distributed storage method based on a multi-layer encryption strategy as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-cloud data processing control method and system

    CN118611948A

  • Method and device for ensuring data security of distributed storage system

    CN118862170A

  • Data privacy protection method based on block chain

    CN119272344A