A destruction-resistant storage system and method for extreme scenarios
Through a semi-decentralized storage system and erasure-correcting code technology, the data security and performance issues of traditional storage systems in extreme scenarios are solved, and efficient data recovery and tamper-proof capabilities are achieved in extreme destruction scenarios.
Patent Information
- Application Number
- CN202511033160.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Traditional storage systems cannot effectively resist all-round malicious attacks in extreme scenarios, have poor data security, unstable performance, high management costs, and cannot meet data recovery needs in extreme destruction scenarios.
It adopts a semi-decentralized storage system, uses super nodes, data nodes, edge nodes and relay servers, combines gRPC and REST API for communication, uses erasure correction code technology for distributed data storage, builds a WORM database to prevent tampering, and establishes a data transmission channel through P2P connection.
It can still work normally even if half of the nodes are damaged, preventing data tampering, improving space utilization, ensuring data security, reducing management costs, and maintaining stable performance.
Smart Images

Figure CN120524534B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage technology, and in particular to a destruction-resistant storage system and method for extreme scenarios. Background Art
[0002] In the traditional storage field, off-site disaster recovery technology solutions usually adopt dual active-active in the same city and off-site synchronization, and then evolve into two-site three-center or three-site five-center. These solutions can solve common natural damage scenarios, but are not capable of resisting all-round malicious attacks.
[0003] The defects of the existing technology are mainly in the following aspects:
[0004] (1) In terms of disaster recovery, a centralized storage system with a small number of nodes cannot cope with extreme destruction scenarios. If a large number of nodes are introduced, the impossibility principle described by the CAP theorem will lead to a sharp increase in management costs and a sharp decline in performance.
[0005] (2) In terms of data security, traditional storage systems do not include data security mechanisms at the bottom layer. Security processing is performed at the application layer, file system layer, network layer, or physical layer. It is difficult to prevent storage owners (non-data owners) from tampering with and obtaining user data.
[0006] (3) In terms of performance, traditional storage may cause long-tail problems of response time delay due to network jitter, resource competition, rapid load changes, software and hardware failures, etc.; at the same time, the recovery indicators RTO and RPO between different locations are heavily dependent on the deployment environment and hardware costs, and cannot meet the needs in extreme scenarios. Summary of the Invention
[0007] In order to solve the above problems, the present invention proposes a storage system and method that can achieve anti-destruction in extreme scenarios. By utilizing the geographical dispersion and random distribution characteristics of massive nodes, data nodes, edge nodes, super nodes and relay servers are combined into a semi-decentralized system, which supports normal operation without obvious performance difference even in the scenario where half of the nodes are damaged.
[0008] In order to achieve the above object, the present invention is implemented through the following technical solutions:
[0009] The present invention provides a survivable storage system for extreme scenarios, comprising:
[0010] Super node clusters are used to store metadata after sharding across regions, manage data nodes, edge nodes, users, and keys, and provide background synchronization services.
[0011] Data node clusters are used to store encrypted user data across regions, provide upload and download functions, and receive task scheduling from super nodes;
[0012] Edge nodes are used to run on the user side, provide interfaces for user and key management, and S3 backend services, and store user metadata locally;
[0013] Relay server, used to establish P2P communication between super nodes, data nodes and edge nodes according to the P2P connection request of edge nodes when super nodes, data nodes and edge nodes cannot communicate directly;
[0014] A further improvement of the present invention is that gRPC and REST API are used for direct communication and data transmission between super nodes, data nodes and edge nodes.
[0015] A further improvement of the present invention is that metadata is stored in the form of a metadata database on a super node. The metadata database is divided into a node sub-database, a user sub-database, a global metadata sub-database, and a user metadata sub-database, wherein:
[0016] The node sub-library is used to store the IDs, public key lists, and status of super nodes and data nodes;
[0017] The user sub-library is used to store user IDs, key ciphertexts, and status;
[0018] The global metadata sub-library is used to store information describing encrypted storage objects, including file objects, data blocks, and data shards.
[0019] The user metadata sub-library is used to store information describing the file, including name, date, size, permissions, ciphertext of random keys, and sharding information. It contains three tables: bucket, file, and file object.
[0020] The node sub-library and user sub-library are both independent sub-libraries, stored on all super nodes, and synchronized in the background when the super node is running;
[0021] The global metadata sub-library is divided into tables according to the ID of the virtual super node corresponding to the super node;
[0022] The user metadata sub-library is divided into sub-libraries according to the ID of the virtual super node corresponding to the super node.
[0023] A further improvement of the present invention is that each super node is composed of a cluster of master-slave systems located locally.
[0024] A further improvement of the present invention is that task scheduling includes a data reconstruction task, data nodes periodically send status to super nodes, and data nodes that fail to send status to super nodes within a timeout are marked as being in a warning state. When the number of data nodes in a warning state reaches a set threshold, a data reconstruction task is triggered. When the data reconstruction task starts, the corresponding data node is marked as being in a damaged state.
[0025] A further improvement of the present invention is that user data is compressed, encrypted, fragmented, and ULRC-encoded on edge nodes and then distributedly stored on data nodes.
[0026] The present invention provides a destruction-resistant storage method for extreme scenarios, comprising:
[0027] The edge nodes shard the user data, calculate the checksum shards using the erasure correction code, and then store them in distributed data nodes across regions.
[0028] Adopting a read-write separated multi-active architecture, metadata is stored in a distributed manner in super nodes across regions in the form of metadata databases. Edge nodes store their own metadata locally as a fast index.
[0029] Super nodes manage data nodes and edge nodes and perform task scheduling, including data reconstruction.
[0030] A further improvement of the present invention is that the storage of user data specifically includes:
[0031] Calculate the hash value of the file content on the edge node and send a query request to the corresponding super node to check whether the file with the corresponding hash value already exists. If so, deduplicate the file.
[0032] Compress user data and split it into several data blocks. Calculate the hash value and deduplication key (DDK) of each data block in sequence. Send a check request to the super node to determine whether there are duplicate data blocks. If so, deduplicate the data blocks. If not, generate a random key (DEK) and use it to encrypt the data blocks.
[0033] Segment the encrypted data block and calculate the check fragment using ULRC coding;
[0034] The data block shards and the corresponding check shards are stored in the corresponding data nodes in a distributed manner.
[0035] The beneficial effects of the present invention are: the present invention has the following advantages:
[0036] 1. Extreme damage resistance: Even if half of the data nodes are destroyed, the system can still operate normally without data loss;
[0037] 2. Data tamper-proofing: This invention utilizes the characteristics of user data storage to construct a database with WORM (write once, read many) properties, effectively preventing tampering;
[0038] 3. High space utilization: The underlying storage uses advanced erasure coding technology to achieve ciphertext deduplication. Compared with traditional copy methods, storage resource utilization is high;
[0039] 4. Data leakage prevention: High-security encryption mechanism, no plaintext data in the entire storage network, no risk of important data leakage, and strong privacy protection capabilities;
[0040] 5. Strong scalability: There is no theoretical limit to storage expansion and it does not affect performance. The expansion of storage scale will not lead to a sharp increase in management costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic diagram of the system architecture in an embodiment of the present invention;
[0042] Figure 2 Schematic diagram of a hash ring of a super node in an embodiment of the present invention;
[0043] Figure 3 is a flowchart of file uploading in an embodiment of the present invention;
[0044] Figure 4 is a flow chart of file downloading in an embodiment of the present invention;
[0045] Figure 5 4 is a flowchart of data reconstruction in an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0047] like Figure 1 As shown, an anti-destruction storage system for extreme scenarios in this embodiment includes: a super node cluster, a data node cluster, an edge node and a relay server. Communication and data transmission between the super node cluster, the data node cluster and the edge node are carried out using gRPC and REST API.
[0048] Super Node: Adopting a read-write split multi-active architecture, metadata is sharded into separate libraries and tables and stored on cross-regional super nodes. Each sharded library and table is deployed on a different node in a master-slave mode, forming an overall multi-active architecture. At the same time, each super node is a cluster with dual-active features constructed in a master-slave mode, mainly providing the following functions:
[0049] Data node management: node registration, allocation, and task scheduling.
[0050] User and key management: user registration, user name change, user deletion, key addition, key deletion, key modification, and password modification. In this embodiment, each user has a pair of login keys (asymmetric), a pair of signature keys (asymmetric), and a pair of encryption keys (symmetric). The login key (ULK) is used for login and asymmetric encryption. The private key is stored locally by the user. The user's public key and private key ciphertext are stored in the user metadata sub-library on the supernode. The ciphertext is encrypted with the encryption key and stored. The signature key (USK) is used to sign sensitive data during transmission. Like the login key, the private key is stored locally. The user information in the supernode stores the public key and private key ciphertext. The encryption key (UEK) is used to encrypt the randomly generated data key DEK when uploading files, as well as the private keys of other asymmetric keys. It is stored locally, that is, on the edge node. The login key and signature key can be lost and can be set to expire. If lost, they can be restored using the UEK. However, if the encryption key UEK is lost, the data cannot be recovered and can only be rebuilt. The UEK can be replaced. Replacing is equivalent to creating a new user with the same ID; the old user's ID will be cleared.
[0051] Metadata management: storage and management of global objects (file objects), blocks (data blocks), and shards (data shards); storage and management of user-related buckets (storage buckets), files (files), and objects (file objects).
[0052] Synchronous management of metadata databases between super nodes.
[0053] Data nodes: Data nodes are a completely decentralized peer-to-peer storage architecture, consisting of a large number of data nodes forming an ultra-large-scale storage cluster. By encrypting and distributing user data across different data nodes across different regions, the high fault tolerance of erasure codes is utilized to ensure the anti-destruction capability of the entire system. The main functions include:
[0054] Store user data shards and provide upload and download functions. Data shards include data block shards and checksum shards.
[0055] Receive task scheduling from super nodes and be responsible for tasks such as data reconstruction, verification, early warning, and status query.
[0056] Edge nodes: run on the user side, mainly providing interfaces for users, key management, and S3 backend services, and implementing functions such as compression, segmentation, encryption, deduplication, and erasure correction code calculation.
[0057] Relay Server: When supernodes, data nodes, and edge nodes are located on the regular consumer internet or behind routers and cannot communicate directly via IP addresses, relay servers assist in establishing connections and P2P communication. Relay servers have fixed, independent public IP addresses, and supernodes, data nodes, and edge nodes communicate directly with them. The relay server identifies all supernodes, data nodes, and edge nodes by node ID, maintains an internal node mapping table, receives P2P connection requests from edge nodes (clients), locates the communicating parties based on the node ID in the P2P connection request, and notifies and assists both parties in establishing the connection.
[0058] The storage principle of the present invention is to compress and split user data to obtain data blocks, shard the encrypted data blocks, calculate the corresponding check shards through erasure correction codes, and disperse the data shards including data block shards and check shards to data nodes in various locations. It fully utilizes the geographical dispersion, Byzantine fault tolerance, and high concurrent transmission characteristics of massive storage to provide theoretical support for anti-destruction in extreme scenarios and provide good performance.
[0059] The ULRC erasure-correcting code algorithm employed in this embodiment is a layered erasure-correcting code scheme based on RS. The underlying layer uses Cauchy RS code, with a default value of 64 raw data shards (k) and a configurable value of 128 for the total number of shards. LRC encoding is then applied above the underlying layer to group and calculate local parity shards. The raw data shards and parity shards are then distributed and stored on data nodes. During download, only half of the data shards are required to calculate the entire data. This ensures that even in the extreme case of half the data nodes being damaged, the data node can still function normally, while completely avoiding data security and long-tail issues in the network.
[0060] The number of data shards is between 128 and 255, which is much smaller than the total number of data nodes. When data nodes are destroyed uniformly and randomly, the loss probability of all shards of a single data to a specific data node group follows a hypergeometric distribution, which is consistent with the mathematical expectation of the loss probability of the entire system. Therefore, the anti-destruction capability of a single data is consistent with the anti-destruction capability of the entire system.
[0061] The data stored in this embodiment is divided into user data and metadata used to describe the distribution of user data and related attributes. The storage of user data is as follows: the user data is compressed, encrypted, fragmented, ULRC encoded, and then distributed and stored on the data nodes, such as Figure 3 As shown, specifically including:
[0062] The edge node calculates the hash value (file CID) of the entire file content and sends a query request to the super node to check whether a file with the hash already exists for file deduplication. That is, if a file with the hash already exists, the corresponding metadata is directly saved.
[0063] The edge node compresses user data and divides it into several data blocks. It then calculates the hash value (block CID) and deduplication key (DDK) for each data block. It then sends a request to the super node to determine whether there are duplicate data blocks for further deduplication. If no duplicates exist, a random key (DEK) is generated and used to encrypt the compressed data block to generate an EBlock. The deduplication key (DDK) is calculated by performing a sha256 operation on the data block's CID and checksum.
[0064] The edge node divides the encrypted data block (EBlock) into 16K fragments and calculates the check fragment using ULRC coding;
[0065] The edge node sends a request to the super node to allocate a data node, and sends the data shards to the allocated data node to complete the actual write operation, that is, to complete the user data storage.
[0066] Encrypt the key DEK using the user encryption key (UEK) to obtain the encrypted key UDEK and save it in the block field of the file object table of the user metadata sub-library;
[0067] Use the deduplication key DDK to encrypt the key DEK, generate the DDEK, and prepare to store it in the DDEK field in the data block table of the global metadata sub-library;
[0068] After all blocks are transferred, the edge node sends a request to the super node to notify the completion of the transfer, and sends a request to the corresponding super node to save the corresponding metadata. The edge node also saves the corresponding metadata locally as a quick index. The metadata includes the file date, size, attributes, user, key ciphertext and data node location of the shard.
[0069] In this embodiment, the user data download process, i.e., the file download process is as follows: Figure 4 Shown, including:
[0070] The edge node sends a request to the super node to read the object file information and read the data block list from it;
[0071] The edge node sends a request to the super node to read the data shard list of the specified data block ID;
[0072] The edge node sends read requests to all data nodes concurrently based on the data shard list information;
[0073] The edge node receives data segments and starts ULRC decoding when the minimum decoding standard is met. After successful decoding, all data segments being downloaded are discarded, effectively eliminating the long tail problem.
[0074] After the edge node successfully decodes, it merges the calculated data fragments into data blocks, decrypts the key UDEK, obtains the key DEK, and then uses the key DEK to decrypt the data blocks. When all data blocks are downloaded and decrypted, the data blocks are spliced together to generate a file to complete the user request.
[0075] The storage of a data node consists of several files or raw partitions, with 16K as a storage unit. The k / v database is used to record the mapping relationship between the hash and offset address of the data shard content, where K is the hash and V is the storage unit address. This mapping table is used for address conversion when reading and writing data shards.
[0076] In this embodiment, data nodes send status to super nodes regularly. Data nodes that fail to send status to super nodes within a timeout period are marked as warning. When the number of data nodes in warning state reaches a specified threshold, a scheduling reconstruction task is triggered. The warning state can be revoked before the scheduling reconstruction task is triggered. When the reconstruction task starts, the corresponding data node is marked as damaged. Figure 5 As shown, the reconstruction process includes:
[0077] Step 1: The super node receives the warning information, scans the data shard information stored in the damaged data node, and creates a rebuild shard table, that is, rebuilds the hash value table;
[0078] Step 2: Allocate N data nodes as reconstruction nodes and evenly distribute the data shards to be reconstructed to the N data nodes for reconstruction;
[0079] Step 3: The reconstruction node downloads the specified data block and calculates the missing data fragments. If only one data fragment is lost, the lost data fragment is directly saved locally and the result is sent to the super node;
[0080] Step 4: If the number of lost data shards is greater than one, upload the data of the corresponding data shards to other data nodes;
[0081] Step 5: Report the reconstruction status to the super node and rewrite the node information of the data shard;
[0082] Repeat steps 3 to 5 until all data shards are rebuilt.
[0083] The metadata storage forms are divided into two types: stored in the form of a database on the super node; the metadata of the user's own files stored locally, that is, the metadata stored by the edge node. The metadata stored in the edge node facilitates rapid local retrieval and serves as a flexible supplement for extreme anti-destruction.
[0084] Regarding the anti-destruction of the metadata database on the super node, this embodiment utilizes the characteristics of user data storage to construct a database with WORM (write once, read many times) properties, and divides the database into several sub-libraries and sub-tables according to the id of the virtual super node, and assigns them to each super node for management. Each super node has its own master library and master table, and multiple copies belonging to other super nodes, that is, slave libraries or slave tables of the super node. The master library and master table can be read and written, while the slave libraries and slave tables can only be read but not written. The metadata stored on the super node is synchronized in the background through each other's slave super nodes. Specifically, it includes: all super nodes, data nodes and edge nodes are distributed on the consistent hash ring according to their id. In a clockwise direction, the data nodes and edge nodes are managed by the nearest super node. The metadata database composed of the metadata corresponding to these nodes exists as an independent sub-library, as the master library and master table managed by the super node (with read and write permissions).
[0085] The background synchronization service runs at startup, queries the time when the data in the slave database and slave table is updated according to the locally configured strategy, and sends synchronization requests to other super nodes based on this time, obtains the master database records of other super nodes after this time, and saves them.
[0086] To further elaborate, Figure 2 As shown, data nodes and edge nodes find the virtual super node in a clockwise direction and accept the management of the virtual super node, as shown in Figure 2 D1 is managed by the virtual super node Snode1, and D2 and D3 are managed by the virtual super node Snode2.
[0087] 1024 virtual super nodes are predefined in the hash ring, and the global metadata sub-library and user metadata sub-library are evenly divided into 1024 sub-libraries. The actual number of super nodes may be less than the number of virtual super nodes, so each super node corresponds to multiple virtual super nodes, such as Figure 2 The super node Snode A corresponds to the virtual nodes Snode 1, 3, and 5, and the super node Snode B corresponds to the virtual super nodes Snode 2, 4, and 6.
[0088] When a new super node is added, the virtual super node will be reallocated to the new super node.
[0089] When a super node is deleted, the virtual super node corresponding to the deleted super node will be released to other super nodes for management.
[0090] The metadata library is divided into four categories: Nodes (node sub-library), Users (user sub-library), UisMeta (global metadata sub-library), and UserMeta (user metadata sub-library). To facilitate application-level synchronization, each record contains a timestamp field.
[0091] The node sub-library is used to record information such as the IDs, public key lists, and status of supernodes and datanodes. The node sub-library contains less data and rarely changes. It is an independent sub-library stored on all supernodes, and each supernode uses the same node sub-library (data may vary). During supernode initialization, a copy of the node sub-library is copied from an existing supernode, and synchronized in the background during runtime.
[0092] The user sub-library is used to record the user's ID, key ciphertext, and status. Super nodes use the same user sub-library and synchronize in the background.
[0093] The global metadata sub-library is used to store information describing encrypted storage objects, including information about file objects, data blocks, and data shards.
[0094] Object file object, used to record the file plaintext hash value, size, data block list, reference count, etc.
[0095] Block data block is used to record the plaintext hash value of the data block, the hash value of the root node of the markle (hash tree) of the data shard, the ciphertext of the random key, and the erasure correction code parameters.
[0096] Shard data sharding is used to record the hash of the data shard, the ID of the data block to which it belongs, and the ID of the data node where it is located;
[0097] The global metadata sub-repository is divided into 1024 sub-tables, corresponding to the IDs of the virtual super nodes corresponding to the super nodes. Each sub-table stores objects and blocks whose IDs fall within the sub-table's range. Super nodes only write to the sub-tables they manage (the master table); other sub-tables (the slave tables) are read-only.
[0098] The user metadata sub-library is used to store information including the name, date, size, permissions, random key ciphertext, and sharding information describing the file. The sharding information includes the number of data shards, encoding algorithm, and other information. The user metadata sub-library contains three tables: bucket, file, and file object:
[0099] Bucket: used to record bucket ID, name, date, etc.
[0100] File: Used to record file names, bucket IDs, and version lists. Each modification to a file generates a new version. Each version contains an object ID corresponding to a storage object.
[0101] File object: It is a reference to the file object in the global metadata database, and a field UDEK is added to store the DEK ciphertext.
[0102] Because the user metadata sub-library is logically independent, it is divided into 1024 sub-libraries based on user IDs, each managed by 1024 virtual super nodes. Similar to the global metadata sub-library, super nodes only write to the sub-library they manage, and read other sub-libraries.
[0103] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless defined similarly as herein, will not be interpreted in an idealized or overly formal sense.
[0104] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A survivable storage system for extreme scenarios, characterized by: include: Super node clusters are used to store metadata after sharding across regions, manage data nodes, edge nodes, users, and keys, and provide background synchronization services. Data node clusters are used to store encrypted user data across regions, provide upload and download functions, and receive task scheduling from super nodes; Edge nodes are used to run on the user side, provide interfaces for user and key management, and S3 backend services, and store user metadata locally; Relay server, used to establish P2P communication between super nodes, data nodes and edge nodes according to the P2P connection request of edge nodes when super nodes, data nodes and edge nodes cannot communicate directly; The metadata is stored in the form of a metadata database on the super node. The metadata database is divided into node sub-database, user sub-database, global metadata sub-database, and user metadata sub-database. The node sub-library is used to store the IDs, public key lists, and status of super nodes and data nodes; The user sub-library is used to store user IDs, key ciphertexts, and status; The global metadata sub-library is used to store information describing encrypted storage objects, including file objects, data blocks, and data shards. The user metadata sub-library is used to store information describing the file, including name, date, size, permissions, ciphertext of random keys, and sharding information. It contains three tables: bucket, file, and file object. The node sub-library and user sub-library are both independent sub-libraries, stored on all super nodes, and synchronized in the background when the super node is running; The global metadata sub-library is divided into tables according to the ID of the virtual super node corresponding to the super node; The user metadata sub-library is divided into sub-libraries according to the ID of the virtual super node corresponding to the super node.
2. The survivable storage system for extreme scenarios according to claim 1, characterized in that: Super nodes, data nodes and edge nodes use gRPC and REST API for direct communication and data transmission.
3. The survivable storage system for extreme scenarios according to claim 1, characterized in that: Each supernode consists of a cluster of master-slave systems located locally.
4. The survivable storage system for extreme scenarios according to claim 1, characterized in that: Task scheduling includes data reconstruction tasks. Data nodes send status to super nodes at regular intervals. Data nodes that fail to send status to super nodes within a timeout period are marked as warning status. When the number of data nodes in the warning status reaches the set threshold, the data reconstruction task is triggered. When the data reconstruction task starts, the corresponding data node is marked as damaged.
5. The survivable storage system for extreme scenarios according to claim 1, characterized in that: User data is compressed, encrypted, fragmented, and ULRC-encoded on edge nodes and then distributedly stored on data nodes.
6. A method for survivable storage in extreme scenarios based on the system according to any one of claims 1 to 5, characterized in that: include: The edge nodes shard the user data, calculate the checksum shards using the erasure correction code, and then store them in distributed data nodes across regions. Adopting a read-write separated multi-active architecture, metadata is stored in a distributed manner in super nodes across regions in the form of metadata databases. Edge nodes store their own metadata locally as a fast index. Super nodes manage data nodes and edge nodes and perform task scheduling, including data reconstruction.
7. The survivable storage method for extreme scenarios according to claim 6, characterized in that: The storage of user data specifically includes: Calculate the hash value of the file content on the edge node and send a query request to the corresponding super node to check whether the file with the corresponding hash value already exists. If so, deduplicate the file. Compress user data and split it into several data blocks. Calculate the hash value and deduplication key (DDK) of each data block in sequence. Send a check request to the super node to determine whether there are duplicate data blocks. If so, deduplicate the data blocks. If not, generate a random key (DEK) and use it to encrypt the data blocks. Slice the encrypted data block to obtain data block fragments, and calculate the check fragments using ULRC coding; The data block shards and the corresponding check shards are stored in the corresponding data nodes in a distributed manner.
Citation Information
Patent Citations
Scalable recursive computation for pattern identification across distributed data processing nodes
US10860622B1
Data Deletion for a Multi-Tenant Environment
US20210110055A1