Data storage method and related equipment
Patent Information
- Application Number
- CN202380081361.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2025-07-18
AI Technical Summary
Existing technologies are difficult to effectively guarantee the integrity and confidentiality of data in distributed storage networks. In particular, there are challenges in how to recover data and protect key security when a node is attacked.
By encrypting the data and splitting it into P shards, storing it in P third nodes, ensuring that at least M shards can recover the data, and using a threshold key sharing algorithm to split and recover the key, Ensure data confidentiality and integrity.
It is realized that even if an individual node is attacked, the data can still be recovered through at least M shards to ensure data integrity, and the threshold key sharing algorithm protects the key security and further ensures the confidentiality of the data.
Smart Images

Figure CN120345210A_ABST
Abstract
Description
Data storage method and related equipment Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a data storage method and related equipment. Background Art
[0002] With the rapid development of communications, the Internet of Things, and edge computing technologies, more and more devices are equipped with the ability to collect and process data. Distributed storage networks store data in a distributed manner across multiple independent devices (or storage nodes). Distributed network storage networks utilize a scalable system architecture, sharing the storage load across multiple devices. This not only improves system reliability, availability, and access efficiency, but also facilitates scalability. However, ensuring the integrity and confidentiality of data stored in distributed storage networks remains a challenge.
[0003] Summary of the Invention
[0004] The present application provides a data storage method and related equipment, which can effectively ensure the integrity and confidentiality of data.
[0005] In a first aspect, the present application provides a data storage method, which is applied to a first node, where the first node has a data storage requirement.
[0006] The data storage method includes the following steps: a first node encrypts first data to obtain encrypted data; and the first node sends a data storage request to a second node. The second node is any node in a distributed storage network. The data storage request includes the encrypted data. The first node receives feedback information. The feedback information indicates that P third nodes have stored P shards, where the P shards are obtained by splitting the encrypted data. At least M of the P shards are used to recover the encrypted data. The P third nodes are nodes in the distributed storage network. P and M are positive integers greater than one, and M is less than P.
[0007] In this solution, the first node first encrypts the first data to generate encrypted data, effectively ensuring data confidentiality. The second node then splits the encrypted data into P shards. These P shards are then stored by P third nodes, with each third node storing one shard. The encrypted data can be recovered based on at least M of the P shards. Therefore, even if individual nodes in the P third nodes are attacked, the encrypted data can still be recovered based on at least M of the P shards, effectively ensuring the integrity of the stored data.
[0008] In conjunction with one possible implementation of the first aspect, the data storage method further includes the following steps: the first node sends P subkeys to P third nodes. The P subkeys are obtained by splitting a first key used to encrypt the first data, wherein at least M of the P subkeys are used to recover the first key.
[0009] Therefore, in this application, when P third nodes have stored P shards, the first node sends P subkeys to the P third nodes, with each third node storing one subkey. If the subkey stored by an individual of the P third nodes is lost, the encrypted data cannot be decrypted using a single subkey. Therefore, the solution of this application can effectively ensure the security of the first key and further protect the confidentiality of the stored data.
[0010] In combination with a possible implementation of the first aspect, the data storage request further includes a data name of the first data, where the data name corresponds to a data digest of the encrypted data.
[0011] Therefore, in this application, the data storage request sent by the first node to the second node may further include the data name of the first data, which corresponds to the data digest of the encrypted data. In this way, the second node can obtain the data digest of the encrypted data based on the data name of the first data. Based on the data digest of the encrypted data, the second node can verify the correctness of the encrypted data it receives, that is, determine whether the encrypted data has been tampered with, thereby ensuring the correctness of the stored data.
[0012] In combination with a possible implementation of the first aspect, the data storage method further includes the following steps: the first node sends a data registration request to the first network, where the data registration request includes a data name of the first data and a data summary of the encrypted data.
[0013] Therefore, in this application, the first node sends a data registration request to the first network. The first network can respond to the data registration request and store the data carried in the request. That is, the first network stores the correspondence between the data name of the first data and the data digest of the encrypted data. In this way, the second node can obtain the corresponding data digest of the encrypted data from the first network based on the data name of the first data.
[0014] In combination with a possible implementation of the first aspect, the data registration request further includes first indication information, where the first indication information is used to determine P, and the first indication information corresponds to the data name.
[0015] Therefore, in the present application, the first network can also store the correspondence between the data name of the first data and the first indication information, so that the second node can obtain the corresponding first indication information according to the data name of the first data, and then determine P according to the first indication information, and then split the encrypted data according to P.
[0016] In one possible implementation of the first aspect, the data storage method further includes the following steps: the first node receives second indication information. The second indication information is used to indicate that the P third nodes have stored the P subkeys. The first node sends third indication information to the first network, the third indication information being used to indicate that the first data storage is complete, and the third indication information corresponds to the data name.
[0017] Therefore, in the present application, after the first node receives the second indication information, it sends the third indication information to the first network, so that the first network can update the storage status of the first data according to the third indication information, and the data requester can learn from the first network that the first data has been stored and is in a requestable state.
[0018] In a second aspect, the present application further provides a data storage method, which is applied to a second node, which is any node in a distributed storage network.
[0019] Specifically, the data storage method includes the following steps: a second node receives a data storage request sent by a first node. The data storage request includes encrypted data, where the encrypted data is obtained by encrypting the first data by the first node. The second node splits the encrypted data into P shards. At least M of the P shards are used to recover the encrypted data, where P and M are positive integers greater than one, and M is less than P. The second node determines P third nodes in a distributed storage network. The second node sends the P shards to the P third nodes, respectively.
[0020] In this solution, the second node splits the encrypted data into P shards. P third nodes then store these P shards, with one third node storing one shard. The encrypted data can be recovered based on at least M of the P shards. Therefore, even if individual nodes in the P third nodes are attacked, the encrypted data can still be recovered based on at least M of the P shards, effectively protecting the integrity of the stored data.
[0021] In one possible implementation of the second aspect, the second node determining P third nodes in the distributed storage network includes: the second node determining the P third nodes based on a first value. The first value includes any of the following: a data digest of encrypted data, a first preset value, a first data digest, or a value stored on a distributed ledger technology (DLT) network, where the first data digest is a data digest of data stored on the DLT network. The first value may also be a digest of other data or other value, as long as the other data or other value is known to each node in the distributed storage network.
[0022] Therefore, in the present application, P third nodes for storing data in the distributed storage network can be determined based on the first value, where P is a positive integer greater than one. The method for determining P third nodes is very convenient and effective.
[0023] In one possible implementation of the second aspect, the ID of the i-th third node among the P third nodes is a hash value of a second value. The second value is the sum of the ID of the i-1th second node and a second preset value, or the sum of the ID of the i-1th second node and a preset function value, where the preset function value is related to i. The ID of the first third node among the P third nodes is the hash value of the first value.
[0024] In one possible implementation of the second aspect, the data storage request further includes a data name of the first data, and the data storage method further includes the following steps: the second node obtains a data digest of the encrypted data corresponding to the data name based on the data name; and the second node verifies the correctness of the encrypted data based on the data digest.
[0025] Therefore, in this application, the data storage request sent by the first node to the second node may further include the data name of the first data, which corresponds to the data digest of the encrypted data. In this way, the second node can obtain the data digest of the encrypted data based on the data name of the first data. Based on the data digest of the encrypted data, the second node can verify the correctness of the encrypted data it receives, that is, determine whether the encrypted data has been tampered with, thereby ensuring the correctness of the stored data.
[0026] In conjunction with a possible implementation of the second aspect, the data storage method further includes the following steps: the second node obtains first indication information corresponding to the data name according to the data name. The first indication information is used to determine P.
[0027] Therefore, in the present application, the second node can obtain the corresponding first indication information according to the data name of the first data, and then can determine P according to the first indication information, and then split the encrypted data according to P.
[0028] In one possible implementation of the second aspect, the second node obtains, based on the data name, a data digest of the encrypted data corresponding to the data name, including: the second node sends a digest request to the first network, the digest request including the data name; and the second node receives the data digest of the encrypted data corresponding to the data name sent by the first network.
[0029] Therefore, in the present application, the second node may send a digest request to the first network to request to obtain the data digest of the encrypted data corresponding to the data name of the first data.
[0030] In combination with a possible implementation of the second aspect, the second node obtains first indication information corresponding to the data name according to the data name, including: the second node obtains the first indication information corresponding to the data name from the first network according to the data name.
[0031] Therefore, in this application, the second node can request the first network to obtain the first indication information corresponding to the data name of the first data. For example, the first network can respond to the digest request and send the data digest of the encrypted data and the first indication information to the second node. For another example, the second node can also initiate another request to the first network to obtain the first indication information.
[0032] In a third aspect, the present application further provides a data storage method, which is applied to a data storage system. The system includes a first node, a second node, and a third node, wherein the first node is a node having a data storage requirement.
[0033] Specifically, the above-mentioned data storage method includes the following steps: the first node encrypts the first data to obtain encrypted data. The first node sends a data storage request to the second node. The data storage request includes the above-mentioned encrypted data, and the second node is any node in the distributed storage network. The second node splits the encrypted data into P shards. Among them, at least M shards among the P shards are used to restore the encrypted data. The second node determines P third nodes in the distributed network and sends the P shards to the P third nodes respectively. P and M are positive integers greater than one, and M is less than P. The P third nodes are nodes in the distributed storage network. Each of the P third nodes stores one shard among the P shards.
[0034] In this solution, the first node first encrypts the first data to generate encrypted data, effectively ensuring data confidentiality. The second node then splits the encrypted data into P shards. These P shards are then stored by P third nodes, with one third node storing one shard. The encrypted data can be recovered based on at least M of the P shards. Therefore, even if individual nodes in the P third nodes are attacked, the encrypted data can still be recovered based on at least M of the P shards, effectively ensuring the integrity of the stored data.
[0035] In one possible implementation of the third aspect, the data storage method further includes the following steps: the first node receives feedback information. The feedback information is used to indicate that the P third nodes have stored the P shards. The first node sends P subkeys to the P third nodes. The P subkeys are obtained by splitting a first key used to encrypt the first data, wherein at least M of the P subkeys are used to recover the first key. Each of the P third nodes stores one of the P subkeys.
[0036] Therefore, in this application, after receiving the feedback information, the first node sends P subkeys to P third nodes, so that each third node stores one subkey. If the subkey stored by an individual of the P third nodes is lost, the encrypted data cannot be decrypted using a single subkey. Therefore, the solution of this application can effectively ensure the security of the first key and further protect the confidentiality of the stored data.
[0037] In one possible implementation of the third aspect, the data storage request further includes a data name of the first data, and the data storage method further includes the following steps: the second node obtains a data digest of the encrypted data corresponding to the data name based on the data name; and the second node verifies the correctness of the encrypted data based on the data digest.
[0038] Therefore, in the present application, the second node can verify the correctness of the encrypted data it receives based on the data summary of the encrypted data, that is, determine whether the encrypted data has been tampered with, thereby ensuring the correctness of the stored data.
[0039] In combination with a possible implementation of the third aspect, the above-mentioned data storage method further includes the following steps: the second node obtains first indication information corresponding to the data name according to the data name, and the first indication information is used to determine P.
[0040] Therefore, in the present application, the second node can obtain the corresponding first indication information according to the data name of the first data, and then can determine P according to the first indication information, and then split the encrypted data according to P.
[0041] In one possible implementation of the third aspect, the data storage system further includes a first network, and the second node obtains a data digest of the encrypted data corresponding to the data name based on the data name, including: the second node sends a digest request to the first network. The digest request includes the data name. The second node receives the data digest of the encrypted data corresponding to the data name sent by the first network.
[0042] Therefore, in the present application, the second node can obtain the data summary of the corresponding encrypted data from the first network according to the data name of the first data.
[0043] In combination with a possible implementation of the third aspect, the second node receives first indication information corresponding to the data name sent by the first network in response to the summary request.
[0044] Therefore, in the present application, the first network may respond to the above digest request and send the data digest of the encrypted data and the first indication information to the second node together.
[0045] In combination with a possible implementation of the third aspect, the data storage method further includes the following steps: the first node sends a data registration request to the first network, where the data registration request includes a data name of the first data and a data summary of the encrypted data.
[0046] Therefore, in the present application, the first node sends a data registration request to the first network, and the first network can respond to the data registration request and store the data carried in the request.
[0047] In combination with a possible implementation of the third aspect, the data registration request further includes first indication information for determining P.
[0048] Therefore, in the present application, the first network may further store the correspondence between the data name of the first data and the first indication information.
[0049] In one possible implementation of the third aspect, the second node determines P third nodes in the distributed storage network, including: the second node determines the P second nodes based on a first value. The first value includes any one of the following: a data digest of encrypted data, a first preset value, a first data digest, or a value stored on a distributed ledger technology (DLT) network, where the first data digest is a data digest of data stored on the DLT network.
[0050] Therefore, in the present application, P third nodes for storing data in the distributed storage network can be determined based on the first value, where P is a positive integer greater than one. The method for determining P third nodes is very convenient and effective.
[0051] In one possible implementation of the third aspect, the ID of the i-th third node among the P third nodes is a hash value of a second value. The second value is the sum of the ID of the i-1th third node and a second preset value, or the second value is the sum of the ID of the i-1th third node and a preset function value, where the preset function value is related to i. The ID of the first third node among the P third nodes is a hash value of the first value.
[0052] In a fourth aspect, the present application also provides a data storage method, which is applied to a distributed storage network.
[0053] Specifically, the above-mentioned data storage method includes the following steps: the second node receives a data storage request sent by the first node. The second node is any node in the distributed storage network, the first node is a node with data storage requirements, and the data storage request includes encrypted data, which is obtained by the first node encrypting the first data. The second node splits the encrypted data into P shards, wherein at least M shards of the P shards are used to restore the encrypted data. The second node determines P third nodes in the distributed network and sends the P shards to the P third nodes respectively, where P and M are positive integers greater than one, and M is less than P. The P third nodes are nodes in the distributed storage network. Each of the P third nodes stores one shard of the P shards.
[0054] In this solution, the second node splits the encrypted data into P shards. P third nodes then store these P shards, with one third node storing one shard. The encrypted data can be recovered based on at least M of the P shards. Therefore, even if individual nodes in the P third nodes are attacked, the encrypted data can still be recovered based on at least M of the P shards, effectively protecting the integrity of the stored data.
[0055] In conjunction with one possible implementation of the fourth aspect, the data storage method further includes the following steps: each of the P third nodes receives one of the P subkeys, and each third node stores one subkey. The P subkeys are obtained by splitting a first key used to encrypt the first data, wherein at least M of the P subkeys are used to recover the first key. The P subkeys are sent by the first node after receiving feedback information, and the feedback information is used to indicate that the P third nodes have stored the P shards.
[0056] In conjunction with one possible implementation of the fourth aspect, the second node determining P third nodes in the distributed storage network includes: the second node determining the P third nodes based on a first value. The first value includes any of the following: a data digest of encrypted data, a first preset value, a first data digest, or a value stored on a distributed ledger technology (DLT) network, where the first data digest is a data digest of data stored on the DLT network. The first value may also be a digest of other data or other value, as long as the other data or other value is known to each node in the distributed storage network.
[0057] In one possible implementation of the fourth aspect, the ID of the i-th third node among the P third nodes is a hash value of a second value. The second value is the sum of the ID of the i-1th second node and a second preset value, or the sum of the ID of the i-1th second node and a preset function value, where the preset function value is related to i. The ID of the first third node among the P third nodes is the hash value of the first value.
[0058] In one possible implementation of the fourth aspect, the data storage request further includes a data name of the first data, and the data storage method further includes the following steps: the second node obtains a data digest of the encrypted data corresponding to the data name based on the data name; and the second node verifies the correctness of the encrypted data based on the data digest.
[0059] In conjunction with a possible implementation of the fourth aspect, the data storage method further includes the following steps: the second node obtains first indication information corresponding to the data name according to the data name. The first indication information is used to determine P.
[0060] In conjunction with a possible implementation of the fourth aspect, the second node obtains, based on the data name, a data digest of the encrypted data corresponding to the data name, including: the second node sends a digest request to the first network. The digest request includes the data name. The second node receives the data digest of the encrypted data corresponding to the data name, sent by the first network in response to the digest request.
[0061] In combination with a possible implementation of the fourth aspect, the second node obtains the first indication information corresponding to the data name according to the data name, including: the second node obtains the first indication information corresponding to the data name from the first network according to the data name.
[0062] In a fifth aspect, the present application also provides a node determination method, which can be applied to a node determination device or a chip in a node determination device.
[0063] Specifically, the node determination method includes the following steps: splitting a first task into P subtasks, where P is a positive integer greater than one; and determining the P nodes in the distributed storage network based on a first value. The first value includes any of the following: a digest corresponding to data to be processed by the first task, a first preset value, a first data digest, or a value stored on a distributed ledger technology (DLT) network, where the first data digest is a data digest of the data stored on the DLT network.
[0064] In this solution, the first value can be used to determine the P nodes in the distributed storage network for executing the P subtasks, where P is a positive integer greater than one. The method for determining the P nodes is very convenient and effective. The first task can include a storage task or a computing task.
[0065] In one possible implementation of the fifth aspect, the identity ID of the i-th node among the P nodes is a hash value of a second value, where the second value is the sum of the ID of the i-1-th node and a second preset value, or the second value is the sum of the ID of the i-1-th node and a preset function value, where the preset function value is related to i. The ID of the first node among the P nodes is the hash value of the first value.
[0066] In combination with a possible implementation of the fifth aspect, the node determination method further includes: sending the P subtasks to the P nodes respectively, so that one node executes one subtask.
[0067] In a sixth aspect, the present application further provides a first node. The first node includes an encryption module, a sending module, and a receiving module. Wherein:
[0068] The encryption module is used to encrypt the first data to obtain encrypted data.
[0069] The sending module is configured to send a data storage request to a second node, where the second node is any node in the distributed storage network, and the data storage request includes encrypted data.
[0070] A receiving module is configured to receive feedback information. The feedback information indicates that P third nodes have stored P shards, where the P shards are obtained by splitting the encrypted data. At least M of the P shards are used to recover the encrypted data. The P third nodes are nodes in a distributed storage network, where P and M are positive integers greater than one, and M is less than P.
[0071] In a seventh aspect, the present application further provides a second node, which is any node in a distributed storage network. The second node includes a receiving module, a splitting module, a determining module, and a sending module, wherein:
[0072] The receiving module is configured to receive a data storage request sent by the first node, wherein the data storage request includes encrypted data, and the encrypted data is obtained by the first node encrypting the first data.
[0073] The splitting module is configured to split the encrypted data into P fragments, wherein at least M fragments of the P fragments are used to recover the encrypted data, where P and M are positive integers greater than one, and M is less than P.
[0074] The determination module is used to determine P third nodes in the distributed storage network.
[0075] The sending module is used to send the P slices to P third nodes respectively.
[0076] In an eighth aspect, the present application further provides a data storage system, which includes a first node, a second node, and a third node.
[0077] The first node is configured to encrypt the first data to obtain encrypted data and send a data storage request to the second node, wherein the data storage request includes the encrypted data. The second node is any node in the distributed storage network.
[0078] The second node is configured to split the encrypted data into P shards. At least M of the P shards are used to recover the encrypted data. P and M are positive integers greater than one, and M is less than P. P third nodes are determined in the distributed storage network, and the P shards are sent to each of the P third nodes.
[0079] The third node is used to store one of the P shards.
[0080] In a ninth aspect, the present application further provides a node determination device, the node determination device including a splitting module and a determination module, wherein:
[0081] The splitting module is used to split the first task into P subtasks, where P is a positive integer.
[0082] The determining module is configured to determine P nodes in the distributed storage network based on a first value. The first value includes any one of the following: a digest corresponding to the data to be processed by the first task, a first preset value, a first data digest, or a value stored on a distributed ledger technology (DLT) network, where the first data digest is a data digest of the data stored on the DLT network.
[0083] In one possible implementation of the ninth aspect, the identity ID of the i-th node among the P nodes is a hash value of a second value, where the second value is the sum of the ID of the i-1-th node and a preset value, or the second value is the sum of the ID of the i-1-th node and a preset function value, where the preset function value is related to i. The ID of the first node among the P nodes is the hash value of the first value.
[0084] In the tenth aspect, the present application also provides a communication device, which includes one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, and the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the communication device performs a method as described in any one of the first aspect, the second aspect, the third aspect, the fourth aspect or the fifth aspect.
[0085] In the eleventh aspect, the present application also provides a computer-readable storage medium, which stores instructions. When the instructions are executed by a processor, the method described in any one of the first aspect, the second aspect, the third aspect, the fourth aspect or the fifth aspect is implemented.
[0086] In the twelfth aspect, the present application also provides a computer program product, including a computer program, which, when the computer program runs on a processor, implements the method described in any one of the first, second, third, fourth or fifth aspects.
[0087] In the thirteenth aspect, the present application provides a chip system, which is applied to a communication device, and the chip system includes one or more processors, and the processors are used to call computer instructions to enable the communication device to execute any method as described in the first aspect, the second aspect, the third aspect, the fourth aspect or the fifth aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] FIG1 is a schematic diagram of the structure of a data storage system provided in an embodiment of the present application;
[0089] FIG2 is a flow chart of a data storage method provided in an embodiment of the present application;
[0090] FIG3A is an interactive flow chart of a data storage method provided in an embodiment of the present application;
[0091] FIG3B is a schematic diagram of an erasure code provided in an embodiment of the present application;
[0092] FIG3C is a schematic diagram of an erasure code in a fault condition provided by an embodiment of the present application;
[0093] FIG3D is a schematic diagram of erasure code operation under a fault condition provided by an embodiment of the present application;
[0094] FIG3E is a schematic diagram of data recovery using erasure codes according to an embodiment of the present application;
[0095] FIG4 is a schematic diagram of a flow chart of a node determination method provided in an embodiment of the present application;
[0096] FIG5 is a schematic structural diagram of a first node provided in an embodiment of the present application;
[0097] FIG6 is a schematic structural diagram of a second node provided in an embodiment of the present application;
[0098] FIG7 is a schematic diagram of the structure of a node determination device provided in an embodiment of the present application;
[0099] FIG8 is a schematic diagram of the structure of a communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0100] The terms used in the following examples of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular expressions "a," "an," "said," "above," "the," and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and encompasses any or all possible combinations of one or more of the listed items.
[0101] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0102] Since the embodiments of the present application involve methods, in order to facilitate understanding, the relevant terms and concepts involved in the embodiments of the present application are first introduced below.
[0103] (1) Distributed storage network
[0104] In an embodiment of the present application, a distributed storage network is a data storage network composed of a plurality of distributed storage nodes (referred to as nodes). For example, a distributed hash table (DHT) network is one type of distributed storage network, wherein the storage nodes of the network can be base stations and / or core network elements. For another example, the storage nodes of the distributed storage network can also be edge servers.
[0105] Distributed network storage network adopts a scalable system structure and uses multiple devices to share the storage load. It not only improves the reliability, availability and access efficiency of the system, but also makes it easy to expand.
[0106] (2) Distributed Ledger Technology (DLT) Network
[0107] A distributed ledger technology (DLT) network is a decentralized data management architecture. The network consists of several nodes, each of which replicates and stores an identical copy of the ledger. When data in the ledger changes, all nodes independently update and use a consensus algorithm to determine the correct copy. Once consensus is reached, all nodes synchronize based on the correct copy. DLT networks also use encryption algorithms and digital signatures to enhance system security. DLT networks can be categorized by the data structure used, such as blockchains and directed acyclic graphs, or by the consensus algorithm used, such as Proof of Work (PoW) and Proof of Stake (PoS). For example, nodes in a DLT network can be base stations and / or core network elements.
[0108] (3) Information dispersion algorithm
[0109] In coding theory, there is a forward error correction (FEC) coding method, also known as erasure coding. This technology can recover k bytes of data lost in the original data from n information containing coded bytes.
[0110] Information dispersal algorithms process the original data using erasure coding to create P fragments, where P = M + N, with M less than or equal to N. The original data can be recovered from any of the M fragments. Information dispersal algorithms include Reed-Solomon (RS) erasure coding or Locally Repairable Codes (LRC). The LRC algorithm is a local parity coding method whose core concept is to divide the parity block into a global parity block and a local reconstruction parity block, which are then grouped for calculation during fault recovery.
[0111] (4) Threshold key sharing algorithm
[0112] The threshold key sharing algorithm processes the original key to obtain P subkeys, and the original key can be solved using a combination of greater than or equal to M subkeys.
[0113] The following describes the methods involved in the embodiments of the present application in combination with the above terms.
[0114] In the prior art, how to ensure the integrity and confidentiality of data stored in a distributed storage network is a problem that needs to be solved. Therefore, an embodiment of the present application provides a data storage method that can effectively ensure the integrity and confidentiality of stored data.
[0115] The method of the embodiment of the present application can be applied to a data storage system. Refer to Figure 1, which is a structural diagram of the data storage system provided by the embodiment of the present application. The data storage system includes a first node 101 and a distributed storage network 102. Optionally, the data storage system also includes a first network 103. There is a wired and / or wireless communication connection between the first node 101, the distributed storage network 102 and the first network 103. Specifically, the distributed storage network 102 includes several nodes with wired and / or wireless communication connections. The distributed storage network 102 includes a second node 104 and P third nodes 105. The second node 104 is any node in the distributed storage network 102, that is, the second node 104 can be one of the P third nodes 105. By way of example, in Figure 1, the nodes in the distributed storage network 102 take base stations as an example, such as the second node 104 and the third node 105. As another example, the node in the distributed storage network 102 may also be a core network element 106. As an example, the core network element 106 includes at least one of the following: an access and mobility management function (AMF) network element, a session management function (SMF) network element, a user plane function (UPF) network element, a mobility management entity (MME) network element, a serving gateway (SGW), or a packet data network gateway (PDN gateway, PGW). The full name of PDN is Packet Data Network.
[0116] The first node 101 is a node with data storage requirements. The specific process of the first node 101 storing data in the distributed storage network 102 can be referred to the specific description of FIG2 , which will not be repeated here.
[0117] The following first introduces an exemplary data storage method provided by an embodiment of the present application.
[0118] FIG2 is a flow chart of a data storage method provided in an embodiment of the present application. The data storage method includes the following steps:
[0119] 201. A first node encrypts first data to obtain encrypted data.
[0120] Specifically, the first node is a node that has a data storage requirement. Optionally, the first node encrypts the first data using a first key to obtain encrypted data. The first key may be an asymmetric key or a symmetric key. Exemplarily, the first node may obtain the first key from another device. In another exemplary embodiment, the first node may randomly generate a symmetric key and then use the symmetric key to encrypt the first data.
[0121] The first node may include a terminal device or an access network device. The terminal device or access network device is not limited to a long term evolution (LTE) system, a long term evolution-advanced (LTE-A) system, an enhanced long term evolution (eLTE) system, a fifth generation (5G) mobile communication system, a new radio (NR) system, a sixth generation (6G) mobile communication system, and other future communication network systems. Terminal devices or access network devices.
[0122] Terminal equipment, also known as user equipment (UE), can include various handheld devices with wireless communication capabilities, vehicle-mounted devices, wearable devices, computing devices, or other processing devices connected to a wireless modem, as well as various forms of terminals, mobile stations (MS), terminals, soft terminals, access terminals, subscriber units, terminal stations, mobile stations, mobile stations (MS), remote stations, remote terminals, mobile devices, terminal agent, terminal devices, etc. For example, water meters, electricity meters, sensors, etc.
[0123] Access network equipment can have any of the following replacement terms: wireless access network, access network (AN), where the access network equipment can be a base station, a further evolved node B (gNB), an evolved node B (eNB), a transmission reception point (TRP), a centralized unit (CU) node, a distributed unit (DU) node, a transmission point (TP), a receiving point (RP), etc., without limitation. In some deployments of access network equipment, the CU node can also be divided into CU-control plane (CP) and CU-user plane (UP), etc. In other deployments of access network equipment, the access network equipment can also be an antenna unit (RU), etc. In some further deployments of access network equipment, the access network equipment can also be an open radio access network (ORAN) architecture, etc. For example, in the ORAN system, CU may also be referred to as open (O)-CU, DU may also be referred to as O-DU, CU-CP may also be referred to as O-CU-CP, CU-UP may also be referred to as O-CU-UP, and RU may also be referred to as O-RU.
[0124] When the first node is an access network device, the access network device receives the first data from the terminal. Exemplarily, the terminal may send a data storage request to the access network device, where the data storage request includes the first data. Thus, upon receiving the data storage request, the terminal may obtain the first data.
[0125] 202. A first node sends a data storage request to a second node, where the second node is any node in a distributed storage network, and the data storage request includes encrypted data.
[0126] Accordingly, referring to FIG3A , FIG3A is an interactive flow chart of a data storage method provided in an embodiment of the present application. The second node receives a data storage request sent by the first node. The second node splits or processes the encrypted data to obtain P shards. At least M of the P shards are used to recover the encrypted data. The second node determines P third nodes in the distributed storage network. The second node sends the P shards to the P third nodes respectively, and accordingly, each third node stores one shard.
[0127] Wherein, P and M are positive integers greater than one, and M is less than P.
[0128] Exemplarily, when the second node uses an information dispersal algorithm to split the encrypted data to obtain P fragments, the encrypted data can be recovered based on at least M fragments.
[0129] For example, the information dispersal algorithm uses RS erasure code technology to process the encrypted data to obtain P (P = M + N) fragments. The erasure code has two parameters M and N, denoted as RS (M, N), where M is the number of source data blocks and N is the number of check blocks. The M source data blocks form a vector D, which is multiplied by a generator matrix B to obtain a data vector, which consists of M data blocks and N check blocks. If a data block is lost, the lost data block can be restored through a series of calculations. RS (M, N) can tolerate the loss of up to N blocks (including data blocks and check blocks).
[0130] Take the erasure code with redundancy level M+N of 5+3 (i.e., M is 5, N is 3) as an example. Referring to FIG3B , FIG3B is a schematic diagram of the erasure code provided by the embodiment of the present application. M Arrange the vectors D by columns and then construct a (M+N)M matrix B, which is called the distribution matrix. Among them, any M row vectors of the matrix B are independent of each other, that is, the MM matrix composed of these M row vectors is reversible. Perform matrix-vector multiplication B*D to obtain N check blocks C1~C N and M source data blocks D1~D M The data vector composed of .
[0131] Assume that each data block in the above data vector is stored in a hard disk. When N hard disks out of (M+N) hard disks fail, that is, data blocks D1, D4, and C2 in FIG3B are lost, it is necessary to recover the source data blocks D1 to D4 from the remaining M data blocks. M . Refer to Figure 3C, which is a schematic diagram of the erasure code under fault conditions provided by an embodiment of the present application. The row vectors corresponding to the remaining data blocks are picked out from matrix B to form a new matrix B'. The result of multiplying B' by vector D is exactly the data vector composed of the data blocks without faults. Because the matrix composed of any M rows of B is invertible, the matrix B' has an inverse matrix, which is recorded as B ’-1 . Refer to FIG3D, which is a schematic diagram of the operation of the erasure code under the fault condition provided by the embodiment of the present application. Multiply the left and right sides of the equation in FIG3C by the matrix B at the same time. ’-1 , where, due to B ’-1 *B'=E, and the unit matrix E multiplied by any matrix is equal to the matrix; therefore, Figure 3E can be obtained from Figure 3D, which is a schematic diagram of erasure code recovery data provided by the embodiment of the present application. That is, according to the inverse matrix B ’-1 The data vector composed of the data blocks without faults can obtain M source data blocks D1~DM , complete data recovery.
[0132] As another example, when the second node uses other methods to split the encrypted data to obtain P fragments, the encrypted data can be recovered based on data of at least M fragments. Other methods, such as the LRC algorithm, are not particularly limited to this.
[0133] After the second node determines the P third nodes, it can obtain network identification information of the P third nodes. The network identification information is, for example, an identity document (ID) of the third node or an IP address of the third node.
[0134] Exemplarily, the second node can send P shards to the corresponding third node based on the network identification information of the third node. For example, the P shards are ten shards S1, S2, S3, S4, S5, S6, S7, S8, S9 and S10, and the P third nodes are ten third nodes J1, J2, J3, J4, J5, J6, J7, J8, J9 and J10. The second node sends shard S1 to the third node J1, sends shard S2 to the third node J2, and sends shard S3 to the third node J3, and so on.
[0135] As another example, the second node can send P shards and network identification information of P third nodes to any third node among the P third nodes, such as the third node J6. After the third node J6 obtains any one of the P shards, it sends the remaining (P-1) shards to any one of the (P-1) third nodes; repeat the previous step until all shards are sent.
[0136] 203. The first node receives feedback information. The feedback information indicates that P third nodes have stored P shards, which are obtained by splitting or processing the encrypted data by the second node. The P third nodes are nodes in the distributed storage network. P and M are positive integers greater than one, and M is less than P.
[0137] Exemplarily, the feedback information consists of P sub-feedback information, each of which indicates that the third node has stored one shard. Each of the P third nodes sends its own sub-feedback information to the first node, informing the first node that the P third nodes have stored P shards. Optionally, when sending the sub-feedback information, the third node may also send its own network identification information. Furthermore, optionally, each of the P third nodes sends its own sub-feedback information to the second node.
[0138] As another example, each of the P third nodes sends its own sub-feedback information to the second node. After the second node receives the sub-feedback information of each of the P third nodes, the second node sends a feedback information to the first node. The feedback information is used to indicate that the P third nodes have stored P shards.
[0139] In one possible embodiment, the second node transmits network identification information of P third nodes to the first node. For example, after receiving the P pieces of sub-feedback information, the second node transmits the network identification information of the P third nodes to the first node. In this way, the first node can send information to the P third nodes based on the network identification information of the P third nodes.
[0140] In this embodiment, the first node first encrypts the first data to obtain encrypted data, effectively ensuring data confidentiality. The second node then splits the encrypted data into P shards, which are then stored by P third nodes, with each third node storing one shard. The encrypted data can be recovered based on at least M of the P shards. Therefore, even if individual nodes in the P third nodes are attacked, the encrypted data can still be recovered based on at least M of the P shards, effectively ensuring the integrity of the stored data.
[0141] In addition, in the prior art, backing up data requires backing up the entire data, while in this application, each node only needs to store one shard to achieve backup, which can effectively save backup costs. Moreover, the solution of the embodiment of this application is a distributed storage solution, which can save communication costs.
[0142] In one possible embodiment, the data storage method further includes the second node calculating the data summary of each shard, i.e., the shard summary. Optionally, the second node sends the shard summary corresponding to the shard to each third node, so that the third node can verify the correctness of the received shard based on the shard summary and confirm whether it has not been tampered with. Optionally, the second node can send P shard summaries to each third node together, so that the third node can find its own corresponding shard summary, and then verify the correctness of the received shard based on the shard summary. Specifically, the third node calculates a shard summary based on the received shard, and then compares the shard summary with the above-mentioned corresponding shard summary to verify the correctness of the received shard. In one possible embodiment, the above-mentioned second node determines P third nodes in the distributed storage network, including:
[0143] The second node determines P third nodes according to the first value.
[0144] The first value includes any of the following: a data digest of the encrypted data, a first preset value, a first data digest, or a value stored on a DLT network. The first data digest is a data digest of the data stored on the DLT network. Exemplarily, the first preset value can be a constant, such as 1, 2, or 3. Furthermore, the first preset value can be an operation of any X items among the data digest of the encrypted data, the preset constant, the first data digest, and the value stored on the DLT network. The operation can be addition, subtraction, or other mathematical operation, without particular limitation. X is greater than or equal to two. The first value can also be a digest of other data or other values, as long as the other data or other values are known to each node in the distributed storage network.
[0145] Therefore, in an embodiment of the present application, P third nodes for storing data in a distributed storage network can be determined based on the first value, where P is a positive integer greater than one. P third nodes can be discovered using a small amount of data, and the method for determining P third nodes is very convenient and effective.
[0146] In one possible implementation, the ID of the i-th third node among the P third nodes is a hash value of a second value. The second value is the sum of the ID of the i-1-th second node and a second preset value, or the sum of the ID of the i-1-th second node and a preset function value, where the preset function value is related to i. The ID of the first third node among the P third nodes is the hash value of the first value.
[0147] The second preset value can be any constant, such as 1, 2, 3, etc. The preset function value can be the corresponding function value of any preset function related to i. The preset function can be a logarithmic function, an exponential function, a linear function, etc. For example, the preset function is log i, or the preset function is 2i+6, or the preset function is 2 i .
[0148] For example, the first value is the data digest HF of the encrypted data, and the second preset value is 1. Assuming P is 10, the IDs of the P third nodes are:
[0149] J1=hash(HF),
[0150] J2=hash(J1+1),
[0151] ……,
[0152] J10=hash(J9+1).
[0153] In a possible implementation, the data storage request further includes a data name of the first data, where the data name corresponds to a data digest of the encrypted data.
[0154] Thus, in an embodiment of the present application, the second node can obtain the data name of the first data from the data storage request, and the data name corresponds to the data digest of the encrypted data. In this way, the second node can obtain the data digest of the encrypted data corresponding to the data name (i.e., the data digest used for verification) based on the data name of the first data. The second node can verify the correctness of the encrypted data it has received based on the data digest of the encrypted data, that is, determine whether the encrypted data has been tampered with, and ensure the correctness of the stored data. Specifically, the second node can calculate the data digest of the received encrypted data, and then compare the calculated data digest with the data digest used for verification. When the former and the latter are the same, it can be confirmed that the encrypted data has not been tampered with.
[0155] In a possible implementation, referring to FIG3A , the data storage method further includes the following steps:
[0156] The first node sends a data registration request to the first network, where the data registration request includes a data name of the first data and a data digest of the encrypted data.
[0157] Thus, in this embodiment of the present application, the first node sends a data registration request to the first network. The first network can respond to the data registration request and store the data included in the request, completing the data registration. That is, the first network stores the correspondence between the data name of the first data and the data digest of the encrypted data. In this way, the second node can obtain the corresponding data digest of the encrypted data from the first network based on the data name of the first data.
[0158] Accordingly, in step 201 , the first node calculates a data digest of the encrypted data.
[0159] The first network mentioned above can be a DLT network or other storage network, which is not particularly limited here.
[0160] In a possible implementation, referring to FIG3A , the second node obtains the data digest of the encrypted data corresponding to the data name according to the data name, including:
[0161] The second node sends a summary request to the first network, where the summary request includes a data name.
[0162] The second node receives a data summary of the encrypted data corresponding to the data name sent by the first network.
[0163] In a possible implementation, the data registration request further includes first indication information, where the first indication information is used to determine P, and the first indication information corresponds to the data name.
[0164] Therefore, in an embodiment of the present application, the first network can also store the correspondence between the data name of the first data and the first indication information, so that the second node can obtain the corresponding first indication information according to the data name of the first data, and then determine P according to the first indication information, and then split the encrypted data according to P.
[0165] In an embodiment of the present application, the second node may request the first network to obtain the first indication information corresponding to the data name of the first data. For example, the first network may respond to the digest request and send the data digest of the encrypted data and the first indication information to the second node. For another example, the second node may also initiate another request to the first network to obtain the first indication information.
[0166] In a possible implementation, referring to FIG3A , the data storage method further includes the following steps:
[0167] The first node sends P subkeys to P third nodes. The P subkeys are obtained by splitting a first key used to encrypt the first data, wherein at least M subkeys of the P subkeys are used to recover the first key.
[0168] Specifically, after receiving the feedback information, the first node splits the first key to obtain P subkeys, and sends the P subkeys to the P third nodes. Exemplarily, when the first node processes the first key using a threshold key sharing algorithm to obtain P subkeys, assuming that the threshold value of the threshold key sharing algorithm is M, the first key can be recovered based on greater than or equal to M subkeys.
[0169] Accordingly, each of the P third nodes stores one of the P subkeys.
[0170] For example, when the threshold key sharing algorithm is used to process the first key, two processes are included: splitting the key and recovering the key. The threshold key sharing algorithm uses an M-1 degree polynomial function to hide the first key. The M-1 degree polynomial function is a(x) = a0 + a1x + a2x 2 +…+a a-1 x M-1 , where the first key is encoded as a constant a0.
[0171] Specifically, the split key is to randomly generate a1 to a M-1 These coefficient values. On this M-1 degree polynomial function curve, randomly select P different points {(x1,y1), (x2,y2),..., (x P ,y P)}, these points are assigned to P third nodes, and each coordinate obtained by the third node is a key shard (secret share), and the key shard is the subkey.
[0172] Recovering the key means aggregating the key fragments by M third nodes and substituting the M coordinate points into the original function to determine a unique curve. Based on the curve, {a0, a1, a2, ... a M-1}, where a0 is the first key.
[0173] In one instance, the first node may send P subkeys to P third nodes respectively based on the network identification information of the P third nodes. For example, the P subkeys are Z1, Z2, Z3, Z4, Z5, Z6, Z7, Z8, Z9 and Z10 respectively. The second node sends subkey Z1 to the third node J1, sends subkey Z2 to the third node J2, sends subkey Z3 to the third node J3, and so on.
[0174] In another example, the first node may first send P subkeys to the second node, and then the second node may send the P subkeys to P third nodes. The method by which the second node sends P subkeys to P third nodes may be the same as the method by which the second node sends P fragments, which will not be elaborated here.
[0175] Therefore, in this embodiment of the present application, when P third nodes have stored P shards, the first node sends P subkeys to the P third nodes, with each third node storing one subkey. If the subkey stored by an individual of the P third nodes is lost, the encrypted data cannot be decrypted using a single subkey. Therefore, the solution of this embodiment of the present application can effectively ensure the security of the first key and further protect the confidentiality of the stored data.
[0176] In a possible implementation, referring to FIG3A , the data storage method further includes the following steps:
[0177] The first node receives second indication information, where the second indication information is used to indicate that the P third nodes have stored the P subkeys.
[0178] The first node sends third indication information to the first network, where the third indication information is used to indicate that storage of the first data is complete, and the third indication information corresponds to a data name.
[0179] Therefore, in an embodiment of the present application, after the first node receives the second indication information, it sends the third indication information to the first network, so that the first network can update the storage status of the first data according to the third indication information. The data requester can learn from the first network that the first data has been stored and is in a requestable state.
[0180] Exemplarily, when the first node is an access network device, after receiving the second indication information, the access network device feeds back the storage result of the first data to the terminal that sends the data storage request.
[0181] In one example, the second indication message includes P sub-indication messages, each of which is used to indicate that the third node has stored a subkey. Each of the P third nodes sends its own sub-indication message to the first node, so that the first node knows that the P third nodes have stored the P subkeys.
[0182] In another example, each of the P third nodes sends its own sub-indication information to the second node. After the second node receives the sub-indication information from each of the P third nodes, the second node sends a second indication information to the first node. The second indication information is used to indicate that the P third nodes have stored P sub-keys.
[0183] The embodiment of the present application further provides a node determination method, which can be applied to a node determination device or a chip in a node determination device. For example, the node determination device can be any node in a distributed storage network.
[0184] Referring to Figure 4, Figure 4 is a flow chart of a node determination method provided in an embodiment of the present application. The node determination method includes the following steps:
[0185] 401. Split the first task into P subtasks.
[0186] Specifically, P is a positive integer greater than 1. The first task may include a storage task or a computing task, where the storage task is used to store data and the computing task is used to process data.
[0187] There are many ways to split the first task. For example, splitting the storage task into P subtasks can be understood as splitting the data to be stored into P shards, such as by using an information dispersal algorithm or other algorithm to split the data. The algorithm for splitting the data is not particularly limited. Another example is splitting the computing task into P sub-computing tasks. Similarly, the algorithm for splitting the computing tasks is not particularly limited.
[0188] 402. Determine P nodes in the distributed storage network according to the first value.
[0189] The first value includes any of the following: a digest of the data to be processed by the first task, a first preset value, a first data digest, or a value stored on a distributed ledger technology (DLT) network. The first data digest is a data digest of the data stored on the DLT network. For a detailed description of the first value, please refer to the description of the first value in the aforementioned data storage method and will not be repeated here.
[0190] In an embodiment of the present application, P nodes in the distributed storage network for executing P subtasks can be determined based on the first value, and the method for determining P nodes is very convenient and effective.
[0191] In one possible implementation, the ID of the i-th node among the P nodes is a hash value of the second value, where the second value is the sum of the ID of the i-1-th node and a second preset value, or the second value is the sum of the ID of the i-1-th node and a preset function value, where the preset function value is related to i. The ID of the first node among the P nodes is the hash value of the first value.
[0192] In a possible implementation, referring to FIG4 , the node determination method further includes:
[0193] 403. Send the P subtasks to the P nodes respectively.
[0194] Specifically, when the first task is a storage task, the node determination device sends P shards to P nodes, respectively, so that each node stores one shard. When the first task is a computation task, the node determination device sends P sub-computation tasks to P nodes, respectively, so that each node processes one sub-computation task. Specifically, the method by which the node determination device sends P shards or P sub-computation tasks can be referred to the description of the second node sending P shards above and will not be repeated here.
[0195] The device provided in the embodiments of the present application is described below.
[0196] Referring to Figure 5, Figure 5 is a schematic diagram of the structure of the first node provided in an embodiment of the present application. The first node 500 includes an encryption module 501, a sending module 502, and a receiving module 503.
[0197] The encryption module 501 is configured to encrypt the first data to obtain encrypted data.
[0198] The sending module 502 is configured to send a data storage request to a second node, where the second node is any node in the distributed storage network, and the data storage request includes encrypted data.
[0199] Receiving module 503 is configured to receive feedback information. The feedback information indicates that P third nodes have stored P shards, where the P shards are obtained by splitting the encrypted data. At least M of the P shards are used to recover the encrypted data. The P third nodes are nodes in the distributed storage network, where P and M are positive integers greater than one, and M is less than P.
[0200] For example, the encryption module 501 may be implemented by a processor, the sending module 502 may be implemented by a transmitter, and the receiving module 503 may be implemented by a receiver. Alternatively, the sending module 502 and the receiving module 503 may be combined into a transceiver.
[0201] In one possible embodiment, the sending module 502 is further configured to send P subkeys to the P third nodes. The P subkeys are obtained by splitting the first key used to encrypt the first data, wherein at least M subkeys of the P subkeys are used to recover the first key.
[0202] In a possible implementation, the data storage request further includes a data name of the first data, where the data name corresponds to a data digest of the encrypted data.
[0203] In a possible implementation, the sending module 502 is further configured to send a data registration request to the first network, where the data registration request includes the data name of the first data and a data digest of the encrypted data.
[0204] In a possible implementation, the data registration request further includes first indication information, where the first indication information is used to determine P, and the first indication information corresponds to the data name.
[0205] In one possible implementation, the receiving module 503 is further configured to receive second indication information. The second indication information is configured to indicate that the P third nodes have stored the P subkeys. The first node sends third indication information to the first network, the third indication information being configured to indicate that the first data storage is complete, and the third indication information corresponds to the data name.
[0206] For a detailed description of the first node, please refer to the description of the first node in the above data storage method, which will not be repeated here.
[0207] An embodiment of the present application also provides a second node, which is any node in the distributed storage network.
[0208] Refer to Figure 6, which is a schematic diagram of the structure of the second node provided in an embodiment of the present application. The second node 600 includes a receiving module 601, a splitting module 602, a determining module 603 and a sending module 604, wherein
[0209] The receiving module 601 is configured to receive a data storage request sent by a first node. The data storage request includes encrypted data, which is obtained by the first node encrypting the first data.
[0210] The splitting module 602 is configured to split the encrypted data into P fragments, wherein at least M fragments of the P fragments are used to recover the encrypted data, where P and M are positive integers greater than one, and M is less than P.
[0211] The determination module 603 is configured to determine P third nodes in the distributed storage network.
[0212] The sending module 604 is configured to send the P slices to P third nodes respectively.
[0213] Exemplarily, the receiving module 601 may be implemented by a receiver, the splitting module 602 and the determining module 603 may be implemented by a processor, and the sending module 604 may be implemented by a transmitter. Optionally, the sending module 604 and the receiving module 601 may also be combined into a transceiver.
[0214] In a possible implementation, the determining module 603 is specifically configured to:
[0215] The P third nodes are determined based on a first value. The first value includes any one of the following: a data digest of encrypted data, a first preset value, a first data digest, or a value stored on a distributed ledger technology (DLT) network, where the first data digest is a data digest of data stored on the DLT network.
[0216] In one possible implementation, the ID of the i-th third node among the P third nodes is a hash value of a second value. The second value is the sum of the ID of the i-1-th second node and a second preset value, or the sum of the ID of the i-1-th second node and a preset function value, where the preset function value is related to i. The ID of the first third node among the P third nodes is a hash value of the first value.
[0217] In a possible implementation, the data storage request further includes a data name of the first data, and the second node further includes:
[0218] The acquisition module is used to obtain the data summary of the encrypted data corresponding to the data name according to the data name.
[0219] The verification module is used to verify the correctness of the encrypted data based on the data digest.
[0220] In a possible implementation, the acquisition module is further configured to acquire first indication information corresponding to the data name according to the data name. The first indication information is used to determine P.
[0221] In a possible implementation, the acquisition module is specifically configured to:
[0222] A summary request is sent to the first network, wherein the summary request includes a data name.
[0223] A data digest of encrypted data corresponding to the data name sent by the first network in response to the digest request is received.
[0224] In a possible implementation, the acquisition module is specifically configured to:
[0225] First indication information corresponding to the data name is obtained from the first network according to the data name.
[0226] For a detailed description of the second node, please refer to the description of the second node in the above data storage method, which will not be repeated here.
[0227] The present application also provides a data storage system. Referring to FIG3A , the system includes a first node, a second node, and a third node.
[0228] The first node is configured to encrypt the first data to obtain encrypted data and send a data storage request to the second node, wherein the data storage request includes the encrypted data. The second node is any node in the distributed storage network.
[0229] The second node is configured to split the encrypted data into P shards. At least M of the P shards are used to recover the encrypted data. P and M are positive integers greater than one, and M is less than P. P third nodes are determined in the distributed network, and the P shards are sent to each of the P third nodes.
[0230] The third node is used to store one of the P shards.
[0231] For a detailed description of the data storage system, please refer to the description of the first node, the second node, and the third node in the above data storage method, which will not be repeated here.
[0232] The embodiment of the present application further provides a distributed storage network, which includes a second node and a third node. Referring to FIG3A , wherein:
[0233] The second node is configured to receive a data storage request sent by the first node. The second node is any node in the distributed storage network, the first node is a node with a data storage requirement, and the data storage request includes encrypted data, which is obtained by encrypting the first data by the first node.
[0234] The second node is further configured to split the encrypted data into P shards, wherein at least M shards among the P shards are used to recover the encrypted data.
[0235] The second node is further configured to determine P third nodes in the distributed network and send the P shards to the P third nodes respectively, where P and M are positive integers greater than one, and M is less than P. The P third nodes are nodes in the distributed storage network.
[0236] Each of the P third nodes is used to store one shard among the P shards.
[0237] For a detailed description of the distributed storage network, please refer to the description of the second node and the third node in the above data storage method, which will not be repeated here.
[0238] The present application also provides a node determination device. Referring to FIG7 , FIG7 is a schematic diagram of the structure of the node determination device provided in the present application. The node determination device 700 includes a splitting module 701 and a determination module 702, wherein:
[0239] The splitting module 701 is configured to split the first task into P subtasks, where P is a positive integer greater than one.
[0240] Determination module 702 is configured to determine P nodes in the distributed storage network based on a first value. The first value includes any one of the following: a digest corresponding to the data to be processed by the first task, a first preset value, a first data digest, or a value stored on a distributed ledger technology (DLT) network, where the first data digest is a data digest of the data stored on the DLT network.
[0241] In one possible embodiment, the ID of the i-th node among the P nodes is a hash value of a second value, where the second value is the sum of the ID of the i-1-th node and a preset value, or the sum of the ID of the i-1-th node and a preset function value, where the preset function value is related to i. The ID of the first node among the P nodes is a hash value of the first value.
[0242] In a possible embodiment, the node determination device 700 further includes:
[0243] The sending module is used to send P subtasks to P nodes respectively. Each node executes one subtask.
[0244] For a detailed description of the node determination device, please refer to the relevant description of the node determination method mentioned above, which will not be repeated here.
[0245] The present application also provides a communication device. Referring to Figure 8, Figure 8 is a schematic diagram of the structure of the communication device provided in the present application. The communication device 800 includes a memory 801, a processor 802, a communication interface 804, and a bus 803. The memory 801, processor 802, and communication interface 804 are interconnected via bus 803. There may be one or more memories 801 and one or more processors 802.
[0246] Exemplarily, the communication device 800 may be a chip or a chip system.
[0247] The memory 801 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 801 may store programs. When the program stored in the memory 801 is executed by the processor 802, the processor 802 is configured to execute the steps of the method described in any of the above embodiments.
[0248] The processor 802 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits to execute relevant programs to implement the method described in any of the above embodiments.
[0249] The processor 802 may also be an integrated circuit chip with signal processing capabilities. During implementation, the various steps of the method described in any embodiment of the present application may be completed by hardware integrated logic circuits or software instructions in the processor 802. The aforementioned processor 802 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method described in conjunction with any embodiment of the present application may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 801, and the processor 802 reads the information in the memory 801 and, in conjunction with its hardware, completes the method described in any of the above embodiments.
[0250] The communication interface 804 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the communication device 800 and other devices or a communication network. For example, the communication device 800 can obtain first data through the communication interface 804.
[0251] The bus 803 may include a path for transmitting information between various components of the communication device 800 (eg, the memory 801 , the processor 802 , and the communication interface 804 ).
[0252] It should be noted that although the communication device 800 shown in FIG8 only shows a memory, a processor, and a communication interface, during specific implementation, those skilled in the art will understand that the communication device 800 also includes other components necessary for normal operation. Furthermore, those skilled in the art will understand that, depending on specific needs, the communication device 800 may also include hardware components that implement other additional functions. Furthermore, those skilled in the art will understand that the communication device 800 may only include the components necessary to implement the embodiments of the present application, and does not necessarily need to include all of the components shown in FIG8 .
[0253] An embodiment of the present application provides a chip system, which is applied to a communication device. The chip system includes one or more processors, and the processors are used to call computer instructions to enable the communication device to execute the method described in any of the above embodiments.
[0254] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0255] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0256] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0257] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).
[0258] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A data storage method, characterized in that: The method comprises the following steps: The first node encrypts the first data to obtain encrypted data; The first node sends a data storage request to a second node, where the second node is any node in the distributed storage network, and the data storage request includes the encrypted data; The first node receives feedback information, and the feedback information is used to indicate that P third nodes have stored P shards, where the P shards are obtained by splitting the encrypted data, wherein at least M of the P shards are used to restore the encrypted data; the P third nodes are nodes in the distributed storage network, and the P and the M are positive integers greater than one, and the M is less than the P.
2. The method according to claim 1, characterized in that The method further comprises the following steps: The first node sends P subkeys to the P third nodes, where the P subkeys are obtained by splitting the first key used to encrypt the first data, wherein at least M subkeys of the P subkeys are used to recover the first key.
3. The method according to claim 1 or 2, characterized in that The data storage request further includes a data name of the first data, where the data name corresponds to a data digest of the encrypted data.
4. The method according to claim 3, characterized in that The method further comprises the following steps: The first node sends a data registration request to the first network, where the data registration request includes a data name of the first data and a data digest of the encrypted data.
5. The method according to claim 4, characterized in that The data registration request further includes first indication information, where the first indication information is used to determine P, and the first indication information corresponds to the data name.
6. The method according to claim 4 or 5, characterized in that The method further comprises the following steps: The first node receives second indication information, where the second indication information is used to indicate that the P third nodes have stored the P subkeys; The first node sends third indication information to the first network, where the third indication information is used to indicate that storage of the first data is complete, and the third indication information corresponds to the data name.
7. A data storage method, characterized in that: The method comprises the following steps: A second node receives a data storage request sent by the first node, where the second node is any node in the distributed storage network, and the data storage request includes encrypted data, where the encrypted data is obtained by encrypting the first data by the first node; The second node splits the encrypted data into P fragments, wherein at least M fragments of the P fragments are used to restore the encrypted data, where P and M are positive integers greater than one, and M is less than P; The second node determines P third nodes in the distributed storage network; The second node sends the P fragments to the P third nodes respectively.
8. The method according to claim 7, characterized in that The second node determines P third nodes in the distributed storage network, including: The second node determines the P third nodes based on a first value, where the first value includes any one of the following: a data summary of the encrypted data, a first preset value, a first data summary, or a value stored on a distributed ledger technology (DLT) network, and the first data summary is a data summary of data stored on the DLT network.
9. The method according to claim 8, characterized in that The identity ID of the i-th third node among the P third nodes is the hash operation value of the second value, where the second value is the sum of the ID of the i-1-th second node and a second preset value, or the second value is the sum of the ID of the i-1-th second node and a preset function value, where the preset function value is related to i; the ID of the first third node among the P third nodes is the hash operation value of the first value.
10. The method according to any one of claims 7 to 9, characterized in that The data storage request further includes a data name of the first data, and the method further includes the following steps: The second node obtains, according to the data name, a data digest of the encrypted data corresponding to the data name; The second node verifies the correctness of the encrypted data according to the data digest.
11. The method according to claim 10, characterized in that The method further comprises the following steps: The second node obtains first indication information corresponding to the data name according to the data name, where the first indication information is used to determine the P.
12. The method according to claim 10 or 11, characterized in that The second node obtains, according to the data name, a data digest of the encrypted data corresponding to the data name, including: The second node sends a summary request to the first network, where the summary request includes the data name; The second node receives a data digest of the encrypted data corresponding to the data name sent by the first network.
13. The method according to claim 12, characterized in that The second node obtains, according to the data name, first indication information corresponding to the data name, including: The second node obtains first indication information corresponding to the data name from the first network according to the data name.
14. A data storage method, characterized in that: The method is applied to a data storage system, the system comprising a first node, a second node and a third node; the method comprises the following steps: The first node encrypts the first data to obtain encrypted data; and sends a data storage request to a second node, the data storage request including the encrypted data, where the second node is any node in the distributed storage network; The second node splits the encrypted data into P shards, wherein at least M of the P shards are used to recover the encrypted data; determines P third nodes in the distributed network, and sends the P shards to the P third nodes respectively, where P and M are positive integers greater than one, and M is less than P; the P third nodes are nodes in the distributed storage network; Each of the P third nodes stores one shard among the P shards.
15. The method according to claim 14, characterized in that The method further comprises the following steps: The first node receives feedback information, where the feedback information is used to indicate that the P third nodes have stored P shards; The first node sends P subkeys to the P third nodes, where the P subkeys are obtained by splitting a first key used to encrypt the first data, wherein at least M subkeys of the P subkeys are used to recover the first key; Each of the P third nodes stores one subkey of the P subkeys.
16. The method according to claim 14 or 15, characterized in that The data storage request further includes a data name of the first data, and the method further includes the following steps: The second node obtains, according to the data name, a data digest of the encrypted data corresponding to the data name; The second node verifies the correctness of the encrypted data according to the data digest.
17. The method according to claim 16, characterized in that The method further comprises the following steps: The second node obtains first indication information corresponding to the data name according to the data name, where the first indication information is used to determine the P.
18. The method according to claim 17, characterized in that The data storage system further includes a first network, and the second node obtains a data summary of the encrypted data corresponding to the data name according to the data name, including: The second node sends a summary request to the first network, where the summary request includes the data name; The second node receives a data digest of the encrypted data corresponding to the data name sent by the first network.
19. The method according to claim 18, characterized in that The second node receives first indication information corresponding to the data name sent by the first network in response to the summary request.
20. The method according to claim 18 or 19, characterized in that The method further comprises the following steps: The first node sends a data registration request to the first network, where the data registration request includes a data name of the first data and a data digest of the encrypted data.
21. The method according to claim 20, characterized in that The data registration request also includes first indication information for determining the P.
22. The method according to any one of claims 14 to 21, characterized in that The second node determines P third nodes in the distributed storage network, including: The second node determines the P second nodes based on a first value, where the first value includes any one of the following: a data summary of the encrypted data, a first preset value, a first data summary, or a value stored on a distributed ledger technology (DLT) network, and the first data summary is a data summary of data stored on the DLT network.
23. The method according to claim 22, characterized in that The identity ID of the i-th third node among the P third nodes is the hash operation value of the second value, where the second value is the sum of the ID of the i-1-th third node and a second preset value, or the second value is the sum of the ID of the i-1-th third node and a preset function value, where the preset function value is related to i; the ID of the first third node among the P third nodes is the hash operation value of the first value.
24. A node determination method, characterized in that: The method comprises the following steps: Split the first task into P subtasks, where P is a positive integer greater than one; P nodes in the distributed storage network are determined based on a first value, where the first value includes any one of the following: a summary corresponding to the data to be processed by the first task, a first preset value, a first data summary, or a value stored on a distributed ledger technology (DLT) network, where the first data summary is a data summary of the data stored on the DLT network.
25. The method according to claim 24, characterized in that The identity ID of the i-th node among the P nodes is the hash operation value of the second value, where the second value is the sum of the ID of the i-1-th node and a second preset value, or the second value is the sum of the ID of the i-1-th node and a preset function value, where the preset function value is related to i; the ID of the first node among the P nodes is the hash operation value of the first value.
26. A first node, characterized in that: include: an encryption module, configured to encrypt the first data to obtain encrypted data; a sending module, configured to send a data storage request to a second node, where the second node is any node in the distributed storage network, and the data storage request includes the encrypted data; A receiving module is used to receive feedback information, where the feedback information is used to indicate that P third nodes have stored P shards, where the P shards are obtained by splitting the encrypted data, wherein at least M of the P shards are used to recover the encrypted data; the P third nodes are nodes in the distributed storage network, where P and M are positive integers greater than one, and M is less than P.
27. A second node, characterized in that: The second node is any node in the distributed storage network, and the second node includes: A receiving module, configured to receive a data storage request sent by a first node, wherein the data storage request includes encrypted data, and the encrypted data is obtained by encrypting the first data by the first node; a splitting module, configured to split the encrypted data into P fragments, wherein at least M fragments of the P fragments are used to recover the encrypted data, where P and M are positive integers greater than one, and M is less than P; A determination module, configured to determine P third nodes in the distributed storage network; A sending module is used to send the P slices to the P third nodes respectively.
28. A data storage system, characterized in that: The system includes a first node, a second node and a third node; The first node is configured to encrypt the first data to obtain encrypted data, and send a data storage request to a second node, the data storage request including the encrypted data, wherein the second node is any node in the distributed storage network; The second node is configured to split the encrypted data into P shards, wherein at least M shards of the P shards are used to recover the encrypted data, where P and M are positive integers greater than one, and M is less than P; determine P third nodes in the distributed storage network, and send the P shards to the P third nodes respectively; The third node is used to store one of the P shards.
29. A node determination device, characterized in that: The device comprises: a splitting module, configured to split the first task into P subtasks, where P is a positive integer greater than one; A determination module is configured to determine P nodes in a distributed storage network based on a first value, where the first value includes any one of the following: a summary corresponding to the data to be processed by the first task, a first preset value, a first data summary, or a value stored on a distributed ledger technology (DLT) network, where the first data summary is a data summary of the data stored on the DLT network.
30. A communication device, characterized in that: The communication device includes one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, and the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the communication device performs the method according to any one of claims 1 to 25.
31. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed by a processor, the method according to any one of claims 1 to 25 is implemented.
32. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 25 when the computer program is run on a processor.