File retrieval method, device and equipment based on decentralized storage network protocol and medium

By constructing a mapping relationship between file index and identifiers in a decentralized storage network, using Merkle tree and member evidence, and combining the single multi-storage node protocol to generate search vectors, the user privacy leakage problem is solved and safe and efficient file retrieval is achieved.

CN120386769AActive Publication Date: 2025-07-29INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD

Patent Information

Application Number
CN202510875766.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In decentralized storage networks, user privacy leaks are prominent, and the existing technology cannot effectively protect users' sensitive information such as video viewing habits and preferences.

Method used

By constructing a mapping relationship between file index and file identifier, using Merkle tree and member evidence to determine file index, combining the privacy information retrieval protocol of single storage nodes or multiple storage nodes, a file search vector is generated for file retrieval, avoiding direct sending of file identifiers, and ensuring that the storage node cannot reverse the file content or category.

Benefits of technology

It improves the security and flexibility of the decentralized network file retrieval process, protects user privacy, prevents sensitive information leakage, and enhances the efficiency and reliability of file retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386769A_ABST
    Figure CN120386769A_ABST
Patent Text Reader

Abstract

The invention discloses a file retrieval method and device based on a decentralized storage network protocol, equipment and a medium, and relates to the field of blockchains, and the method comprises the following steps: constructing a mapping relation between a file index and a corresponding file identifier; when file retrieval is carried out based on a preset privacy information retrieval protocol supporting a single storage node, the single storage node is used for obtaining a file retrieval vector so that the single storage node can construct a target database according to a mapping relation between a file index and a corresponding file identifier; determining a target retrieval result based on the target database and the file retrieval vector, and sending the target retrieval result to the user side; when file retrieval is carried out based on a preset privacy information retrieval protocol supporting multiple storage nodes, the multiple storage nodes are used for obtaining file retrieval vectors so that the multiple storage nodes can construct target databases corresponding to the multiple storage nodes respectively, and target retrieval results are determined based on the target databases and the file retrieval vectors; and sending the target retrieval result to the user side.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of blockchain, and particularly to a file retrieval method, device, equipment and medium based on a decentralized storage network protocol. Background Art

[0002] A decentralized storage network (DSN) is an innovative storage system that combines a peer-to-peer storage network with blockchain technology. The DSN aggregates the idle storage resources from storage providers (i.e., storage nodes) around the world to create a fault-tolerant decentralized network. In the DSN, a client uploads a file to a storage node and retrieves the file when needed. In recent years, as a key infrastructure for Web 3.0, the DSN has provided support for the storage requirements of applications such as non-fungible tokens (NFTs) and decentralized artificial intelligence (AI). However, privacy leakage remains a long-standing problem in the DSN, mainly involving data privacy and user privacy. Currently, research on user privacy in the DSN is still limited. In the current DSN, a user retrieves a file by sending a file identifier (FID) to a storage node. However, this method exposes the user's habits and preferences to the storage node, posing a risk of sensitive information leakage. For example, if the DSN is used as the basic storage facility for a media transmission system such as Douyin, sensitive information such as the user's video viewing habits and interests may be leaked to the storage node when the user retrieves a short video file. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a file retrieval method, device, equipment and medium based on a decentralized storage network protocol, which can improve the DSN's ability to protect user privacy and effectively enhance the security and resilience of the file retrieval process in the decentralized network. The specific solutions are as follows:

[0004] In a first aspect, this application provides a file retrieval method based on a decentralized storage network protocol, including:

[0005] Determine a target vector corresponding to a file identifier saved in a vector set; wherein, the vector set includes each file vector, and each file vector corresponds to a Merkle tree; the Merkle tree is used to save the file identifier corresponding to the target file that has been stored;

[0006] Using the membership evidence corresponding to the file identifier, and starting from the next vector of the target vector in the vector set until the entire vector set is traversed, to determine the file index corresponding to the file identifier based on a preset index determination process, and construct a mapping relationship between the file index and the file identifier; the membership evidence is used to represent the path of the file identifier in the corresponding Merkle tree;

[0007] When performing file retrieval based on a preset private information retrieval protocol that supports a single storage node, use a single storage node to obtain the file retrieval vector sent by the client, so that the single storage node constructs a target database according to the mapping relationship between the file index and the corresponding file identifier, and determines the target retrieval file corresponding to the file retrieval vector from the target database, determines the target retrieval result based on the target retrieval file, and sends the target retrieval result to the client; the file retrieval vector is a vector determined based on the file index;

[0008] When performing file retrieval based on a preset private information retrieval protocol that supports multiple storage nodes, use multiple storage nodes to obtain the file retrieval vector sent by the client, so that the multiple storage nodes respectively construct their own corresponding target databases according to the mapping relationship between the file index and the corresponding file identifier, and determine the target retrieval file corresponding to the file retrieval vector from the target database, determine the target retrieval result based on the target retrieval file, and send the target retrieval result to the client.

[0009] Optionally, the using the membership evidence corresponding to the file identifier, and starting from the next vector of the target vector in the vector set until the entire vector set is traversed, to determine the file index corresponding to the file identifier based on a preset index determination process, includes:

[0010] Initialize the target counter to 1, and start traversing from the next vector of the target vector in the vector set;

[0011] If the currently traversed vector is not empty, determine the vector number corresponding to the currently traversed vector, and increment the current target counter based on the vector number until the entire vector set is traversed;

[0012] Determine the membership evidence corresponding to the file identifier, traverse the position information corresponding to each target element in the membership evidence, if the current position information indicates left, increment the current target counter based on the element number corresponding to the current position information until all target elements in the membership evidence are traversed, and use the value corresponding to the current target counter as the file index corresponding to the file identifier.

[0013] Optionally, after constructing the mapping relationship between the file index and the file identifier, the method further includes:

[0014] Using the storage node storing the target file corresponding to the file identifier to construct a file upload proof based on the vector set, the target vector, the membership evidence, the file identifier, the file index, and the status hash value corresponding to the file identifier, and using the storage node to submit the file upload proof to the blockchain so that the blockchain determines whether the file upload proof passes the verification; the status hash value is the previous status hash value generated by the storage node during the last file storage or file deletion;

[0015] Correspondingly, using the blockchain to determine whether the file upload proof passes the verification includes:

[0016] Using the blockchain to determine whether the file identifier exists in the Merkle tree corresponding to the target vector and whether the mapping relationship between the file index and the file identifier is correct based on the vector set, the target vector, the membership evidence, and the file identifier in the file upload proof, so as to obtain a first verification result;

[0017] Using the blockchain to determine a first target file index based on the vector set, the target vector, and the membership evidence and according to the preset index determination process, and to determine whether the first target file index matches the file index, so as to obtain a second verification result;

[0018] Using the blockchain to compare the current preset cryptographic accumulator state with the previous preset cryptographic accumulator state based on the vector set, the membership evidence, and the status hash value, so as to verify whether the state can be transferred from the previous preset cryptographic accumulator state to the current preset cryptographic accumulator state, so as to obtain a third verification result; the preset cryptographic accumulator includes a vector set and a file identifier and is used to construct a mapping relationship between a file index and the corresponding file identifier;

[0019] If the first verification result, the second verification result, and the third verification result all indicate passing, the file upload proof passes the verification, and the file upload proof is saved to the blockchain.

[0020] Optionally, the file retrieval method based on the decentralized storage network protocol further includes:

[0021] In the file deletion stage, delete the file identifier corresponding to the target deleted file from the preset cryptographic accumulator, generate a corresponding file deletion proof for the file identifier corresponding to the target deleted file, and upload the file deletion proof to the blockchain so that the blockchain can verify whether the file deletion proof passes the verification;

[0022] If the file deletion proof passes the verification, save the file deletion proof to the blockchain.

[0023] Optionally, when using the preset private information retrieval protocol that supports a single storage node, the process of file storage and file deletion by the storage node is the same as the process of file storage and file deletion by the storage node when using the preset private information retrieval protocol that supports multiple storage nodes.

[0024] Optionally, the file retrieval method based on the decentralized storage network protocol further includes:

[0025] Obtain the target file upload proof corresponding to the file to be retrieved from the blockchain to determine the file index of the file identifier corresponding to the file to be retrieved in the corresponding target storage node from the target file upload proof;

[0026] Determine the database size corresponding to the target storage node based on the maximum file index of the file identifiers stored in the target storage node;

[0027] Generate a file retrieval vector based on the file index of the file identifier corresponding to the file to be retrieved in the target storage node and the database size corresponding to the target storage node so that the client can send the file retrieval vector to the target storage node;

[0028] Correspondingly, when the target storage node receives the file retrieval vector, it further includes:

[0029] Construct a target database, save the second target file index corresponding to the file identifier saved by itself to the target database, and save the file corresponding to the second target file index to the target database; wherein, if the file corresponding to the second target file index is deleted, save a blank file to the target database.

[0030] Optionally, when performing file retrieval based on the preset private information retrieval protocol that supports multiple storage nodes, it further includes:

[0031] Control the honest nodes in the storage node to maintain a consistent database state based on the Byzantine fault-tolerant state machine replication protocol;

[0032] Utilize the Byzantine-robust private information retrieval protocol and perform file retrieval based on the preset private information retrieval protocol that supports multiple storage nodes.

[0033] In a second aspect, the present application provides a file retrieval device based on a decentralized storage network protocol, including:

[0034] A vector determination module, configured to determine a target vector corresponding to a file identifier saved in a vector set; wherein, the vector set includes each file vector, and each file vector corresponds to a Merkle tree; the Merkle tree is used to save the file identifier corresponding to the target file that has been stored;

[0035] An index determination module, configured to use the membership proof corresponding to the file identifier and traverse from the next vector of the target vector in the vector set until the vector set is traversed, so as to determine the file index corresponding to the file identifier based on a preset index determination process, and construct a mapping relationship between the file index and the file identifier; the membership proof is used to represent the path of the file identifier in the corresponding Merkle tree;

[0036] A first result sending module, configured to, when performing file retrieval based on a preset private information retrieval protocol that supports a single storage node, use a single storage node to obtain a file retrieval vector sent by a user terminal, so that the single storage node constructs a target database according to the mapping relationship between the file index and the corresponding file identifier, determine a target retrieval file corresponding to the file retrieval vector from the target database, determine a target retrieval result based on the target retrieval file, and send the target retrieval result to the user terminal; the file retrieval vector is a vector determined based on the file index;

[0037] A second result sending module, configured to, when performing file retrieval based on a preset private information retrieval protocol that supports multiple storage nodes, use multiple storage nodes to obtain a file retrieval vector sent by a user terminal, so that the multiple storage nodes respectively construct their own corresponding target databases according to the mapping relationship between the file index and the corresponding file identifier, determine a target retrieval file corresponding to the file retrieval vector from the target database, determine a target retrieval result based on the target retrieval file, and send the target retrieval result to the user terminal.

[0038] In a third aspect, the present application provides an electronic device, including:

[0039] A memory, configured to save a computer program;

[0040] A processor, configured to execute the computer program to implement the foregoing file retrieval method based on a decentralized storage network protocol.

[0041] Fourthly, the present application provides a computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, the foregoing file retrieval method based on a decentralized storage network protocol is implemented.

[0042] In the present application, a target vector corresponding to a file identifier stored in a vector set is determined; wherein the vector set includes each file vector, and each file vector corresponds to a Merkle tree; the Merkle tree is used to store the file identifier corresponding to the stored target file; by using the membership proof corresponding to the file identifier and traversing from the next vector of the target vector in the vector set until the vector set is traversed, a file index corresponding to the file identifier is determined based on a preset index determination process, and a mapping relationship between the file index and the file identifier is constructed; the membership proof is used to represent the path of the file identifier in the corresponding Merkle tree; when file retrieval is performed based on a preset privacy information retrieval protocol supporting a single storage node, a single storage node is used to obtain a file retrieval vector sent by a client, so that the single storage node constructs a target database according to the mapping relationship between the file index and the corresponding file identifier, determines a target retrieval file corresponding to the file retrieval vector from the target database, determines a target retrieval result based on the target retrieval file, and sends the target retrieval result to the client; the file retrieval vector is a vector determined based on the file index; when file retrieval is performed based on a preset privacy information retrieval protocol supporting multiple storage nodes, multiple storage nodes are used to obtain a file retrieval vector sent by a client, so that the multiple storage nodes respectively construct their own corresponding target databases according to the mapping relationship between the file index and the corresponding file identifier, determine a target retrieval file corresponding to the file retrieval vector from the target databases, determine a target retrieval result based on the target retrieval file, and send the target retrieval result to the client. As can be seen from the above, in the present application, the client does not directly send the FID, but generates a file retrieval vector through the file index and sends the file retrieval vector to the storage node, so that the storage node cannot reverse-infer the file content or category corresponding to the FID from the file index, ensuring user privacy and effectively improving the security of the decentralized network file retrieval process. Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0044] Figure 1 Flowchart of a file retrieval method based on a decentralized storage network protocol disclosed in this application;

[0045] Figure 2 Schematic diagram of a specific pre - set privacy information retrieval protocol architecture supporting a single storage node disclosed in this application;

[0046] Figure 3 Schematic diagram of a specific system implementation framework disclosed in this application;

[0047] Figure 4 Schematic diagram of a specific test result of file upload time overhead disclosed in this application;

[0048] Figure 5 Schematic diagram of a specific test result of file deletion time overhead disclosed in this application;

[0049] Figure 6 Schematic diagram of a specific test result of proof verification time overhead disclosed in this application;

[0050] Figure 7 Schematic diagram of a specific test result of file retrieval time overhead disclosed in this application;

[0051] Figure 8 (a) Schematic diagram of a specific test result of file upload throughput disclosed in this application;

[0052] Figure 8 (b) Schematic diagram of a specific test result of file upload latency disclosed in this application;

[0053] Figure 9 (a) Schematic diagram of a specific test result of file deletion throughput disclosed in this application;

[0054] Figure 9 (b) Schematic diagram of a specific test result of file deletion latency disclosed in this application;

[0055] Figure 10 (a) Schematic diagram of a specific test result of file retrieval throughput disclosed in this application;

[0056] Figure 10 (b) Schematic diagram of a specific test result of file retrieval latency disclosed in this application;

[0057] Figure 11 Schematic diagram of the structure of a file retrieval device based on a decentralized storage network protocol disclosed in this application;

[0058] Figure 12 Schematic diagram of the structure of an electronic device disclosed in this application. Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0060] Currently, the research on DSN in terms of user privacy is still limited. In the current DSN, users retrieve files by sending file identifiers to storage nodes. However, this method will expose the user's habits and preferences to the storage nodes, posing a risk of sensitive information leakage. For this reason, the present application provides a file retrieval method based on a decentralized storage network protocol, which can improve the ability of DSN in user privacy protection and effectively improve the security and elasticity of the decentralized network file retrieval process.

[0061] See Figure 1 As shown, the embodiments of the present application disclose a file retrieval method based on a decentralized storage network protocol, including:

[0062] Step S11: Determine a target vector corresponding to the file identifier stored in the vector set; wherein, the vector set includes each file vector, and each file vector corresponds to a Merkle tree; the Merkle tree is used to store the file identifier corresponding to the target file that has been stored.

[0063] In this embodiment, the storage node can perform file storage (i.e., file upload), file deletion on the file data sent by the user end, or file retrieval based on the file retrieval vector sent by the user end; wherein, the storage node can perform file retrieval based on a preset privacy information retrieval protocol supporting a single storage node, or based on a preset privacy information retrieval protocol supporting multiple storage nodes. The above-mentioned preset privacy information retrieval protocol supporting a single storage node is applicable to any scenario of retrieving files from a single storage node, and the above-mentioned preset privacy information retrieval protocol supporting multiple storage nodes is applicable to the scenario where the client stores redundant file copies in multiple storage nodes to enhance storage elasticity. In this scenario, the preset privacy information retrieval protocol supporting multiple storage nodes has a higher retrieval efficiency than the preset privacy information retrieval protocol supporting a single storage node. These two protocols together cover most of the DSN file retrieval requirements.

[0064] It should be noted that the processes of file storage and file deletion by the storage node when using the preset private information retrieval protocol supporting a single storage node are the same as those of the storage node when using the preset private information retrieval protocol supporting multiple storage nodes.

[0065] When the storage node performs file storage, the file identifier can be mapped to a compact index range to build the mapping relationship between the file index and the file identifier. First, it is necessary to determine the target vector corresponding to the file identifier of the file to be stored in the vector set, that is , indicating that the FID is inserted into the Merkle tree rooted at ; among them, the vector set can be expressed as .

[0066] Step S12: Utilize the membership evidence corresponding to the file identifier, and start traversing from the next vector of the target vector in the vector set until the entire vector set is traversed, so as to determine the file index corresponding to the file identifier based on the preset index determination process, and build the mapping relationship between the file index and the file identifier; the membership evidence is used to represent the path of the file identifier in the corresponding Merkle tree.

[0067] In this embodiment, first, the target counter can be initialized to 1, and start traversing from the next vector of the target vector in the vector set; if the currently traversed vector is not empty, determine the vector number corresponding to the currently traversed vector, and increase the current target counter based on the vector number until the entire vector set is traversed; then determine the membership evidence corresponding to the file identifier, traverse the position information corresponding to each target element in the membership evidence, if the current position information is characterized as left, increase the current target counter based on the element number corresponding to the current position information until all the target elements in the membership evidence are traversed, and use the value corresponding to the current target counter as the file index corresponding to the file identifier; finally, build the mapping relationship between the file index and the corresponding file identifier.

[0068] After building the mapping relationship between the file index and the file identifier, it may further include: using the storage node storing the target file corresponding to the file identifier to build a file upload proof based on the vector set, target vector, membership evidence, file identifier, file index, and status hash value corresponding to the file identifier, and using the storage node to submit the file upload proof to the blockchain, so that the blockchain can determine whether the file upload proof passes the verification; where the status hash value is the previous status hash value generated by the storage node during the previous file storage or file deletion.

[0069] Correspondingly, using the blockchain to determine whether a file upload proof passes verification may include: First, using the blockchain to determine whether the file identifier exists in the Merkle tree corresponding to the target vector and whether the mapping relationship between the file index and the file identifier is correct based on the vector set, target vector, membership evidence, and file identifier in the file upload proof, so as to obtain a first verification result; Then, using the blockchain to determine a first target file index based on the vector set, target vector, and membership evidence and according to a preset index determination process, so as to judge whether the first target file index matches the file index, so as to obtain a second verification result; Finally, using the blockchain to compare the current preset cryptographic accumulator state with the previous preset cryptographic accumulator state based on the vector set, membership evidence, and state hash value, so as to verify whether the state can be transferred from the previous preset cryptographic accumulator state to the current preset cryptographic accumulator state, so as to obtain a third verification result; Wherein, the preset cryptographic accumulator includes a vector set and a file identifier, and is used to construct a mapping relationship between the file index and the corresponding file identifier; If the first verification result, the second verification result, and the third verification result all indicate passing, the file upload proof passes verification, and the file upload proof is saved to the blockchain.

[0070] It can be understood that only when all the above three verification results pass, the blockchain will accept the file upload proof. By verifying the file upload proof, on the one hand, it can ensure that each FID corresponds to a unique and correct index, providing a reliable mapping basis for subsequent privacy retrieval. For example, if a storage node maliciously inserts FID2 into the index position occupied by FID1, the verification will fail due to index conflict, thereby preventing FID2 from being inserted into the index position occupied by FID1 and preventing data chaos; On the other hand, the file upload proof verification process is public to all storage nodes, and any participant can query and verify the legality of the file upload operation through the blockchain, establishing decentralized trust and avoiding the risk of "single-point trust" in traditional centralized systems, ensuring that storage nodes cannot forge or tamper with the mapping relationship; At the same time, when the FID index is incorrect, the retrieval vector generated by the user side will point to the wrong position, resulting in retrieval failure or privacy exposure. At this time, the verification mechanism can ensure that the index is strictly bound to the FID. For example, the retrieval vector generated by the user side can accurately hit the target file, preventing the storage node from misleading the user through the wrong index.

[0071] When a storage node deletes a file, it can delete the file identifier corresponding to the target deleted file from the preset cryptographic accumulator, then generate a corresponding file deletion proof for the file identifier corresponding to the target deleted file, and upload the file deletion proof to the blockchain so that the blockchain can verify whether the file deletion proof passes verification; If the file deletion proof passes verification, the file deletion proof is saved to the blockchain.

[0072] By verifying the file deletion proof, on the one hand, only legitimate deletion requests initiated by the user or the authorized party will be accepted, preventing storage nodes from privately deleting valid files; on the other hand, after verification, the preset cryptographic accumulator removes the deleted FIDs, releases the occupied index space, and allows subsequent FIDs to reuse this space, ensuring that the preset cryptographic accumulator always reflects the true file storage status and avoiding space waste or index confusion caused by "zombie indexes".

[0073] Taking the preset privacy information retrieval protocol supporting a single storage node as an example, the processes of file storage and file deletion are elaborated below.

[0074] In the file storage stage, the mapping relationship between file indexes and corresponding file identifiers can be established through a preset Asynchronous Cryptographic Accumulator (ACA). The ACA includes a vector set and n FIDs (structured as a Merkle forest); among them, each element can be either the root of a Merkle tree or a placeholder . If , then its leaf nodes either store FIDs or placeholders ; if , it means that all its leaf nodes store . Through the ACA, n FIDs can be mapped to this compact index range. In addition, the ACA can be used to provide membership witnesses to generate proofs to confirm the validity of the mapping relationship between FIDs and the corresponding indexes, and record them on a bulletin board, such as a blockchain, to achieve public verification of the correctness of the mapping relationship.

[0075] See Figure 2 shown, the initial state of the ACA vector set is , where . When inserting an FID, the following two situations may occur: 1. All leaf nodes in the ACA have been occupied by FIDs. In this case, the ACA needs to be extended to accommodate additional FIDs; 2. At least one leaf node stores , indicating the existence of an empty space. In this case, the new FID will be inserted into the first available empty space. For this purpose, the ACA vector set can be traversed from to find , and this The element satisfies or and The leaf nodes of the associated Merkle tree contain . Specifically, if no that meets the conditions can be found, it means the ACA is full. At this time, a new Merkle tree can be created to provide additional index space. The new FID will be merged with the to leaf nodes of the Merkle tree to form a complete Merkle tree containing leaf nodes. The root of this tree is stored in , and the new FID is inserted into the rightmost leaf node of this tree. Subsequently, to to values and all merged leaf nodes are set to . If exists, there are two possible cases: 1. , in which case the FID is inserted into the leftmost leaf node of the Merkle tree associated with ; 2. , in which case, the leaf nodes of the associated Merkle tree can be traversed using depth - first search (DFS, Depth First Search), and the FID is inserted into the first leaf node that stores .

[0076] After the FID is correctly inserted, update the ACA and perform the index allocation process. Assume the FID is inserted into the Merkle tree with as the root. The index corresponding to the FID can be determined by the following three parameters: , and the membership proof of the FID . The membership proof represents the Merkle path of the FID, where each element consists of a hash value and position information (left or right). The index calculation method is as follows: First, initialize the counter index = 1; Second, traverse the elements in the ACA vector set from to . If , then index is incremented by ; Finally, traverse each element of the membership proof . If the position information of a certain is left, then index is incremented by .

[0077] After the mapping is established, the storage node that inserts the FID can submit a proof to the blockchain to verify the correctness of the mapping. The proof can be expressed as: Where h represents the hash value of the previous state proof generated by the storage node during the last file storage or file deletion operation. The following three steps are required: First, use and FID to ensure that the FID exists in the ACA. Specifically, it is necessary to verify Is it The kth element in , and whether it can be obtained by and FID are calculated. Secondly, use Confirm the correctness of the index assignment by rerunning the index calculation process and ensuring that the calculation result matches the index. , Compare the current ACA status with h Compared to the previous ACA status (Use h to obtain from the blockchain). By treating ACA as a state machine, it is possible to efficiently verify whether the state can be obtained from Transfer to , ensuring that newly inserted FIDs do not cause index conflicts.

[0078] During the file deletion phase, see Figure 2 As shown, the FID is deleted from the ACA, the ACA is updated and a certificate is generated for each deleted FID, and the certificate is recorded on the blockchain. This method ensures that the storage nodes in the DSN can publicly verify the integrity of the mapping deletion. Specifically, to delete a FID, first locate the leaf node containing the FID and the Merkle tree to which it belongs. , the file deletion process depends on Whether other FIDs are included can be divided into the following two situations: 1. Still contains other FIDs. At this time, the membership evidence of FID can be generated , and update the value stored in the leaf node where FID is located to . Then, recalculate the root of the Merkle tree and store the updated root value in ; 2. Does not include other FIDs. Membership evidence of FID is also generated at this time (at this time ) and updates the value stored in the leaf node where FID is located to At this time, due to All leaf nodes are , The value of is also set to .

[0079] After the mapping deletion, the storage node that deletes the FID can submit a proof to the blockchain to verify the correctness of the mapping deletion process. This proof can be expressed as: . Verification The following two steps need to be completed: First, use and h to confirm that the FID existed in the previous state of the ACA before deletion, and the mappings of other FIDs were not affected by the deletion operation. Specifically, the previous ACA state corresponding to h can be retrieved from the blockchain , verify whether it can be calculated through and the FID, so as to verify the existence of the FID in the previous ACA state and ensure that the other FID mappings remain unchanged. At the same time, check and whether the other elements except the k-th element are exactly the same to confirm that the FID mappings in other Merkle trees have not been modified; Second, use and the FID to verify whether the mapping deletion operation is correctly executed. Specifically, it can be checked whether is equal to . If , it indicates that still contained other FIDs before deletion. At this time, verify whether is the k-th element in , and whether it can be recalculated through and . If , it means that only contained this FID before deletion. In this case, can be further checked to ensure the correctness of the deletion operation.

[0080] Step S13: When retrieving a file based on a preset privacy information retrieval protocol that supports a single storage node, use a single storage node to obtain a file retrieval vector sent by the client, so that the single storage node constructs a target database according to the mapping relationship between the file index and the corresponding file identifier, determines a target retrieval file corresponding to the file retrieval vector from the target database, determines a target retrieval result based on the target retrieval file, and sends the target retrieval result to the client; the file retrieval vector is a vector determined based on the file index.

[0081] In this embodiment, the preset privacy information retrieval protocol supporting a single storage node allows the client to retrieve files from a single storage node through the Single-Server Private Information Retrieval protocol. Compared with the file retrieval technology in the existing DSN, the method in this embodiment enhances user privacy during the retrieval process. Specifically, this embodiment can determine the database size corresponding to the target storage node based on the maximum file index of the file identifier stored in the target storage node; then generate a file retrieval vector according to the file index of the file identifier corresponding to the file to be retrieved in the target storage node and the database size corresponding to the target storage node, so that the client can send the file retrieval vector to the target storage node.

[0082] It should be noted that when a user queries a file, a file retrieval vector needs to be generated and sent to the storage node. To generate a file retrieval vector for a specific FID, the following two key parameters are required: 1. The index of the FID in the storage node, which is used to determine the location of the retrieved file; 2. The size of the database held by the storage node, which is used to determine the dimension (i.e., size) of the file retrieval vector. Both of these parameters can be obtained from the blockchain. Among them, the index corresponding to the FID is stored in the proof submitted by the storage node when inserting the FID, and the database size is determined by the maximum index value when the storage node stores the FID, and this value can also be obtained from the proof submitted by the storage node. After obtaining these parameters, the client can use the Query algorithm of single-server private information retrieval to construct a file retrieval vector, send it to the storage node, and wait for the storage node to respond. If the client retrieves the FID deletion proof submitted by the storage node on the blockchain, it will look for other storage nodes storing the file.

[0083] For example Figure 2 As shown, the client obtains the index of FID1 and the size of the retrieval vector through the on-chain proof to determine the retrieval vector q, sends the retrieval vector q to the storage node, the storage node calculates the retrieval result ans through single-server private information retrieval technology, and returns the retrieval result ans to the client, and the client reconstructs the required target file using the retrieval result ans.

[0084] Correspondingly, when the target storage node receives the file retrieval vector, it may further include: constructing a target database, saving the second target file index corresponding to the file identifier stored by itself into the target database, and saving the file corresponding to the second target file index into the target database; wherein, if the file corresponding to the second target file index is deleted, a blank file is saved into the target database.

[0085] It should be noted that the storage node indexes according to the FID and places the corresponding file in the corresponding entry of the target database. Due to the existence of file deletion operations, some entries in the target database may be empty. For these empty entries, the storage node can fill them with blank files. After constructing the target database, the storage node can use the Answer algorithm of single-server private information retrieval to determine the target retrieval file according to the target database and the file retrieval vector, determine the target retrieval result based on the target retrieval file, and send the target retrieval result to the client. After receiving the target retrieval result, the client can use the Decrypt algorithm of single-server private information retrieval to obtain the corresponding file and verify the file integrity using the FID to check whether the storage node provides an incorrect query result.

[0086] Step S14, when retrieving a file based on a preset privacy information retrieval protocol that supports multiple storage nodes, use multiple storage nodes to obtain the file retrieval vector sent by the client, so that the multiple storage nodes respectively construct their corresponding target databases according to the mapping relationship between the file index and the corresponding file identifier, determine the target retrieval file corresponding to the file retrieval vector from the target database, determine the target retrieval result based on the target retrieval file, and send the target retrieval result to the client.

[0087] In this embodiment, the preset privacy information retrieval protocol that supports multiple storage nodes allows the client to retrieve files from multiple storage nodes through the multi-server private information retrieval (MPIR-DSN; MPIR, i.e., Multi-Server Private Information Retrieval) protocol.

[0088] In DSN, files are usually redundantly stored on multiple storage nodes to enhance security, availability, and fault tolerance. Although users can use the SPIR-DSN protocol (i.e., single-server private information retrieval protocol) to retrieve files from a single storage node, the multi-storage-node file privacy retrieval protocol further extends this function, enabling the client to retrieve files from any subset of the set of storage nodes where the files are stored. Different from single-server private information retrieval which mainly relies on computationally intensive homomorphic multiplication, the multi-server private information retrieval uses more efficient computations, such as bitwise exclusive OR operations and simple multiplication operations, thus enabling better retrieval efficiency.

[0089] It should be noted that the multi-server private information retrieval protocol runs in the following environment: Suppose there is a subnet composed of N storage nodes for processing multi-server private file retrieval, where at most N / 3 storage nodes may be malicious. All storage nodes in the subnet simultaneously process the user's file storage, deletion, and retrieval requests.

[0090] The main problem to be solved in the multi-server private information retrieval protocol is how to ensure that all honest nodes maintain a consistent database state after any file is stored or deleted, and a consistent database state is a prerequisite for correctly retrieving files using the multi-server private information retrieval protocol. Among them, the honest nodes in the storage nodes can be controlled to maintain a consistent database state based on the Byzantine fault-tolerant state machine replication protocol, and the Byzantine-robust private information retrieval protocol can be used to retrieve files based on the multi-server private information retrieval protocol.

[0091] For example, the consistency of the database of storage nodes within a subnet can be achieved based on the BFT-SMR protocol (Byzantine Fault-Tolerant State Machine Replication). BFT-SMR ensures that all honest nodes can reach a consensus in the order of processing user requests even in the presence of Byzantine nodes (i.e., malicious nodes). In the multi-server private information retrieval protocol, the Hotstuff protocol can be adopted. The Hotstuff protocol is widely used for efficient consensus in a partially synchronous environment. In addition, the multi-server private information retrieval protocol is not limited to Hotstuff, and other BFT-SMR protocols such as Tendermint or PBFT (Practical Byzantine Fault Tolerance) can also be selected according to actual needs to achieve the above-mentioned database state replication.

[0092] Specifically, broadcast the file storage request or file deletion request to all storage nodes within the subnet. Each storage node processes these requests locally and conducts consensus according to the following steps to ensure the consistency of the user request order: First, the storage nodes within the subnet take turns to act as the primary node. The primary node sorts the user requests received from the client, updates its local ACA, and generates corresponding proofs. These proofs are packaged into a block and broadcast to other storage nodes in the subnet. After receiving the block sent by the primary node, other storage nodes verify each proof in the block and confirm that they have received the corresponding user requests. If the verification passes, the storage node broadcasts a PREPARE message. Finally, when any storage node receives at least 2 / 3N PREPARE messages from other nodes, it broadcasts a COMMIT message; when any node receives at least 2 / 3N COMMIT messages, it finally determines the block and processes the user requests in order. If no storage node starts to process the user requests within the preset specified time, at this time, trigger the view change protocol of BFT-SMR to select a new primary node to ensure the liveness of the consensus process.

[0093] During the file retrieval phase, a Byzantine-Robust Private Information Retrieval (BR-PIR) protocol can be integrated. BR-PIR is a protocol specifically designed for multi-server private information retrieval. This protocol has three important parameters ; among them represents the number of storage nodes participating in the retrieval, represents the maximum number of colluding storage nodes that the protocol can guarantee to not leak the user's privacy, represents the maximum number of storage nodes that the protocol can identify to provide incorrect responses or no responses. In this embodiment, it is required that the BR-PIR protocol needs to satisfy , which enables the user to send queries to all N storage nodes, retrieve the correct file privately, and detect the storage nodes that do not respond or return incorrect results when the subnet can reach a consensus through the database state replication method. For this purpose, the BR-PIR scheme proposed by Goldberg can be selected. Because this scheme can tolerate at most malicious storage nodes, and its Byzantine fault tolerance exceeds N / 3 required by this embodiment.

[0094] Specifically, the user can first obtain the index and database size of the file from the proofs stored on the blockchain, and then call the Query algorithm of the selected BR-PIR scheme to generate N file retrieval vectors, and distribute the file retrieval vectors to N storage nodes. Each storage node that receives the file retrieval vector constructs the target database, calculates the target retrieval result using the Answer algorithm of the BR-PIR scheme, and returns the target retrieval result to the user side. After the user receives all the retrieval results or after a preset time, the user calls the Reconstruct algorithm in the BR-PIR scheme to reconstruct the target file, and identify which storage nodes have returned incorrect retrieval results or have not responded, so as to detect and punish malicious storage nodes.

[0095] As can be seen from the above, in this embodiment, the storage node generates a proof when processing a file upload request or a file deletion request. These proofs can ensure that when the storage node processes a file upload request, the newly generated mapping of the FID of the uploaded file will not conflict with the existing mapping; when the storage node performs a file deletion operation, through the mechanism stipulated by the protocol, it is ensured that the mapping relationship between the FID and the index is correctly removed or updated, and this process can be verified by other nodes, enhancing the publicly verifiable property. On the other hand, in this embodiment, based on the preset private information retrieval protocol supporting a single storage node and the preset private information retrieval protocol supporting multiple storage nodes, the client does not directly send the FID, but generates a file retrieval vector through the file index and sends the file retrieval vector to the storage node. Even if the storage node can obtain all the query information sent by the client during the file retrieval process, it cannot reverse-infer the file content or category corresponding to the FID, protecting the user privacy. At the same time, when the storage node applies any polynomial-time algorithm for speculation, the probability of successfully speculating the FID of the file the user wants to query is not better than random guessing, improving the privacy of file retrieval. Meanwhile, when using the single-server private information retrieval protocol to retrieve a file from a single storage node, it is ensured that the client either successfully retrieves the correct file or can detect that the storage node provides a wrong answer or fails to respond; when using the multi-server private information retrieval protocol to retrieve a file from multiple storage nodes, it is ensured that the client can not only successfully retrieve the correct file, but also identify which storage nodes return wrong results or fail to respond, improving the retrievability auditability.

[0096] See Figure 3 As shown, taking Filecoin as an example below, the technical solution in this application will be described.

[0097] This embodiment implements the file retrieval based on the decentralized storage network protocol proposed in this application on top of Filecoin. Among them, Filecoin is the most widely used DSN implementation at present. Filecoin is developed using Golang, and the architecture of the protocol is as Figure 3 shown, where the file upload and deletion module, node authentication module, message pool, and gateway are components inherited from Filecoin, the file retrieval module, file storage module, and consensus module are components modified based on Filecoin, and the private file retrieval module is a newly developed component.

[0098] On the user side (i.e., user nodes), the file retrieval module is modified to enable users to obtain the index of the target file from the blockchain and call the newly introduced privacy file retrieval module, which is responsible for generating file retrieval vectors and reconstructing files based on the target retrieval results. The single-server PIR and BR-PIR schemes use SealPIR provided by Microsoft and percy++ developed by Goldberg, respectively. On the storage node side, the file storage module is modified to implement the mapping construction and mapping deletion methods proposed in this application. In addition, the file retrieval module is adapted to handle the file retrieval vectors generated by users and call the newly introduced privacy file retrieval module, which implements the Answer algorithm of the PIR protocol. On the public service side, the consensus module is modified to integrate the database state replication method based on HotStuff; the Hotstuff algorithm is provided by Resilient Systems Lab.

[0099] As can be seen from the above, in this embodiment, by modifying the original file retrieval module, file storage module, and consensus module of Filecoin, and adding the privacy file retrieval module, the ability of DSN in user privacy protection is improved, and the security and resilience of the decentralized network file retrieval process are effectively enhanced.

[0100] The technical solutions in this application will be described below with a specific experimental process.

[0101] Experimental environment:

[0102] The experiment is divided into two parts. In the first part of the experiment, the time overheads of file upload, file deletion, proof verification, and retrieval processes are measured. Since this part of the experiment is less sensitive to bandwidth, it can be tested in a local environment. Specifically, on a DELL PowerEdge R740 server (running Ubuntu 22.04 LTS), systems built based on the methods in this application (i.e., Ours_SPIR and Ours_MPIR), Filecoin, Sia, and Storj are deployed. The server is equipped with 2 12-core CPUs (Central Processing Unit), 16GB of memory, and 300GB of SSD (Solid State Disk), and runs the Ubuntu 22.04 LTS operating system. Each DSN contains four storage nodes and one user node. When testing MPIR-DSN, the four storage nodes are regarded as a subnet. The experimental dataset consists of 300 randomly generated text files, each with a size of 7.5MB, created through the Linux dd command. To ensure fair comparison, the above four DSN systems all store only a single copy of each file and generate proofs for all the data of the file.

[0103] In the second part of the experiment, the throughput and latency of file upload, deletion, and retrieval are evaluated. These metrics are greatly affected by bandwidth. Therefore, systems built based on the methods in this application, Filecoin, Sia, and Storj can be deployed in a wide area network environment. Specifically, 125 SA5.MEDIUM8 instances are deployed, each equipped with 1 2-core CPU, 8GB of memory, and 50GB of SSD, and runs the Ubuntu 22.04 LTS operating system. Among these 125 instances, 100 instances are used as storage nodes, and the remaining 25 are used as user nodes. When testing MPIR-DSN, the 100 storage nodes are divided into 25 subnets, each subnet containing 4 storage nodes, and the experimental dataset is the same as in the first part. The request rate at the user end is controlled to vary between 1 and 5 requests per second, and each request is either sent to 4 storage nodes or to a MPIR-DSN subnet. In addition, in the file retrieval experiment, it is ensured that the storage nodes in all DSN systems store 50 files.

[0104] Time overhead of the file operation process:

[0105] In terms of file upload, as Figure 4As shown, the upload times of SPIR-DSN and MPIR-DSN are basically the same as that of Filecoin, indicating that the additional overhead introduced by the system built based on this application is extremely small. Specifically, the mapping establishment only adds 0.04 to 0.21 milliseconds, and the database state replication in MPIR-DSN only consumes an additional 0.03 seconds. These additional overheads are almost negligible compared to the 33 seconds required to generate storage proofs in Filecoin. The file upload time overheads of the two systems, Sia and Storj, are relatively small, but their security is lower than that of the system built based on this application.

[0106] In terms of file deletion, as Figure 5 shown, the deletion time of SPIR-DSN is slightly longer than that of Filecoin, mainly because the mapping deletion introduces an additional overhead of 0.11 to 1.56 milliseconds. The deletion time of MPIR-DSN is about 0.03 seconds longer than that of SPIR-DSN and Filecoin, which is due to the additional overhead of database state replication. Overall, the system built based on this application only introduces a very small additional overhead during the file deletion process. The deletion overhead of Storj is slightly higher than that of the system built based on this application, while in the deletion process of Sia, the file is not actually deleted at the storage node. Instead, only the file identifier is deleted on the user side, making the file unretrievable. Therefore, the time overhead of its deletion process is very small.

[0107] In terms of proof verification, as Figure 6 shown, the system built based on this application has the shortest verification time during the file deletion process, indicating that the overhead of mapping deletion verification is extremely small. The proof verification time of the system built based on this application during the file upload stage is slightly longer than that of Filecoin because it is necessary to verify both the storage proof and the mapping establishment proof at the same time. This also shows that the additional overhead of mapping establishment verification is extremely small and has a limited impact on the overall verification time. Sia and Storj need to split the file into multiple 256KB sectors, so multiple proofs are required to verify the storage of the complete file, and their proof verification time overheads are relatively large.

[0108] In terms of file retrieval, as Figure 7As shown, the retrieval time of the system constructed based on this application grows linearly with the number of stored files. This is because when the storage node executes the Answer algorithm of the PIR protocol, it needs to process the entire database, resulting in a proportional increase in the computational workload with the number of stored files. The retrieval time of SPIR-DSN is higher than that of MPIR-DSN because SealPIR relies on computationally intensive homomorphic multiplication, causing a large computational overhead. The retrieval time of the system constructed based on this application is higher than that of Filecoin because the storage nodes of Filecoin directly return the files without additional computation. In addition, the retrieval method of Sia is similar to that of Filecoin, and its file retrieval time is basically the same as that of Filecoin. Since Storj uses AES-GCM encryption and sets the file block size for encryption and decryption to 1KB, the decryption process of the file is slow during retrieval, increasing the additional overhead.

[0109] Throughput and latency of the file operation process:

[0110] In terms of file upload, as Figure 8 shown in (a) and Figure 8 shown in (b), the upload throughput of the system constructed based on this application is about 1.19 times higher than that of Filecoin, while the latency remains similar, indicating that the additional overhead of the mapping establishment process is extremely small. Since both Filecoin and the system constructed based on this application involve computationally intensive packaging processes, when the storage node processes 5 files simultaneously, the throughputs of both tend to converge. The throughput of Sia grows linearly and the latency is stable, indicating good scalability. Storj relies on satellite nodes to distribute and audit sectors, its throughput remains unchanged, while the latency grows linearly.

[0111] In terms of file deletion, as Figure 9 shown in (a) and Figure 9 shown in (b), the deletion throughput of the system constructed based on this application grows linearly and is about 1.97 times that of Filecoin, while the deletion latency is almost constant and comparable to that of Filecoin. This indicates that the file deletion process of the system constructed based on this application has good scalability and the additional overhead of mapping deletion is extremely small. Storj relies on satellite nodes to manage deletion requests, so the throughput remains constant while the deletion latency grows linearly. Since Sia does not perform actual file deletion at the storage node, it does not participate in this experiment.

[0112] In terms of file retrieval, as Figure 10 shown in (a) and Figure 10As shown in (b), the throughput of MPIR-DSN is slightly lower than that of Filecoin because multi-server PIR requires additional computations, while the throughput of SPIR-DSN is even lower because the computational overhead of single-server PIR is greater. In addition, the data transfer volume of both MPIR-DSN and SPIR-DSN is higher than that of Filecoin. When retrieving a file from 4 storage nodes, the transfer volume of MPIR-DSN is approximately 4 times the file size, and that of SPIR-DSN is approximately 5.7 times the file size. This results in the fact that although the throughput of Filecoin is 1.02 times higher than that of MPIR-DSN and approximately 2.17 times higher than that of SPIR-DSN, the latency of MPIR-DSN and SPIR-DSN is approximately 4.51 times and 13.16 times higher than that of Filecoin, respectively. The throughput of Sia is approximately 32.78 Mbps, and its retrieval latency linearly increases from 1.86 seconds to 10.13 seconds. Due to poor parameter configuration, the throughput of Storj is approximately 14.77 Mbps, and its latency linearly increases from 17.26 seconds to 77.49 seconds.

[0113] As can be seen from the above, through two parts of experiments in this embodiment, the system constructed based on the method in this application is compared with other systems in terms of time overhead, throughput, and latency in aspects of file deletion, proof verification, and retrieval process, demonstrating the superiority of the comprehensive performance of the system constructed based on the method in this application.

[0114] See Figure 11 As shown, the embodiment of this application also discloses a file retrieval device based on a decentralized storage network protocol, including:

[0115] A vector determination module 11, configured to determine a target vector corresponding to a file identifier saved in a vector set; wherein, the vector set includes each file vector, and each file vector corresponds to a Merkle tree; the Merkle tree is used to save the file identifier corresponding to the stored target file;

[0116] An index determination module 12, configured to use the membership proof corresponding to the file identifier and traverse from the next vector of the target vector in the vector set until the entire vector set is traversed, so as to determine the file index corresponding to the file identifier based on a preset index determination process, and construct a mapping relationship between the file index and the file identifier; the membership proof is used to represent the path of the file identifier in the corresponding Merkle tree;

[0117] The first result sending module 13 is configured to, when retrieving a file based on a preset privacy information retrieval protocol supporting a single storage node, obtain a file retrieval vector sent by the client using a single storage node, so that the single storage node constructs a target database according to the mapping relationship between the file index and the corresponding file identifier, determine a target retrieval file corresponding to the file retrieval vector from the target database, determine a target retrieval result based on the target retrieval file, and send the target retrieval result to the client; the file retrieval vector is a vector determined based on the file index.

[0118] The second result sending module 14 is configured to, when retrieving a file based on a preset privacy information retrieval protocol supporting multiple storage nodes, obtain a file retrieval vector sent by the client using multiple storage nodes, so that the multiple storage nodes respectively construct their own corresponding target databases according to the mapping relationship between the file index and the corresponding file identifier, determine a target retrieval file corresponding to the file retrieval vector from the target database, determine a target retrieval result based on the target retrieval file, and send the target retrieval result to the client.

[0119] As can be seen from the above, in this application, the client does not directly send the FID, but generates a file retrieval vector through the file index and sends the file retrieval vector to the storage node, so that the storage node cannot reverse-infer the file content or category corresponding to the FID from the file index, ensuring user privacy and effectively improving the security of the decentralized network file retrieval process.

[0120] In some specific embodiments, the index determination module 12 includes:

[0121] An initialization unit configured to initialize the target counter to 1 and start traversing from the next vector of the target vector in the vector set.

[0122] A first traversal unit configured to, if the currently traversed vector is not empty, determine the vector serial number corresponding to the currently traversed vector, and increment the current target counter based on the vector serial number until the vector set is traversed.

[0123] A second traversal unit configured to determine the membership evidence corresponding to the file identifier, traverse the position information corresponding to each target element in the membership evidence, if the current position information indicates left, increment the current target counter based on the element serial number corresponding to the current position information until all target elements in the membership evidence are traversed, and use the value corresponding to the current target counter as the file index corresponding to the file identifier.

[0124] In some specific embodiments, the index determination module 12 further includes:

[0125] A proof submission sub-module, configured to construct a file upload proof based on the vector set, the target vector, the membership evidence, the file identifier, the file index, and the status hash value corresponding to the file identifier by using a storage node that stores the target file corresponding to the file identifier, and submit the file upload proof to a blockchain by using the storage node, so that the blockchain determines whether the file upload proof passes verification; the status hash value is a previous status hash value generated by the storage node during the last file storage or file deletion;

[0126] Correspondingly, the proof submission sub-module includes:

[0127] A first verification result determination unit, configured to use the blockchain to determine whether the file identifier exists in the Merkle tree corresponding to the target vector and whether the mapping relationship between the file index and the file identifier is correct based on the vector set, the target vector, the membership evidence, and the file identifier in the file upload proof, so as to obtain a first verification result;

[0128] A second verification result determination unit, configured to use the blockchain to determine a first target file index based on the vector set, the target vector, and the membership evidence and according to the preset index determination process, so as to determine whether the first target file index matches the file index, so as to obtain a second verification result;

[0129] A third verification result determination unit, configured to use the blockchain to compare the current preset cryptographic accumulator state with the previous preset cryptographic accumulator state based on the vector set, the membership evidence, and the status hash value, so as to verify whether the state can be transferred from the previous preset cryptographic accumulator state to the current preset cryptographic accumulator state, so as to obtain a third verification result; the preset cryptographic accumulator includes a vector set and a file identifier and is used to construct a mapping relationship between a file index and the corresponding file identifier;

[0130] A first proof storage unit, configured to, if the first verification result, the second verification result, and the third verification result all indicate passing, then the file upload proof passes verification, and save the file upload proof to the blockchain.

[0131] In some specific embodiments, the file retrieval device based on the decentralized storage network protocol further includes:

[0132] A proof verification unit, which is used to, in the file deletion phase, delete the file identifier corresponding to the target deleted file from the preset cryptographic accumulator, generate a corresponding file deletion proof for the file identifier corresponding to the target deleted file, and upload the file deletion proof to the blockchain so that the blockchain can verify whether the file deletion proof passes the verification;

[0133] A second proof storage unit, which is used to, if the file deletion proof passes the verification, save the file deletion proof to the blockchain.

[0134] In some specific embodiments, when using the preset private information retrieval protocol that supports a single storage node, the process of file storage and file deletion by the storage node is the same as the process of file storage and file deletion by the storage node when using the preset private information retrieval protocol that supports multiple storage nodes.

[0135] In some specific embodiments, the file retrieval device based on the decentralized storage network protocol further includes:

[0136] An index determination unit, which is used to obtain the target file upload proof corresponding to the file to be retrieved from the blockchain, so as to determine the file index of the file identifier corresponding to the file to be retrieved in the corresponding target storage node from the target file upload proof;

[0137] A database size determination unit, which is used to determine the database size corresponding to the target storage node based on the maximum file index of the file identifiers stored in the target storage node;

[0138] A vector sending unit, which is used to generate a file retrieval vector according to the file index of the file identifier corresponding to the file to be retrieved in the target storage node and the database size corresponding to the target storage node, so that the client can send the file retrieval vector to the target storage node;

[0139] Correspondingly, when the target storage node receives the file retrieval vector, it further includes:

[0140] A database construction unit, which is used to construct a target database, save the second target file index corresponding to the file identifier it stores to the target database, and save the file corresponding to the second target file index to the target database; wherein, if the file corresponding to the second target file index is deleted, a blank file is saved to the target database.

[0141] In some specific embodiments, the second result sending module 14 further includes:

[0142] A database status unification unit for controlling honest nodes in storage nodes to maintain a consistent database status based on the Byzantine Fault Tolerant State Machine Replication protocol;

[0143] A file retrieval unit for utilizing the Byzantine Robust Private Information Retrieval protocol and performing file retrieval based on the preset private information retrieval protocol supporting multiple storage nodes.

[0144] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 12 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation to the scope of use of the present application.

[0145] Figure 12 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the file retrieval method based on the decentralized storage network protocol disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0146] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0147] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0148] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the file retrieval method based on the decentralized storage network protocol executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0149] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, it implements the aforementioned file retrieval method based on the decentralized storage network protocol. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0150] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method part.

[0151] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0152] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0153] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0154] The above has introduced the technical solution provided by the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A file retrieval method based on a decentralized storage network protocol, characterized in that, Including: Determine a target vector corresponding to a file identifier stored in a vector set; wherein, the vector set includes each file vector, and each file vector corresponds to a Merkle tree; the Merkle tree is used to store the file identifier corresponding to the stored target file; Utilize the membership proof corresponding to the file identifier, and start traversing from the next vector of the target vector in the vector set until the entire vector set is traversed, so as to determine the file index corresponding to the file identifier based on a preset index determination process, and construct a mapping relationship between the file index and the file identifier; the membership proof is used to represent the path of the file identifier in the corresponding Merkle tree; When performing file retrieval based on a preset private information retrieval protocol that supports a single storage node, use a single storage node to obtain a file retrieval vector sent by a client, so that the single storage node constructs a target database according to the mapping relationship between the file index and the corresponding file identifier, and determines a target retrieval file corresponding to the file retrieval vector from the target database, determines a target retrieval result based on the target retrieval file, and sends the target retrieval result to the client; The file retrieval vector is a vector determined based on the file index; When performing file retrieval based on a preset private information retrieval protocol that supports multiple storage nodes, use multiple storage nodes to obtain a file retrieval vector sent by a client, so that the multiple storage nodes respectively construct their own corresponding target databases according to the mapping relationship between the file index and the corresponding file identifier, and determine a target retrieval file corresponding to the file retrieval vector from the target database, determine a target retrieval result based on the target retrieval file, and send the target retrieval result to the client.

2. The file retrieval method based on the decentralized storage network protocol according to claim 1, wherein The step of utilizing the membership proof corresponding to the file identifier, and starting to traverse from the next vector of the target vector in the vector set until the entire vector set is traversed, so as to determine the file index corresponding to the file identifier based on a preset index determination process, includes: Initialize a target counter to 1, and start traversing from the next vector of the target vector in the vector set; If the currently traversed vector is not empty, determine the vector number corresponding to the currently traversed vector, and increment the current target counter based on the vector number until the entire vector set is traversed; Determine the membership proof corresponding to the file identifier, traverse the position information corresponding to each target element in the membership proof, if the current position information is characterized as left, increment the current target counter based on the element number corresponding to the current position information until all target elements in the membership proof are traversed, and use the value corresponding to the current target counter as the file index corresponding to the file identifier.

3. The file retrieval method based on the decentralized storage network protocol according to claim 1, characterized in that, After constructing the mapping relationship between the file index and the file identifier, it further includes: The storage node storing the target file corresponding to the file identifier constructs a file upload proof based on the vector set, the target vector, the membership evidence, the file identifier, the file index, and the status hash value corresponding to the file identifier, and submits the file upload proof to the blockchain by using the storage node so that the blockchain determines whether the file upload proof passes the verification; the status hash value is the previous status hash value generated by the storage node during the last file storage or file deletion. Correspondingly, using the blockchain to determine whether the file upload proof passes the verification includes: Using the blockchain to determine whether the file identifier exists in the Merkle tree corresponding to the target vector and whether the mapping relationship between the file index and the file identifier is correct based on the vector set, the target vector, the membership evidence, and the file identifier in the file upload proof, so as to obtain a first verification result; Using the blockchain to determine a first target file index based on the vector set, the target vector, and the membership evidence, and according to the preset index determination process, so as to determine whether the first target file index matches the file index, so as to obtain a second verification result; Using the blockchain to compare the current preset cryptographic accumulator state with the previous preset cryptographic accumulator state based on the vector set, the membership evidence, and the status hash value, so as to verify whether the state can be transferred from the previous preset cryptographic accumulator state to the current preset cryptographic accumulator state, so as to obtain a third verification result; the preset cryptographic accumulator includes a vector set and a file identifier, and is used to construct a mapping relationship between a file index and a corresponding file identifier; If the first verification result, the second verification result, and the third verification result all indicate passing, the file upload proof passes the verification, and the file upload proof is saved to the blockchain.

4. The file retrieval method based on the decentralized storage network protocol according to claim 3, characterized in that, It further includes: In the file deletion stage, delete the file identifier corresponding to the target deleted file from the preset cryptographic accumulator, generate a corresponding file deletion proof for the file identifier corresponding to the target deleted file, and upload the file deletion proof to the blockchain so that the blockchain verifies whether the file deletion proof passes the verification; If the file deletion proof passes the verification, the file deletion proof is saved to the blockchain.

5. The file retrieval method based on the decentralized storage network protocol according to claim 1, wherein When using the preset privacy information retrieval protocol supporting a single storage node, the process of file storage and file deletion by the storage node is the same as the process of file storage and file deletion by the storage node when using the preset privacy information retrieval protocol supporting multiple storage nodes.

6. The file retrieval method based on the decentralized storage network protocol according to claim 3, wherein It further includes: Obtain the target file upload proof corresponding to the file to be retrieved from the blockchain, so as to determine the file index of the file identifier corresponding to the file to be retrieved in the corresponding target storage node from the target file upload proof; Determine the database size corresponding to the target storage node based on the maximum file index of the file identifiers stored in the target storage node; Generate a file retrieval vector based on the file index of the file to be retrieved corresponding to the file identifier in the target storage node and the database size corresponding to the target storage node, so that the client can send the file retrieval vector to the target storage node; Correspondingly, when the target storage node receives the file retrieval vector, it further includes: Construct a target database, save the second target file index corresponding to the file identifier saved by itself into the target database, and save the file corresponding to the second target file index into the target database; wherein, if the file corresponding to the second target file index is deleted, a blank file is saved into the target database.

7. The file retrieval method based on the decentralized storage network protocol according to any one of claims 1 to 6, characterized in that, When performing file retrieval based on a privacy information retrieval protocol that supports multiple storage nodes preset, it further includes: Control the honest nodes in the storage nodes to maintain a consistent database state based on the Byzantine fault-tolerant state machine replication protocol; Utilize the Byzantine robust privacy information retrieval protocol and perform file retrieval based on the preset privacy information retrieval protocol that supports multiple storage nodes.

8. A file retrieval device based on a decentralized storage network protocol, characterized in that, It includes: A vector determination module, configured to determine a target vector corresponding to a file identifier saved in a vector set; wherein the vector set includes each file vector, and each file vector corresponds to a Merkle tree; the Merkle tree is used to save the file identifier corresponding to the stored target file; An index determination module, configured to use the membership proof corresponding to the file identifier and traverse from the next vector of the target vector in the vector set until the vector set is traversed, so as to determine the file index corresponding to the file identifier based on a preset index determination process, and construct a mapping relationship between the file index and the file identifier; the membership proof is used to represent the path of the file identifier in the corresponding Merkle tree; A first result sending module, configured to, when performing file retrieval based on a preset privacy information retrieval protocol that supports a single storage node, use a single storage node to obtain a file retrieval vector sent by the client, so that the single storage node constructs a target database according to the mapping relationship between the file index and the corresponding file identifier, determine a target retrieval file corresponding to the file retrieval vector from the target database, determine a target retrieval result based on the target retrieval file, and send the target retrieval result to the client; the file retrieval vector is a vector determined based on the file index; A second result sending module, configured to, when performing file retrieval based on a preset privacy information retrieval protocol that supports multiple storage nodes, use multiple storage nodes to obtain a file retrieval vector sent by the client, so that the multiple storage nodes respectively construct their own corresponding target databases according to the mapping relationship between the file index and the corresponding file identifier, determine a target retrieval file corresponding to the file retrieval vector from the target databases, determine a target retrieval result based on the target retrieval file, and send the target retrieval result to the client.

9. An electronic device, characterized in that, It includes: A memory, configured to save a computer program; A processor for executing the computer program to implement the file retrieval method based on the decentralized storage network protocol according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program which, when executed by a processor, implements the file retrieval method based on the decentralized storage network protocol according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Search encryption method supporting dynamic updating and multi-keyword safe ranking

    CN104615692A

  • Cloud storage-oriented multi-keyword ciphertext retrieval method and system

    CN108628867A

  • Data encryption and search method capable of protecting file privacy in cloud environment

    CN108768951A

  • Decentralized storage system, method and equipment for multi-version files and storage medium

    CN117056297A

  • File management method based on Byzantine fault-tolerant decentralized storage network

    CN118018561A

Cited By

  • MPT-based block chain database system, data access method, terminal and medium

    CN121278137A