Distributed storage system anti-collusion privacy information retrieval method and system

The query strategy is constructed through disguise and compression methods, and the problem of insufficient code rate in existing private information retrieval methods is solved, and more efficient private information retrieval is achieved. It is suitable for more parameters and multi-file retrieval scenarios, improving the information retrieval rate.

CN120336264APending Publication Date: 2025-07-18SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510298876.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing private information retrieval methods need to improve in terms of adapting information retrieval rate, especially in terms of MDS-TPIR capacity, which cannot effectively adapt to more parameters and multi-file retrieval scenarios.

Method used

Using the disguise and compression method, by constructing query vectors and encoded symbols, the query strategy for each file is designed, so that any set of T conspiracy servers cannot distinguish between the requested file and the non-requested file, and in the compression stage, analyzing the server's response to query saves the number of bits of the non-requested file, and eliminates redundant symbols to improve the bit rate.

Benefits of technology

It improves the PIR code rate, is better than existing research results, adapts to more parameters and multi-file retrieval scenarios, and enhances the security and efficiency of private information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336264A_ABST
    Figure CN120336264A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of privacy information retrieval, and provides an anti-collusion privacy information retrieval method and system for a distributed storage system, the technical scheme mainly comprises two stages of camouflage and compression, the camouflage stage aims to protect privacy, and the query, compression and compression of each file are designed; therefore, any group of T collusion servers cannot distinguish the requested file from the non-requested file. In the compression stage, the number of bits of non-requested files that can be saved by each server in response to a query by applying a combinatorial policy is analyzed. And more redundancy about non-retrieval file symbols can be compressed, so that the code rate is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of private information retrieval, and in particular relates to a method and system for anti-collusion private information retrieval in a distributed storage system. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] The concept of Private Information Retrieval (PIR) was first proposed by Chor et al. In recent years, PIR based on Maximum Distance Separable (MDS) - coded databases and allowing server collusion (abbreviated as MDS - TPIR) has received extensive attention. In the MDS - TPIR setting, M files are stored on N servers, and each file is independently stored using the same (N, K) - MDS code. A user wishes to retrieve a single file while not revealing any information about the index of the requested file to any up to T colluding servers. The measure of the efficiency of an MDS - TPIR scheme is called the MDS - TPIR rate, which is the ratio of the size of the retrieved file to the total download volume. The maximum achievable rate is called the MDS - TPIR capacity.

[0004] For the degenerate parameter cases K = 1 or T = 1, the MDS - TPIR capacity has been completely determined. For the non - degenerate case K≥2 and T≥2, Zhang and Ge proposed a scheme with a rate of (1 + ρ + … + ρ M-1 ) -1 , where Using the Schur product of generalized Reed - Solomon codes, Freij - Hollanti et al. proposed a scheme with a rate of . Combining this result with the known capacity expressions in the degenerate cases, Freij - Hollanti et al. conjectured that the MDS - TPIR capacity is given by the following formula: where Using the methods of refinement and lifting, D’Oliveira and El Rouayheb proposed a scheme whose PIR rate is exactly consistent with the above conjecture. In addition, when the PIR scheme is further required to be linear and full support-rank, the conjecture is proven to be correct. However, without such restrictions, the conjecture is proven to be wrong by Sun and Jafar. They constructed a PIR scheme with parameters (M,N,T,K)=(2,4,2,2), achieving a rate of 3 / 5, exceeding the conjectured value of 4 / 7. In addition, for the specific case of K=N - 1 and M = 2, Sun and Jafar gave an exact expression for the capacity. They also pointed out that in the MDS-TPIR capacity expression, the roles of K and T are not necessarily symmetric, which is different from the expectation of previously known results.

[0005] Therefore, it can be seen that the information retrieval rate adapted by existing private information retrieval methods needs to be improved. Summary of the Invention

[0006] To solve at least one of the technical problems existing in the above background art, the present invention provides a method and system for anti-collusion private information retrieval in a distributed storage system, and its general case can be adapted to more parameters, and can also be adapted to other private information retrieval scenarios such as multi-file retrieval and arbitrary collusion patterns.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The first aspect of the present invention provides a method for anti-collusion private information retrieval in a distributed storage system, including the following steps:

[0009] Store a number of files composed of independently and identically distributed and uniform symbols into multiple servers;

[0010] Construct a query based on the query vectors of the retrieved files and non-retrieved files, take the inner product of the query vectors with the symbols of each corresponding file to obtain encoded symbols;

[0011] Analyze the redundant symbols existing in the non-retrieved files, and eliminate the redundant symbols to obtain the maximum dimension of the non-retrieved file symbols;

[0012] Construct an effective file download policy for each server when responding to the query based on the maximum dimension of the non-retrieved file symbols;

[0013] Compress the encoded symbols according to the effective file download policy, recover all the symbols in the retrieved files based on the compressed encoded symbols, and parse to obtain the files to be retrieved.

[0014] Furthermore, each file consists of L independent symbols over a sufficiently large finite field F p where K is the number of data packets contained in each file, N is the number of servers, and each file is stored in N servers through a (N, K)-MDS code.

[0015] Furthermore, the query vector of the file to be retrieved is The construction process is as follows:

[0016] Denote the set of all p full-rank matrices over F as Randomly and privately select two matrices S, S′ independently and uniformly from Let Label the rows of S as Divide the rows of S′ into two parts, labeled as {Z : i ∈ [σ]} and i respectively, where each

[0017] contains row vectors of S, and any K sets in share exactly one common row vector of S, and any two sets share exactly common row vectors;

[0018] The query vector of the non-retrieved file is The construction process is as follows:

[0019] Put {Z i : i ∈ [σ]} into all sets Let

[0020] and and assign to

[0021] Furthermore, the query vector sent to the server is:

[0022]

[0023] where is the query sent to the nth server, W m is the file to be retrieved, is the non-retrieved file;

[0024] After receiving the query vector, the nth server takes the inner product between these vectors and the symbols of each file it stores to calculate 2L / N coded symbols ​​​​

[0025] Furthermore, the symbols of the retrieved files are independent;

[0026] The symbols of the non-retrieved files have a dimension of at most

[0027] Furthermore, the effective file download policy of each server in response to a query includes:

[0028] According to the set order full-rank matrix C n , map the symbols of the retrieved files and the symbols of the non-retrieved files to {X1,…,X L / N ,Y1,…,Y L / N};

[0029] The effective download policy is: compress these symbols to form a part of isolated symbols and some pairwise sums

[0030] The second aspect of the present invention provides a distributed storage system anti-collusion privacy information retrieval system, including:

[0031] A file storage module that stores several files composed of independently and identically distributed and uniform symbols into multiple servers;

[0032] A query construction module that constructs a query based on the query vectors of the retrieved files and non-retrieved files, takes the inner product of the query vectors and the symbols of the corresponding files to obtain encoded symbols;

[0033] A redundancy extraction module that analyzes the redundant symbols existing in the non-retrieved files and eliminates the redundant symbols to obtain the maximum dimension of the non-retrieved file symbols;

[0034] An effective download module that constructs the effective file download policy of each server in response to a query based on the maximum dimension of the non-retrieved file symbols; compresses the encoded symbols according to the effective file download policy, restores all the symbols in the retrieved files based on the compressed encoded symbols, and parses to obtain the file to be retrieved.

[0035] The third aspect of the present invention provides a computer-readable storage medium.

[0036] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in an anti-collusion privacy information retrieval method for a distributed storage system as described above.

[0037] The fourth aspect of the present invention provides a computer device.

[0038] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a method for retrieving privacy information against collusion in a distributed storage system as described above are implemented.

[0039] The fifth aspect of the present invention provides a program product.

[0040] A program product, which is a computer program product, includes a computer program. When the computer program is executed by a processor, the steps in a method for retrieving privacy information against collusion in a distributed storage system as described in the first aspect are implemented.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] 1. The present invention proposes a method for retrieving privacy information against collusion in a distributed storage system. By using the methods of disguise and compression, for each file query, any group of T colluding servers cannot distinguish the requested file from the non-requested files. In the compression stage, the number of bits of non-requested files that can be saved by each server applying a combination strategy when responding to the query is analyzed; more redundancy of non-retrieved file symbols can be compressed, thereby further improving the code rate.

[0043] 2. When the parameters are (M, N, T, K) = (2, N, 2, K), the PIR code rate of the present invention is better than the existing research results. Its generalized situation can be adapted to more parameters and can also be adapted to other PIR scenarios such as multi-file retrieval and arbitrary collusion patterns.

[0044] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0046] Figure 1 It is a flowchart of a method for retrieving privacy information against collusion in a distributed storage system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention will be further described below in conjunction with the drawings and embodiments.

[0048] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention pertains.

[0049] It should be noted that the terms used herein are merely for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0050] Regarding the problem of private information retrieval (MDS-TPIR) from an MDS-coded database in the presence of colluding servers. Specifically, the problem consists of M files and N servers, and each file is stored using an (N, K)-MDS code. A user wishes to retrieve one of the files without revealing any information about the index of the requested file to up to T colluding servers. Regarding the MDS-TPIR capacity, for the degenerate cases K = 1 or T = 1, the MDS-TPIR capacity has been determined. For the non-degenerate cases K ≥ 2 and T ≥ 2, Freij-Hollanti et al. proposed a conjecture where the conjecture was proven wrong by Sun and Jafar. They constructed a scheme with parameters (M, N, T, K) = (2, 4, 2, 2). Based on this counterexample, the present invention refined the ideas of disguise and compression therein and constructed a new class of anti-collusion PIR schemes for distributed storage systems. Using the methods of disguise and squeeze, the disguise stage aims to protect privacy and designs queries for each file such that any group of T colluding servers cannot distinguish the requested file from the non-requested files. In the compression stage, the number of bits of non-requested files that can be saved by each server applying a combination strategy when responding to queries is analyzed.

[0051] Embodiment 1

[0052] As Figure 1 shown, this embodiment provides a method for anti-collusion private information retrieval in a distributed storage system. The problem of private information retrieval (MDS-TPIR) from an MDS-coded database consists of M files and N servers, and each file is stored using an (N, K)-MDS code. A user wishes to retrieve one of the files without revealing any information about the index of the requested file to up to T colluding servers.

[0053] Specifically, it includes the following steps:

[0054] Step 1: Distributively store several files on multiple servers;

[0055] In this embodiment, taking (M, N, T, K) = (2, N, 2, K) as an example for discussion, its generalized case can be adapted to more parameters, and can also be adapted to other PIR scenarios such as multi-file retrieval and arbitrary collusion patterns.

[0056] Assume that each file consists of p independent symbols over a sufficiently large finite field F. Let:

[0057] W1 = (a1, …, a K ), W2 = (b1, …, b K ),

[0058] where W1 is the first file and W2 is the second file; is a vector composed of independent and identically distributed and uniform symbols from F p . K is the number of data packets contained in each file, N is the number of servers, and each file is stored in N servers through (N, K)-MDS codes, specifically expressed as: (W 11 , …, W 1N ) = W1G, (W 21 , …, W 2N ) = W2G. Among them, the generating matrix G comes from Lemma 2;

[0059] Among them, Lemma 1: Let G = [I K A] be a K×N matrix, where I K is the identity matrix and A is a K×(N - K) matrix; then G has the MDS property, that is, the necessary and sufficient condition for any K columns of G to be linearly independent is that any sub-square matrix of A is full-rank.

[0060] Lemma 2: Let G = [I K g1 … g N-K be a K×N matrix with the MDS property, where When p is sufficiently large, there exists a matrix:

[0061]

[0062] with the row-MDS property (any two rows are linearly independent), where

[0063] ​Given a generating matrix $G$ of an MDS storage code, the corresponding matrix $H$ plays an important role in designing queries. In fact, when $K = N - 2$ and $G=[I K g1g2]$ has the MDS property, then according to Lemma 1, we can take to have the row-MDS property.

[0064] Step 2: For the query of each file, ensure that any set of $T$ colluding servers cannot distinguish between the requested file and non-requested files;

[0065] Specifically, it includes the following steps:

[0066] Step 201: Construct the query vectors of the files to be retrieved

[0067] For $n\in[N]$ and $1\leq m\leq2$, to privately retrieve the $m$-th file, the user sends a query set to the $n$-th server. This set consists of two parts: Each part contains row vectors, called query vectors.

[0068] Denote the set of all p full-rank matrices over as The user privately and independently selects two matrices $S,S'$ uniformly from .

[0069] Let Label the rows of $S$ as Divide the rows of $S'$ into two parts, labeled as $\{Z i :i\in[\sigma]\}$ and

[0070] Define the set as follows:

[0071] Each contains row vectors of $S$, and any $K$ sets in share exactly one common row vector of $S$. Thus, any two sets share exactly common row vectors. Essentially, considering the linear space spanned by each set

[0072] Step 202: Construct the query vectors of the non-retrieved files Specifically, it includes:

[0073] First, for $\{Z i: {i ∈ [σ]} into all sets among them.

[0074] Let and

[0075]

[0076] and assign to among them.

[0077] The matrix H here comes from Lemma 2 and has the row-MDS property, such that for any n, n′ ∈ [N], The subspace spanned by and the intersection of the subspaces spanned by

[0078] Equation (1) is well-defined because Considering each The dimension of the intersection of any two spaces is exactly σ. This is the key property for disguising the retrieval file and the non-retrieval file query vectors.

[0079] Step 203: Construct a query based on the query vectors of the retrieval file and the non-retrieval file, take the inner product of the query vectors with the symbols of each file to obtain the coded symbols;

[0080] Let W m be the file to be retrieved, and

[0081]

[0082] where π n and π n ′ are independent and uniformly random permutations on the symmetric group .

[0083] After receiving the query vector, the n-th server will take the inner product between and the symbols of each file it stores, and calculate to obtain 2L / N coded symbols

[0084] Step 3: Analyze the redundant symbols existing in the non-retrieval file, and eliminate the redundant symbols to obtain the maximum dimension of the non-retrieval file symbols;

[0085] The symbols of the retrieval file are independent: Because in the (N, K)-MDS system, each query vector V i is sent to exactly K servers, so the user can retrieve V i W m. Since S is full rank, the retrieval is equivalent to retrieving file W m . Therefore, as long as the user collects this scheme is effective.

[0086] The dimension of the non-retrieval file symbols is at most For each i ∈ [σ], there are N - K redundant symbols in .

[0087] Let For each j, 1 ≤ j ≤ μ, due to the relationship between the two matrices G and H in Lemma 2, it can be guaranteed that This indicates that in there is at least one redundant symbol. Therefore, in there are at least σ(N - K) + μ redundant symbols in total. That is, the dimension of the non-retrieval file symbols is at most

[0088] Step 4: Based on the maximum dimension of the non-retrieval file symbols, construct the effective file download strategy for each server when responding to a query. Compress the encoded symbols according to the effective file download strategy, and recover all the symbols in the retrieval file based on the compressed encoded symbols, and parse to obtain the file to be retrieved;

[0089] When the nth server responds to the query vector, it will adopt a combination strategy to map to a smaller number of encoded symbols.

[0090] This idea comes from Sun and Jafar. A preset full-rank matrix C of order n is used to map and to {X1, …, X L / N , Y1, …, Y L / N}.

[0091] Then a combination function compresses these symbols to form a part of isolated symbols and some pairwise sums

[0092] The key of the combination strategy is to ensure that ∑ n∈[N] I n = I, and the union of the isolated symbols of the non-retrieval file can linearly generate all the symbols in, which will enable the user to eliminate the interference terms appearing in the paired sums and finally recover All L symbols in are solved for W m The C that satisfies the above conditions n The selection of is very non-trivial (need to consider (L / N)! N cases of permutations), but its existence can be guaranteed when the field is large enough.

[0093] Theorem 1: For (M, N, T, K) = (2, N, 2, K), there exists a PIR scheme with a code rate of where N - K ≥ 2, K ≥ 2.

[0094] Proof: The correctness of the scheme has been explained, the construction of the query vectors and and the permutation {π n , π n ′: n ∈ [N]} can ensure that for any two servers holds (the probability distributions of the two sets of queries are the same). The code rate of the scheme can be calculated by Substituting obtains the theorem.

[0095] For the parameters (M, N, T, K) = (2, N, 2, K), the PIR scheme designed by the present invention has a code rate of where N - K ≥ 2, K ≥ 2. Note that when (M, N, T, K) = (2, N, 2, K), the originally conjectured capacity is Therefore, when 2N - K - 1 > (N - K) 2 the results of the present invention provide more counterexamples and are also superior to the currently known PIR schemes. For (M, N, T, K) = (2, 5, 2, 2), by further carefully designing the query vectors, more redundancy regarding non-retrieved file symbols can be compressed to further improve the code rate.

[0096] The implementation process of a collusion-resistant private information retrieval method in this embodiment will be described in detail below from several aspects including files and storage, construction of queries, compression stage, and effective download.

[0097] Overall scheme example: (M, N, T, K) = (2, 5, 2, 3)

[0098] 1. Files and storage:

[0099] Assume that each file consists of L = 30 independent symbols over a sufficiently large finite field F p Let:

[0100] W1 = (a1, a2, a3) W2 = (b1, b2, b3),

[0101] where is a vector composed of independent and identically distributed uniform symbols from F p . Each file is stored in the system through a (5,3)-MDS code. Let

[0102]

[0103] both G and H have the MDS property. The storage scheme is set as:

[0104] (W 11 , …, W 15 ) = W1G(W 21 , …, W 25 ) = W2G;

[0105] 2. Disguise stage (construction of queries)

[0106] Let S and S′ be two matrices independently and uniformly selected from the set of all 10×10 full-rank matrices. Label the row vectors of S as {V i , i ∈

[10] }, and the row vectors of S′ as {Z1, Z2, Z3, U j , j ∈ [1:7]}. Define:

[0107]

[0108] where

[0109] Let W m be the file to be retrieved, be the non-retrieved files. For n ∈ [5], the query sent to the nth server consists of two parts:

[0110]

[0111] where π n and π′ n are independent and uniformly random permutations on the symmetric group S6.

[0112] After receiving the query, the nth server takes the inner product of these vectors with the symbols it stores to obtain 12 coded symbols

[0113] 3. Compression stage (extracting redundancy)

[0114] The symbols of the retrieved file are independent: because in a (5,3)-MDS system, each query vector V i is sent to exactly K = 3 servers, so the user can retrieve V i W m. Since S is full rank, retrieving {V i W m : i ∈

[10] } is equivalent to retrieving file W m . Therefore, as long as the user collects this scheme is effective.

[0115] Symbols of non-retrieved files have a dimension of at most For each i ∈ [3], there are 2 redundant symbols in . Let For each j, 1 ≤ j ≤ 3,

[0116] (U j + U j+3 )x1 + (U j + 2U j+3 )x2 + (U j + 3U j+3 )x3

[0117] = U j (x1 + x2 + x3) + U j+3 (x1 + 2x2 + 3x3)

[0118] This shows that there is at least one redundant symbol in . Therefore, there are at least 9 redundant symbols in total in . That is, the dimension of the symbols of non-retrieved files is at most

[0119] 4. Effective Download (Combination Strategy)

[0120] Suppose L = 30 non-retrieved file symbols can be expressed as a linear combination of the symbols in set s, where s is a set containing I = I1 + … + I5 = 5 + 4 + 4 + 4 + 4 = 21 symbols, and the n-th server for n ∈ [5] contains I n symbols in s. The symbols in set s can generate all the symbols in

[0121] Define the function

[0122]

[0123] The first I n symbols are downloaded directly, and the last 6 - I n symbols are downloaded in combination. The symbols of retrieved files and non-retrieved files are used to generate the answer in the following combined form.

[0124]

[0125] where C n is a deterministic 6×6 matrix that is required to satisfy the following two properties (denote the first I n rows of C n as ).

[0126] P1. All C n , n ∈ [5] are full rank.

[0127] P2. For all (6!) 5 different realizations of π n ′, n ∈ [5], there is a bijection between the linear combination of the symbols of the 21 non-retrievable files downloaded directly, i.e., and s.

[0128] Using the Schwartz-Zippel lemma, it can be shown that when the finite field F p is sufficiently large, a C n satisfying the above two conditions exists. When the size of the field tends to infinity, the probability that a C n matrix randomly selected uniformly satisfies these properties approaches 1.

[0129] 5. Correctness and privacy of the scheme

[0130] That C n satisfies the two properties can guarantee correctness. The construction of the query vector can guarantee privacy.

[0131] The code rate is exceeding the capacity conjecture value

[0132] 6. Extracting more redundancy Example: (M, N, T, K) = (2, 5, 2, 2)

[0133] 1) Files and storage:

[0134] Assume that each file consists of L = 20 independent symbols over a sufficiently large finite field F p . Let

[0135] W1 = (a1, a2), W2 = (b1, b2),

[0136] where is a vector composed of independent and identically distributed uniform symbols from F p . Each file is stored in the system through a (5, 2)-MDS code. Let

[0137]

[0138] Both G and H have the MDS property. The storage scheme is set as

[0139] (W 11 ,…,W 15 ) = W1G(W 21 ,…,W 25 ) = W2G,

[0140] 2) Disguise Phase (Construction of Queries)

[0141] Let S and S′ be two matrices independently and uniformly selected from the set of all 10×10 full-rank matrices. Label the row vectors of S as {V i , i ∈

[10] }, and the row vectors of S′ as {Z1, U j , j ∈ [1:9]}. Define

[0142]

[0143]

[0144] where

[0145]

[0146] Let W m be the retrieved file, and be the non-retrieved file. For n ∈ [5], the query sent to the nth server consists of two parts:

[0147]

[0148] where π n and π n ′ are independent and uniformly random permutations on the symmetric group S4.

[0149] After receiving the query, the nth server takes the inner product of these vectors with the symbols it stores to obtain 8 coded symbols

[0150] 3) Compression Phase (Extracting Redundancy)

[0151] The symbols of the non-retrieved files have dimensions of at most There are 3 redundant symbols in For each j, 1 ≤ j ≤ 3,

[0152]

[0153] The rank of the above digital matrix is 3. Thus, for each 1 ≤ j ≤ 3, There are two redundancies in In total, there are 9 redundant symbols in the non-retrieved symbols.

[0154] It should be emphasized that here we further compress the extra redundancy by using a better retrieval vector design method, thus being superior to the general description of Theorem 1.

[0155] Through the combination strategy, the code rate of the scheme can reach exceeding the value of Theorem 1 and the conjectured value

[0156] Example 2

[0157] This embodiment provides a distributed storage system anti-collusion private information retrieval system, including:

[0158] A file storage module that stores several files composed of independently and identically distributed and uniform symbols into multiple servers;

[0159] A query construction module that constructs a query based on the query vectors of the retrieved files and the non-retrieved files, takes the inner product of the query vectors with the symbols of the corresponding files to obtain encoded symbols;

[0160] A redundancy extraction module that analyzes the redundant symbols existing in the non-retrieved files and eliminates the redundant symbols to obtain the maximum dimension of the non-retrieved file symbols;

[0161] An effective download module that constructs an effective file download policy for each server when responding to a query based on the maximum dimension of the non-retrieved file symbols; compresses the encoded symbols according to the effective file download policy, restores all the symbols in the retrieved file based on the compressed encoded symbols, and parses to obtain the file to be retrieved.

[0162] It should be noted that the specific implementation manner of a distributed storage system anti-collusion private information retrieval system in the embodiments of the present invention is similar to the specific implementation manner of a distributed storage system anti-collusion private information retrieval method in the embodiments of the present invention. For details, please refer to the description in the method part. To reduce redundancy, it will not be elaborated here.

[0163] Example 3

[0164] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a distributed storage system anti-collusion private information retrieval method as described above.

[0165] Example 4

[0166] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a distributed storage system anti-collusion private information retrieval method as described above.

[0167] Example 5

[0168] This embodiment provides a program product, which is a computer program product and includes a computer program. When the computer program is executed by a processor, it implements the steps in a distributed storage system anti-collusion privacy information retrieval method as described above.

[0169] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0170] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0171] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0173] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0174] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for retrieving privacy information against collusion in a distributed storage system, characterized in that, Including the following steps: Storing several files composed of independently and identically distributed and uniform symbols into multiple servers; Constructing a query based on the query vectors of the retrieved files and the non-retrieved files, taking the inner product of the query vectors with the symbols of the corresponding files to obtain coded symbols; Analyzing the redundant symbols existing in the non-retrieved files, removing the redundant symbols to obtain the maximum dimension of the symbols of the non-retrieved files; Constructing the effective file download strategy for each server when responding to the query based on the maximum dimension of the symbols of the non-retrieved files; Compressing the coded symbols according to the effective file download strategy, restoring all the symbols in the retrieved files based on the compressed coded symbols, and parsing to obtain the files to be retrieved.

2. The anti-collusion private information retrieval method for a distributed storage system according to claim 1, wherein Each file consists of L independent symbols over a finite field F that is large enough, where p K is the number of data packets contained in each file, and N is the number of servers. Each file is stored in N servers through a (N, K)-MDS code. K is the number of data packets contained in each file, N is the number of servers, and each file is stored in N servers through a (N, K)-MDS code.

3. A method for retrieving private information against collusion in a distributed storage system according to claim 1, characterized in that, The query vector of the file to be retrieved is The construction process is as follows: Denote F p All of the The set of full-rank matrices is Independently and uniformly privately select two matrices S, S′ from ; Label the rows of S as Divide the rows of S′ into two parts, labeled as {Z i : i ∈ [σ]} and Each contains row vectors in S, and any K sets in share exactly one common row vector in S, and any two sets share exactly The query vector of the non-retrieval document is The construction process is as follows: Put {Z i : i ∈ [σ]} into all sets . Let and and assign to in.

4. The anti-collusion private information retrieval method for a distributed storage system according to claim 3, wherein, The query vector sent to the server is: Among them, is the query sent to the nth server, W m is the file to be retrieved, is the non-retrievable file; After receiving the query vectors, the nth server computes the inner products between these vectors and the symbols of each file it stores to obtain 2L / N coded symbols 5. The anti-collusion privacy information retrieval method for a distributed storage system according to claim 4, characterized in that Symbol of the retrieved document is independent; Symbol for non-search document has a dimension of at most 6. A method for retrieving privacy information in a distributed storage system against collusion, as described in claim 4, wherein The effective file download strategy for each server when responding to the query includes: According to the set full-rank matrix C of order n , map the symbols of the retrieved documents and the symbols of the non-retrieved documents to {X1,…,X L / N ,Y1,…,Y L / N}; Effective download strategy is: Compress these symbols to form a part of isolated symbols and some pairwise sums 7. A distributed storage system anti-collusion privacy information retrieval system, characterized in that, Including: A file storage module for storing several files composed of independently and identically distributed and uniform symbols into multiple servers; A query construction module for constructing a query based on the query vectors of the retrieved files and the non-retrieved files, taking the inner product of the query vectors with the symbols of the corresponding files to obtain coded symbols; A redundancy extraction module for analyzing the redundant symbols existing in the non-retrieved files, removing the redundant symbols to obtain the maximum dimension of the symbols of the non-retrieved files; An effective download module for constructing the effective file download strategy for each server when responding to the query based on the maximum dimension of the symbols of the non-retrieved files; compressing the coded symbols according to the effective file download strategy, restoring all the symbols in the retrieved files based on the compressed coded symbols, and parsing to obtain the files to be retrieved.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a method for retrieving private information in a distributed storage system against collusion as described in any one of claims 1-6.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a method for retrieving private information in a distributed storage system against collusion as described in any one of claims 1-6.

10. A program product, which is a computer program product and includes a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps in a method for retrieving private information in a distributed storage system against collusion as described in any one of claims 1-6.