Coding method, confidentiality measurement method, storage medium and device for confidentiality measurement of similarity of archival information based on Chebyshev distance

Through the archive information similarity density measurement method based on Chebishev distance, combined with the NTRU encryption algorithm and hash function, a secure computing protocol is designed, which solves the problems of high complexity and low efficiency of Chebishev distance calculation, and realizes efficient and secure similarity measurement and clustering without leaking confidential archive data.

CN118013556BActive Publication Date: 2025-07-01INNER MONGOLIA UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410276779.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-07-01
Estimated Expiration
2044-03-12

AI Technical Summary

Technical Problem

The existing Chebishev distance calculation solution has high computational complexity and low efficiency without leaking the privacy of confidential archive data, and the growth rate is high as the number of elements increases, and lacks effective anti-malicious attack methods.

Method used

Using the archive information similarity density measurement method based on Chebishev distance, confidential archives are mapped into two-dimensional vectors through feature extraction and dimensionality reduction technology, combined with NTRU encryption algorithm and hash function, a secure computing protocol under semi-honest and malicious models is designed, and digital commitments are used to ensure the accuracy and security of the calculation results.

Benefits of technology

Without leaking the privacy of confidential archive data, it significantly reduces the computational complexity and execution time, improves computing efficiency, can resist quantum attacks, and ensures the accuracy and security of computational results under malicious models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118013556B_ABST
    Figure CN118013556B_ABST
Patent Text Reader

Abstract

Coding method, confidentiality measurement method, storage medium and device for confidentiality measurement of file information similarity based on Chebyshev distance, belonging to the technical field of file information processing. In order to solve the problems of high computational complexity and poor computational efficiency in the existing document distance measurement scheme without revealing the privacy of confidential file data, the present invention extracts features from the information in the confidential files to obtain feature vectors, and then maps the feature vectors into two-dimensional vectors through dimensionality reduction technology, that is, maps the information in the first confidential file and the second confidential file into private points on a plane and encodes them. Then, based on the NTRU encryption algorithm, a secure Chebyshev distance calculation protocol under a semi-honest model and a secure Chebyshev distance calculation protocol under a malicious model are proposed, which can transform the confidential calculation of the Chebyshev distance into the confidential calculation of the inner product of vectors, thereby realizing the confidentiality measurement of file information similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of archival information processing, and particularly relates to a method for measuring the similarity and confidentiality of archival information. Background Art

[0002] With the rapid development of new generation information technologies such as cloud computing and edge computing, the global data volume has increased explosively. Data has become an important strategic resource affecting global competition. However, at the current stage, a vast amount of data is distributed in different organizations and information systems, and cross-departmental, cross-regional, and cross-system data sharing needs to be realized in order to fully utilize the value of data. However, data security and compliance issues pose many challenges to data sharing. Secure multi-party computation (MPC), as the core technology of privacy computing, provides a way to break the deadlock for ensuring the value of data under the premise of security and compliance. It is an interdisciplinary technical system covering many fields such as cryptography, artificial intelligence, and blockchain.

[0003] Secure multi-party computation allows participants to jointly perform a certain operation using their private data confidentially without disclosing their own private data. The earliest secure multi-party computation problem was the millionaire problem proposed by Professor Yao Qizhi, a computer scientist, in 1982. Subsequently, research scholars such as Goldreich conducted in-depth research on it, and the research field of secure multi-party computation has been continuously expanded, including secure data mining, secure computational geometry and set problems, secure scientific computing, secure statistical analysis problems, and secure database query problems, etc. These studies have continuously promoted the development of secure multi-party computation and effectively solved many practical problems.

[0004] "Secure manhattan distance computation" studied the problem of securely computing the Manhattan distance between two points, that is, computing the Manhattan distance MD = |x1 - x2| + |y1 - y2| between two points P(x1, y1) and Q(x2, y2) under privacy protection. Protocol 1 in it uses a coding method and the Goldwasser-Micali public key encryption algorithm to transform the problem into securely computing the Hamming distance between two bit strings; Protocol 2 in it combines another coding method with the Paillier encryption algorithm to securely compute the Manhattan distance between two points and can prevent malicious participants from cheating in key links. However, this method has high computational complexity, poor computational efficiency, and a high growth rate corresponding to the increase in the number of elements.

[0005] The Chebyshev distance is an important distance metric. The greater the Chebyshev distance, the greater the difference between file information, which can be used in machine learning tasks such as clustering, classification, and anomaly detection. The Chebyshev distance originated from the movement of the king in chess. The Chebyshev distance between two positions on a chessboard refers to the number of steps the king needs to take to move from one position to another. Since the king can move one square diagonally forward or backward, it can reach the destination square more efficiently. Figure 1 is the Chebyshev distance from all positions on the chessboard to the position (3, 4). The Chebyshev distance is also known as the L∞ distance, which is a metric between two n-dimensional vectors, representing the maximum value of the differences in each dimension of these two vectors. Let point A(x1, y1) and point B(x2, y2), the Chebyshev distance between the two points is defined as

[0006] d(x, y) = max i |x i - y i | = max(|x1 - x2|, |y1 - y2|).

[0007] where x and y represent two n-dimensional vectors respectively, and x i and y i represent their values in the i-th dimension respectively. For example, if there are two two-dimensional vectors x = (1, 2) and y = (3, 5), then the Chebyshev distance between them is: d(x, y) = max(|1 - 3|, |2 - 5|) = 3. In this example, the maximum value of the differences between the two vectors is 3, so the Chebyshev distance between them is 3.

[0008] In the file management system, the Chebyshev distance is mainly used for similarity measurement, classification, and clustering of file data, as well as evaluating the degree of data compression and dimensionality reduction. By representing the characteristics of classified file data as vectors and calculating the Chebyshev distance between different data, the similarity degree between them can be judged, which helps with information retrieval, clustering analysis, anomaly detection, etc. By calculating the Chebyshev distance between files, similar files can be grouped into one category, thus realizing the classification of files; at the same time, the Chebyshev distance metric can also be used for clustering analysis to divide files into different clusters for better management and organization of files. In order to compress and reduce the dimensionality of a large amount of file data to save storage space and improve processing efficiency, the Chebyshev distance can be used as a distance metric method to evaluate the effect of data compression and dimensionality reduction, so as to select the optimal compression and dimensionality reduction method. Calculating the Chebyshev distance while protecting the privacy of classified files has important theoretical significance and application value. Securely calculating the Chebyshev distance can better manage and organize file data, and improve the query, sharing efficiency, and security of classified files. However, there is currently a lack of a secure calculation scheme for the Chebyshev distance against malicious adversaries. Summary of the Invention

[0009] The present invention aims to solve the problems of high computational complexity, poor computational efficiency in the existing document distance measurement scheme without revealing the privacy of classified file data, and the high growth rate corresponding to the increase in the number of elements.

[0010] An encoding method for secure measurement of file information similarity based on the Chebyshev distance. For the first classified file and the second classified file, the information in the classified file is extracted to obtain a feature vector, and then the feature vector is mapped to a two-dimensional vector through dimensionality reduction technology, that is, the information in the first classified file and the second classified file is respectively mapped to a private point on a plane. Let the information in the first classified file and the second classified file be mapped to points S(x1, y1) and T(x2, y2) respectively. Let the universal set U of coordinates = {u1,..., u n}, where u1,..., u n are n consecutive integers, satisfying u1 <... < u n ; Let the points S(x1, y1) and T(x2, y2) satisfy x1, y1, x2, y2 ∈ U;

[0011] Encoding method 1: For the encoding of x1 in S(x1, y1), construct an n-dimensional array A1 = (a 11 ,..., a 1n ) according to x1 and the universal set U. Construction method: Assume x1 = u k , k ∈ [1, n] = {1,..., n}, then make the first k elements of the array 0 and the last n - k elements 1, that is, make a 11=,…,=a 1k =0,a 1(k+1) =,…,=a 1n =1;

[0012] In the same way, the array constructed by y1 is A′1 = (a′ 11 ,…,a′ 1n );

[0013] In the same way, the array constructed by x2 is B1, and the array constructed by y2 is B1′;

[0014] Coding method 2: For the coding of x1 in S(x1, y1), construct an n-dimensional array A2 = (a 21 ,…,a 2n ) according to x1 and the universal set U. The construction method is as follows: Assume x1 = u k , k ∈ [1, n] = {1,…, n}, then make the first k elements of the array be 1 and the last n - k elements be 0, that is, make a 21 =,…,=a 2k =1,a 2(k+1) =,…,=a 2n =0;

[0015] In the same way, the array constructed by y1 is A2′ = (a2′1,…,a′ 2n );

[0016] In the same way, the array constructed by x2 is B2, and the array constructed by y2 is B2′;

[0017] Based on coding method 1 and coding method 2, for the first classified file and the second classified file that respectively have private points S(x1, y1) and T(x2, y2), perform coding:

[0018] First code x1 according to coding method 1, and then code x1 according to coding method 2. Finally, splice the two coding sequences in order to obtain the corresponding vector denoted as A; code y1 in the same coding method to obtain A′;

[0019] First code x2 according to coding method 2, and then code x2 according to coding method 1. Finally, splice the two coding sequences in order to obtain the corresponding vector denoted as B; code y1 in the same coding method to obtain B′;

[0020] Or,

[0021] First code x1 according to coding method 2, and then code x1 according to coding method 1. Finally, splice the two coding sequences in order to obtain the corresponding vector denoted as A; code y1 in the same coding method to obtain A′;

[0022] First, encode x2 according to encoding method 1, and then encode x2 according to encoding method 2. Finally, splice the two encoding sequences to obtain the corresponding vector denoted as B; encode y1 using the same encoding method to obtain B'.

[0023] The method for measuring the confidentiality of the similarity of file information based on the Chebyshev distance includes the following steps:

[0024] S100. For the first classified file and the second classified file, use the encoding method for measuring the confidentiality of the similarity of file information based on the Chebyshev distance described above for encoding. For the convenience of representation, denote the vectors corresponding to x1 and y1 in S(x1, y1) of the first classified file as F = (a 11 , …, a 1n , a 21 , …, a 2n ) and G = (a1′1, …, a1′ n , a2′1, …, a2′ n ), and denote the vectors corresponding to x2 and y2 in T(x2, y2) of the second classified file as P = (b 11 , …, b 1n , b 21 , …, b 2n ) and Q = (b1′1, …, b1′ n , b2′1, …, b2′ n ); The first classified file runs the NTRU encryption scheme to generate a public key / private key pair pk / sk, and sends the public key pk to the second classified file;

[0025] S101. The first classified file encrypts each element in vectors F and G using the public key pk to obtain

[0026] E(F) = (E(a 11 ), …, E(a 1n ), E(a 21 ), …, E(a 2n ))

[0027] E(G) = (E(a1′1), …, E(a1′ n ), E(a2′1), …, E(a2′ n ))

[0028] and sends E(F) and E(G) to the second classified file;

[0029] where E(·) is the ciphertext calculated during the encryption process of the NTRU encryption algorithm;

[0030] S102. The second classified file performs the following calculations using E(F), E(G) and its own vectors P, Q

[0031] W1 = (E(a 11 )b 11 , …, E(a 1n )b 1n , E(a 21 )b 21 , …, E(a 2n )b 2n )

[0032] W2 = (E(a1′1)b1′1, …, E(a1′ n )b1′ n , E(a2′1)b2′1, …, E(a′ 2n )b2′ n )

[0033] Then, randomly permute the 2n elements in W1 and W2 respectively to obtain new vectors, denoted as and and send and to the first classified file;

[0034] S103. The first classified file decrypts and using the private key sk to obtain

[0035]

[0036]

[0037] where D(·) is the plaintext calculated during the decryption process of the NTRU encryption algorithm;

[0038] Calculate z1 = d 11 + … + d 1n + d 21 , + … + d 2n , z2 = d1′1 + … + d1′ n + d2′1, + … + d2′ n , then let f(S, T) = max(z1, z2) and publish f(S, T);

[0039] The obtained Chebyshev distance f(S, T) is the measurement result of the confidentiality of the file information similarity.

[0040] A computer storage medium for measuring the confidentiality of file information similarity based on the Chebyshev distance, wherein a computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the method for measuring the confidentiality of file information similarity based on the Chebyshev distance.

[0041] An information similarity confidentiality measurement device for archives based on Chebyshev distance, the device includes a processor and a memory, and a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the information similarity confidentiality measurement method for archives based on Chebyshev distance.

[0042] An information similarity confidentiality measurement method for archives based on Chebyshev distance, comprising the following steps:

[0043] S200. For the first classified archive and the second classified archive, encode them using the encoding method for information similarity confidentiality measurement of archives based on Chebyshev distance. For the convenience of representation, the vectors corresponding to x1 and y1 in S(x1, y1) corresponding to the first classified archive are respectively denoted as F = (a 11 , …, a 1n , a 21 , …, a 2n ) and G = (a1′1, …, a1′ n , a2′1, …, a2′ n ); the vectors corresponding to x2 and y2 in T(x2, y2) corresponding to the second classified archive are denoted as P = (b 11 , …, b 1n , b 21 , …, b 2n ) and Q = (b1′1, …, b1′ n , b2′1, …, b2′ n ); both parties agree on a hash function Hash(); the first classified archive runs the NTRU encryption scheme to generate a public key / private key pair pk1 / sk1, and sends the public key pk1 to the second classified archive; the second classified archive runs the NTRU encryption scheme to generate a public key / private key pair pk2 / sk2, and sends the public key pk2 to the first classified archive;

[0044] S201. The first classified archive encrypts each element in the vectors F and G with the public key pk1 to obtain

[0045]

[0046]

[0047] and sends and to the second classified archive;

[0048] wherein, is the ciphertext calculated during the encryption process of the NTRU encryption algorithm;

[0049] S202. The second classified archive encrypts each element in the vectors P and Q with the public key pk2 to obtain

[0050]

[0051]

[0052] and send and to the first classified file;

[0053] Among them, is the ciphertext calculated during the encryption process of the NTRU encryption algorithm;

[0054] S203. The first classified file selects random numbers s1 and s2 to calculate h1 = Hash(s1) and h2 = Hash(s2), and further calculates

[0055]

[0056] The first classified file sends h1, h2, W1, and W2 to the second classified file;

[0057] S204. The second classified file selects random numbers t1 and t2 to calculate h3 = Hash(t1) and h4 = Hash(t2), and further calculates

[0058]

[0059] The second classified file sends h3, h4, W3, and W4 to the first classified file;

[0060] S205. The first classified file decrypts W3 and W4 with the private key sk1 to obtain w3 = D(W3) and w4 = D(W4), and sends w3 and w4 to the second classified file;

[0061] Among them, D(·) is the plaintext calculated during the decryption process of the NTRU encryption algorithm;

[0062] S206. The second classified file decrypts W1 and W2 with the private key sk2 to obtain w1 = D(W1) and w2 = D(W2), and sends w1 and w2 to the first classified file;

[0063] S207. The first classified file calculates z1 = w1 / s1 and z2 = w2 / s2, and sends z1 and z2 to the second classified file;

[0064] S208. The second classified file calculates z3 = w3 / t1 and z4 = w4 / t2, and sends z3 and z4 to the first classified file;

[0065] S209. Verify whether Hash(w3 / z3) = h3 and Hash(w4 / z4) = h4 hold for the first classified file; if not, reject z3 and z4; if so, let f1(S, T) = max(z3, z4) and publish it.

[0066] S210. Verify whether Hash(w1 / z1) = h1 and Hash(w2 / z2) = h2 hold for the second classified file; if not, reject z1 and z2; if so, let f2(S, T) = max(z1, z2) and publish it.

[0067] S211. If f1(S, T) = f2(S, T), it proves that the calculation result is correct, and the obtained Chebyshev distance f(S, T) is the measurement result of the confidentiality of the file information similarity.

[0068] A computer storage medium for measuring the confidentiality of file information similarity based on the Chebyshev distance, wherein a computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the method for measuring the confidentiality of file information similarity based on the Chebyshev distance.

[0069] A device for measuring the confidentiality of file information similarity based on the Chebyshev distance, the device includes a processor and a memory, and a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method for measuring the confidentiality of file information similarity based on the Chebyshev distance.

[0070] Beneficial effects:

[0071] (1) The present invention first proposes a protocol for securely calculating the Chebyshev distance under a semi - honest model, analyzes the correctness of the protocol, and proves the security under the semi - honest model by using the simulation paradigm. At the same time, aiming at the possible malicious behaviors in the semi - honest model protocol, with the help of cryptographic tools such as digital commitments and the method of jointly participating in decryption by both parties, the present invention proposes a protocol for securely calculating the Chebyshev distance under the malicious model; and analyzes the correctness of the protocol under the malicious model, and proves that the protocol is secure under the malicious model by using the ideal - real paradigm method. Therefore, the present invention can perform similarity measurement, classification and clustering, improving the security of classified file query and sharing. In addition, the present invention can also resist quantum attacks.

[0072] (2) The present invention is based on the NTRU encryption algorithm and adapts the vector coding method. Based on the present invention, the execution time for measuring the document distance without revealing the privacy of classified file data is shorter, with better operation efficiency and lower computational complexity, especially for the case where the number of elements increases, and the growth rate of the present invention is lower. Description of the drawings

[0073] Figure 1 Schematic diagram of the Chebyshev distance from the king's position to other positions on the chessboard.

[0074] Figure 2 Flowchart of Protocol 1 of the present invention.

[0075] Figure 3 Flowchart of Protocol 2 of the present invention.

[0076] Figure 4 Variation law of the execution time of Protocol 1 in the present invention and Protocol 1 in "Secure manhattan distance computation" with the increase of the n value.

[0077] Figure 5 Variation law of the execution time of Protocol 2 in the present invention and Protocol 2 in "Secure manhattan distance computation" with the increase of the n value.

[0078] Figure 6 Relationship between the latency time of Protocol 1 in the present invention and Protocol 1 in "Secure manhattan distance computation" and n.

[0079] Figure 7 Relationship between the latency time of Protocol 2 in the present invention and Protocol 2 in "Secure manhattan distance computation" and n. Detailed implementation manners

[0080] Detailed implementation manner 1: This implementation manner is a method for measuring the confidentiality of the similarity of archive information based on the Chebyshev distance. This implementation manner provides two methods for measuring the confidentiality of the similarity of archive information based on the Chebyshev distance.

[0081] Before giving a specific description, first, the relevant content involved in the present invention will be described.

[0082] (A), NTRU encryption algorithm

[0083] The NTRU encryption algorithm is considered to be strongly related to the problem of finding the shortest vector in a lattice. It can resist quantum attacks, and has fast speed and high security. The NTRU algorithm is defined on the ring R = Z[X] / (X N -1), and its elements can be represented in the form of polynomials or vectors. For example, F ∈ R,

[0084] Key Generation: Generate an integer triple (N, p, q) (public) according to the given security parameter k, satisfying: N is a prime number, gcd(p, q) = 1 and q >> p; then generate 4 sets of integer coefficient polynomials of degree N - 1, L f , L g , L r , L m , satisfying: the set L m selected for the plaintext m L f = L(d f , d f -1), L g = L(d g , d g ); L r = L(d r , d r ), that is, 3 positive integers d f , d g , d r can determine the sets L f , L g , L r , where: L(d1, d2) = {H ∈ R: H has d1 coefficients = 1, d2 coefficients = -1, and the remaining coefficients = 0}.

[0085] First, randomly select polynomials f ∈ L f , g ∈ L g , and f has inverses modulo p and modulo q, denoted as F p and F q , respectively. Then F p *f = 1 mod p, F q *f = 1 mod q. Calculate the public key h = p·F q *g (mod q) and the private key (f, F p ), and F p , F q , g are stored secretly as the private key. gcd represents the greatest common divisor, and mod represents the modulo function.

[0086] Encryption: Encode the plaintext as a plaintext polynomial M ∈ L m , select a random polynomial r ∈ L r , and calculate the ciphertext C = E(M) = r*h + M (mod q).

[0087] Decryption: First calculate the intermediate polynomial d = f*C (mod q), where the coefficients of d belong to the set Then calculate the plaintext polynomial M = D(C) = F p *d (mod p), and the plaintext data can be obtained according to the plaintext encoding rule.

[0088] Additive homomorphism:

[0089] E(M1)+E(M2)=[r1*h+M1(mod q)]+[r2*h+M2(mod q)]

[0090] =(r1+r2)*h+(M1+M2)(mod q)

[0091] =E(M1+M2)

[0092] (B) Security definition under the honest model

[0093] Semi-honest model: In a semi-honest protocol, participants will fully follow every step of the protocol, will not provide false information, will not stop the protocol midway, and will not collude with other participants to attack the protocol. However, they will record the public information in the protocol in order to attempt to deduce the information of other participants. For details, please refer to "Symmetric cryptographic solution to Yao’s Millionaires’ problem and an evaluation of secure multiparty computations".

[0094] Suppose the two parties participating in the secure computation are Steven and Tom respectively. Steven has data x and Tom has data y. They hope to cooperatively compute the probabilistic polynomial-time function f(x,y)=(f1(x,y),f2(x,y)) while ensuring the privacy of x and y. Let π be the protocol for computing f. denotes the sequence of information obtained by Steven during the execution of the protocol, where r1 represents the random number selected by Steven; denotes the i-th information received by Steven. After executing the protocol π, the output result obtained by Steven is f1(x,y). The sequence of information obtained by Tom can be defined similarly.

[0095] Definition 1: Suppose f(x,y)=(f1(x,y),f2(x,y)) is a two-party computation function, x and y represent the variables input by the two participants, f1(x,y) and f2(x,y) represent the function values obtained by the two participants respectively, and π is a two-party computation protocol for computing f. If there exist probabilistic polynomial-time algorithms S1 and S2 such that:

[0096]

[0097]

[0098] Then it is said that the protocol π securely computes the function f. Among them represents computational indistinguishability.

[0099] (C) Security Definition under the Malicious Model

[0100] The malicious model is a more practical secure multi-party computation model. To prove that a protocol is secure under the malicious model, it must be proven that it satisfies the security conditions under the malicious model, that is, an ideal protocol with the help of a trusted third party: For example, Alice and Bob have data x and y, and they use a trusted third party (Trusted Third Party, TTP) to compute the function f(x, y) = (f1(x, y), f2(x, y)). They obtain f1(x, y) and f2(x, y) respectively without leaking x and y. The ideal protocol is described as follows:

[0101] (1) Send the input information to the TTP: Honest participants will provide the correct data x or y to the TTP, and malicious participants may not execute the protocol or provide false data x' or y' to the TTP.

[0102] (2) The TTP sends the result data to Alice: After receiving the data (x, y), the TTP independently computes f(x, y) and sends f1(x, y) to Alice. Otherwise, it sends a special symbol ⊥ to Alice to indicate the termination of the protocol.

[0103] (3) The TTP sends the result data to Bob: If Alice is a malicious participant and terminates the protocol after receiving the data f1(x, y) in the second step, in this case, the TTP sends a special symbol ⊥ to Bob. Otherwise, it sends f2(x, y) to Bob.

[0104] Because the participants can only obtain the result f i (x, y) from the TTP and cannot obtain any other information, the ideal protocol is the most secure protocol. If a practical protocol can have the same security as the ideal protocol, then this practical protocol is secure.

[0105] Let F: {0, 1} * × {0, 1} * → {0, 1} * × {0, 1} * be a probabilistic polynomial-time function, and F1(x, y), F2(x, y) represent the first and second elements of F(x, y). Let be a pair of probabilistic polynomial-time algorithms representing the strategies of the participants in the ideal protocol. If during the execution of the protocol, at least one B i (i ∈ {1, 2}) satisfies B i(u, z, r) = u, B i (u, z, r, v) = v, where u is the input of B i z is its auxiliary input, r is the randomly selected number by it, and v is the partial output F obtained from a trusted third party i (), then such is acceptable. In the ideal model, the participants jointly calculate F(x, y) according to the auxiliary information z and with the strategy The process is Defined as the adversary uniformly selects a random number r, and let

[0106]

[0107] where γ(x, y, z, r) is defined as follows:

[0108] · If Alice is honest, then there is

[0109] γ(x, y, z, r) = (f1(x, y′), B2(y, z, r, f2(x, y′))),

[0110] where y′ = B2(y, z, r).

[0111] · If Bob is honest,

[0112]

[0113] In both cases, x′ = B1(x, z, r).

[0114] Let П be a two-party protocol for computing F. Are two probabilistic polynomial-time algorithms representing the strategies of the participants in the real model. If at least one A i (i ∈ {1, 2}) is consistent with the strategy specified by П, then is acceptable with respect to П. In particular, this A i Ignores its auxiliary input. When the input is (x, y), its auxiliary input is z, and the process of executing the protocol П in the real model with the strategy is denoted as REAL П,A(z) (x, y), defined as the output pair generated by the interaction between A1(x, z) and A2(y, z).

[0115] Definition 2: Security in the malicious model

[0116] If an acceptable strategy pair can be found in the actual protocol There exists an acceptable strategy pair in the ideal model such that

[0117]

[0118] Then the protocol Π securely computes F, where x, y, z ∈ {0, 1} * such that |x| = |y| and |z| = poly(|x|).

[0119] It should be noted that: The security definition in the malicious model implies that in the case of two-party participation in the computation, at least one party must be honest to ensure the feasibility of the protocol. If both parties are malicious, it is impossible to design a secure computation protocol.

[0120] (D), One-way hash function

[0121] A one-way hash function, also known as a one-way Hash function, is a function that transforms an input message string of any length into an output string of a fixed length and it is difficult to obtain the input string from the output string.

[0122] For a function f, if there exists a polynomial-time algorithm A such that A(x) = f(x), but there does not exist a polynomial-time algorithm B such that B(y) = x' ∧ f(x') = y, then f is called a one-way function.

[0123] If a one-way function f satisfies |f(x1)| ≠ |f(x2)| for any |x1| ≠ |x2|, it is called a one-way hash function.

[0124] Properties of one-way hash functions: Given a message m, it is easy to calculate h = hash(m); Given h = hash(m), it is difficult to calculate the message m from h; Given a message m, it is difficult to find another message m' such that hash(m) = hash(m'); It is difficult to find two random messages m and m' such that hash(m) = hash(m'); If m is slightly changed, even if only one bit is changed, the value of hash(m) will change significantly.

[0125] (E), Digital commitment

[0126] Digital commitment is a basic tool in cryptography. It is divided into two stages, the commitment stage and the stage of revealing the commitment.

[0127] (1) In the commitment stage, the committer selects a random number r according to his own message m, calculates the commitment function c = c(m, r) and sends it to the receiver, which is equivalent to putting the message into a sealed envelope. In this stage, it is required that any receiver cannot obtain any information about the committed value m within probabilistic polynomial time. Even if the receiver tries to deceive, he cannot obtain the corresponding information. This property is called hiding.

[0128] (2) In the revealing commitment stage, the committer provides the receiver with a random number r and his own commitment value m. The receiver calculates the commitment function c′ = c(m, r). If c = c′, the original commitment is valid; otherwise, it is rejected. This stage is equivalent to opening the envelope face to face. In this stage, it is required that any committer cannot find m ≠ m′ and r ≠ r′ such that c(m, r) = c(m′, r′) within probabilistic polynomial time. This property is called bindingness.

[0129] Example 1. Secure Computation of Chebyshev Distance Protocol under the Semi - honest Model

[0130] Problem description: Suppose Steven and Tom map the classified file data to the private points S(x1, y1) and T(x2, y2) on the plane. They want to securely compute the Chebyshev distance between the two points, that is, securely compute the function f(S, T)=max(|x1 - x2|, |y1 - y2|) for analysis or judgment, and the computation will not reveal any information about their respective private points.

[0131] In the following text, the points S(x1, y1) and T(x2, y2) are regarded as two two - dimensional vectors.

[0132] To securely compute the Chebyshev distance, two encoding methods are designed as follows:

[0133] Suppose the universal set U = {u1,…,u n}, where u1,…,u n are n consecutive integers satisfying u1 < … < u n . Let the points S(x1, y1) and T(x2, y2) satisfy x1, y1, x2, y2 ∈ U.

[0134] Coding method 1: Taking the encoding of x1 in S(x1, y1) as an example, the encoding method is introduced as follows. According to x1 and the universal set U, construct an n - dimensional array A1=(a 11 ,…,a 1n ), and the construction method is as follows:

[0135] Suppose x1 = u k , k ∈ [1, n]={1,…,n}, then let the first k elements of the array be 0, and the last n - k elements be 1, that is, let a 11 =,…,=a 1k = 0, a 1(k+1) =,…,=a 1n = 1.

[0136] For y1, construct its corresponding array in the same way, denoted as A′1=(a′ 11 ,…,a′ 1n ).

[0137] Example: Assume the universal set is U = {1, 2, 3, 4, 5}. For the point S(2, 4), according to Encoding Method 1, the abscissa x1 = 2 is encoded as A1 = (0, 0, 1, 1, 1). The ordinate y1 = 4 is encoded as A′1 = (0, 0, 0, 0, 1).

[0138] Similarly, for the point T(1, 5), according to Encoding Method 1, x2 = 1 can be encoded as B1 = (0, 1, 1, 1, 1), and y2 = 5 can be encoded as B′1 = (0, 0, 0, 0, 0).

[0139] Encoding Method 2: Still taking the encoding of x1 in S(x1, y1) as an example to introduce the encoding method. Construct an n-dimensional array A2 = (a 21 , …, a 2n ) according to x1 and the universal set U. The construction method is as follows:

[0140] Assume x1 = u k , k ∈ [1, n] = {1, …, n}, then let the first k elements of the array be 1 and the last n - k elements be 0, that is, let a 21 =, …, = a 2k = 1, a 2(k+1) =, …, = a 2n = 0.

[0141] Construct its corresponding array for y1 in the same way, denoted as A′2 = (a′ 21 , …, a′ 2n ).

[0142] Example: Assume the universal set is U = {1, 2, 3, 4, 5}. For the point S(2, 4), according to Encoding Method 2, the abscissa x1 = 2 is encoded as A2 = (1, 1, 0, 0, 0). The ordinate y1 = 4 is encoded as A′2 = (1, 1, 1, 1, 0).

[0143] Similarly, for the point T(1, 5), according to Encoding Method 2, x2 = 1 can be encoded as B2 = (1, 0, 0, 0, 0), and y2 = 5 can be encoded as B′2 = (1, 1, 1, 1, 1).

[0144] Encoding Method 3: Assume the universal set is U = {u1, …, u n}, Steven and Tom respectively have private points S(x1, y1) and T(x2, y2), satisfying x1, y1, x2, y2 ∈ U. Under the universal set U, first encode x1 (or x2) according to Encoding Method 1 (or Encoding Method 2) to get (a 11 , …, a 1n ) (or (b 11 , …, b 1n), and then encode x1 (or x2) according to encoding method 2 (or encoding method 1) to obtain (a 21 , …, a 2n )(or (b 21 , …, b 2n ). Finally, the corresponding vector is denoted as

[0145] A = (a 11 , …, a 1n , a 21 , …, a 2n )(or B = (b 11 , …, b 1n , b 21 , …, b 2n ).

[0146] For y1 and y2, use the same encoding method as above to encode, and obtain the vector

[0147] A′ = (a′ 11 , …, a′ 1n , a′ 21 , …, a′ 2n )(or B′ = (b′ 11 , …, b′ 1n , b′ 21 , …, b′ 2n ).

[0148] Example: Let the universal set be U = {1, 2, 3, 4, 5}, taking the encoding of x1 and x2 as an example.

[0149] It should be noted that: the order of the two encoding methods adopted by x1 is opposite to the order of the two encoding methods adopted by x2, that is, when x1 adopts encoding method 1 - encoding method 2, x2 adopts encoding method 2 - encoding method 1.

[0150] First, encode x1 = 2 as (0, 0, 1, 1, 1) according to encoding method 1, then encode x1 = 2 as (1, 1, 0, 0, 0) according to encoding method 2, and then let A = (0, 0, 1, 1, 1, 1, 1, 0, 0, 0).

[0151] First, encode x2 = 1 as (1, 0, 0, 0, 0) according to encoding method 2, then encode x2 = 1 as (0, 1, 1, 1, 1) according to encoding method 1, and then let B = (1, 0, 0, 0, 0, 0, 1, 1, 1, 1).

[0152] Regarding the calculation of |x1 - x2| and |y1 - y2|, the following conclusions hold:

[0153] Proposition 1: The value of |x1 - x2| is equal to the inner product of vectors A and B, and the value of |y1 - y2| is equal to the inner product of vectors A' and B'. That is, the following formula holds

[0154]

[0155] Based on the above calculation principle and combined with the NTRU encryption algorithm, the present invention designs a method for measuring the confidentiality of file information similarity based on the Chebyshev distance in the semi - honest model, as Figure 2 shown

[0156]

[0157]

[0158] Correctness Analysis

[0159] In step (2) of the protocol, combined with the additive homomorphism of the NTRU encryption algorithm, it can be known that

[0160]

[0161] Similarly, E(a i ′ j )b i ′ j =E(a i ′ j b i ′ j ). It can also be known that

[0162] D(W1)=(a 11 b 11 ,…,a 1n b 1n ,a 21 b 21 ,…,a 2n b 2n )

[0163] D(W2)=(a1′1b1′1,…,a1′ n b1′ n ,a2′1b2′1,…,a2′ n b2′ n )

[0164] Since and are obtained by randomly permuting the elements in W1 and W2, therefore, and are also obtained by randomly permuting the corresponding elements in D(W1) and D(W2). That is, there is

[0165]

[0166]

[0167] According to Proposition 1, it is proved that y1 = |x1 - x2| and y2 = |y1 - y2|. That is, Protocol 1 is correct.

[0168] Security Analysis

[0169] Regarding the security of the protocol, the following conclusion holds:

[0170] Theorem 1: Protocol 1 can securely compute the Chebyshev distance.

[0171] Proof: This theorem is proved by constructing simulators S1 and S2 that satisfy equations (1) and (2).

[0172] The process of S1's simulation is as follows:

[0173] (1) After receiving the inputs (x1, f1(x1, x2)) and (y1, f1(y1, y2)), S1 randomly selects x′2 ∈ U and y′2 ∈ U such that f1(x1, x′2) = f1(x1, x2) and f1(y1, y′2) = f1(y1, y2), and encodes x1, x′2, and y1, y′2 according to Coding Method 3 to obtain the corresponding vectors F = (a 11 , …, a 1n , a 21 , …, a 2n ), P′ = (b′ 11 , …, b′ 1n , b′ 21 , …, b′ 2n ), and G = (a′ 11 , …, a′ 1n , a′ 21 , …, a′ 2n ), Q′ = (b″ 11 , …, b″ 1n , b″ 21 , …, b″ 2n ).

[0174] (2) S1 encrypts the vectors F and G to obtain E(F) and E(G), and calculates

[0175] W1′ = (E(a 11 )b′ 11 , …, E(a 1n )b′ 1n , E(a 21 )b′ 21 , …, E(a 2n )b′ 2n )

[0176] W′2 = (E(a′ 11 )b″ 11 , …, E(a′ 1n )b″ 1n , E(a′ 21 )b″ 21 , …, E(a′ 2n )b″ 2n )

[0177] Then, randomly permute the elements in W1′ and W′2 to obtain and

[0178] (3) S1 decrypts and to obtain

[0179]

[0180]

[0181] (4) S1 calculates

[0182] z′1 = d′ 11 + … + d′ 1n + d′ 21 , + … + d′ 2n

[0183] z′2 = d″ 11 + … + d″ 1n + d″ 21 , + … + d″ 2n

[0184] Then let f′(S, T) = max(z′1, z′2).

[0185] During the execution of the protocol, while

[0186]

[0187] and are obtained by Tom through randomly permuting 2n elements in W1 and W2 and sent to Steven. Although Steven has the private key sk to decrypt and but Steven can only know the decrypted vectors and (the elements in the vectors consist of 0 and 1), and cannot know which specific ciphertexts in W1 and W2 decrypt to 0 (or 1). Therefore, there are and Also, since \(f_1(x_1,x'_2)=f_1(x_1,x_2)\) and \(f_1(y_1,y'_2)=f_1(y_1,y_2)\), we have

[0188]

[0189] Next, simulate the execution process of S2.

[0190] (1) After S2 receives the inputs \((x_2,f_2(x_1,x_2))\) and \((y_2,f_2(y_1,y_2))\), randomly select \(x'_1\in U\) and \(y'_1\in U\) such that \(f_2(x'_1,x_2)=f_2(x_1,x_2)\) and \(f_2(y'_1,y_2)=f_2(y_1,y_2)\), and encode \(x'_1,x_2\) and \(y'_1,y_2\) according to encoding method 3 to obtain the corresponding vectors \(F'=(a' 11 ,…,a' 1n ,a' 21 ,…,a' 2n ), P=(b 11 ,…,b 1n ,b 21 ,…,b 2n ), and \(G'=(a'' 11 ,…,a'' 1n ,a'' 21 ,…,a'' 2n ), Q=(b' 11 ,…,b' 1n ,b' 21 ,…,b' 2n ).

[0191] (2) S2 encrypts the vectors \(F'\) and \(G'\) to obtain \(E(F')\) and \(E(G')\), and calculates

[0192] W1'=(E(a' 11 )b 11 ,…,E(a' 1n )b 1n ,E(a' 21 )b 21 ,…,E(a' 2n )b 2n ),

[0193] W'2=(E(a'' 11 )b' 11 ,…,E(a'' 1n )b' 1n ,E(a'' 21 )b' 21 ,…,E(a'' 2n )b' 2n ),

[0194] Randomly permute the elements in W1′ and W′2 to obtain and

[0195] (3) S2 decryption and to obtain

[0196]

[0197] (4) S2 calculation

[0198] z′1 = d′ 11 + … + d′ 1n + d′ 21 , + … + d′ 2n , z′2 = d″ 11 + … + d″ 1n + d″ 21 , + … + d″ 2n .

[0199] Let f′(S,T) = max(z′1,z′2). During the execution of the protocol, while

[0200] S2(x2,f2(x1,x2)) = {x2,E(F′),f2(x′1,x2)},

[0201] S2(y2,f2(y1,y2)) = {y2,E(G′),f2(y′1,y2)}.

[0202] E(F) and E(G) are encrypted by Steven using the NTRU encryption algorithm. Tom does not have the private key. According to the semantic security of the NTRU encryption algorithm, for Tom, there is and Since f2(x1′,x2) = f2(x1,x2) and f2(y1′,y2) = f2(y1,y2), so

[0203]

[0204]

[0205] Therefore, Protocol 1 can securely compute the Chebyshev distance.

[0206] Example 2. Secure Chebyshev Distance Computation Protocol under the Malicious Model

[0207] In Protocol 1 under the semi - honest model, the malicious behaviors that a malicious party may implement include:

[0208] (1) In step (2) of Protocol 1, participant Tom may provide false ciphertext to Steven.

[0209] (2) If one of the participants has a public-private key pair while the other can only passively wait for the result, there may be a situation where the party with the public-private key pair tells the other party a wrong result. In step (3) of Protocol 1, Steven may tell Tom a wrong result, preventing Tom from obtaining the correct result.

[0210] Solution ideas:

[0211] (1) Both parties need to have public-private key pairs. In the protocol, Steven and Tom decrypt the calculation results respectively to obtain the Chebyshev distance. Eventually, both parties can calculate the correct results separately.

[0212] (2) Combine the NTRU encryption scheme and the digital commitment method to design and construct a secure computing protocol in an anti-cheating scenario, ensuring that the two participants obtain the same calculation result (if one party attempts to cheat, the other party can detect it). The so-called digital commitment can be simply understood as a two-phase protocol involving two parties, namely the committer and the receiver. Through this protocol, the committer can bind himself to a number. This binding should satisfy confidentiality and determinacy. Confidentiality means that after the committer makes a commitment, the receiver cannot obtain any knowledge about the number committed by the committer; determinacy means that the receiver only accepts the legal number sent by the committer. If the committer cheats, the receiver can detect it and refuse to accept.

[0213] Based on the foregoing content, the present invention designs a method for measuring the confidentiality of the similarity of archive information based on the Chebyshev distance in a malicious model, as Figure 3 shown.

[0214]

[0215]

[0216]

[0217] Correctness analysis

[0218] (1) In steps (3) and (4) of the protocol, combined with the additive homomorphism of the NTRU encryption algorithm, it can be known that

[0219]

[0220]

[0221]

[0222]

[0223] (2) In steps (5) and (6) of the protocol, Steven decrypts W3 and W4 to obtain and Tom decrypts W1 and W2 to obtain and From Steven's calculation in step (7), and From Tom's calculation in step (8), and By Proposition 1, the correctness is proven.

[0224] Security Analysis

[0225] In Protocol 2, Steven and Tom decrypt respectively to obtain the calculation results, thus avoiding the malicious behavior of one participant tampering with the results. To prevent the malicious behavior of the participants, in step (3) of the protocol, Steven has sent the Hash values h1 and h2 of the random numbers s1 and s2 to Tom, thus making a commitment to W1 and W2; in step (4) of the protocol, Tom has sent the Hash values h3 and h4 of the random numbers t1 and t2 to Steven, thus making a commitment to W3 and W4. According to the one-way property of the one-way hash function, that is, it is easy to calculate the hash value from the message, but it is impossible to reverse-calculate the message from the hash value, forcing the participants to finally only announce the true results. If participant Steven wants to modify the results, then Tom will discover Steven's deception behavior when verifying Hash(w1 / z1) = h1 and Hash(w2 / z2) = h2 in step (10) of the protocol; if participant Tom wants to modify the results, then Steven will discover Tom's deception behavior when verifying Hash(w3 / z3) = h3 and Hash(w4 / z4) = h4 in step (9) of the protocol. Therefore, the security of Protocol 2 is guaranteed through digital commitment. In addition, the participants decrypt respectively to obtain the calculation results and verify at the end of the protocol, ensuring that the two participants obtain the same calculation results.

[0226] The following uses the ideal-real paradigm to prove the security of Protocol 2.

[0227] Theorem 2: Protocol 2 (denoted by Π) is secure in the malicious model.

[0228] Proof By Definition 2, for protocol Π to securely compute function F, the participating parties need to find an acceptable strategy pair in the actual protocol and the strategy pair in the ideal model to be indistinguishable. Then it can be proven that the protocol is secure.

[0229] At least one of A1 and A2 in the protocol is honest, so there are two cases.

[0230] (1) A1 is honest and A2 is dishonest.

[0231] When executing protocol Π with A1 being honest, there exists:

[0232]

[0233] When A1 is the honest party, he will execute the protocol honestly, and then B1 is determined. What needs to be proved is that A2 in the actual protocol is indistinguishable from B2 in the ideal model. Therefore, a strategy pair in the ideal model needs to be found whose output is indistinguishable from that in the actual model During the execution of the protocol, the actual executor is A2, and the correctness of the protocol needs to be verified according to the behavior A2(T) of A2.

[0234] (1) In the actual protocol, since A1 is the honest party, B1 is also honest and will send the real information S to the TTP by imitating A1's behavior.

[0235] (2) In the actual protocol, because A2 is dishonest, B2 is also dishonest. The information it sends to the TTP depends on B2's strategy, which is the same as A2's strategy. Then the input information that B2 sends to the TTP is A2(T).

[0236] (3) The input information obtained by the TTP is (S, A2(T)), and it calculates F(S, A2(T)).

[0237] (4) B2 gets F(S, A2(T)) from the TTP and uses F(S, A2(T)) to obtain a computation indistinguishable from that obtained by A2 when actually executing the protocol and gives to A2 to get A2's output. Then the simulator B2 executes the protocol according to its own input and the result of the protocol, assuming the input value of the other party that satisfies the result, that is, B2 selects S′ to simulate the protocol and makes F(S′, A2(T)) = F(S, A2(T)). The specific execution process of B2 is as follows:

[0238] ① B2 sends the information required in step (1) of the protocol and to A2;

[0239] ② In step (3) of the protocol, B2 calculates h1′, h2′, and then sends W1′ and W2′ to A2;

[0240] ③ In step (5) of the protocol, B2 decrypts W3′ and W4′ and sends the decryption results w3′ and w4′ to A2;

[0241] ④ In step (7) of the protocol, B2 calculates and publishes z1′ and z′2;

[0242] ⑤ In step (9) of the protocol, verify that Hash(w3′ / z′3) = h3′ and Hash(w4′ / z′4) = h4′, and then publish f1(S′,T) = max(z3′,z′4).

[0243] B2 uses to call A2. Output In this way, we get:

[0244]

[0245] In the protocol, since the protocol uses the NTRU encryption algorithm with additive homomorphic property, therefore The digital commitment guarantees Then there is:

[0246]

[0247] (2) A1 is dishonest and A2 is honest

[0248] Then there are the following two cases:

[0249] (1) If Steven ignores the TTP after obtaining the information, the TTP will send ⊥ to Tom, then:

[0250]

[0251] (2) Otherwise, if Steven publishes the result and passes the proof of digital commitment, then the TTP will send F(A1(S),T) to Tom, then:

[0252]

[0253] Since A2 is honest, A2 will execute the protocol as required, then B2 is determined. What needs to be proved is that A1 in the actual protocol is indistinguishable from B1 in the ideal model, so as to find a strategy pair in the ideal model whose output is indistinguishable from that in the actual model During the execution of the protocol, the actual executor is A1. Therefore, during the proof process, the correctness of the protocol must be verified according to the behavior A1(S) of A1.

[0254] (1) In the actual protocol, A1 is dishonest. Therefore, B1 is also dishonest, and the information it sends to TTP depends on B1's strategy, which is the same as A1's strategy. Then B1 will send A1(S) to TTP.

[0255] (2) In the actual protocol, A2 is honest. Therefore, B2 is also honest, and it sends the true input information T to TTP.

[0256] (3) The input information obtained by TTP is (A1(S), T), and it calculates F(A1(S), T).

[0257] (4) B1 uses F(A1(S), T) obtained from TTP to get and should be computationally indistinguishable from what is obtained by executing the actual protocol with A1, and delivers to A1 to obtain A1's output. Then B1 executes the protocol by assuming the input of the other party that satisfies the result based on its own input and the calculation result, that is, B1 selects T′ to simulate the protocol and makes F(A1(S), T′) = F(A1(S), T). The specific execution process of B1 is as follows:

[0258] ① B1 sends the information required in step (2) of the protocol and to A1;

[0259] ② In step (4) of the protocol, B1 calculates h3′, h4′, and then sends W3′ and W4′ to A1;

[0260] ③ In step (6) of the protocol, B1 decrypts W1′ and W2′ and sends the decryption results w1′ and w2′ to A1;

[0261] ④ In step (8) of the protocol, B1 calculates and publishes z3′ and z′4;

[0262] ⑤ In step (10) of the protocol, it verifies Hash(w1′ / z1′) = h1′ and Hash(w2′ / z′2) = h2′, and then publishes f2(S, T′) = max(z1′, z′2).

[0263] During the process of B1 executing the protocol, there are two situations that may occur:

[0264] (1) If A1 ignores TTP after obtaining the information, then we get:

[0265]

[0266] (2) Otherwise, we get

[0267]

[0268] In either case, the outputs of A2 and B2 in the actual protocol and the ideal model are the same. It suffices to prove that and are computationally indistinguishable. In the protocol, since the NTRU encryption algorithm with additive homomorphic property is adopted in the protocol, the digital commitment guarantees that Then we have:

[0269]

[0270] In summary, for any acceptable probabilistic polynomial-time strategy pair in the actual protocol there exists an acceptable probabilistic polynomial-time strategy pair in the ideal model such that and are computationally indistinguishable. Therefore, Protocol 2 is secure in the malicious model.

[0271] Performance Analysis and Comparison of the Protocol

[0272] (a) Computational Complexity Analysis

[0273] Protocol 1 in "Secure manhattan distance computation" uses the Goldwasser-Micali encryption algorithm for encryption and decryption. One of the participants encrypts 2n times (n is the number of elements in the universal set U) and decrypts 2n times; the other participant needs to perform 2n encryption operations and 2n modular multiplication operations. Encryption once requires 2 modular multiplication operations, and decryption once requires lgp modular multiplication operations. Therefore, Protocol 1 therein needs to perform 10n + 2nlgp modular multiplication operations.

[0274] Protocol 2 in "Secure manhattan distance computation" uses the Paillier encryption algorithm for encryption and decryption. One of the participants encrypts 4n times (n is the number of elements in the universal set U) and decrypts 1 time; the other participant needs to perform at most 4n modular multiplication operations. Therefore, Protocol 2 therein needs to perform 2(6n + 1) modular multiplication operations.

[0275] Computational complexity of Protocol 1 in the embodiments of the present invention: In Protocol 1, the NTRU encryption algorithm is adopted. Steven needs to encrypt 4n times and decrypt 4n times (n is the number of elements included in the universal set U). Since the computational speed of multiplying by 0 or 1 is extremely fast, the time required for the calculation of multiplying by 0 or 1 can be ignored. Encrypting once using the NTRU encryption algorithm requires 1 convolution operation, and decrypting once requires 2 convolution operations. Therefore, Protocol 1 requires 12n convolution operations.

[0276] Computational complexity of Protocol 2 in the embodiments of the present invention: In Protocol 2, the NTRU encryption algorithm is adopted. Steven needs to encrypt 4n times and decrypt 2 times (n is the number of elements included in the universal set U), and Tom needs to encrypt 4n times and decrypt 2 times. Since the hashing operation speed is extremely fast, the time taken for the hashing operation can be ignored. Therefore, Protocol 2 requires 8n + 8 convolution operations.

[0277] (b) Communication complexity

[0278] The communication complexity of the protocol is measured by the number of communication rounds.

[0279] Both Protocol 1 and Protocol 2 in "Secure manhattan distance computation" require 2 rounds of communication.

[0280] Communication complexity of Protocol 1 in the embodiments of the present invention: In Protocol 1, Steven and Tom need to conduct 2 rounds of communication.

[0281] Communication complexity of Protocol 2 in the embodiments of the present invention: In Protocol 2, in order to further improve security by adopting the digital commitment method and the method of each of the participating parties decrypting the ciphertext to obtain the calculation result and verifying the results of both parties, Steven and Tom need to conduct 5 rounds of communication.

[0282] As shown in Table 1, for solving the problem of securely computing the Chebyshev distance in the semi - honest model (where n is the number of elements included in the universal set U, that is, the cardinality of the universal set U), the computational complexity of Protocol 1 in the embodiments of the present invention is lower than that of Protocol 1 in "Securemanhattan distance computation", and the computational efficiency is higher. For solving the problem of securely computing the Chebyshev distance in the malicious model, the computational complexity of Protocol 2 in the embodiments of the present invention is lower than that of Protocol 2 in "Secure manhattandistance computation", and the computational efficiency is higher. In addition, due to the adoption of the NTRU encryption algorithm, Protocol 1 and Protocol 2 in the embodiments of the present invention can resist quantum attacks.

[0283] Table 1 Performance comparison

[0284]

[0285] "1" in the above table represents the present invention, and "2" in the above table represents "Secure manhattan distancecomputation".

[0286] Experimental simulation

[0287] To further evaluate the efficiency of the protocol of the present invention, experimental simulations were carried out on Protocol 1 and Protocol 2 using the Python language on the PyCharm platform and compared with existing solutions.

[0288] Experimental environment: Window10 64-bit system, Intel(R) Core(TM) i5-8400 CPU @ 2.80GHz, 16GB RAM.

[0289] Experimental method: Randomly select two points S and T and set the number of elements contained in the universal set U to n. In the experiment, n takes different values in turn (n = 5, 7,..., 23), and 1000 simulation experiment tests are carried out for each n, and the average value of the protocol execution time is statistically calculated.

[0290] Experimental parameter settings: In the experiment, the encryption key lengths of the Goldwasser-Micali encryption algorithm, Paillier encryption algorithm, and NTRU are 512 bits, and the length of the randomly selected number is 64 bits.

[0291] Figure 4 Describes the variation law of the execution time of Protocol 1 (Protocol 1) in the present invention and Protocol 1 (

[27] -Protocol 1) in "Secure manhattan distancecomputation" with the increase of the n value. As Figure 4 shown, in the semi-honest model, the execution times of Protocol 1 and Protocol 1 in "Secure manhattan distance computation" increase with the increase of the number n of elements contained in the universal set U. The execution time of Protocol 1 shows a linear growth. Compared with Protocol 1 in "Secure manhattan distance computation", under the condition that the number of elements n is the same, the execution time of Protocol 1 in the present invention is shorter, the growth rate is lower, and it has better operation efficiency.

[0292] Figure 5Describes the variation law of the execution time of Protocol 2 in the present invention and Protocol 2 in "Secure manhattan distancecomputation" (

[27] - Protocol 2) with the increase of the value of n. As Figure 5 shown, the execution time of Protocol 2 under the malicious model and Protocol 2 in "Secure manhattan distance computation" increases with the increase of the number n of elements contained in the universal set U. The execution time of Protocol 2 shows a linear growth. Compared with Protocol 2 in "Secure manhattan distance computation", under the condition that the number n of elements is the same, the execution time of Protocol 2 in the present invention is shorter, the growth rate is lower, and it has better operation efficiency.

[0293] Next, a communication experiment is carried out to further evaluate the performance of the proposed scheme of the present invention. A simulation experiment is carried out using a Python program (bandwidth is 100 Mbps) on the Pycharm platform to determine the possible delay time when executing Protocol 1 and Protocol 2. In actual situations, the delay time between different networks will be different, which will affect the performance of the protocol operation. This kind of influencing factor is not considered in the performance evaluation. The communication experiment results of Protocol 1 are as Figure 6 shown, and the communication results of Protocol 2 are as Figure 7 shown.

[0294] The experimental results show that the delay times of Protocol 1 and Protocol 2 in the present invention are generally low, increase with the increase of the number n of elements contained in the universal set U, show a linear growth, and the growth rate is low, and the communication efficiency is high; under the condition that the number n of elements is certain, the delay time of Protocol 2 under the malicious model is slightly greater than the delay time of Protocol 1 under the semi - honest model.

[0295] The Chebyshev distance is an important distance metric. In the archive management system, calculating the Chebyshev distance without revealing the privacy of confidential archive data can perform similarity measurement, classification and clustering, and improve the security of confidential archive query and sharing. The present invention proposes a secure Chebyshev distance calculation protocol under the semi - honest model based on the NTRU encryption algorithm with additive homomorphic property and a vector coding method, and on this basis, proposes a secure Chebyshev distance calculation protocol under the malicious model. While ensuring fairness, it can also effectively prevent attacks from malicious participants, and at the same time proves the security of the protocol through the ideal - real paradigm. Compared with the existing schemes, the protocol proposed by the present invention is more efficient and has practical value. Specific Embodiment 2:

[0297] This embodiment is a computer storage medium for the secure measurement of the similarity of file information based on the Chebyshev distance. This embodiment provides two computer storage media for the secure measurement of the similarity of file information based on the Chebyshev distance.

[0298] Example 3: A computer storage medium for the secure measurement of the similarity of file information based on the Chebyshev distance.

[0299] Stored in the computer storage medium for the secure measurement of the similarity of file information based on the Chebyshev distance described in Example 3 is a computer program, which is loaded and executed by a processor to implement the method for the secure measurement of the similarity of file information based on the Chebyshev distance corresponding to the secure Chebyshev distance calculation protocol under the semi - honest model in Example 1.

[0300] Example 4: A computer storage medium for the secure measurement of the similarity of file information based on the Chebyshev distance.

[0301] Stored in the computer storage medium for the secure measurement of the similarity of file information based on the Chebyshev distance described in Example 4 is a computer program, which is loaded and executed by a processor to implement the method for the secure measurement of the similarity of file information based on the Chebyshev distance corresponding to the secure Chebyshev distance calculation protocol under the malicious model in Example 2.

[0302] It should be understood that the storage medium described in this embodiment includes but is not limited to magnetic storage media and optical storage media; the magnetic storage media include but are not limited to RAM, ROM, and other storage media such as hard disks and USB flash drives. Specific Embodiment 3:

[0304] This embodiment is a device for the secure measurement of the similarity of file information based on the Chebyshev distance. This embodiment provides two devices for the secure measurement of the similarity of file information based on the Chebyshev distance.

[0305] Example 5: A device for the secure measurement of the similarity of file information based on the Chebyshev distance.

[0306] The device for the secure measurement of the similarity of file information based on the Chebyshev distance described in Example 5 includes a processor and a memory. Stored in the memory is a computer program, which is loaded and executed by the processor to implement the method for the secure measurement of the similarity of file information based on the Chebyshev distance corresponding to the secure Chebyshev distance calculation protocol under the semi - honest model in Example 1.

[0307] Example 6: A device for the secure measurement of the similarity of file information based on the Chebyshev distance.

[0308] A confidentiality measurement device for the similarity of file information based on the Chebyshev distance described in Embodiment 6 includes a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the confidentiality measurement method for the similarity of file information based on the Chebyshev distance corresponding to the secure calculation of the Chebyshev distance protocol under the malicious model in Embodiment 2.

[0309] It should be understood that the device described in this embodiment includes, but is not limited to, a device including a processor and a memory, and may also include other devices corresponding to units or modules with information collection, information interaction, and control functions. For example, the device may also include a signal acquisition device, etc. The device includes, but is not limited to, a PC, a workstation, a mobile device, etc.

[0310] The present invention may also have various other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.

Claims

1. A coding method for archival information similarity confidentiality measurement based on Chebyshev distance, characterized in that: For the first confidential file and the second confidential file, the information in the confidential file is feature extracted to obtain a feature vector, and then the feature vector is mapped into a two-dimensional vector by using the dimensionality reduction technology, that is, the information in the first confidential file and the second confidential file is mapped to a private point on a plane, and the information in the first confidential file and the second confidential file is respectively mapped to points S(x1,y1) and T(x2,y2), and the full set of coordinates U = {u1,…,u n }, where u1,…,u n are n consecutive integers, satisfying u1<… n ; Let points S(x1,y1) and T(x2,y2) satisfy x1,y1,x2,y2∈U;​ Coding method 1: For the encoding of x1 in S(x1,y1), construct an n-dimensional array A1=(a 11 ,…,a 1n ), construction method: Assume x1 = u k , k∈[1,n]={1,…,n}, then let the first k elements of the array be 0 and the last nk elements be 1, that is, let a 11 =,…,=a 1k =0,a 1(k+1) =,…,=a 1n =1; In the same way, the array constructed by y1 is A′1=(a′ 11 ,…,a′ 1n ); In the same way, the array constructed by x2 is B1, and the array constructed by y2 is B′1; Coding method 2: For the encoding of x1 in S(x1,y1), construct an n-dimensional array A2=(a 21 ,…,a 2n ), construction method: Assume x1 = u k , k∈[1,n]={1,…,n}, then let the first k elements of the array be 1, and the last nk elements be 0, that is, let a 21 =,…,=a 2k =1,a 2(k+1) =,…,=a 2n =0; In the same way, the array constructed by y1 is A′2=(a′ 21 ,…,a′ 2n ); In the same way, the array constructed by x2 is B2, and the array constructed by y2 is B′2; Based on encoding method 1 and encoding method 2, the first confidential file and the second confidential file have private points S(x1,y1) and T(x2,y2) respectively, which are encoded: First encode x1 according to encoding method 1, then encode x1 according to encoding method 2, and finally concatenate the two encoding sequences to obtain the corresponding vector denoted as A; use the same encoding method to encode y1 to obtain A′; First encode x2 according to encoding method 2, then encode x2 according to encoding method 1, and finally concatenate the two encoding sequences to obtain the corresponding vector denoted as B; use the same encoding method to encode y1 to obtain B′; or, First encode x1 according to encoding method 2, then encode x1 according to encoding method 1, and finally concatenate the two encoding sequences to obtain the corresponding vector denoted as A; use the same encoding method to encode y1 to obtain A′; First encode x2 according to encoding method 1, then encode x2 according to encoding method 2, and finally concatenate the two encoding sequences to obtain the corresponding vector denoted as B; use the same encoding method to encode y1 to obtain B′.

2. A confidentiality measurement method for archival information similarity based on Chebyshev distance, characterized by: The following steps are involved: S100, for the first confidential file and the second confidential file, the encoding method of the file information similarity confidentiality measurement based on Chebyshev distance described in claim 1 is used for encoding, and for the convenience of representation, the vectors corresponding to x1 and y1 in S(x1,y1) corresponding to the first confidential file are recorded as F=(a 11 ,…,a 1n ,a 21 ,…,a 2n ) and G=(a1′1,…,a1′ n ,a2′1,…,a2′ n ), the vector corresponding to x2 and y2 in T(x2,y2) corresponding to the second confidential file is recorded as P = (b 11 ,…,b 1n ,b 21 ,…,b 2n ) and Q=(b1′1,…,b1′ n ,b2′1,…,b2′ n ); The first confidential file runs the NTRU encryption scheme to generate a public key / private key pair pk / sk, and sends the public key pk to the second confidential file; S101. The first confidential file encrypts each element of the vectors F and G with the public key pk, and obtains E(F)=(E(a 11 ),…,And(a 1n ),And(a 21 ),…,And(a 2n )) E(G)=(E(a1′1),…,E(a1′ n ),E(a2′1),…,E(a2′ n )) and send E(F) and E(G) to the second confidential file; Where E(·) is the ciphertext calculated during the encryption process of the NTRU encryption algorithm; S102, the second confidential file uses E(F), E(G) and its own vectors P and Q to perform the following calculations W1=(E(a 11 )b 11 ,…,E(a 1n )b 1n ,E(a 21 )b 21 ,…,E(a 2n )b 2n ) W2=(E(a′ 11 )b′ 11 ,…,E(a′ 1n )b′ 1n ,E(a′ 21 )b′ 21 ,…,E(a′ 2n )b′ 2n ) Then randomly permute the 2n elements in W1 and W2 to obtain new vectors, denoted as and and will and Sent to the first confidential file; S103, the first confidential file is decrypted using the private key sk and get Where D(·) is the plaintext calculated during the decryption process of the NTRU encryption algorithm; Calculate z1 = d 11 +…+d 1n +d 21 ,+…+d 2n , z2=d1′1+…+d1′ n +d2′1,+…+d2′ n , then let f(S,T)=max(z1,z2) and publish f(S,T); The obtained Chebyshev distance f(S,T) is the confidentiality measurement result of the archival information similarity.

3. A computer storage medium for measuring the confidentiality of archival information similarity based on Chebyshev distance, wherein a computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the method for measuring the confidentiality of archival information similarity based on Chebyshev distance as described in claim 2.

4. A device for measuring the confidentiality of archival information similarity based on Chebyshev distance, the device comprising a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the method for measuring the confidentiality of archival information similarity based on Chebyshev distance as described in claim 2.

5. A confidentiality measurement method for archival information similarity based on Chebyshev distance, characterized in that: The following steps are involved: S200, for the first confidential file and the second confidential file, the encoding method of the file information similarity confidentiality measurement based on Chebyshev distance described in claim 1 is used for encoding, and for the convenience of representation, the vectors corresponding to x1 and y1 in S(x1,y1) corresponding to the first confidential file are respectively recorded as F=(a 11 ,…,a 1n ,a 21 ,…,a 2n ) and G=(a1′1,…,a1′ n ,a2′1,…,a2′ n ); the vector corresponding to x2 and y2 in T(x2,y2) corresponding to the second confidential file is recorded as P = (b 11 ,…,b 1n ,b 21 ,…,b 2n ) and Q=(b1′1,…,b1′ n ,b2′1,…,b2′ n ); The two parties agree on a hash function Hash(); The first confidential file runs the NTRU encryption scheme to generate a public / private key pair pk1 / sk1, and sends the public key pk1 to the second confidential file; The second confidential file runs the NTRU encryption scheme to generate a public / private key pair pk2 / sk2, and sends the public key pk2 to the first confidential file; S201, the first confidential file uses the public key pk1 to encrypt each element in the vector F and G, and obtains and will and Send to the second confidential file; in, It is the ciphertext calculated during the encryption process of the NTRU encryption algorithm; S202, the second confidential file uses the public key pk2 to encrypt each element in the vectors P and Q, and obtains and will and Sent to the first confidential file; in, It is the ciphertext calculated during the encryption process of the NTRU encryption algorithm; S203, the first confidential file selects random numbers s1 and s2 to calculate h1=Hash(s1) and h2=Hash(s2), and further calculates The first confidential file sends h1, h2, W1 and W2 to the second confidential file; S204, the second confidential file selects random numbers t1 and t2 to calculate h3 = Hash (t1) and h4 = Hash (t2), and further calculates The second confidential file sends h3, h4, W3 and W4 to the first confidential file; S205, the first confidential file uses the private key sk1 to decrypt W3 and W4 to obtain w3=D(W3) and w4=D(W4), and sends w3 and w4 to the second confidential file; Where D(·) is the plaintext calculated during the decryption process of the NTRU encryption algorithm; S206, the second confidential file uses the private key sk2 to decrypt W1 and W2 to obtain w1=D(W1) and w2=D(W2), and sends w1 and w2 to the first confidential file; S207, the first confidential file calculates z1=w1 / s1 and z2=w2 / s2, and sends z1 and z2 to the second confidential file; S208, the second confidential file calculates z3=w3 / t1 and z4=w4 / t2, and sends z3 and z4 to the first confidential file; S209, the first confidential file verifies whether Hash(w3 / z3)=h3 and Hash(w4 / z4)=h4 are established; if not, z3 and z4 are rejected; if established, f1(S,T)=max(z3,z4) is set and announced; S210, the second confidential file verifies whether Hash(w1 / z1)=h1 and Hash(w2 / z2)=h2 are established; if not, z1 and z2 are rejected; if established, f2(S,T)=max(z1,z2) is set and announced; S211. If f1(S,T)=f2(S,T), it is proved that the calculation result is correct, and the obtained Chebyshev distance f(S,T) is the confidentiality measurement result of the archive information similarity.

6. A computer storage medium for measuring the confidentiality of archival information similarity based on Chebyshev distance, wherein a computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the method for measuring the confidentiality of archival information similarity based on Chebyshev distance as described in claim 5.

7. A device for measuring the confidentiality of archival information similarity based on Chebyshev distance, the device comprising a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the method for measuring the confidentiality of archival information similarity based on Chebyshev distance as described in claim 5.

Citation Information

Patent Citations

  • Multi-channel ship radiation noise feature extraction method based on entropy

    CN113869289A

  • Graph data similarity method based on node-level embedded feature three-dimensional relation reconstruction

    CN114511708A