Method, apparatus, device, and medium for estimating the plaintext similarity of encrypted strings

By modeling plaintext data sets and ciphertext data sets and applying Bayesian statistical models, the distribution of decryption functions is estimated to calculate the plaintext similarity of encrypted strings, which solves the problem of lack of effective methods in the prior art to estimate the similarity of plaintext data before encryption, and realizes the effect of estimating the similarity and association relationship of plaintext data while protecting data privacy.

CN114117487BActive Publication Date: 2025-06-10SHANGHAI PARAVIEW SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111402823.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-06-10
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

There is a lack of effective methods in the prior art to estimate the similarity of plaintext data before encryption using encrypted ciphertext data, especially from the data security perspective.

Method used

By obtaining the plaintext data set and encrypting it using a preset encryption algorithm to obtain the ciphertext data set. Then, the plaintext data set and the ciphertext data set are modeled based on multiple distributions to obtain their respective estimated distributions. Based on Bayesian statistical model, the estimated distribution of the decryption function is estimated based on the estimated distribution of the plaintext data set and the ciphertext data set, and the plaintext similarity between different target encryption strings is finally estimated based on the estimated distribution of the decryption function.

Benefits of technology

The similarity of the plaintext data before encryption is realized through the encrypted ciphertext data, and while protecting the privacy of plaintext data, it is possible to estimate the relationship between multiple plaintext data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114117487B_ABST
    Figure CN114117487B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose a method, apparatus, device, and medium for estimating the plaintext similarity of encrypted strings. The method includes: obtaining a plaintext data set, performing an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set; respectively modeling the plaintext data set and the ciphertext data set based on a multinomial distribution to obtain an estimated distribution corresponding to the plaintext data set and an estimated distribution corresponding to the ciphertext data set; based on a Bayesian statistical model, estimating the estimated distribution corresponding to the decryption function according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; estimating the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function. By adopting the above technical solution, the similarity of the plaintext data before encryption can be estimated through the encrypted ciphertext data, and while protecting the privacy of the plaintext data, the correlation relationship between multiple plaintext data can also be estimated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer technology, and in particular, to a method, device, equipment and medium for estimating the plaintext similarity of encrypted strings. Background Art

[0002] With the rapid development of informatization, people's demand for information security is also increasing. In information security and data confidentiality applications, data encryption is a basic application technology for protecting information.

[0003] In the prior art, there are many methods for judging string similarity, but in terms of data security, the methods of using encrypted ciphertext data as algorithm input and outputting the similarity of plaintext data before encryption are not very rich. Summary of the Invention

[0004] The embodiments of the present invention provide a method, device, equipment and medium for estimating the plaintext similarity of encrypted strings, which can optimize the existing related solutions for estimating plaintext data.

[0005] In a first aspect, the embodiments of the present invention provide a method for estimating the plaintext similarity of encrypted strings, including:

[0006] Obtain a plaintext data set, and perform an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set;

[0007] Based on the multinomial distribution, model the plaintext data set and the ciphertext data set respectively to obtain the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set;

[0008] Based on the Bayesian statistical model, estimate the estimated distribution corresponding to the decryption function according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set;

[0009] Estimate the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function, where the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm.

[0010] In a second aspect, the embodiments of the present invention provide an apparatus for estimating the plaintext similarity of encrypted strings, including:

[0011] An encryption operation module, configured to obtain a plaintext data set and perform an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set;

[0012] An estimated distribution obtaining module, configured to model the plaintext data set and the ciphertext data set respectively based on the multinomial distribution to obtain the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set;

[0013] An estimated distribution calculation module, configured to estimate the estimated distribution corresponding to the decryption function based on the Bayesian statistical model, according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set;

[0014] A plaintext similarity estimation module, configured to estimate the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function, where the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm.

[0015] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for estimating the plaintext similarity of encrypted strings provided by the embodiment of the present invention is implemented.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method for estimating the plaintext similarity of encrypted strings provided by the embodiment of the present invention is implemented.

[0017] In the solution for estimating the plaintext similarity of encrypted strings provided by the embodiment of the present invention, first, a plaintext data set is obtained, and the plaintext data set is encrypted using a preset encryption algorithm to obtain a ciphertext data set; then, the plaintext data set and the ciphertext data set are respectively modeled based on the multinomial distribution to obtain the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; then, based on the Bayesian statistical model, the estimated distribution corresponding to the decryption function is estimated according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; finally, the plaintext similarity between different target encrypted strings is estimated according to the estimated distribution corresponding to the decryption function, where the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm. By adopting the above technical solution, the similarity of the plaintext data before encryption can be estimated through the encrypted ciphertext data, and while protecting the privacy of the plaintext data, the correlation between multiple plaintext data can also be estimated. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic flowchart of a method for estimating the plaintext similarity of encrypted strings provided by an embodiment of the present invention;

[0019] Figure 2 It is a schematic flowchart of another method for estimating the plaintext similarity of encrypted strings provided by an embodiment of the present invention;

[0020] Figure 3 It is a structural block diagram of a device for estimating the plaintext similarity of encrypted strings provided by an embodiment of the present invention;

[0021] Figure 4 A structural block diagram of a computer device provided by an embodiment of the present invention. Specific embodiments

[0022] The technical solution of the present invention will be further described below with reference to the accompanying drawings and through specific embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0023] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0024] Embodiment 1

[0025] Figure 1 A schematic flowchart of a method for estimating the plaintext similarity of encrypted strings provided by an embodiment of the present invention. This method can be executed by a device for estimating the plaintext similarity of encrypted strings, where the device can be implemented by software and / or hardware and is generally integrated in a computer device such as a server. As Figure 1 shown, the method includes:

[0026] S110. Obtain a plaintext data set, and perform an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set.

[0027] The plaintext data is a string that has not been encrypted. This string can be obtained from characters with different lengths and / or combinations, and the combination includes at least one of numbers, letters, and symbols.

[0028] Correspondingly, the plaintext data set is a set of strings with different lengths and / or combinations including a preset number. Among them, the preset number can be 10,000 or 20,000, etc., which is determined according to the training samples required by developers and is not limited here. The purpose of performing an encryption operation on the plaintext data set is to ensure the confidentiality of the data during transmission.

[0029] Correspondingly, the ciphertext data set is obtained by performing an encryption operation on the strings in the plaintext data set using a preset encryption algorithm. The preset encryption algorithm can be a symmetric algorithm (Data Encryption Standard, abbreviated as DES), International Data Encryption Algorithm (abbreviated as IDEA), Digital Signature Algorithm (abbreviated as DSA), etc., which is not limited herein.

[0030] Correspondingly, the ciphertext data set after the encryption operation includes a preset number of encrypted strings.

[0031] S120. Model the plaintext data set and the ciphertext data set respectively based on the multinomial distribution to obtain the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set.

[0032] The multinomial distribution is a generalization of the binomial distribution, which is induced by the joint distribution of two or more random variables X 1 , X 2 , …, X k (where k ≥ 2).

[0033] Since the plaintext data provided in the embodiments of the present invention consists of strings, according to the conventional input method, for example, there are a total of m different inputs (including numbers, letters, and symbols), then when modeling the plaintext data set using the multinomial distribution, the predicted distribution corresponding to the plaintext data set can be expressed by the following expression:

[0034] Multinomial(n 1 , n 2 , …, n m , p 1 , p 2 , …, p m ) (1)

[0035] In the formula, Multinomial represents the multinomial distribution, m represents the total dimension of the character types, n represents the mean vector of the corresponding dimension variables, and p represents the covariance vector of the corresponding dimension variables.

[0036] Specifically, according to the conventional input method of the keyboard, 94 different inputs can be counted, then the predicted distribution corresponding to the plaintext data set can be specifically expressed as:

[0037] Multinomial(n 1 , n 2 , …, n94 , p 1 , p 2 , …, p 94 ) (2)

[0038] A total of 94 input methods can be used to obtain the mean vector and covariance vector in 94 dimensions.

[0039] Correspondingly, the ciphertext data set is obtained by encrypting the plaintext data set. When modeling the ciphertext data set based on the multinomial distribution, the prediction distribution expression corresponding to the ciphertext data set is the same as the prediction distribution expression corresponding to the plaintext data set.

[0040] S130. Based on the Bayesian statistical model, estimate the predicted distribution corresponding to the decryption function according to the predicted distribution corresponding to the plaintext data set and the predicted distribution corresponding to the ciphertext data set.

[0041] The decryption function is a way to decrypt the encrypted ciphertext data set, which can be understood as the inverse function of the above encryption algorithm.

[0042] When determining the predicted distribution corresponding to the decryption function, it is necessary to model the m different input methods obtained in step S120. The m methods correspond to m dimensions. Then, the big data law and the multivariate normal distribution can be used to determine the predicted distribution corresponding to the decryption function, which can be expressed by the following expression:

[0043] N(μ m , Σ m ) (3)

[0044] In the formula, μ m represents the mean vector corresponding to the decryption function in the total dimension, and Σ m represents the variance matrix corresponding to the decryption function in the total dimension.

[0045] Corresponding to the 94 different inputs statistically obtained, the predicted distribution corresponding to the decryption function can be expressed as:

[0046] N(μ 94 , Σ 94 ) (4)

[0047] Furthermore, take the distributions corresponding to the obtained plaintext data set and the ciphertext data set as the input of the Bayesian statistical model, so that the output is the predicted distribution corresponding to the decryption function.

[0048] According to the predicted distribution expression corresponding to the decryption function, it can be seen that based on the Bayesian statistical model, the predicted distribution of the mean vector parameter μ corresponding to the current dimension and the predicted distribution of the variance matrix parameter ∑ corresponding to the current dimension are mainly estimated.

[0049] S140. Estimate the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function.

[0050] In the embodiments of the present invention, the target encrypted string can be understood as an encrypted string for which the plaintext similarity needs to be estimated, and the specific number can be more than two. The encryption algorithm corresponding to the target encrypted string is a preset encryption algorithm.

[0051] After determining the estimated distribution corresponding to the decryption function, the target encrypted string can be decrypted according to the estimated distribution corresponding to the current decryption function, so as to obtain the plaintext string estimated for the corresponding encrypted string. After obtaining the estimated plaintext strings corresponding to the target encrypted strings, the similarity between any two estimated plaintext strings can be calculated, so as to estimate the plaintext similarity between different target encrypted strings.

[0052] In the method for estimating the plaintext similarity of encrypted strings provided in the embodiments of the present invention, first obtain a plaintext data set, perform an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set; then model the plaintext data set and the ciphertext data set respectively based on the multinomial distribution to obtain the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; then, based on the Bayesian statistical model, estimate the estimated distribution corresponding to the decryption function according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; finally, estimate the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function, where the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm. By adopting the above technical solution, the similarity of the plaintext data before encryption can be estimated through the encrypted ciphertext data, and while protecting the privacy of the plaintext data, the correlation between multiple plaintext data can also be estimated.

[0053] Embodiment 2

[0054] The embodiments of the present invention are further optimized on the basis of the above embodiments. The optimization of estimating the estimated distribution corresponding to the decryption function according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set based on the Bayesian statistical model includes: converting the estimated distribution corresponding to the plaintext data set into a posterior distribution with respect to the Bayesian statistical model, and converting the estimated distribution corresponding to the ciphertext data set into a prior distribution with respect to the Bayesian statistical model; estimating the likelihood distribution of the Bayesian statistical model based on the posterior distribution and the prior distribution, where the likelihood distribution of the Bayesian statistical model is the estimated distribution corresponding to the decryption function. The advantage of this setting is to transform the decryption function problem into a Bayesian statistical model problem, which is convenient for calculation.

[0055] The step of further optimizing the estimation of the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function includes: selecting a first target encrypted string and a second target encrypted string; respectively inputting the first target encrypted string and the second target encrypted string into the estimated distribution corresponding to the decryption function to obtain first estimated plaintext data corresponding to the first target encrypted string and second estimated plaintext data corresponding to the second target encrypted string; calculating the similarity between the first estimated plaintext data and the second estimated plaintext data to obtain the plaintext similarity corresponding to the first target encrypted string and the second target encrypted string. The advantage of such a setting is that by obtaining the estimated distribution corresponding to the decryption function, the corresponding plaintext data estimated for the encrypted string is obtained, so as to predict the similarity of the plaintext, ensuring data security during the data transmission process.

[0056] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another method for estimating the plaintext similarity of encrypted strings provided by an embodiment of the present invention; specifically, the method includes the following steps:

[0057] S210. Obtain a plaintext data set, and use a preset encryption algorithm to perform an encryption operation on the plaintext data set to obtain a ciphertext data set.

[0058] If the obtained plaintext data set is denoted as x, the encryption algorithm is denoted as f(·), and the ciphertext data set is denoted as y, then the relational expression y = f(x) can be obtained.

[0059] In order to implement the method for estimating the plaintext similarity of encrypted strings based on ciphertexts provided by an embodiment of the present invention, it is necessary to find the decryption function n(·) of the encryption algorithm f(·), where n(·) is the generalized inverse function of f(·). Generally, as long as the decryption result obtained by the found decryption function does not affect the similarity evaluation of the decrypted characters.

[0060] When obtaining the plaintext data set x, a certain number of strings with different lengths and different combinations can be randomly generated as the plaintext data set. Exemplarily, 20000 strings are selected.

[0061] Then, use the preset encryption algorithm f(·) to perform an encryption operation on each plaintext string in the plaintext data set to obtain the corresponding ciphertext data set y.

[0062] S220. Based on the multinomial distribution, model the plaintext data set and the ciphertext data set respectively to obtain the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set.

[0063] Further, model the plaintext data set based on the multinomial distribution, and denote the estimated distribution corresponding to the plaintext data set as X; correspondingly, model the ciphertext data set based on the multinomial distribution, and denote the estimated distribution corresponding to the ciphertext data set as Y. Then, according to S210, the relational expression Y = f(X) can be obtained.

[0064] S230. Convert the estimated distribution corresponding to the plaintext data set into a posterior distribution with respect to the Bayesian statistical model, and convert the estimated distribution corresponding to the ciphertext data set into a prior distribution with respect to the Bayesian statistical model.

[0065] The method for estimating the plaintext similarity of an encrypted string provided by the embodiment of the present invention can transform the problem of calculating the decryption function into a Bayesian statistical model problem. Among them, denote the distribution corresponding to the plaintext data set as X, denote the distribution corresponding to the ciphertext data set as Y, and convert the estimated distribution X corresponding to the plaintext data set into a posterior distribution with respect to the Bayesian statistical model. The posterior distribution can be understood as the probability distribution of estimating the cause from the known result; the distribution Y corresponding to the ciphertext data set can be regarded as a prior distribution in the Bayesian statistical model, and the prior distribution can be understood as the probability distribution of determining the cause prior to the result.

[0066] S240. Based on the posterior distribution and the prior distribution, estimate the likelihood distribution of the Bayesian statistical model, where the likelihood distribution of the Bayesian statistical model is the estimated distribution corresponding to the decryption function.

[0067] Further, denote the decryption function as q, and denote the estimated distribution corresponding to the decryption function as Q. Then, the estimated distribution X corresponding to the plaintext data set, the distribution Y corresponding to the ciphertext data set, and the estimated distribution Q corresponding to the decryption function have the following relationship:

[0068] X ∝ QY (5)

[0069] That is, the distribution X corresponding to the plaintext data set is proportional to the product of the distribution Y corresponding to the ciphertext data set and the estimated distribution Q corresponding to the decryption function.

[0070] In formula (5), the estimated distribution corresponding to the decryption function can be regarded as a likelihood distribution in the Bayesian statistical model. The likelihood distribution can be understood as the probability distribution of estimating the result based on the determined cause.

[0071] Further, before estimating the likelihood distribution of the Bayesian statistical model based on the posterior distribution and the prior distribution, it further includes: using the law of large numbers and the multivariate normal distribution to estimate the parameters of the estimated distribution corresponding to the decryption function.

[0072] Since the estimated distribution corresponding to the decryption function is related to the input dimension of the plaintext data set, when the input dimension is 94, using the law of large numbers and the multivariate normal distribution to determine the estimated distribution corresponding to the decryption function can be expressed as:

[0073] N(μ 94 ,Σ 94 )

[0074] where μ 94 represents the mean vector corresponding to the decryption function in the total dimension, and Σ 94 represents the variance matrix corresponding to the decryption function in the total dimension.

[0075] Therefore, before determining the estimated distribution corresponding to the decryption function, it is also necessary to determine the parameters μ 94 and Σ 94 of the estimated distribution corresponding to the decryption function.

[0076] Specifically, a multivariate normal distribution is selected as the estimated distribution corresponding to the decryption function, and the Bayesian statistical model after the problem transformation is used, that is, using X and Y, to estimate the parameters μ 94 and Σ 94 in the estimated distribution corresponding to the decryption function, so as to determine the estimated distribution Q corresponding to the decryption function.

[0077] Furthermore, the law of large numbers and the multivariate normal distribution can be used to estimate the parameters of the estimated distribution corresponding to the decryption function.

[0078] The law of large numbers discusses the law that the arithmetic mean of a sequence of random variables converges to the arithmetic mean of the mathematical expectations of the random variables. The multivariate normal distribution is a generalization of the univariate normal distribution to multiple dimensions. For a variable subject to a normal distribution, as long as its mean and standard deviation are known, the frequency proportion within any value range can be estimated according to the formula. Then, the embodiments of the present invention use the law of large numbers and the multivariate normal distribution to solve the parameter problem in the decryption function.

[0079] Correspondingly, step S240 can further be based on obtaining the posterior distribution X (i.e., the estimated distribution corresponding to the plaintext data set) and the prior distribution Y (i.e., the estimated distribution corresponding to the ciphertext data set) of the Bayesian statistical model, and the parameters μ 94 and Σ 94 of the estimated distribution corresponding to the decryption function. After that, the likelihood distribution of the Bayesian statistical model can be estimated, and the likelihood distribution of the Bayesian statistical model is the estimated distribution corresponding to the decryption function.

[0080] S250. Select the first target encrypted string and the second target encrypted string.

[0081] The first target encrypted string and the second target encrypted string are ciphertext strings obtained by performing encryption operations using an encryption algorithm, and it is necessary to estimate the similarity of the plaintexts corresponding to the first target encrypted string and the second target encrypted string.

[0082] S260. Input the first target encrypted string and the second target encrypted string into the estimated distribution corresponding to the decryption function respectively, to obtain the first estimated plaintext data corresponding to the first target encrypted string and the second estimated plaintext data corresponding to the second target encrypted string.

[0083] According to the obtained estimated distribution Q corresponding to the decryption function, the first target encrypted string can be denoted as y_new 1 , and the second target encrypted string is denoted as y_new 2 . Input y_new 1 and y_new 2 into the estimated distribution Q corresponding to the decryption function respectively, and the first estimated plaintext data corresponding to y_new 1 can be obtained, denoted as x_new 1 , and the second estimated plaintext data corresponding to y_new 2 can be obtained, denoted as x_new 2 .

[0084] S270. Calculate the similarity between the first estimated plaintext data and the second estimated plaintext data to obtain the similarity of the plaintexts corresponding to the first target encrypted string and the second target encrypted string.

[0085] The method for calculating the similarity between the first estimated plaintext data and the second estimated plaintext data can be: calculating the cosine similarity, calculating the Euclidean distance, or calculating the Mahalanobis distance, etc. The specific calculation method is not limited here.

[0086] When the method for estimating the similarity of the plaintexts of the encrypted string provided by the embodiment of the present invention uses ciphertext data to estimate the similarity of the corresponding plaintexts, in the field of data security, while protecting privacy data or plaintext data, the correlation relationship between plaintext data can be calculated. It is also possible to obtain the similarity of the plaintexts according to the known encryption algorithm and the ciphertext relationship; at the same time, the calculation method is a generalized decryption model, which transforms the decryption problem into a Bayesian statistical model. This method is theoretically applicable to various encryption algorithms, and there is no need to perform precise decryption calculations for the encryption algorithm, which greatly reduces the decryption cost on the basis of a certain accuracy.

[0087] Embodiment III

[0088] Figure 3The following is a structural block diagram of an apparatus for estimating the plaintext similarity of encrypted strings provided by an embodiment of the present invention. The apparatus can be implemented by software and / or hardware, and is generally integrated in a computer device such as a server. It can estimate the plaintext similarity of encrypted strings by executing the method for estimating the plaintext similarity of encrypted strings. As Figure 3 shown, the apparatus includes: an encryption operation module 31, an estimated distribution obtaining module 32, an estimated distribution calculation module 33, and a plaintext similarity estimation module 34, where:

[0089] The encryption operation module 31 is configured to obtain a plaintext data set, and perform an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set;

[0090] The estimated distribution obtaining module 32 is configured to model the plaintext data set and the ciphertext data set respectively based on a multinomial distribution to obtain an estimated distribution corresponding to the plaintext data set and an estimated distribution corresponding to the ciphertext data set;

[0091] The estimated distribution calculation module 33 is configured to estimate an estimated distribution corresponding to a decryption function based on a Bayesian statistical model according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set;

[0092] The plaintext similarity estimation module 34 is configured to estimate the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function, where the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm.

[0093] In the apparatus for estimating the plaintext similarity of encrypted strings provided by the embodiment of the present invention, first, a plaintext data set is obtained, and an encryption operation is performed on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set; then, the plaintext data set and the ciphertext data set are respectively modeled based on a multinomial distribution to obtain an estimated distribution corresponding to the plaintext data set and an estimated distribution corresponding to the ciphertext data set; then, based on a Bayesian statistical model, an estimated distribution corresponding to a decryption function is estimated according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; finally, the plaintext similarity between different target encrypted strings is estimated according to the estimated distribution corresponding to the decryption function, where the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm. By adopting the above technical solution, the similarity of the plaintext data before encryption can be estimated through the encrypted ciphertext data, and while protecting the privacy of the plaintext data, the correlation relationship between multiple plaintext data can also be estimated.

[0094] Optionally, the plaintext data set includes a preset number of strings with different lengths and / or different combinations, and the combination includes at least one of numbers, letters, and symbols.

[0095] Optionally, the estimated distribution calculation module 33 includes: an estimated distribution conversion unit and an estimated distribution calculation unit, where:

[0096] The estimated distribution conversion unit is configured to convert the estimated distribution corresponding to the plaintext data set into a posterior distribution with respect to the Bayesian statistical model, and convert the estimated distribution corresponding to the ciphertext data set into a prior distribution with respect to the Bayesian statistical model;

[0097] The estimated distribution calculation unit is configured to estimate the likelihood distribution of the Bayesian statistical model based on the posterior distribution and the prior distribution, where the likelihood distribution of the Bayesian statistical model is the estimated distribution corresponding to the decryption function.

[0098] Optionally, the estimated distribution calculation unit further includes: a parameter estimation sub-unit;

[0099] The parameter estimation sub-unit is configured to estimate the parameters of the estimated distribution corresponding to the decryption function using the law of large numbers and the multivariate normal distribution.

[0100] Accordingly, the estimated distribution calculation unit is further configured to estimate the likelihood distribution of the Bayesian statistical model based on the posterior distribution, the prior distribution, and the parameters of the estimated distribution corresponding to the decryption function.

[0101] Optionally, the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set are represented by the following expressions:

[0102] Multinomial(n 1 ,n 2 ,…,n m ,p 1 ,p 2 ,…,p m )

[0103] In the formula, Multinomial represents the multinomial distribution, m represents the total dimension of the character types, n represents the mean vector of the corresponding dimension variables, and p represents the covariance vector of the corresponding dimension variables.

[0104] Optionally, the estimated distribution corresponding to the decryption function is represented by the following expression:

[0105] N(μ m ,Σ m )

[0106] In the formula, μ m represents the mean vector corresponding to the decryption function in the total dimension, and Σ m represents the variance matrix corresponding to the decryption function in the total dimension.

[0107] Optionally, the plaintext similarity estimation module 34 includes: a target encrypted string selection unit, a target encrypted string input unit, and a plaintext similarity estimation unit, where:

[0108] The target encrypted string selection unit is configured to select a first target encrypted string and a second target encrypted string;

[0109] The target encrypted string input unit is configured to respectively input the first target encrypted string and the second target encrypted string into the estimated distribution corresponding to the decryption function, so as to obtain first estimated plaintext data corresponding to the first target encrypted string and second estimated plaintext data corresponding to the second target encrypted string;

[0110] The plaintext similarity estimation unit is configured to calculate the similarity between the first estimated plaintext data and the second estimated plaintext data, so as to obtain the plaintext similarity corresponding to the first target encrypted string and the second target encrypted string.

[0111] The plaintext similarity estimation device for encrypted strings provided by the embodiments of the present invention can execute the plaintext similarity estimation method for encrypted strings provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0112] Embodiment 4

[0113] The embodiments of the present invention provide a computer device, and the plaintext similarity estimation device for encrypted strings provided by the embodiments of the present invention can be integrated in the computer device. Figure 4 It is a structural block diagram of a computer device provided by the embodiments of the present invention. The computer device 400 may include: a memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor. When the processor 402 executes the computer program, it implements the plaintext similarity estimation method for encrypted strings as described in the embodiments of the present invention.

[0114] The computer device provided by the embodiments of the present invention can execute the plaintext similarity estimation method for encrypted strings provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0115] Embodiment 5

[0116] The embodiments of the present invention further provide a storage medium including computer-executable instructions, and when the computer-executable instructions are executed by a computer processor, they are used to execute the plaintext similarity estimation method for encrypted strings. The method includes:

[0117] Obtain a plaintext data set, and perform an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set;

[0118] Model the plaintext data set and the ciphertext data set respectively based on the multinomial distribution to obtain the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set;

[0119] Based on the Bayesian statistical model, estimate the estimated distribution corresponding to the decryption function according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set;

[0120] Estimate the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function, where the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm.

[0121] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROMs, floppy disks or tape drives; computer system memory or random access memory such as DRAM, DDRRAM, SRAM, EDORAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. The storage medium may also include other types of memory or combinations thereof. Additionally, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system may provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media that may reside in different locations (such as in different computer systems connected via a network). The storage medium may store program instructions (such as embodied as a computer program) executable by one or more processors.

[0122] Of course, the storage medium containing computer-executable instructions provided by the embodiments of the present invention is not limited to the operation of estimating the plaintext similarity of the encrypted string as described above, and may also execute related operations in the method for estimating the plaintext similarity of the encrypted string provided by any embodiment of the present invention.

[0123] The device, equipment, and storage medium for estimating the plaintext similarity of the encrypted string provided in the above embodiments can execute the method for estimating the plaintext similarity of the encrypted string provided by any embodiment of the present invention, and have corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the above embodiments can be found in the method for estimating the plaintext similarity of the encrypted string provided by any embodiment of the present invention.

[0124] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for estimating the plaintext similarity of encrypted strings, characterized in that, it includes: Obtain a plaintext data set, and perform an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set; Based on the multinomial distribution, model the plaintext data set and the ciphertext data set respectively to obtain the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; Based on the Bayesian statistical model, estimate the estimated distribution corresponding to the decryption function according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; Estimate the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function, wherein the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm; Among them, the estimating the estimated distribution corresponding to the decryption function according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set based on the Bayesian statistical model includes: Convert the estimated distribution corresponding to the plaintext data set into a posterior distribution with respect to the Bayesian statistical model, and convert the estimated distribution corresponding to the ciphertext data set into a prior distribution with respect to the Bayesian statistical model; Use the law of large numbers and the multivariate normal distribution to estimate the parameters of the estimated distribution corresponding to the decryption function; Based on the posterior distribution, the prior distribution, and the parameters of the estimated distribution corresponding to the decryption function, estimate the likelihood distribution of the Bayesian statistical model, wherein the likelihood distribution of the Bayesian statistical model is the estimated distribution corresponding to the decryption function.

2. The method according to claim 1, characterized in that, the plaintext data set includes a preset number of strings with different lengths and / or combinations, and the combination includes at least one of numbers, letters, and symbols.

3. The method according to claim 1, characterized in that, the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set are represented by the following expressions: Multinomial(n 1 ,n 2 ,…,n m ,p 1 ,p 2 ,…,p m ) In the formula, Multinomial represents the multinomial distribution, m represents the total dimension of the character types, n represents the mean vector of the corresponding dimension variables, and p represents the covariance vector of the corresponding dimension variables.

4. The method according to claim 1, characterized in that, the estimated distribution corresponding to the decryption function is represented by the following expression: N(μ m ,Σ m ) where μ m represents the mean vector corresponding to the decryption function in the total dimension, and Σ m represents the variance matrix corresponding to the decryption function in the total dimension.

5. The method according to claim 1, characterized in that, the estimating the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function includes: Select a first target encrypted string and a second target encrypted string; Input the first target encrypted string and the second target encrypted string into the estimated distribution corresponding to the decryption function respectively to obtain the first estimated plaintext data corresponding to the first target encrypted string and the second estimated plaintext data corresponding to the second target encrypted string; Perform a similarity calculation on the first estimated plaintext data and the second estimated plaintext data to obtain the plaintext similarity corresponding to the first target encrypted string and the second target encrypted string.

6. An apparatus for estimating the plaintext similarity of encrypted strings, characterized in that, Comprising: An encryption operation module, configured to obtain a plaintext data set, and perform an encryption operation on the plaintext data set using a preset encryption algorithm to obtain a ciphertext data set; An estimated distribution obtaining module, configured to model the plaintext data set and the ciphertext data set respectively based on a multinomial distribution to obtain an estimated distribution corresponding to the plaintext data set and an estimated distribution corresponding to the ciphertext data set; An estimated distribution calculation module, configured to estimate an estimated distribution corresponding to a decryption function based on a Bayesian statistical model according to the estimated distribution corresponding to the plaintext data set and the estimated distribution corresponding to the ciphertext data set; A plaintext similarity estimation module, configured to estimate the plaintext similarity between different target encrypted strings according to the estimated distribution corresponding to the decryption function, wherein the encryption algorithm corresponding to the target encrypted string is the preset encryption algorithm; Wherein, the estimated distribution calculation module includes: an estimated distribution conversion unit, a parameter estimation sub-unit, and an estimated distribution calculation unit; The estimated distribution conversion unit is configured to convert the estimated distribution corresponding to the plaintext data set into a posterior distribution with respect to the Bayesian statistical model, and convert the estimated distribution corresponding to the ciphertext data set into a prior distribution with respect to the Bayesian statistical model; The parameter estimation sub-unit is configured to estimate the parameters of the estimated distribution corresponding to the decryption function using the law of large numbers and a multivariate normal distribution; The estimated distribution calculation unit is configured to estimate the likelihood distribution of the Bayesian statistical model based on the posterior distribution, the prior distribution, and the parameters of the estimated distribution corresponding to the decryption function.

7. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, the method described in any one of claims 1-5 is implemented.

8. A computer-readable storage medium, on which a computer program is stored, wherein, when the program is executed by the processor, the method described in any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Cipher text JPEG image retrieval method based on tree-form BoW model

    CN108600573A

  • KNN classification service system and method supporting privacy protection

    CN110011784A