Privacy intersection method and system based on national cryptographic algorithm

Through the privacy interception method based on the Guomi algorithm, fast encryption and local feature hash computing are used to construct a hash correlation matrix to filter the intersection data, solving the problem of large-scale sensitive data interception and leakage, and achieving efficient and secure data intersection determination.

CN119885283BActive Publication Date: 2025-07-11SHENZHEN OLYM INFORMATION SECURITY TECHOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510386716.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-11
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing data interchange method is computationally expensive and easy to leak original data when facing large-scale sensitive data, resulting in inefficiency and insufficient security.

Method used

The privacy interception method based on the national secret algorithm is adopted, and the data set is quickly encrypted, local feature hash computing, hash correlation matrix construction and similarity screening are filtered out, and the potential intersection data pairs are compared in the encrypted state to determine the final intersection data.

Benefits of technology

On the premise of ensuring data privacy and security, it significantly reduces the amount of computing and leakage risks, and efficiently and accurately finds the intersection data between data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119885283B_ABST
    Figure CN119885283B_ABST
Patent Text Reader

Abstract

The present invention provides a privacy intersection method and system based on national cryptographic algorithms, including: quickly encrypting a first data set based on national cryptographic algorithms to obtain a first encrypted data set; quickly encrypting a second data set based on national cryptographic algorithms to obtain a second encrypted data set; performing local feature hashing calculation to obtain corresponding first and second hash sequences; using the hash data in the first hash sequence as row elements and the hash data in the second hash sequence as column elements; calculating the similarity between the row elements and the column elements and filling it in a matrix to obtain a hash association matrix; screening out the hash data corresponding to the row-column intersection positions where the similarity is greater than a threshold as potential intersection data pairs; comparing the original data corresponding to the hash data in the potential intersection data pairs; if they are consistent, they are the final intersection data. In the present invention, the calculation amount during data intersection is reduced, and at the same time, the risk of data leakage is also reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a privacy intersection method and system based on a national secret algorithm. Background Art

[0002] In an era of rapid development of digital information, data has become a key factor in promoting innovation and improving efficiency in various fields. Different institutions and enterprises often need to cooperate on data to tap into more value, but this process inevitably involves the interaction of sensitive data.

[0003] On the one hand, enterprises and institutions must strictly protect user data privacy and prevent data leakage risks. On the other hand, traditional data intersection methods have many disadvantages when facing large-scale sensitive data. In the past, most common practices were to directly perform hash operations on the entire data set. This method has a large amount of calculations, especially in massive data scenarios, which consumes a lot of computing resources, resulting in low efficiency and hindering the timeliness of data interaction.

[0004] Meanwhile, in the prior art, direct comparison is usually performed when finding data intersection, which not only requires large amount of comparison calculation, but also easily leaks the original data. Summary of the invention

[0005] The main purpose of the present invention is to provide a privacy intersection method and system based on a national secret algorithm, aiming to overcome the defects of easy leakage of original data and large amount of calculation when performing data intersection.

[0006] To achieve the above purpose, the present invention provides a privacy intersection method based on a national secret algorithm, comprising the following steps:

[0007] The first data set involved in the intersection is quickly encrypted based on the national secret algorithm to obtain a first encrypted data set; the second data set involved in the intersection is quickly encrypted based on the national secret algorithm to obtain a second encrypted data set;

[0008] Performing local feature hash calculation on the first encrypted data set and the second encrypted data set to obtain a corresponding first hash sequence and a second hash sequence; wherein, during the local feature hash calculation, only a key data segment in each encrypted data is selected for hash calculation;

[0009] Constructing a matrix, using the hash data in the first hash sequence as row elements and the hash data in the second hash sequence as column elements; calculating the similarity between the row elements and the column elements and filling them in the matrix to obtain a hash association matrix;

[0010] Filtering out hash data corresponding to row and column intersection positions with similarity greater than a threshold from the hash association matrix as potential intersection data pairs;

[0011] Compare the original data corresponding to the hash data in the potential intersection data pair; if they are the same, the corresponding original data is the final intersection data.

[0012] Further, during fast encryption, only the basic structure information of the original data is retained, and the metadata is removed.

[0013] Further, calculating the similarity between row elements and column elements and filling it in the matrix includes:

[0014] Based on the similarity evaluation algorithm, calculate the similarity between the row elements and column elements corresponding to each row-column intersection position respectively, and fill it in the row-column intersection position in the matrix; the similarity evaluation algorithm combines the number of bit differences and the data size ratio of the hash data on the row elements and column elements to calculate the similarity between the two.

[0015] Further, before performing local feature hashing calculation on the first encrypted data set and the second encrypted data set, it also includes:

[0016] Scramble and reorganize each encrypted data in the first encrypted data set and the second encrypted data set respectively.

[0017] Further, after determining that if they are the same, the corresponding original data is the final intersection data, it includes:

[0018] Store each original data in the first data set and the second data set in a data table in the database;

[0019] Generate a query code based on the first data set and the second data set;

[0020] Establish a mapping relationship between the final intersection data and the data table, and configure a query permission code for the final intersection data based on the query code.

[0021] Further, generating a query code based on the first data set and the second data set includes:

[0022] Extract the multi-dimensional first feature information of the first data set, and generate a closed figure based on the first feature information;

[0023] Extract the multi-dimensional second feature information of the second data set, and generate a two-dimensional curve based on the second feature information;

[0024] Evenly divide the data table into four sub-regions, overlay the closed figure on each sub-region, and obtain the original data located at the center of the closed figure in each sub-region as the target data;

[0025] Generate a multi - digit number based on the feature information of each target data; substitute the multi - digit number as the abscissa into the two - dimensional curve for calculation to obtain the corresponding ordinate value as the query code.

[0026] Further, generate a query code based on the first data set and the second data set, including:

[0027] Add the characters in the first data set to the matrix one by one in sequence to generate a first matrix; add the characters in the second data set to the matrix one by one in sequence to generate a second matrix;

[0028] Perform an exclusive - OR calculation on the characters in the same matrix positions of the first matrix and the second matrix, and add the calculation results to a new matrix to obtain an exclusive - OR matrix; the exclusive - OR matrix includes a first number and a second number;

[0029] Generate the query code based on the first number in the exclusive - OR matrix.

[0030] Further, generate the query code based on the first number in the exclusive - OR matrix, including:

[0031] Select the first numbers that meet the preset position conditions from the exclusive - OR matrix in sequence as target numbers, and connect the target numbers in sequence to form a target curve;

[0032] Extract the curve feature information of multiple dimensions of the target curve, and generate the query code based on the curve feature information of multiple dimensions.

[0033] Further, generate the query code based on the first number in the exclusive - OR matrix, including:

[0034] Extract multiple features of the first number in the exclusive - OR matrix; the features at least include the number of the first numbers, row positions, column positions, and the maximum interval number;

[0035] Generate the query code based on the multiple features of the first number.

[0036] The present invention also provides a privacy intersection system based on the national cryptographic algorithm, including:

[0037] An encryption unit for quickly encrypting the first data set participating in the intersection based on the national cryptographic algorithm to obtain a first encrypted data set; quickly encrypting the second data set participating in the intersection based on the national cryptographic algorithm to obtain a second encrypted data set;

[0038] A hash unit for performing local feature hashing calculations on the first encrypted data set and the second encrypted data set to obtain corresponding first and second hash sequences; wherein, when performing local feature hashing, only the key data segments in each encrypted data are selected for hashing calculation;

[0039] A construction unit for constructing a matrix, using the hash data in the first hash sequence as row elements and the hash data in the second hash sequence as column elements; calculating the similarity between the row elements and the column elements and filling it in the matrix to obtain a hash association matrix;

[0040] A screening unit for screening out the hash data corresponding to the row-column intersection positions with similarity greater than the threshold from the hash association matrix as potential intersection data pairs;

[0041] An intersection finding unit for comparing the original data corresponding to the hash data in the potential intersection data pairs; if they are the same, the corresponding original data is the final intersection data.

[0042] The present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0043] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0044] The privacy intersection method and system based on the national cryptographic algorithm provided by the present invention include: quickly encrypting the first data set participating in the intersection based on the national cryptographic algorithm to obtain a first encrypted data set; quickly encrypting the second data set participating in the intersection based on the national cryptographic algorithm to obtain a second encrypted data set; performing local feature hashing calculation on the first encrypted data set and the second encrypted data set to obtain corresponding first and second hash sequences; wherein, when performing local feature hashing, only select the key data segments in each encrypted data for hashing calculation; construct a matrix, use the hash data in the first hash sequence as row elements, and use the hash data in the second hash sequence as column elements; calculate the similarity between the row elements and the column elements and fill it in the matrix to obtain a hash association matrix; screen out the hash data corresponding to the row-column intersection positions with a similarity greater than the threshold from the hash association matrix as potential intersection data pairs; compare the original data corresponding to the hash data in the potential intersection data pairs; if they are the same, the corresponding original data is the final intersection data. In the present invention, intersection calculation is performed after encrypting and hashing the original data, avoiding data leakage; local feature hashing calculation is adopted, significantly reducing the calculation amount; the similarity screening is performed in the way of a hash association matrix, reducing the calculation amount during data intersection and also reducing the risk of data leakage. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a schematic diagram of the steps of the privacy intersection method based on the national cryptographic algorithm in an embodiment of the present invention;

[0046] Figure 2 is a structural block diagram of the privacy intersection system based on the national cryptographic algorithm in an embodiment of the present invention;

[0047] Figure 3 is a schematic structural diagram of a computer device in an embodiment of the present invention.

[0048] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0050] Refer to Figure 1 , an embodiment of the present invention provides a privacy intersection method based on the national cryptographic algorithm, including the following steps:

[0051] Step S1, quickly encrypt the first data set participating in the intersection based on the national cryptographic algorithm to obtain the first encrypted data set; quickly encrypt the second data set participating in the intersection based on the national cryptographic algorithm to obtain the second encrypted data set;

[0052] Step S2, perform local feature hashing calculation on the first encrypted data set and the second encrypted data set to obtain the corresponding first hash sequence and second hash sequence; among them, when performing local feature hashing, only select the key data segments in each encrypted data for hashing calculation;

[0053] Step S3, construct a matrix, use the hash data in the first hash sequence as row elements, and use the hash data in the second hash sequence as column elements; calculate the similarity between the row elements and the column elements and fill it in the matrix to obtain the hash association matrix;

[0054] Step S4, screen out the hash data corresponding to the row-column intersection positions with a similarity greater than the threshold from the hash association matrix as potential intersection data pairs;

[0055] Step S5, compare the original data corresponding to the hash data in the potential intersection data pairs; if they are the same, the corresponding original data is the final intersection data.

[0056] In this embodiment, as described in step S1 above, the national cryptographic algorithm is a cryptographic algorithm system with high security. In this step, the national cryptographic algorithm is used to quickly encrypt the first data set to protect the privacy and integrity of the data. Specifically, cryptographic algorithms such as the national cryptographic SM4 symmetric encryption algorithm can be selected (other suitable national cryptographic algorithms can also be selected according to actual needs). During the encryption process, each piece of data or data block in the first data set is encrypted through a specific key, converting the original data that may contain sensitive information into ciphertext form, so that even if the data is illegally obtained during transmission or storage, the attacker cannot directly interpret the content, and finally the first encrypted data set is obtained. This encryption process emphasizes quickness while ensuring security, so as to efficiently process large-scale data and reduce the time cost brought by encryption operations.

[0057] Quickly encrypt the second data set participating in the intersection based on the national cryptographic algorithm to obtain the second encrypted data. The principle and purpose are the same as those for encrypting the first data set. The national cryptographic algorithm is also used, and according to the established encryption rules and appropriate keys, the second data set is encrypted, converting the data in it into ciphertext state to form the second encrypted data set. This ensures that the two data sets participating in the intersection are under encryption protection from the beginning, laying the foundation for subsequent intersection operations under the premise of privacy and security, and maintaining the symmetry of the processing flow, which is convenient for subsequent unified operations and comparative analysis.

[0058] As described in step S2 above, in traditional hash calculation, the entire data is often processed. However, this method innovatively adopts local feature hash calculation. Only the key data segments in each encrypted data are selected for hash calculation, which is considered from multiple aspects. On the one hand, the overall encrypted data is relatively long and complex. If all are subjected to hash operations, it will increase unnecessary computational effort. Focusing on the key data segments can not only reflect the core features of the data but also effectively reduce the consumption of computing resources and improve the overall efficiency. For example, for an encrypted structured data record, the key data segments are specific fields (such as the field representing the unique identifier of the data subject, the important timestamp field, etc.). After selecting these key segments, using a suitable hash algorithm (such as the national cryptography SM3 hash algorithm, etc.), corresponding hash values are generated for each data in the first encrypted data set. These hash values arranged in order form the first hash sequence. Similarly, the same operation is performed on the second encrypted data set to obtain the second hash sequence. Through such local feature hash calculation, the two data sets are transformed into a more representative hash sequence form that is convenient for subsequent comparison operations, and at the same time, the complete content of the original data is further hidden, enhancing the privacy protection effect.

[0059] As described in step S3 above, constructing a matrix is a key intermediate link in this private intersection method. By taking each hash data in the first hash sequence as the row elements of the matrix in turn and each hash data in the second hash sequence as the column elements of the matrix in turn, a two-dimensional structure is built. The dimension size of this matrix depends on the number of elements in the two hash sequences, and its significance lies in providing a clear and regular framework for comprehensively comparing the hash features of the two data sets later, facilitating the systematic examination of the correlation between different hash data.

[0060] After constructing the matrix framework, it is necessary to further calculate the similarity between each row element (i.e., a certain hash data in the first hash sequence) and each column element (i.e., each hash data in the second hash sequence). The calculation method of similarity can adopt a variety of suitable strategies, such as measuring based on the number of bit differences of the hash values (the smaller the bit difference, the higher the similarity), or through some more complex functions that comprehensively consider factors such as the length of the hash value and the data distribution characteristics to determine the similarity value. For each pair of row and column elements, after calculating their similarity value, fill this value into the corresponding row-column intersection position of the matrix. After the calculation and filling operations for all element pairs, the hash correlation matrix is finally obtained. The above hash correlation matrix fully reflects the similarity degree of each element between the two hash sequences, providing an intuitive basis for screening potential intersections later.

[0061] As described in step S4 above, a reasonable threshold is set. The determination of this threshold usually needs to consider factors such as the actual application scenario, data characteristics, and requirements for the accuracy and recall rate of intersection calculation. For example, in the scenario of calculating intersections for financial data with extremely high requirements for data accuracy, the threshold will be set relatively high to minimize misjudgment situations. From the constructed hash association matrix, traverse the similarity values at the intersection positions of each row and column, and extract the hash data corresponding to the positions where the similarity is greater than the set threshold. Since these hash data come from the first hash sequence and the second hash sequence respectively, they form potential intersection data pairs in pairs. The above data pairs represent data with high similarity at the hash level and are very likely to be the same or related data in the original dataset, and are important candidate objects for further verification to determine the final intersection.

[0062] As described in step S5 above, the potential intersection data pairs screened in the previous steps are only preliminary judgments based on the hash feature level, and it is necessary to further verify whether they are truly consistent at the original data level. Since in the previous operations, the data was first encrypted and then subjected to hash calculation and other processes, here it is necessary to first restore the data corresponding to the hash data in the potential intersection data pairs through corresponding decryption means (such as decrypting using the corresponding key during previous encryption), and then perform a detailed comparison operation on these restored original data word by word, field by field, or according to the established data comparison rules. Only when the comparison result shows complete consistency can it be determined that these original data are the true intersection part of the two datasets, and then they are recognized as the final intersection data. Such a final verification process ensures the accuracy of the intersection calculation result, avoids misjudgment situations caused by hash collisions and other reasons, and makes the finally obtained intersection data meet the actual requirements and be reliable.

[0063] In this embodiment, by performing intersection calculation after encrypting and hashing the original data, data leakage is avoided; by using local feature hashing calculation, the calculation amount is significantly reduced; by using the method of hash association matrix for similarity screening, the calculation amount during data intersection calculation is reduced, and at the same time, the risk of data leakage is also reduced. Through the close cooperation and orderly execution of the above steps, the privacy intersection method based on the national cryptography algorithm can efficiently and accurately find the intersection data between two datasets while ensuring data privacy and security, meeting the actual requirements in many scenarios involving sensitive data interaction.

[0064] In one embodiment, during fast encryption, only the basic structure information of the original data is retained for encryption, and the metadata is removed without encryption.

[0065] In this embodiment, during the data encryption process, the basic structure information of the original data refers to the organization form and architecture of the data, which defines how the data is arranged and associated. For example, in a database table, the basic structure information includes the field definitions of the table (such as field names, data types, relationships between fields, etc.), the association relationships between tables, etc. For a file system, the basic structure information may include the directory structure, the hierarchical relationship of files, etc. The purpose of retaining the basic structure information is to still be able to recognize and understand the overall organization method of the data after encryption, so as to facilitate subsequent processing and analysis.

[0066] Metadata provides descriptive and interpretive information about the original data. Metadata can include the source of the data, the creation time, the author, the meaning of the data, the quality information of the data, etc. Removing metadata in this encryption scenario is mainly for privacy protection and improving encryption efficiency. Removing metadata can reduce the amount of data to be encrypted, thereby accelerating the encryption speed, and at the same time avoiding the privacy leakage risk that metadata may bring, because metadata sometimes may contain sensitive information, such as the identity of the data owner, the data collection location, etc.

[0067] In one embodiment, calculating the similarity between row elements and column elements and filling it in a matrix includes:

[0068] Based on the similarity evaluation algorithm, calculate the similarity between the row elements and column elements corresponding to each row-column intersection position respectively, and fill it in the row-column intersection position in the matrix; the similarity evaluation algorithm combines the number of bit differences of the hash data on the row elements and column elements and the data size ratio to calculate the similarity between the two.

[0069] In this embodiment, in the entire method process of private set intersection, the similarity evaluation algorithm is a key link, which undertakes the important task of quantifying the similarity degree between two different hash data (that is, row elements and column elements). Through this algorithm, the association situation between the relatively abstract data that originally only exists in hash form can be numerically represented specifically, and then provide a clear and measurable basis for screening potential intersection data pairs in the subsequent process.

[0070] In the previously constructed matrix, the row elements come from the first hash sequence, and the column elements come from the second hash sequence. Each row-column intersection position represents a corresponding relationship between a certain hash data in the first hash sequence and a certain hash data in the second hash sequence.

[0071] For each row element (denoted as hash data A) and column element (denoted as hash data B) corresponding to such a row-column intersection position, a similarity evaluation algorithm is used to calculate the similarity between them. This calculation process needs to comprehensively consider various factors. It is not a simple direct comparison, but an in-depth analysis of the characteristics of the hash data itself to obtain a reasonable similarity value.

[0072] Hash data is usually a string of numbers (or characters) presented in binary form. The number of bit differences refers to the total number of different bits in the corresponding binary bits of two hash data (A and B). The number of bit differences reflects the similarity between the two hash data to a certain extent. The fewer the number of different bits, the closer the two hash data are, which means that the corresponding original data is more likely to be the same or highly correlated. Therefore, this is an important dimension for measuring similarity.

[0073] The above data size ratio refers to the proportional relationship between two hash data (A and B) in terms of length or the data scale they represent, etc. Hash data of different sizes may have differences in generation mechanisms, data characteristics they represent, etc. Even if the number of bit differences is the same, but if the data size ratio is different, the actual meaning of the similarity degree may also be different. For example, a shorter hash data and a very long hash data, even if the number of bit differences is small, may have relatively low similarity due to the large difference in overall scale; while two hash data with similar lengths are more likely to have higher similarity when the number of bit differences is small. By comprehensively considering this proportional relationship, the true similarity between two hash data can be measured more comprehensively and accurately.

[0074] The similarity evaluation algorithm combines these two factors, the number of bit differences and the data size ratio, for calculation (such as weighted calculation), and finally obtains a specific similarity value. The range of this value is usually set between 0 and 1, where 0 means completely dissimilar, 1 means completely the same, and intermediate values represent different degrees of similarity. For example, through a weighted summation method, appropriate weights are assigned to the number of bit differences and the data size ratio respectively (the determination of the weights can be determined in advance based on experiments, data characteristic analysis, etc.), and then the similarity value is calculated.

[0075] In one embodiment, before performing local feature hashing calculation on the first encrypted data set and the second encrypted data set, it further includes:

[0076] Randomly shuffle and reorganize each encrypted data in the first encrypted data set and the second encrypted data set respectively.

[0077] In one embodiment, if they are consistent, after the corresponding original data is the final intersection data, it includes:

[0078] Store each piece of original data in the first dataset and the second dataset in a data table of the database;

[0079] Generate a query code based on the first dataset and the second dataset;

[0080] Establish a mapping relationship between the final intersection data and the data table, and configure a query permission code for the final intersection data based on the query code.

[0081] In this embodiment, after completing the private intersection operation and determining the final intersection data, storing each piece of original data in the first dataset and the second dataset in a data table of the database is mainly for various considerations such as data management, subsequent use, and traceability. As an efficient data storage and management tool, the database can organize data in a structured manner, facilitating operations such as querying, updating, and analyzing data.

[0082] Generating a query code is for effective access control and permission management of the stored data, especially the final intersection data. By creating a unique query code, corresponding query permissions can be configured based on this code later. Only authorized users or systems with the correct query code can obtain and view the relevant data, which can greatly enhance data security, prevent unauthorized access to data, and protect data privacy, especially important in scenarios involving sensitive information.

[0083] There can be various ways to specifically generate the query code, and it is necessary to comprehensively consider the overall characteristics of the first dataset and the second dataset to ensure the uniqueness and effectiveness of the query code. One possible idea is to first extract features from the two datasets. For example, for numerical data, statistical features such as the data distribution range, mean, and variance can be calculated; for text data, semantic features such as keywords and topic words can be extracted. Then, these extracted features are combined according to certain rules, which can be concatenation, encryption, or through specific mathematical transformations (such as hash operations, encoding conversions, etc.). Finally, a query code that uniquely identifies the overall characteristics of the two datasets is generated. For example, after concatenating the numerical features of the first dataset and the text keywords of the second dataset, and then encrypting the result through an encryption algorithm (such as a symmetric encryption algorithm, using a specific key), the resulting ciphertext is used as the query code, which can not only integrate the key information of the two datasets but also ensure its confidentiality and uniqueness.

[0084] Establishing the mapping relationship between the final intersection data and the data table is to clearly determine the specific location and association of the final intersection data in the entire stored data system. A large amount of original data is stored in the data table, and the final intersection data is the part with specific value selected from these original data. Through the mapping relationship, the specific record rows or field positions of these intersection data in the data table can be quickly located, facilitating subsequent data extraction, display, and permission-based access control operations, etc. This helps improve the efficiency of data management and ensures that the use and access of data can accurately correspond to the actual stored data content.

[0085] Configuring the query permission code for the final intersection data based on the previously generated query code is the core step in implementing data access control. The query permission code can be understood as a "key" derived from the query code and used to verify whether a user has the permission to access the final intersection data. For example, a role-based permission management mechanism can be adopted to assign different query permission codes to different roles (such as administrators, ordinary business personnel, etc.). When a user attempts to access the final intersection data, they will be required to enter the query permission code, and then the entered code will be compared and verified with the corresponding permission code configured based on the query code. If the match is successful, it means the user has the corresponding access permission and can obtain the final intersection data; otherwise, access will be denied. The above permission configuration method can flexibly manage data access in a refined manner according to actual business needs, ensuring data security while ensuring that only legitimate and authorized users can access sensitive final intersection data.

[0086] In one embodiment, generating a query code based on the first data set and the second data set includes:

[0087] Extract the multi-dimensional first feature information of the first data set and generate a closed figure based on the first feature information;

[0088] Extract the multi-dimensional second feature information of the second data set and generate a two-dimensional curve based on the second feature information;

[0089] Evenly divide the data table into four sub-regions, overlay the closed figure on each sub-region, and obtain the original data located at the center of the closed figure in each sub-region as the target data;

[0090] Generate a multi-digit number based on the feature information of each target data; substitute the multi-digit number as the abscissa into the two-dimensional curve for calculation to obtain the corresponding ordinate value as the query code.

[0091] In this embodiment, the first data set often contains rich and diverse data content. Extracting its feature information from multiple dimensions is to comprehensively and synthetically grasp the characteristics of the data set, so as to subsequently generate representative identification elements (i.e., closed figures). The multiple dimensions here can cover multiple aspects. For example, for numerical data, statistical dimension features such as the maximum and minimum values, mean, variance, and the interval range of data distribution are considered; for text data, semantic and structural dimension features such as the frequency of keyword occurrences, the theme category of the text, and the mean sentence length are extracted.

[0092] After obtaining the first feature information, using these features to generate a closed figure is an innovative representation method. For example, the size parameters (such as perimeter, area, etc.) of the figure can be determined according to the numerical features, the shape category of the figure can be determined according to the data distribution (for example, if the data features are evenly distributed, a circle may be generated; if the data features are biased, an ellipse may be generated, etc.), and some textures or filling styles inside the figure can be determined through text-related features (of course, this may be an abstract representation for distinguishing different semantic features), etc. The closed figure generated in this way is like a visual fingerprint of the first data set, presenting complex data features in an intuitive and unique graphical form, providing a unique basis for subsequent further association and operation.

[0093] Similar to extracting the feature information of the first data set, feature mining should also be carried out from multiple dimensions for the second data set. Similarly, it will involve the key dimensions corresponding to different types of data, such as various statistics of numerical data, semantic key elements of text data, etc. By comprehensively analyzing these dimensions, the second feature information that can accurately describe the characteristics of the second data set is summarized. This step ensures a comprehensive description of the second data set and prepares for generating a corresponding two-dimensional curve subsequently.

[0094] Based on the extracted second feature information to generate a two-dimensional curve, the principle is to convert these features into relevant parameters of the curve. For example, a certain trend feature of the data can be used as the change of the abscissa with time or other variables, and the correlation feature of the data can be used as the value of the ordinate. Through a specific functional relationship or fitting algorithm, a two-dimensional curve is drawn. This curve is like a unique identifier of the second data set, which can reflect the dynamic changes and internal correlations of the second data set under multi-dimensional features, making it different from other different data sets and laying a foundation for subsequent association operations with the first data set.

[0095] Dividing the data table evenly into four sub-regions is a strategy for partitioning the storage data space, aiming to screen and correlate data from a more refined local perspective. This partitioning method can be based on the average distribution of the number of rows and columns in the data table. For example, it can be partitioned according to the row index range and column index range, so that each sub-region has a relatively independent and similar-scale data part.

[0096] Overlay the closed figures generated based on the first data set onto each sub-region. By finding the original data located at the center of the closed figure in each sub-region, a mapping relationship between the first data set and the specific original data in the data table is actually established. Extracting this data as the target data filters out some key data with strong correlation to the first data set in the entire data table, further focusing on the key elements for generating the query code finally.

[0097] Extracting the feature information of each obtained target data again is to further refine the key attributes of these key data, and then generate a multi-digit number based on this feature information. For example, different features of the target data can be combined according to a certain weight and converted into a unique multi-digit number representation through mathematical operations (such as weighted summation, coding conversion, etc.). This multi-digit number carries the core features of the target data and becomes an important intermediate identifier connecting the first data set, the target data, and subsequent operations.

[0098] Substituting the generated multi-digit number as the abscissa into the two-dimensional curve generated based on the second data set for calculation is a key step in associating the previous two data sets to generate the query code. Since the two-dimensional curve represents the features of the second data set and the multi-digit number carries the features of the target data related to the first data set, through such substitution calculation and using the function relationship contained in the curve, the corresponding ordinate value is obtained, and this ordinate value is determined as the query code. The query code generated in this way skillfully integrates the respective features of the two data sets and the association established between them through the target data, not only ensuring the uniqueness of the query code but also making it closely related to the two data sets, providing a reliable and creative basis for subsequent operations such as permission configuration based on the query code, and realizing effective access control and privacy protection for the final intersection data.

[0099] In one embodiment, generating a query code based on the first data set and the second data set includes:

[0100] Sequentially adding the characters in the first data set to the matrix one by one to generate a first matrix; sequentially adding the characters in the second data set to the matrix one by one to generate a second matrix;

[0101] Perform an exclusive OR calculation on the characters at the same matrix positions in the first matrix and the second matrix, and add the calculation results to a new matrix to obtain an exclusive OR matrix; the exclusive OR matrix includes a first number and a second number;

[0102] Generate the query code based on the first number in the exclusive OR matrix.

[0103] In this embodiment, the first data set contains a large number of data elements. Here, the focus is on the character information among them (for example, if the data set is in text form, the characters are the text content itself; if the data set is of a mixed type, the part that can be represented in character form can be extracted). These characters are sequentially added to the matrix one by one, that is, in a certain order, such as starting from the first piece of data in the data set, and each character is filled into the respective element positions of the matrix in turn. For example, fixed number of rows and columns can be set. If the number of characters exceeds the matrix capacity, appropriate strategies such as truncation, block division, etc. can be adopted to ensure that the character information of the first data set can be presented in the matrix in an orderly manner, and then the first matrix is generated.

[0104] In the same principle and operation mode as constructing the first matrix, for the second data set, extract the character information therein, and then add each character to another matrix one by one in the established order to generate the second matrix.

[0105] Exclusive OR (XOR) calculation is a logical operation. Its rule is that for two participating binary bits, when the two bits are the same (both are 0 or both are 1), the result is 0; when the two bits are different (one is 0 and the other is 1), the result is 1. In this step, for the characters in the same positions in the first matrix and the second matrix, perform the XOR operation bit by bit.

[0106] The first number and the second number (i.e., 0, 1) contained in the exclusive OR matrix are the results reflected after the above exclusive OR operation.

[0107] Selecting the first number in the exclusive OR matrix to generate the query code is because this number has a key representative or unique identification role in the information represented by the entire exclusive OR matrix. By generating the query code in this way, the difference information between the characters of the two data sets is cleverly utilized. Through the refinement of the exclusive OR operation and the transformation based on the specific number, the query code can be closely related to the characteristics of the two data sets, and at the same time has a certain degree of confidentiality and uniqueness, providing strong support for subsequent effective permission configuration and other operations on the final intersection data based on the query code.

[0108] In one embodiment, generating the query code based on the first number in the exclusive OR matrix includes:

[0109] Select the first numbers that meet the preset position conditions from the XOR matrix in sequence as target numbers, and connect the target numbers in sequence to form a target curve;

[0110] Extract the curve feature information of multiple dimensions of the target curve, and generate the query code based on the curve feature information of the multiple dimensions.

[0111] In this embodiment, the XOR matrix is a result matrix obtained by performing XOR operations on the corresponding position characters in the first matrix and the second matrix, which contains rich digital information reflecting the difference features of the two data sets. The preset position condition here is a screening rule designed to select numbers with specific representativeness or relevance as target numbers from numerous numbers. For example, the preset position condition can be to select the numbers on the main diagonal of the XOR matrix, or it can also be set to select a number every fixed number of rows and columns. According to such a preset position condition, the corresponding first numbers are sequentially selected from the XOR matrix as target numbers, and the set of these target numbers becomes the basic elements for constructing the subsequent target curve.

[0112] Connect the selected target numbers in sequence to form a target curve, which is a way to transform discrete digital information into a continuous and geometric feature representation form. Specifically, each target number can be regarded as the ordinate of a point in the plane rectangular coordinate system, and then these points are connected in sequence, and a curve can be drawn on the plane, that is, the target curve. This target curve carries the information contained in the target numbers selected from the XOR matrix, and shows the correlation features of the two data sets after specific processing in a visual (abstract sense) and geometric form, providing a new intermediate data structure for further extracting features to generate the query code.

[0113] After the target curve is formed, the feature information it contains can be mined and extracted from multiple dimensions. From the geometric dimension, it can include the length of the curve, the degree of bending of the curve (measured by, for example, curvature-related indicators, and curvature can reflect the speed of bending change at each point of the curve), the area enclosed by the curve (if the curve is closed or can enclose a certain area with the coordinate axes), etc.; from the data distribution dimension, it will involve the position range of the curve in the coordinate system (such as the value range of the abscissa, the value range of the ordinate), the density of data points (reflecting whether the distribution of target numbers in terms of values is relatively concentrated or dispersed), etc.; it can also be analyzed from the function fitting dimension, for example, trying to fit this curve with different functions (such as polynomial functions, trigonometric functions, etc.) and extracting the relevant parameters of the fitting function (such as the coefficients of each term of the polynomial function) as the curve feature information. By comprehensively considering the curve feature information of these different dimensions, the internal connection and unique properties between the two data sets represented by the target curve can be comprehensively and deeply grasped.

[0114] After obtaining multi-dimensional curve feature information, it is necessary to convert this information into a query code. A feasible approach is to arrange and combine the feature information of each dimension in a specific order. For example, first sort by dimension importance or a fixed order (such as first geometric dimension, then data distribution dimension, and finally function fitting dimension), and concatenate the corresponding feature values of each dimension to form a long string. Then, further processing can be performed on this string, such as using a hash algorithm (such as a common hash function to map the string to a hash value of a fixed length), encoding conversion (such as converting the string to hexadecimal encoding, etc.), or through specific mathematical operations (such as taking the integer after weighted summation of each feature value, etc.), ultimately generating a unique query code that can be used for subsequent permission management and other purposes. The query code generated in this way fully integrates the correlation features of the two data sets extracted and transformed from the XOR matrix, not only ensuring its close connection with the original data set but also possessing the characteristics of being easy to identify, verify, and used for access control, providing a reliable basis for the permission configuration link in the entire data privacy intersection and subsequent data management process.

[0115] In one embodiment, generating the query code based on the first number in the XOR matrix includes:

[0116] Extracting multiple features of the first number in the XOR matrix; the features at least include the quantity of the first number, row position, column position, and maximum interval quantity;

[0117] Generating the query code based on the multiple features of the first number.

[0118] In this embodiment, in the data structure of the XOR matrix, counting the quantity of the first number is to count the frequency of its occurrence in the entire XOR matrix. This quantity can reflect the distribution density of the first number in the matrix.

[0119] Recording the row position information of the first number in the XOR matrix, that is, clarifying which row of the matrix each first number is in. The row position is an important spatial position attribute, and together with the column position, it can determine the specific coordinates of the first number in the two-dimensional structure of the matrix.

[0120] Similar to the row position, the column position is used to determine the horizontal coordinate position of the first number in the XOR matrix. Considering the column position feature and combining it with the row position information can more accurately locate the specific distribution of the first number in the matrix, providing more detailed information support for mining the deep association between the two data sets and generating a unique and representative query code.

[0121] The maximum interval number refers to the maximum interval distance between adjacent first digits in the XOR matrix (the interval here can be measured according to the row and column order of the matrix, for example, how many rows and columns are there between two adjacent first digits). This feature reflects the dispersion degree of the first digits in the matrix. A larger maximum interval number indicates that the distribution of the first digits in the matrix is relatively dispersed; while a smaller maximum interval number means the distribution is relatively concentrated, reflecting a certain regularity.

[0122] Generating a query code based on multiple features of the extracted first digits is a process of integrating and transforming this discrete information that describes the first digits from different perspectives into a unique identifier. One approach is to first determine the representation form and order of each feature. For example, the quantity of the first digits can be represented by decimal digits, and the row position and column position can be respectively represented by binary codes with a fixed number of bits (determine the appropriate number of encoding bits according to the scale of the XOR matrix to ensure that all possible positions can be accurately represented), and the maximum interval number is also represented in a suitable numerical form. Then, in a predetermined order, such as first the quantity, then the row position, then the column position, and finally the maximum interval number, the codes representing each feature are concatenated together in sequence to form a string.

[0123] After that, in order to make this string more in line with the requirements of the query code (usually, it is expected that the query code has certain confidentiality, uniqueness, and is convenient for storage and verification, etc.), it can be further processed. For example, use a hash algorithm (such as the national secret SM3 hash algorithm or other common hash algorithms) to perform a hash operation on the concatenated string to obtain a hash value with a fixed length, and this hash value can be used as the final query code; or use an encryption algorithm to encrypt the concatenated string with a specific key, and the obtained ciphertext is used as the query code; or through some mathematical transformations, such as performing weighted summation and modulo operations on the concatenated string, and converting the result into a query code in a specific format. In this way, making full use of multiple features of the first digits in the XOR matrix, the generated query code can closely relate to the specific situation after the XOR operation of the two data sets, reflecting both the internal connection between them and playing an important role in subsequent data access control, permission management, etc., ensuring that only users or systems with corresponding permissions can access relevant data resources with the correct query code, thus protecting the security and privacy of the data.

[0124] In one embodiment, generating a query code based on the first data set and the second data set includes:

[0125] Classify and group the first data set to obtain multiple groups; among them, according to the obvious characteristic differences presented by the data itself, it is divided into multiple different groups. For example, according to factors such as the business field involved in the data, the time period when the data is generated, or the geographical scope corresponding to the data, the first data set is divided into corresponding groups.

[0126] For each group, calculate the key indicators of the overall characteristics of the group; for example, a certain average value of the data within the group, which can reflect the central tendency of the data in the group, and the degree of dispersion of the data within the group, so as to understand the dispersion of the data.

[0127] Classify and group the second data set, and calculate the corresponding key indicators for each group.

[0128] Based on the key indicators of each group of the first data set, construct a graph; the above graph can intuitively display its characteristic changes. For example, take the average values of different groups as the ordinates of the points on the graph, and take the sorting of the groups as the abscissa, and connect these points in turn to form a unique line trajectory.

[0129] Based on the key indicators of each group of the second data set, construct a line trajectory.

[0130] Perform transformation processing on the graph and the line trajectory; mainly transform them from the original shape representation into another characteristic form that can reflect their internal laws, and extract the undulating characteristics, flat characteristics, and trend characteristics at different stages during the change process of the line trajectory.

[0131] Compare the characteristics after transformation and find out the differences between them. These differences reflect the differences in the internal laws of the two data sets.

[0132] According to the above differences, generate a unique identifier, and use the identifier as the final query code. For example, if the difference shows that the graph has an obvious upward trend at a certain stage and the line trajectory is flat at the corresponding stage, then encode this characteristic combination into a specific symbol or a short character sequence, and this character sequence is the query code. Subsequently, accurate permission control of relevant data can be carried out by relying on this query code.

[0133] Refer to Figure 2 , in another embodiment of the present invention, a privacy intersection system based on the national secret algorithm is further provided, including:

[0134] An encryption unit for quickly encrypting the first data set participating in the intersection based on the national secret algorithm to obtain a first encrypted data set; quickly encrypting the second data set participating in the intersection based on the national secret algorithm to obtain a second encrypted data set;

[0135] A hash unit for performing local feature hashing calculations on the first encrypted data set and the second encrypted data set to obtain corresponding first and second hash sequences; wherein, when performing local feature hashing, only the key data segments in each encrypted data are selected for hashing calculations;

[0136] A construction unit for constructing a matrix, using the hash data in the first hash sequence as row elements and the hash data in the second hash sequence as column elements; calculating the similarity between the row elements and the column elements and filling them in the matrix to obtain a hash association matrix;

[0137] A screening unit for screening out the hash data corresponding to the row-column intersection positions with similarity greater than the threshold from the hash association matrix as potential intersection data pairs;

[0138] An intersection finding unit for comparing the original data corresponding to the hash data in the potential intersection data pairs; if they are the same, the corresponding original data is the final intersection data.

[0139] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to that described in the above method embodiment and will not be elaborated here.

[0140] Refer to Figure 3 , in the embodiment of the present invention, a computer device is further provided. This computer device can be a server, and its internal structure can be as Figure 3 shown. This computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of this computer design is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store the corresponding data in this embodiment. The network interface of this computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the above method.

[0141] Those skilled in the art can understand that Figure 3 the structure shown in

[0142] is only a block diagram of a part of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0143] In summary, the privacy intersection method and system based on the national cryptographic algorithm provided in the embodiments of the present invention include: quickly encrypting the first data set participating in the intersection based on the national cryptographic algorithm to obtain a first encrypted data set; quickly encrypting the second data set participating in the intersection based on the national cryptographic algorithm to obtain a second encrypted data set; performing local feature hashing calculation on the first encrypted data set and the second encrypted data set to obtain corresponding first and second hash sequences; wherein, when performing local feature hashing, only the key data segments in each encrypted data are selected for hashing calculation; constructing a matrix, using the hash data in the first hash sequence as row elements and the hash data in the second hash sequence as column elements; calculating the similarity between the row elements and the column elements and filling it in the matrix to obtain a hash association matrix; screening out the hash data corresponding to the row-column intersection positions with a similarity greater than the threshold from the hash association matrix as potential intersection data pairs; comparing the original data corresponding to the hash data in the potential intersection data pairs; if they are the same, the corresponding original data is the final intersection data. In the present invention, intersection calculation is performed after encrypting and hashing the original data, avoiding data leakage; local feature hashing calculation is adopted, significantly reducing the calculation amount; the similarity screening is performed in the form of a hash association matrix, reducing the calculation amount during data intersection and also reducing the risk of data leakage.

[0144] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database or other media provided in the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0145] It should be noted that in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements that are inherent to such process, apparatus, article or method. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.

[0146] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A privacy intersection method based on national cryptographic algorithms, characterized in that, Including the following steps: Quickly encrypt the first data set participating in the intersection based on the national secret algorithm to obtain a first encrypted data set; quickly encrypt the second data set participating in the intersection based on the national secret algorithm to obtain a second encrypted data set; Perform local feature hashing calculation on the first encrypted data set and the second encrypted data set to obtain corresponding first hash sequence and second hash sequence; when performing local feature hashing, only select the key data segments in each encrypted data for hashing calculation; Construct a matrix, use the hash data in the first hash sequence as row elements, and use the hash data in the second hash sequence as column elements; calculate the similarity between the row elements and the column elements and fill it in the matrix to obtain a hash association matrix; Select the hash data corresponding to the row-column intersection positions with similarity greater than the threshold from the hash association matrix as potential intersection data pairs; Compare the original data corresponding to the hash data in the potential intersection data pairs in the first data set and the second data set; if they are the same, the corresponding original data is the final intersection data.

2. The privacy intersection method based on the national cryptographic algorithm according to claim 1, wherein, When performing quick encryption, only retain the basic structure information of the original data and remove the metadata.

3. The privacy intersection method based on the national cryptographic algorithm according to claim 1, wherein Calculating the similarity between the row elements and the column elements and filling it in the matrix includes: Based on the similarity evaluation algorithm, calculate the similarity between the row elements and the column elements corresponding to each row-column intersection position respectively, and fill it in the row-column intersection position in the matrix; the similarity evaluation algorithm calculates the similarity between the two by combining the number of bit differences and the data size ratio of the hash data on the row elements and the column elements.

4. The privacy intersection method based on the national cryptographic algorithm according to claim 1, characterized in that Before performing local feature hashing calculation on the first encrypted data set and the second encrypted data set, it further includes: Reshuffle each encrypted data in the first encrypted data set and the second encrypted data set respectively.

5. The privacy intersection method based on the national cryptographic algorithm according to claim 1, characterized in that, After if they are the same, the corresponding original data is the final intersection data, it includes: Store each original data in the first data set and the second data set in a data table of the database; Generate a query code based on the first data set and the second data set; Establish a mapping relationship between the final intersection data and the data table, and configure a query permission code for the final intersection data based on the query code.

6. The privacy intersection method based on the national cryptographic algorithm according to claim 5, wherein Generating a query code based on the first data set and the second data set includes: Extract the multi-dimensional first feature information of the first data set, and generate a closed figure based on the first feature information; Extract the multi-dimensional second feature information of the second data set, and generate a two-dimensional curve based on the second feature information; Evenly divide the data table into four sub-regions, overlay the closed figure on each sub-region, and obtain the original data located at the center of the closed figure in each sub-region as target data; Generate a multi-digit number based on the feature information of each target data; substitute the multi-digit number as the abscissa into the two-dimensional curve for calculation, and obtain the corresponding ordinate value as the query code.

7. The privacy intersection method based on the national cryptographic algorithm according to claim 5, wherein Generating a query code based on the first data set and the second data set includes: Add the characters in the first dataset to the matrix one by one in sequence to generate a first matrix; add the characters in the second dataset to the matrix one by one in sequence to generate a second matrix; Perform an exclusive OR calculation on the characters at the same matrix positions in the first matrix and the second matrix, and add the calculation results to a new matrix to obtain an exclusive OR matrix; the exclusive OR matrix includes a first number and a second number; Generate the query code based on the first number in the exclusive OR matrix.

8. The privacy intersection method based on the national cryptographic algorithm according to claim 7, characterized in that, Generating the query code based on the first number in the exclusive OR matrix includes: Select the first numbers that meet the preset position conditions from the exclusive OR matrix in sequence as target numbers, and connect the target numbers in sequence to form a target curve; Extract the curve feature information of multiple dimensions of the target curve, and generate the query code based on the curve feature information of multiple dimensions.

9. The privacy intersection method based on the national cryptographic algorithm according to claim 7, characterized in that Generating the query code based on the first number in the exclusive OR matrix includes: Extract multiple features of the first number in the exclusive OR matrix; the features at least include the quantity of the first number, row position, column position, and maximum interval quantity; Generate the query code based on the multiple features of the first number.

10. A privacy intersection system based on national cryptographic algorithms, characterized in that, Includes: An encryption unit for quickly encrypting the first dataset participating in the intersection based on the national cryptography algorithm to obtain a first encrypted dataset; Quickly encrypt the second dataset participating in the intersection based on the national cryptography algorithm to obtain a second encrypted dataset; A hash unit for performing local feature hashing calculations on the first encrypted dataset and the second encrypted dataset to obtain corresponding first hash sequences and second hash sequences; wherein, when performing local feature hashing, only select the key data segments in each encrypted data for hashing calculation; A construction unit for constructing a matrix, using the hash data in the first hash sequence as row elements and the hash data in the second hash sequence as column elements; calculate the similarity between the row elements and the column elements and fill it in the matrix to obtain a hash association matrix; A screening unit for screening out the hash data corresponding to the row-column intersection positions with a similarity greater than the threshold from the hash association matrix as potential intersection data pairs; An intersection unit for comparing the original data corresponding to the hash data in the potential intersection data pairs in the first dataset and the second dataset; if they are consistent, the corresponding original data is the final intersection data.

Citation Information

Patent Citations

  • Approximate data processing method and apparatus, medium and electronic device

    WO2021143016A1

  • Private information retrieval method supporting multiple parties

    WO2024239592A1