Data retrieval method based on information security

By encrypting user data requests and searching in encrypted data index using standardized keyword collections and encryption clustering algorithms, the problem of insufficient efficiency and accuracy in ensuring data security is solved, and efficient, accurate and secure encrypted data retrieval is achieved.

CN120197206APending Publication Date: 2025-06-24SHENZHEN SHENMI XINAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411054152.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditional data retrieval methods are difficult to achieve efficient retrieval while ensuring data security when processing sensitive data, and the application of encryption technology in the retrieval process may affect efficiency and accuracy.

Method used

A data retrieval method based on information security is proposed. By encrypting user data requests, and using standardized keyword sets to search in encrypted data index, and clustering analysis is performed in combination with encrypted clustering algorithms to ensure that the data remains encrypted throughout the process.

Benefits of technology

It realizes efficient and accurate encrypted data retrieval under high data security guarantees, balancing the needs of security and retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197206A_ABST
    Figure CN120197206A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information security and data retrieval, and particularly relates to a data retrieval method based on information security. According to the method, a user data request is received and encrypted, and the encrypted request is analyzed to extract feature information. A keyword set is formed, retrieval is carried out in the encrypted data index, and a preliminary candidate data set is generated; a security candidate data set is generated by excluding non-compliant data items. And performing similarity calculation and screening on the encrypted feature values of the data items to form a preliminary retrieval result. The results are sorted after being classified and subjected to safety weight calculation, and the sorted data items are subjected to subset division and feature distribution calculation; and based on the subset feature index and the security weight, generating a comprehensive feature vector, performing clustering analysis, and identifying a data cluster. And finally, calculating feature distribution of the clustering clusters, evaluating safety quality parameters of data retrieval, and generating and outputting an encrypted retrieval result. According to the method, the security and the accuracy of data retrieval are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of information security and data retrieval, and particularly relates to a data retrieval method based on information security. Background Art

[0002] In the current fields of information security and data processing, data retrieval technology has become a crucial part. With the explosion of data volume and the strictness of privacy protection regulations, traditional data retrieval methods are facing severe challenges. Especially when dealing with sensitive data, how to perform efficient retrieval while ensuring data security has become a hot issue in the industry.

[0003] Traditional data retrieval methods mainly rely on plaintext processing, that is, the data is not encrypted during storage and retrieval, which exposes certain data security risks to a certain extent. Once the database is illegally accessed, the stored data may be stolen or misused. In addition, even if encryption technology is adopted, in most cases, it is only limited to encryption during the data storage stage, and decryption is still required during the retrieval process, which may also lead to the leakage of sensitive information.

[0004] In recent years, with the development of encryption technology, some encryption-based data retrieval methods have emerged. These methods improve security by keeping the data encrypted during storage and retrieval. However, these methods often involve complex encryption and decryption processes, which may seriously affect the efficiency and accuracy of retrieval. In addition, how to ensure the integrity and availability of data while guaranteeing the retrieval effect is also a technical difficulty.

[0005] Therefore, it is particularly urgent to develop a new type of data retrieval method based on information security. Summary of the Invention

[0006] (1) Technical Problems to be Solved

[0007] The present invention mainly aims at the above problems and proposes a data retrieval method based on information security, aiming to solve how to improve the efficiency and accuracy of encrypted data retrieval while ensuring a high degree of data security.

[0008] (2) Technical Solutions

[0009] To achieve the above object, the first aspect of the present invention provides a data retrieval method based on information security, and the data retrieval method includes the following steps:

[0010] Receive a user data request and perform encryption processing on it;

[0011] Parse the encrypted user data request, extract feature information and perform preprocessing;

[0012] Extract keywords and perform semantic analysis on the preprocessed feature information to form a standardized keyword set;

[0013] Use the standardized keyword set to retrieve in the encrypted data index to generate a preliminary candidate data set;

[0014] Perform security filtering on the preliminary candidate data set, exclude data items that do not meet the security requirements according to the preset security policy to obtain a secure candidate data set;

[0015] In the secure candidate data set, calculate the similarity according to the encrypted feature values of the data items, and screen out highly matching data items through the set similarity threshold to generate a preliminary retrieval result;

[0016] Classify the data items in the preliminary retrieval result, calculate the security weight of each data item and sort them to form a sorted list;

[0017] Divide the data items in the sorted list into several subsets according to a predetermined rule, calculate the data feature distribution of each subset to generate an encrypted feature sequence;

[0018] Calculate the contrast between subsets based on the standard deviation of the encrypted feature sequence and the encrypted similarity distance between features to generate a feature index for each data item;

[0019] Combine the security weight and feature index of the data item to generate a comprehensive feature vector;

[0020] Use the comprehensive feature vector to perform clustering analysis through an encrypted clustering algorithm to identify and divide data clustering clusters;

[0021] Count the total amount of encrypted data in each clustering cluster and calculate the encrypted feature distribution of the clustering cluster;

[0022] Calculate the security quality parameter of data retrieval according to the encrypted feature distribution of the clustering cluster, generate the final encrypted retrieval result and decrypt and output it.

[0023] Furthermore, the specific method for encrypting the user data request is: use a symmetric encryption algorithm to encrypt the user data request to generate a ciphertext request.

[0024] Furthermore, the specific method for parsing the encrypted user data request is: decrypt the encrypted user data request through a parsing algorithm, extract feature information and perform data preprocessing, including data cleaning and format standardization.

[0025] Furthermore, the specific method for using the standardized keyword set to retrieve in the encrypted data index is: match the standardized keyword set with the encrypted data index and retrieve through an encrypted search algorithm to generate a preliminary candidate data set.

[0026] Furthermore, the specific method for calculating the similarity based on the encrypted eigenvalue of the data item is as follows:

[0027] In the received secure candidate dataset, perform feature encoding processing on each data item, encrypt each data item using an asymmetric encryption algorithm, and retain its encrypted eigenvalue;

[0028] Perform feature decomposition on each encrypted data item, split each encrypted data item into basic feature units, and generate an encryption identifier for each feature unit;

[0029] Perform feature analysis on the basic feature units of each data item, calculate the correlation between the feature units through a neural network, and generate a feature correlation graph;

[0030] Based on the feature correlation graph, use graph theory algorithms to calculate the encrypted feature similarity between data items and determine the similarity relationship between each data item;

[0031] Set a similarity threshold, and based on the calculated encrypted feature similarity, screen out the data item pairs higher than the set threshold to form a preliminary set of similar data items.

[0032] Furthermore, the method for dividing data items into several subsets based on a sorted list and generating an encrypted feature sequence includes the following steps:

[0033] Receive the sorted list generated according to the similarity calculation, divide the data items in the sorted list into several subsets according to a predetermined rule, and each subset contains several data items;

[0034] Perform feature encoding on the data items of each subset, encrypt the eigenvalue of each data item using an encryption algorithm to generate an encrypted eigenvalue;

[0035] Calculate the data feature distribution of each subset, count the distribution of the encrypted eigenvalues of all data items in each subset to form a feature distribution matrix;

[0036] Perform normalization processing on the feature distribution matrix to convert the values in the feature distribution matrix into normalized feature values;

[0037] Combine the normalized feature values in sequence to form an encrypted feature sequence;

[0038] Perform sequence analysis on the encrypted feature sequence,

[0039] According to the results of the sequence analysis, classify the encrypted feature sequences of each subset to form an encrypted feature sequence classification table;

[0040] Compare the encrypted feature sequence classification table with the sorted list of original data items to verify the accuracy and integrity of the encrypted feature sequence, and generate the final output of the encrypted feature sequence.

[0041] Further, the specific method for calculating the contrast between subsets is as follows:

[0042] Receive the encrypted feature sequence classification table generated according to the encrypted feature sequence, group the subsets in the classification table according to a predetermined rule, and each group contains several subsets:

[0043] Calculate the feature vector of each subset. Based on the encrypted feature values of all data items within the subset, use the weighted average method to generate the subset feature vector;

[0044] Conduct a differential analysis of the feature vectors between subsets. Use the vector space model to calculate the Euclidean distance between each pair of subset feature vectors to obtain the preliminary contrast matrix between subsets;

[0045] Normalize the preliminary contrast matrix to form a standardized contrast matrix;

[0046] Perform dimensionality reduction on the standardized contrast matrix, and map the high-dimensional contrast data into a two-dimensional or three-dimensional space;

[0047] In the dimensionality reduction space, calculate the density distribution between subsets, and use the density clustering algorithm to identify the density peak regions between subsets;

[0048] Based on the density peak regions, calculate the final contrast value between each pair of subsets to form a contrast map between subsets;

[0049] According to the contrast map between subsets, determine the similarity and difference between subsets, and generate an optimized subset classification table.

[0050] Further, when encrypting the user data request, the symmetric encryption algorithms used include AES, DES, or 3DES.

[0051] Further, when parsing the encrypted user data request, the parsing algorithms used include RSA, ECC, or DSA.

[0052] (III) Beneficial effects

[0053] Compared with the prior art, a data retrieval method based on information security provided by the present invention first receives and encrypts the user's data request to ensure the security of the data during transmission and processing. Then, an algorithm is used to parse the encrypted data request, extract key feature information and perform preprocessing. This step improves the accuracy of data processing by optimizing keyword extraction and semantic analysis methods. Next, a standardized keyword set is used to perform efficient retrieval in the encrypted data index to generate a candidate data set. In addition, a security filtering mechanism is used to exclude data that does not meet security standards, enhancing data security protection. Finally, by calculating the similarity of the encrypted feature values of data items and combining clustering analysis techniques, highly relevant data clusters are accurately identified and classified, thereby improving the relevance and accuracy of retrieval results. Throughout the process, the data remains encrypted, effectively balancing the requirements of security and retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 FIG. is a data processing flowchart disclosed in the present application.

[0055] Figure 2 FIG. is a system architecture diagram disclosed in the present application.

[0056] Figure 3 FIG. is a schematic diagram of an encryption and decryption method disclosed in the present application.

[0057] Figure 4 FIG. is a flowchart of feature processing and analysis disclosed in the present application.

[0058] Figure 5 FIG. is a schematic diagram of a clustering and classification method disclosed in the present application.

[0059] Figure 6 FIG. is an analysis diagram of an encrypted feature sequence and a feature vector disclosed in the present application.

[0060] Figure 7 FIG. is a diagram of calculating and optimizing the contrast between subsets disclosed in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The present invention will be described in detail below with reference to the accompanying drawings. The technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] An embodiment of the present invention provides a data retrieval method based on information security. Through a series of steps, this method ensures data security while improving retrieval efficiency. Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0063] As Figure 1 , Figure 2 shown, the implementation details of steps S1 to S13 are as follows:

[0064] Step S1: Receive the user data request and encrypt it:

[0065] In this step, the system first receives the data request submitted by the user through the interface. The request information input by the user is encrypted through a secure encryption protocol (such as the RSA encryption algorithm) to ensure the data security during the transmission process.

[0066] Step S2: Parse the encrypted user data request, extract the feature information and preprocess it:

[0067] After the encrypted data request is transmitted to the server, a dedicated parsing tool is used to decrypt and parse the data to extract the key feature information. This feature information is then preprocessed, including data normalization and error correction, to prepare the data basis for the subsequent steps.

[0068] Step S3: Extract keywords and perform semantic analysis on the preprocessed feature information:

[0069] Using natural language processing technology, keywords are extracted and semantic analysis is performed on the preprocessed data to generate a standardized keyword set, Figure 4 describing steps such as the extraction, preprocessing, keyword generation, feature encoding, feature decomposition, and feature association graph generation of the feature information.

[0070] Step S4: Use the standardized keyword set to retrieve in the encrypted data index and generate a preliminary candidate data set:

[0071] The extracted standardized keywords are used to search the encrypted data index, and the index structure is constructed based on the inverted index technology, which can quickly locate the data sets related to the keywords.

[0072] Step S5: Perform security filtering on the preliminary candidate data set:

[0073] The preliminarily generated candidate data set needs to go through the security filtering step to remove those data items that do not conform to the preset security policies, such as sensitive information or access-restricted data.

[0074] Step S6, calculate the similarity of encrypted eigenvalues:

[0075] In the candidate dataset that meets the security requirements, calculate the similarity based on the encrypted eigenvalue of each data item, and filter out the highly matching data items through a preset threshold to form a preliminary retrieval result.

[0076] Step S7, classify and sort the data items in the preliminary retrieval result:

[0077] Classify the retrieval result and calculate the security weight of each data item, and sort according to the weight for the user to browse and select more effectively.

[0078] Step S8, calculate the data feature distribution:

[0079] Divide the sorted data items into several subsets according to a predetermined rule, calculate the data feature distribution of each subset, and generate a feature sequence, which helps to further analyze and process the data.

[0080] Step S9, calculate the contrast between subsets:

[0081] Based on the generated encrypted feature sequence, calculate the standard deviation between subsets and the encrypted similarity distance between features, so as to obtain the contrast of each subset for evaluating the distribution uniformity of the data.

[0082] Step S10, generate a comprehensive feature vector:

[0083] Combine the security weight and feature index of the data item to generate a feature vector representing the comprehensive attributes of the data, providing a basis for clustering analysis.

[0084] Step S11, perform clustering analysis using the comprehensive feature vector:

[0085] Such as Figure 5 , apply the encrypted clustering algorithm to analyze the data and identify the clustering clusters between the data, which helps to identify the patterns and associations in the data.

[0086] Step S12, count the total amount of encrypted data in the clustering cluster:

[0087] Statistically analyze each clustering cluster, calculate the total amount and feature distribution of its encrypted data, and provide a quantitative basis for the final data processing.

[0088] Step S13, generate the final encrypted retrieval result and decrypt and output it:

[0089] Calculate the security quality parameter of data retrieval according to the feature distribution of each clustering cluster, and generate the final encrypted retrieval result. On the premise of ensuring data security, decrypt the result and display it to the user.

[0090] A data retrieval method based on information security disclosed by the above embodiments shows that the method ensures data security through full - process encryption. First, the method receives and encrypts the user's data request, then parses the encrypted data to extract key feature information. This information undergoes further processing and analysis to form a standardized keyword set for effective retrieval in the encrypted data index. Through a series of security filtering and similarity calculations, the method can screen out the data sets most relevant to the query. Then, these data are further analyzed and classified, the security weights and feature indices of each are calculated, and the data are more meticulously organized and classified through clustering analysis techniques. Finally, the security quality parameter of the retrieval is calculated based on the clustering results, and the final retrieval result is decrypted and output. This process not only improves the accuracy and efficiency of the retrieval, but also strengthens the security of the data during the retrieval process.

[0091] The embodiments described above are only exemplary, and those skilled in the art can modify or substitute the embodiments without departing from the principles of the present invention.

[0092] In step S1, in order to ensure the security of the user data request during transmission and processing, a symmetric encryption algorithm is used for encryption. The symmetric encryption algorithm is a commonly used encryption method, where the encryption and decryption processes use the same key. Specifically, the system converts the user's original data request into ciphertext, that is, an encrypted format that cannot be directly read, and only the system or individual with the correct key can decrypt the ciphertext to access the original data.

[0093] In step S2, the system first uses a parsing algorithm to decrypt the encrypted user data request. This process involves converting the received ciphertext data back to its original readable format to ensure that subsequent steps can effectively process the data. After decryption, the system continues to extract the key feature information in the data to ensure that the subsequent retrieval can be carried out specifically. After extracting the features, the data will go through a pre - processing stage, including data cleaning and format standardization. Data cleaning mainly removes irrelevant information, incorrect data, or duplicate data in the data, while format standardization ensures that the data formats are consistent for effective analysis and retrieval.

[0094] In step S4, the system uses the formed standardized keyword set to retrieve the encrypted data index. In this process, the system matches the keyword set with the encrypted entries in the data index. By using an encrypted search algorithm, the system can accurately identify the data entries related to the keywords while keeping the data in an encrypted state. After successful matching, the system will generate a preliminary candidate data set, which contains all the encrypted data entries that may be related to the user's query.

[0095] In step S6, the specific method for calculating the similarity based on the encrypted eigenvalue of the data item is as follows:

[0096] Step S6-1: In the received secure candidate dataset, perform feature encoding processing on each data item, encrypt each data item using an asymmetric encryption algorithm, and retain its encrypted eigenvalue.

[0097] "Performing feature encoding processing on each data item" means converting the feature information (such as text content, numerical data, etc.) of each data item into a coding format, and these codes can represent the key features and attributes of the data item. For example, the text content can be converted into a series of numbers and symbols through a hash function or other coding algorithms.

[0098] "Retaining its encrypted eigenvalue" means that on the basis of feature encoding, further encrypt these coded values using an asymmetric encryption algorithm to generate an encrypted eigenvalue. This encrypted eigenvalue is used to protect the privacy of the original data and ensure that even if the data features are accessed during the retrieval process, the security of the data will not be threatened. For example, encrypt the coded eigenvalue through the RSA algorithm to generate a ciphertext that can only be decrypted by the corresponding private key.

[0099] Step S6-2: Perform feature decomposition on each encrypted data item, split each encrypted data item into basic feature units, and generate an encrypted identifier for each feature unit.

[0100] Further refine and decompose the encrypted eigenvalue of each data item into smaller components, and each component represents a basic aspect of the data feature. For example, for a text data segment, its encrypted eigenvalue can be split into units representing different semantic or statistical attributes, such as keyword frequency, grammatical structure, etc.

[0101] "Generating an encrypted identifier" is to assign a unique encrypted flag to each basic feature unit so that these feature units can be securely and uniquely identified and referenced in subsequent processes. These encrypted identifiers are usually generated by a security algorithm to ensure that the identifier of each feature unit is both unique and difficult to crack. For example, a hash function can be used to calculate a hash value for the content of each feature unit as its encrypted identifier.

[0102] Step S6-3: Perform feature analysis on the basic feature units of each data item, calculate the correlation between the feature units through a neural network, and generate a feature correlation graph.

[0103] Use a neural network to analyze the correlations between basic feature units and represent these relationships graphically, where nodes represent feature units and edges represent the degree of association between these feature units. For example, if two feature units often appear in similar data environments, the connection (edge) between them will be identified in the feature association graph, indicating that they have a high degree of correlation. Such a graphical representation can help to more intuitively understand the complex relationships between data features, thereby more effectively identifying and utilizing these relationships in subsequent analysis.

[0104] Step S6-4: Based on the feature association graph, use graph theory algorithms to calculate the encrypted feature similarity between data items and determine the similarity relationships between data items.

[0105] "Graph theory algorithms" refer to using the principles of graph theory in mathematics to analyze the connection strength and patterns between nodes (feature units) in the feature association graph, thereby calculating the similarity between data items. For example, cosine similarity is commonly used to calculate the similarity between two non-binary vectors and can be applied to the comparison of encrypted feature values. The calculation formula is:

[0106]

[0107] where A and B are the feature vectors of two data items in the feature association graph, A·B represents the dot product of the two vectors, and ||A|| and ||B|| are the magnitudes of the two vectors respectively.

[0108] Step S6-5: Set a similarity threshold, and based on the calculated encrypted feature similarity, filter out data item pairs with similarity higher than the set threshold to form a preliminary set of similar data items.

[0109] As Figure 6 shown, in step S8, the method of dividing data items into several subsets and generating an encrypted feature sequence based on the sorted list includes the following steps:

[0110] Step S8-1: Receive the sorted list generated according to similarity calculation, and divide the data items in the sorted list into several subsets according to a predetermined rule, with each subset containing several data items.

[0111] Organize a sorted data list into different groups or sets according to classification criteria (such as similarity level, data type, user-defined classification criteria, etc.), with each group containing data items with similar features or meeting specific conditions. For example, the entire sorted list can be divided into subsets of high, medium, and low security levels according to the security weight or feature similarity score of the data items, with each level containing data items with similar security scores.

[0112] Step S8-2: Perform feature encoding on the data items of each subset, and use an encryption algorithm to encrypt the feature values of each data item to generate encrypted feature values;

[0113] Extract features from the data items in each subset and convert these features into a numerical or symbolic encoding form so that they can be processed by a computer. For example, if the data items in the subset are text documents, feature encoding includes extracting keywords, counting word frequencies, and then mapping these keywords or word frequencies to numerical codes for subsequent encryption processing and data analysis.

[0114] Step S8-3: Calculate the data feature distribution of each subset, and count the distribution of the encrypted feature values of all data items in each subset to form a feature distribution matrix;

[0115] Count the set conditions of the encrypted feature values of all data items in each subset, including statistical attributes such as frequency and range, and organize these statistical data into a matrix form, where each row represents the feature value of a data item and each column represents the statistical results of different features. If the subset contains different types of documents, the feature distribution matrix includes the frequencies of keywords appearing in each document, each row represents a document, and each column represents the frequency of a keyword, thus forming a matrix that comprehensively describes the characteristics of the subset.

[0116] Step S8-4: Perform standardization processing on the feature distribution matrix to convert the values in the feature distribution matrix into standardized feature values;

[0117] Adjust the original values in the feature distribution matrix through statistical methods (such as Z-score standardization, min-max standardization, etc.) to make them comparable under different scales or measurement units, thereby eliminating the influence of dimensions and improving the accuracy of subsequent analysis. For example, if the feature distribution matrix contains data with different scales (such as age and income), by converting these values into standardized Z-scores (subtracting the mean and dividing by the standard deviation), it can be ensured that each feature has equal weight when comparing and analyzing.

[0118] Step S8-5: Combine the standardized feature values into an encrypted feature sequence in order;

[0119] For example, the standardized feature values within a subset (such as the time series standard values of user behavior data) can be arranged in chronological order to form an encrypted feature sequence representing the behavior trend of the subset, which is convenient for time series analysis or other forms of data mining.

[0120] Step S8-6: Perform sequence analysis on the encrypted feature sequence;

[0121] Step S8-7: Classify the encrypted feature sequences of each subset according to the results of sequence analysis to form a classification table of encrypted feature sequences;

[0122] Organize the classification results into a table, which will show the results of dividing the encrypted feature sequences of different subsets according to the classification criteria. For example, they can be classified into different groups according to the similarity patterns of the sequences, and a unique identifier is assigned to each group in the classification table.

[0123] Step S8-8: Compare the classification table of encrypted feature sequences with the sorted list of original data items to verify the accuracy and integrity of the encrypted feature sequences and generate the final output of encrypted feature sequences.

[0124] As Figure 7 shown, in Step S9, the specific method for calculating the contrast between subsets is as follows:

[0125] Step S9-1: Receive the classification table of encrypted feature sequences generated according to the encrypted feature sequences, group the subsets in the classification table according to a predetermined rule, and each group contains several subsets:

[0126] The table obtained previously by analyzing the encrypted feature sequences, which contains the classification information of each subset and its encrypted feature sequences. These subsets are further organized into larger groups according to specific criteria (such as the similarity of feature sequences or specific attributes), and each group includes multiple subsets with similar characteristics or meeting specific conditions. For example, subsets with similar data density or feature distribution can be grouped together.

[0127] Step S9-2: Calculate the feature vector of each subset. Based on the encrypted feature values of all data items within the subset, use the weighted average method to generate the subset feature vector;

[0128] When calculating the feature vector of each subset, the weighted average method is adopted, based on the encrypted feature values of all data items within the subset. The specific calculation formula is as follows:

[0129] Suppose subset S contains n data items, and each data item i has an encrypted feature value υ i and the corresponding weight ω i . The feature vector V S of subset S can be calculated through the following weighted average method:

[0130]

[0131] where, is the sum of weighted feature values, and is the sum of all weights.

[0132] Step S9-3: Conduct a difference analysis of the eigenvectors between subsets. Use the vector space model to calculate the Euclidean distance between each pair of subset eigenvectors to obtain a preliminary contrast matrix between subsets.

[0133] Suppose there are two subsets of eigenvectors V A and V B , and each vector is composed of eigenvalues in d dimensions, that is, V A =(a1, a2, …, a d ) and V B =(b1, b2, …, b d ). The Euclidean distance D(V A , V B ) between these two subsets can be calculated by the following formula:

[0134]

[0135] Here, represents the sum of the squares of the differences of the corresponding eigenvalues between V A and V B . The smaller the Euclidean distance, the more similar the two subsets are in terms of features; the larger the distance, the more significant the difference.

[0136] Step S9-4: Normalize the preliminary contrast matrix to form a standardized contrast matrix.

[0137] Step S9-5: Perform dimensionality reduction on the standardized contrast matrix to map the high-dimensional contrast data into a two-dimensional or three-dimensional space.

[0138] It can be understood that if the original contrast matrix contains dozens of dimensions, through PCA, it can be converted into two or three main components, which not only makes the visualization and analysis of the data easier, but also helps to more clearly identify and interpret the main differences between the data.

[0139] Step S9-6: In the dimensionality-reduced space, calculate the density distribution between subsets, and use the density clustering algorithm to identify the density peak regions between subsets.

[0140] After analysis by the density clustering algorithm, it can be found that some data subsets form obvious dense regions in the two-dimensional or three-dimensional space. These regions are the density peak regions, which reflect the aggregation trend of the data under specific eigenvector conditions and help to identify and analyze data subsets with similar features or behaviors.

[0141] Step S9-7: Based on the density peak regions, calculate the final contrast value between each pair of subsets to form a contrast map between subsets.

[0142] Step S9-8: Determine the similarities and differences between subsets based on the contrast map between subsets, and generate an optimized subset classification table.

[0143] In this embodiment, as Figure 3 shown, the encryption and decryption processes of data requests are presented. When encrypting user data requests, symmetric encryption algorithms such as AES, DES, or 3DES can be used to ensure data security; while when parsing the encrypted user data requests, asymmetric decryption algorithms such as RSA, ECC, or DSA can be adopted to securely and effectively decrypt and extract the key information of the data requests.

[0144] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of this application is defined by the appended claims rather than the above description. Thus, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within this application. Any reference signs in the claims should not be construed as limiting the claims concerned. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the apparatus claims can also be implemented by the same unit or device through software or hardware. First, second, etc. are used to denote names and do not represent any particular order.

[0145] The above embodiments are only used to illustrate the technical solutions of this application and not to limit them. Although this application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of this application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of this application.

Claims

1. A data retrieval method based on information security, characterized in that: The data retrieval method comprises the following steps: Receive user data requests and encrypt them; Parse the encrypted user data request, extract feature information and perform preprocessing; Perform keyword extraction and semantic analysis on the preprocessed feature information to form a standardized keyword set; Use standardized keyword sets to search in the encrypted data index to generate preliminary candidate data sets; Perform security filtering on the preliminary candidate data set, exclude data items that do not meet security requirements according to the preset security policy, and obtain a secure candidate data set; In the security candidate data set, similarity is calculated based on the encrypted feature values ​​of the data items, and highly matching data items are screened out through the set similarity threshold to generate preliminary retrieval results; Classify the data items in the preliminary search results, calculate the security weight of each data item and sort them to form a sorted list; Divide the data items in the sorted list into several subsets according to a predetermined rule, calculate the data feature distribution of each subset, and generate an encrypted feature sequence; According to the standard deviation of the encrypted feature sequence and the encrypted similarity distance between features, the contrast between subsets is calculated to generate the feature index of each data item; Combine the security weight and feature index of the data item to generate a comprehensive feature vector; Using comprehensive feature vectors, cluster analysis is performed through encrypted clustering algorithms to identify and divide data clusters; Count the total amount of encrypted data in each cluster and calculate the encrypted feature distribution of the cluster; According to the encrypted feature distribution of clusters, the security quality parameters of data retrieval are calculated, the final encrypted retrieval results are generated and decrypted for output.

2. A data retrieval method based on information security according to claim 1, characterized in that: The specific method for encrypting the user data request is: encrypting the user data request using a symmetric encryption algorithm to generate a ciphertext request.

3. A data retrieval method based on information security according to claim 1, characterized in that: The specific method for parsing the encrypted user data request is: decrypting the encrypted user data request through a parsing algorithm, extracting feature information and performing data preprocessing, including data cleaning and format standardization.

4. The data retrieval method based on information security according to claim 1, characterized in that: The specific method of using a standardized keyword set to search in an encrypted data index is: matching the standardized keyword set with the encrypted data index, and retrieving and generating a preliminary candidate data set through an encrypted search algorithm.

5. The data retrieval method based on information security according to claim 1, characterized in that: The specific method for calculating similarity based on the encrypted feature value of the data item is: In the received security candidate data set, feature encoding processing is performed on each data item, each data item is encrypted using an asymmetric encryption algorithm, and its encrypted feature value is retained; Perform feature decomposition on each encrypted data item, split each encrypted data item into basic feature units, and generate an encryption identifier for each feature unit; Perform feature analysis on the basic feature units of each data item, calculate the correlation between feature units through a neural network, and generate a feature correlation graph; Based on the feature association graph, the graph theory algorithm is used to calculate the encrypted feature similarity between data items and determine the similarity relationship between each data item; A similarity threshold is set, and based on the calculated encrypted feature similarity, data item pairs with values ​​higher than the set threshold are screened out to form a preliminary set of similar data items.

6. A data retrieval method based on information security according to claim 5, characterized in that: The method for dividing data items into several subsets based on a sorted list and generating an encrypted signature sequence comprises the following steps: Receiving a sorted list generated according to similarity calculation, dividing the data items in the sorted list into a plurality of subsets according to a predetermined rule, each subset containing a plurality of data items; Perform feature encoding on the data items of each subset, encrypt the feature value of each data item using an encryption algorithm, and generate an encrypted feature value; Calculate the data feature distribution of each subset, count the distribution of encrypted feature values ​​of all data items in each subset, and form a feature distribution matrix; Standardize the feature distribution matrix and convert the values ​​in the feature distribution matrix into standardized feature values; Combining the standardized feature values ​​into an encrypted feature sequence in order; Performing sequence analysis on the encrypted signature sequence; According to the results of sequence analysis, the encrypted feature sequences of each subset are classified to form an encrypted feature sequence classification table; The encrypted feature sequence classification table is compared with the sorted list of original data items to verify the accuracy and completeness of the encrypted feature sequence and generate the final encrypted feature sequence output.

7. A data retrieval method based on information security according to claim 6, characterized in that: The specific method for calculating the contrast between subsets is: Receive the encrypted feature sequence classification table generated according to the encrypted feature sequence, and group the subsets in the classification table according to a predetermined rule, each group containing a number of subsets: Calculate the eigenvector of each subset, and generate the subset eigenvector using the weighted average method based on the encrypted eigenvalues ​​of all data items in the subset; Perform difference analysis of feature vectors between subsets, use vector space model to calculate the Euclidean distance between feature vectors of each pair of subsets, and obtain the preliminary contrast matrix between subsets; Normalizing the preliminary contrast matrix to form a standardized contrast matrix; Perform dimensionality reduction on the standardized contrast matrix to map the high-dimensional contrast data into a two-dimensional or three-dimensional space; In the dimensionality reduction space, the density distribution between subsets is calculated, and the density peak area between subsets is identified using the density clustering algorithm; Based on the density peak area, the final contrast value between each pair of subsets is calculated to form a contrast map between subsets; According to the contrast map between subsets, the similarities and differences between subsets are determined, and an optimized subset classification table is generated.

8. The data retrieval method based on information security according to claim 1, characterized in that: When encrypting user data requests, the symmetric encryption algorithms used include AES, DES, or 3DES.

9. The data retrieval method based on information security according to claim 1, characterized in that: When parsing the encrypted user data request, the adopted parsing algorithm includes RSA, ECC or DSA.

10. The data retrieval method based on information security according to claim 1, characterized in that: When searching using a standardized keyword set, the encrypted search algorithms used include Bloom filters, fully homomorphic encryption, or partially homomorphic encryption.