A data backup method and system based on encryption algorithms
By using regional growth algorithms and clustering algorithms in differential privacy protection technology to evaluate the leakage probability and confidentiality needs of data items, combined with random noise and encryption technology, the problem of difficult to balance data availability and privacy protection in the existing technology is solved, and more efficient data security and privacy protection is achieved.
Patent Information
- Application Number
- CN202510352224.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Existing differential privacy protection technologies are difficult to balance data availability and protection of personal information, and it is difficult to determine the parameters for adding random noise.
By obtaining the data to be backed up by users, using the region growth algorithm and clustering algorithm, the first leakage probability and confidentiality requirements of each data item are evaluated, the privacy budget is determined based on similarity characteristics, random Laplace noise is added to the word vectors and numbers in the data unit, and encrypted backups are performed.
It achieves a more accurate assessment and meeting the privacy protection needs of each data unit, enhances data security and user privacy protection, and avoids the risk of personal privacy leakage caused by data breaches.
Smart Images

Figure CN119885242B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a data backup method and system based on an encryption algorithm. Background Art
[0002] With the rapid development of informatization, a large amount of production and life data is generated daily in various industries. For example, in a trading platform, a large amount of sensitive personal data such as transaction records and user account information is involved. For security reasons, the platform usually needs to back up the database regularly. However, by deeply analyzing, mining, and performing machine learning modeling on this massive amount of data, the platform can extract the potential value in the data, thereby gaining insights into user behavior and providing more personalized services for users.
[0003] Differential privacy is a technique that interferes with user data by introducing a privacy protection algorithm, usually executed in a trusted execution environment. In this method, by adding an appropriate amount of noise to the statistical analysis results, the sensitive information of individuals can be effectively blurred, making it impossible to trace back to a specific personal identity. This approach can not only protect the privacy data of users but also ensure that when performing data analysis, general information with statistical significance can still be extracted, meeting the needs of the platform in data mining and user behavior analysis while avoiding potential data leakage risks.
[0004] However, for existing differential privacy protection technologies, it is difficult to select the key parameters of differential privacy, that is, it is difficult to balance the data availability and the degree of protection of personal information; when the added random noise falls within a very small interval, attackers can easily analyze and obtain accurate user sensitive data through differential attacks and probabilistic inference attacks. Summary of the Invention
[0005] In view of the above, it is necessary to provide a data backup method and system based on an encryption algorithm to solve the above problems.
[0006] The first aspect of this application provides a data backup method based on an encryption algorithm, and the method includes:
[0007] Obtain the data to be backed up of several users, where the data to be backed up includes several data units; each data unit includes several data items, and obtain a C-dimensional vector for the content corresponding to each data item; where C is a preset value;
[0008] Apply the region growing algorithm to the set composed of the C-dimensional vectors corresponding to each data item of each data unit of all users, and combine the number of elements in the set to obtain the first leakage probability of each data item of each data unit of each user;
[0009] Classify all data items of each user by combining the clustering algorithm according to the degree of chaos of the C-dimensional vectors corresponding to each data item in each data unit among all the data to be backed up by users; combine every two data items in all data units of each user, and obtain the confidentiality requirements of each data unit of each user based on the first leakage probability and classification information of the data items in all combinations.
[0010] According to the similarity characteristics between all data items of any two data units belonging to different users, combine the clustering algorithm and the confidentiality requirements to obtain the privacy budget of each data unit of each user; add random Laplace noise variables to the word vectors and numbers in the data unit based on the privacy budget and restore them, and encrypt and back up the restored data shards.
[0011] Among them, the process of obtaining the C-dimensional vector corresponding to the content of each data item is specifically as follows:
[0012] If the content of the data item is a number, place the number in the first dimension of the C-dimensional vector, and fill the remaining dimensions with 0; if the content of the data item is text, perform word segmentation on the text to obtain the word vector corresponding to the text, reduce each word vector to a C-dimensional vector, fill with 0 at the end if it is less than C dimensions, and add the C-dimensional vectors of all word vectors as the C-dimensional vector corresponding to the data item.
[0013] Among them, the process of obtaining the first leakage probability of each data item in each data unit of each user is as follows:
[0014] Denote the set composed of all C-dimensional vectors corresponding to the same data item in each data unit of all users as the value range of the data item corresponding to each data unit;
[0015] Apply the region growing algorithm to the value range of each data item in each data unit to obtain several regions;
[0016] The formula form of the first leakage probability of the j-th data item in the i-th data unit of the e-th user is: ; where represents the first leakage probability of the j-th data item in the i-th data unit of the e-th user; represents the number of users in the region where the j-th data item in the i-th data unit is located; represents the number of users; represents the number of C-dimensional vectors in the value range of the j-th data item in the i-th data unit.
[0017] Among them, the steps of applying the region growing algorithm to the value range of each data item in each data unit to obtain several regions are as follows:
[0018] Take the C-dimensional vector corresponding to the j-th data item of the i-th data unit of the user who is obtained first as the growth seed;
[0019] The growth condition of the region growth algorithm is that the cosine similarity between the C-dimensional vector corresponding to the j-th data item of the i-th data unit of the user and the growth seed is greater than a preset value.
[0020] Among them, the specific method of classifying all data items of each user by combining the clustering algorithm is as follows:
[0021] Obtain the information entropy of the C-dimensional vector corresponding to the content of each data item among all data items of all users;
[0022] Use the clustering algorithm to cluster all data items of each user to obtain a preset number of clustering clusters;
[0023] Obtain the mean value of the information entropy corresponding to all data items in each clustering cluster, and record all data items in the clustering cluster with the largest mean value of the information entropy as commodity-type data items; record all data items in the clustering cluster with the smallest mean value of the information entropy as personal-type data items.
[0024] Among them, the specific method of obtaining the confidentiality requirement of each data unit of each user is as follows:
[0025] Analyze the first leakage probability and the category information of the data items in each combination to obtain the second leakage probability and the interception probability of each combination of each data unit of each user;
[0026] Take the product of the second leakage probability and the interception probability of each combination of each data unit of each user as the confidentiality requirement information of each combination of each data unit of each user;
[0027] Take the maximum value of the confidentiality requirement information in each data unit as the confidentiality requirement of each data unit.
[0028] Among them, the formula form of the second leakage probability is specifically as follows: ; the formula form of the interception probability is specifically as follows: ; where represents the second leakage probability of the b-th combination data item of the i-th data unit of the e-th user; represents the number of commodity-type data items in the b-th combination data item of the i-th data unit of the e-th user, represents the number of personal-type data items in the b-th combination data item of the i-th data unit of the e-th user; u represents the number of all data items included in a user; represents the first leakage probability of all commodity-type data items in the b-th combination data item of the i-th data unit of the e-th user; represents the first leakage probability of all personal data items in the b-th combined data item of the i-th data unit of the e-th user; represents a normalization function; represents the interception probability of the b-th combined data item of the i-th data unit of the e-th user; represents the preset type of attack required for each combined data item of each data unit of each user.
[0029] Among them, the process of obtaining the privacy budget of each data unit of each user is specifically as follows:
[0030] According to the similarity characteristics between the C-dimensional vectors corresponding to all data items of any two data units belonging to different users, obtain the difference degree between any two data units belonging to different users;
[0031] Take the difference degree as the distance between data units in the clustering algorithm, cluster the data units of all users, calculate the average value of the confidentiality requirements of all data units in each cluster, and take the product of the normalized result of the average value of the confidentiality requirements and the preset parameter value as the privacy budget of all data units in each cluster.
[0032] Among them, the process of obtaining the difference degree between any two data units belonging to different users is specifically as follows:
[0033] Calculate the cumulative sum of the cosine similarities between the C-dimensional vectors corresponding to each data item of any two data units belonging to different users, and take the negative correlation mapping result of all the cumulative sums of the any two data units as the difference degree between the any two data units.
[0034] In a second aspect, an embodiment of the present application further provides a data backup system based on an encryption algorithm, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0035] The present application has at least the following beneficial effects:
[0036] In the embodiments of the present application, in view of the problems in the existing differential privacy protection technology that it is difficult to balance data availability and the degree of protection of personal information, and it is difficult to determine the addition of random noise, the data to be backed up of platform users is analyzed. First, for the set of C-dimensional vectors corresponding to each data item of each data unit of all users, the region growing algorithm is used, and combined with the number of elements in the set, the first leakage probability of each data item of each data unit of each user is obtained, which helps to evaluate the privacy leakage risk of each data item of the user; further, since the personal information of the user changes less in the data of various platforms, while the transaction information changes at any time, so according to the degree of chaos of the C-dimensional vectors corresponding to each data item of each data unit in the data to be backed up of all users, combined with the clustering algorithm, all data items of each user are classified, which helps to evaluate the privacy risk and leakage possibility of the data according to the change characteristics of the data item information. This evaluation can help to better understand the sensitivity of the data item and provide a basis for taking more targeted privacy protection measures; also because different combinations of data items can release more information, for each pair of all data items in each data unit of each user, based on the first leakage probability and classification information of the data items in all combinations, the confidentiality requirements of each data unit of each user are obtained, which helps to accurately evaluate the privacy protection requirements of each data unit and facilitate subsequent determination of relevant parameters in the differential privacy protection technology according to the leakage probability and the sensitivity of the data combination, and perform encryption to further ensure the security of the data, and then back up the data, so that even if the data is leaked, personal privacy data cannot be obtained, which is beneficial to improving the security of the backed-up data in the database; this method directly introduces a privacy protection mechanism on the user device, ensuring that the data does not expose the sensitive information of individuals during storage and processing, effectively preventing possible privacy attacks, and at the same time being able to retain the analysis value of the data. This method significantly enhances data security and user privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of the steps of a data backup method based on an encryption algorithm provided by an embodiment of the present application;
[0038] Figure 2 It is a flowchart for obtaining the privacy budget provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary", "or", "for example" is intended to present relevant concepts in a specific manner.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application.
[0041] In addition, it should be noted that the terms "first" and "second" in this application and the accompanying drawings are used to distinguish similar objects and are not used to describe a specific order or sequence. For the methods disclosed in the embodiments of this application or the methods shown in the flowcharts, including one or more steps for implementing the methods, without departing from the scope of protection of this application, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs.
[0043] The following specifically describes the specific solutions of a data backup method and system based on an encryption algorithm provided by this application with reference to the accompanying drawings.
[0044] Please refer to Figure 1 , which shows a flowchart of the steps of a data backup method based on an encryption algorithm provided by an embodiment of this application. The method includes the following steps:
[0045] The first step: Obtain the data to be backed up of several users. The data to be backed up contains several data units. Each data unit contains several data items, and a C-dimensional vector is obtained for the content corresponding to each data item.
[0046] In the process of this application using a privacy algorithm to interfere with user data, a balance is found between ensuring the availability and effectiveness of the data and protecting personal privacy. Based on this, obtain the data to be backed up of several users. The data to be backed up contains several data units, and each data unit contains several data items.
[0047] In this embodiment, taking a trading platform as an example, the data to be backed up of users includes personal information and several transaction information. Both personal information and transaction information are a data unit. Personal information includes several data items such as name, gender, and account ID; transaction information includes several data items such as product name, order number, shipping address, and receiving address.
[0048] Since the content corresponding to the data item contains text data, in this embodiment, the maximum matching word segmentation algorithm is used to segment the text data, and any vocabulary is converted into a word vector through the Word2Vec model. Among them, converting text data into word vectors is a well-known prior art, and this application will not elaborate on this.
[0049] Obtain a C-dimensional vector for the content corresponding to each data item: If the content of the data item is a number, place the number in the first dimension of the C-dimensional vector and fill the remaining dimensions with 0; if the content of the data item is text, after word segmentation to obtain multiple word vectors, first reduce the dimension of each word vector to a C-dimensional vector using principal component analysis. If it is less than C dimensions, fill 0 at the end, and then add the C-dimensional vectors corresponding to each word vector as the vector corresponding to the data item. Since there are some extremely high-dimensional word vectors in the content of the data item, such as product names, the value of C should not be too small. In this embodiment, the value of C is 3000, and the implementer can adjust it according to the actual situation. It should be understood that in this embodiment, one data item corresponds to one C-dimensional vector.
[0050] The second step: Apply the region growing algorithm to the set composed of the C-dimensional vectors corresponding to each data item of each data unit of all users, and combine the number of elements in the set to obtain the first leakage probability of each data item of each data unit of each user.
[0051] In the differential privacy protection of this application, the ideal goal of privacy protection is carried out in units of data units: Even if a certain data unit is leaked, it is impossible to accurately identify the user to which the data unit belongs. For example, in a trading platform, when using the Laplace mechanism for protection, since the privacy requirements of data items such as gender, account ID, product name, order number, etc. are different, the risk of leaking user privacy when different data items are queried will also be different. In addition, different types of products also have different privacy risks. For example, there are differences in the confidentiality requirements between daily necessities and niche products. The latter has a higher risk of leaking user information due to its stronger directivity, and the consequences of leakage are more serious. Therefore, it is necessary to set different privacy budgets for each data unit according to the privacy risks of different data units.
[0052] First, analyze the value range of each data item of each data unit. For example, the number of elements contained in the value range of the data item gender is 2, the value range of the product name is relatively wide, and the value range of the account ID is related to the number of users; the value range represents the richness of the data item. The smaller the value range, the more monotonous the data item, and the smaller the user's selection range. Then, when the data item is obtained, the possibility of leaking user information is smaller. However, at the same time, although the value range of the account ID is large, it has directivity, and each user's account ID is unique. Then, when the data item is obtained, the possibility of leaking user information is very large.
[0053] In this application, a set composed of all C-dimensional vectors corresponding to the same data item in each data unit for all users is obtained, denoted as the value range of the data item corresponding to each data unit; the region growing algorithm is applied to the value range of each data item in each data unit to obtain several regions, where the C-dimensional vector corresponding to the j-th data item of the i-th data unit of the user obtained first is used as the growth seed; the growth condition of the region growing algorithm is that the cosine similarity between the C-dimensional vector corresponding to the j-th data item of the i-th data unit of the user and the growth seed is greater than a preset value. It should be noted that when the growth terminates, the C-dimensional vector corresponding to the j-th data item of the i-th data unit of the user obtained first that has not grown is used as the growth seed to continue growing until all users have completed growth; if there are users that do not belong to any region, they are incorporated into the cluster where the user with the highest cosine similarity to them is located.
[0054] The vector distribution in the region can reflect the data distribution characteristics of the data item. For a cluster with more users, its first leakage probability is lower. On this basis, the weaker the directivity of the data item, the lower the first leakage probability. Based on this, according to the number of elements in the region where the C-dimensional vector corresponding to each data item in each data unit of each user is located, combined with the number of elements in the value range of each data item, the first leakage probability of each data item in each data unit of each user is obtained. The formula for the first leakage probability of the j-th data item of the i-th data unit of the e-th user is: ; where represents the first leakage probability of the j-th data item of the i-th data unit of the e-th user; represents the number of users in the region where the j-th data item of the i-th data unit of the e-th user is located; represents the number of users; represents the number of C-dimensional vectors in the value range of the j-th data item of the i-th data unit, The closer it is to 1, the stronger the directivity.
[0055] The third step: Classify all data items of each user according to the degree of chaos of the C-dimensional vectors corresponding to each data item in each data unit in the data to be backed up by all users, combined with the clustering algorithm; for each pair of all data items in each data unit of each user, based on the first leakage probability and classification information of the data items in all combinations, the confidentiality requirements of each data unit of each user are obtained.
[0056] When analyzing the data items in each data unit, it is necessary to consider the distribution characteristics of the data items. For example, for a data item such as a product name, its first leakage probability is closely related to the popularity and demand degree of the product. If a product is common and in wide demand, such as daily necessities, many people will buy it, then the risk of leaking user privacy is relatively low because it has strong non - directivity. However, for some special products, such as experimental materials or large musical instruments, their purchases often have stronger directivity and seasonality. Therefore, once obtained, the risk of leaking user information will be higher. In other words, the universality and personalization degree of a product directly affect its first leakage probability. The privacy leakage risk of common products is relatively low, while that of personalized products is relatively high.
[0057] According to the degree of chaos of the content of each data item in each data unit in all users' data to be backed up, combined with the clustering algorithm, classify all data items of each user: obtain the information entropy of the C - dimensional vector corresponding to the content of each data item in all users' data items; use the clustering algorithm to cluster all data items of each user to obtain a preset number of clustering clusters; obtain the average value of the information entropy corresponding to all data items in each clustering cluster, and record all data items in the clustering cluster with the largest average value of the information entropy as product - type data items; record all data items in the clustering cluster with the smallest average value of the information entropy as personal - type data items. In this embodiment, the preset number takes a value of 2, which is used to divide all data items of a user into two categories.
[0058] The first leakage probability of a data unit cannot be simply calculated by adding the first leakage probabilities of each data item. The reason is that: the amount of information revealed by different combinations of data items is different. For example, when data items such as products, prices, and delivery addresses are combined, the combination of products and prices may more reflect information related to the products themselves, while the combination of products and delivery addresses may reveal more details about user behavior. For example, if the product and the delivery address are combined, it can be inferred about the user's purchase habits and interests. Especially when it comes to daily necessities, it may indicate that the user tends to buy products on the online platform; if it is other niche products, the combined information can more clearly reflect the user's personalized needs and behavior patterns. Therefore, different combinations of data items can release more information than just the content exposed by individual data items.
[0059] In addition, the difficulty or probability of intercepting different combinations of data items is not a simple linear relationship. Taking differential attacks and repeated query attacks as examples, the difference between two data units of the same user lies in the commodity-related information. Therefore, only a small number of differential attacks are required to determine the commodity-related privacy. The repeated part of two data units of the same user lies in the personal-related information, such as account nickname, mobile phone number, delivery address, etc. These information can achieve the purpose of protecting privacy by adding a small amount of noise. However, if the added noise is within a small range, the relevant privacy of the user can be obtained through repeated queries.
[0060] For each pair of combinations of all data items in each data unit of each user, based on the first leakage probability and the category information of the data items in each combination, the second leakage probability and the interception probability of each combination of each data unit of each user are obtained. Among them, the formula form of the second leakage probability is: , where represents the second leakage probability of the b-th combination of data items in the i-th data unit of the e-th user; represents the number of commodity-type data items in the b-th combination of data items in the i-th data unit of the e-th user, represents the number of personal-type data items in the b-th combination of data items in the i-th data unit of the e-th user; u represents the number of all data items included in a user; represents the first leakage probability of all commodity-type data items in the b-th combination of data items in the i-th data unit of the e-th user; represents the first leakage probability of all personal-type data items in the b-th combination of data items in the i-th data unit of the e-th user; represents the normalization function.
[0061] The formula form of the interception probability is: ; where represents the interception probability of the b-th combination of data items in the i-th data unit of the e-th user; represents the number of commodity-type data items in the b-th combination of data items in the i-th data unit of the e-th user, represents the number of personal-type data items in the b-th combination of data items in the i-th data unit of the e-th user; represents the preset number of attack types required for each combination of data items in each data unit of each user. In this embodiment, the value is 2, and the implementer can adjust it according to the actual situation; exp() represents the exponential function with the natural constant as the base.
[0062] Take the product of the second leakage probability and the interception probability of each combination of each data unit of each user as the confidentiality requirement factor of each combination of each data unit of each user. The degree of data security lies in the short board. Therefore, taking the maximum value of the confidentiality requirement information in each data unit as the confidentiality requirement of each data unit can ensure that all combined data items in the data unit can meet the confidentiality requirement.
[0063] The fourth step: According to the similarity characteristics between all data items of any two data units belonging to different users, combined with the clustering algorithm and the confidentiality requirements, obtain the privacy budget of each data unit of each user; based on the privacy budget, add random Laplace noise variables to the word vectors and numbers in the data unit and restore them, and encrypt and back up the restored data shards.
[0064] The above steps only analyze a single user and do not consider the differences between all users, which may lead to uneven distribution of privacy budgets among users. Therefore, it is necessary to analyze the number of data units of different users and construct a unified privacy budget allocation model.
[0065] First, according to the similarity characteristics between the C-dimensional vectors corresponding to all data items of any two data units belonging to different users, obtain the difference degree between any two data units belonging to different users: calculate the cumulative sum of the cosine similarities between the C-dimensional vectors corresponding to each data item of any two data units belonging to different users, and take the negative correlation mapping result of all the cumulative sums of the any two data units as the difference degree between the any two data units. In this embodiment, the cumulative sum of the cosine similarities between the C-dimensional vectors corresponding to each data item of any two data units belonging to different users is denoted as A, and the formula form of the difference degree between the any two data units is: .
[0066] There are similar purchased goods among different users, and similar goods should have similar confidentiality requirements. Take the difference degree between data units as the distance, and use the DBSCAN algorithm to cluster the data units of all users. In this embodiment, the neighborhood radius is , and the minimum number of sample points is 2; obtain several clustering clusters, denote the data units included in the same clustering cluster as similar data units, and denote the mean value of the confidentiality requirements of the similar data units as the confidentiality requirement of the belonging clustering cluster; calculate the privacy budget of the data units in each clustering cluster , and its formula form is: , where: represents the confidentiality requirement of the data units of each cluster class, represents the preset parameter value, and the value in this embodiment is , which means controlling the range of the privacy budget between 0 and 10, and the implementer can adjust it according to the specific situation; represents a normalization function. It should be understood that the privacy budgets of data units in the same clustering cluster are equal. Among them, the flowchart for obtaining the privacy budget is as Figure 2 shown.
[0067] After obtaining the privacy budget , add random Laplace noise variables to the word vectors and numbers in the data unit, and the noise , and then restore the noisy word vector to the word with the maximum cosine similarity to obtain the noisy data unit; among them, represents an independent and identically distributed random Laplace noise variable that follows the scale parameter is , which is a well-known prior art, and this application will not elaborate on it.
[0068] Use the AES algorithm to perform sharding encryption on the noisy data unit, complete the encryption to obtain the encrypted data unit, create a backup file storage database, and export the encrypted data unit to the backup file, which is a well-known prior art and will not be elaborated here.
[0069] This application adds noise and encrypts the data to be backed up, so that even if the backup data is intercepted, the original data cannot be obtained.
[0070] Based on the same inventive concept as the above method, an embodiment of this application also provides a data backup system based on an encryption algorithm, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods of a data backup method based on an encryption algorithm.
[0071] In summary, in view of the problems in the existing differential privacy protection technology, namely, it is difficult to balance the data availability and the degree of protection of personal information, and it is difficult to determine the addition of random noise, the present application analyzes the data to be backed up of platform users. First, for the set of C-dimensional vectors corresponding to each data item in each data unit of all users, the region growing algorithm is adopted, and combined with the number of elements in the set, the first leakage probability of each data item in each data unit of each user is obtained, which helps to evaluate the privacy leakage risk of each data item of the user; further, since the personal information of the user changes little in the data of various platforms, while the transaction information changes at any time, according to the degree of chaos of the C-dimensional vectors corresponding to each data item in each data unit in the data to be backed up of all users, and combined with the clustering algorithm, all data items of each user are classified, which helps to evaluate the privacy risk and leakage possibility of the data according to the change characteristics of the data item information. This evaluation can help better understand the sensitivity of the data item and provide a basis for taking more targeted privacy protection measures; also, because different combinations of data items can release more information, for each pair of all data items in each data unit of each user, based on the first leakage probability of the data items and the classification information in all combinations, the confidentiality requirement of each data unit of each user is obtained, which helps to accurately evaluate the privacy protection requirement of each data unit and is convenient for subsequent determination of relevant parameters in the differential privacy protection technology according to the leakage probability and the sensitivity of the data combination, and encryption is performed to further ensure the security of the data, and then the data is backed up, so that even if the data is leaked, personal privacy data cannot be obtained, which is beneficial to improving the security of the backed-up data in the database; this method directly introduces a privacy protection mechanism on the user device, ensuring that the sensitive information of the individual will not be exposed during the storage and processing of the data, effectively preventing possible privacy attacks, and at the same time being able to retain the analysis value of the data. This method significantly enhances data security and user privacy protection.
[0072] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the description corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0073] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the basic characteristics of the present application. Therefore, from any point of view, the above-described embodiments of the present application should be considered exemplary and non-restrictive; modifications to the technical solutions described in the foregoing embodiments, or equivalent replacements of some of the technical features, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A data backup method based on encryption algorithm, characterized in that: The method comprises the following steps: Acquire data to be backed up of several users, wherein the data to be backed up includes several data units; the data units include several data items, and obtain a C-dimensional vector for the content corresponding to each data item; wherein C is a preset value; A region growing algorithm is used on a set consisting of C-dimensional vectors corresponding to each data item in each data unit of all users, and a first leakage probability of each data item in each data unit of each user is obtained in combination with the number of elements in the set; According to the disorder degree of the C-dimensional vector corresponding to each data item of each data unit in the data to be backed up by all users, all data items of each user are classified in combination with the clustering algorithm; all data items in each data unit of each user are combined in pairs, and based on the first leakage probability of the data items in all combinations and the classification information, the confidentiality requirement of each data unit of each user is obtained; According to the similarity characteristics between the C-dimensional vectors corresponding to all data items of any two data units belonging to different users, the difference between any two data units belonging to different users is obtained, and the difference is used as the metric of the clustering algorithm for clustering. The overall distribution characteristics of confidentiality requirements in each cluster cluster are analyzed to obtain the privacy budget of each data unit of each user; based on the privacy budget, random Laplace noise variables are added to the word vectors and numbers in the data units and restored, and the restored data fragments are encrypted and backed up.
2. A data backup method based on encryption algorithm as claimed in claim 1, characterized in that: The C-dimensional vector is obtained for the content corresponding to each data item, specifically: If the content of the data item is a number, the number is placed in the first dimension of the C-dimensional vector, and the remaining dimensions are filled with 0; if the content of the data item is text, the text is segmented to obtain the word vector corresponding to the text, and each word vector is reduced to a C-dimensional vector. If it is less than C dimensions, 0 is added at the end, and the C-dimensional vectors of all word vectors are added together to obtain the C-dimensional vector of the corresponding data item.
3. The data backup method based on encryption algorithm as claimed in claim 1, characterized in that: The process of obtaining the first leakage probability of each data item of each data unit of each user is as follows: The set of all C-dimensional vectors corresponding to the same data item of all users in each data unit is recorded as the value range of the data item corresponding to each data unit; A region growing algorithm is used for the value range of each data item of each data unit to obtain several regions; The formula form of the first leakage probability of the jth data item of the i-th data unit of the e-th user is: ;in, represents the first leakage probability of the jth data item of the ith data unit of the eth user; The number of users in the region where the jth data item of the ith data unit of the eth user is located; Indicates the number of users; The number of C-dimensional vectors representing the value range of the j-th data item of the i-th data unit.
4. A data backup method based on encryption algorithm as claimed in claim 3, characterized in that: The steps of using a region growing algorithm for the range of each data item of each data unit to obtain a plurality of regions are as follows: The C-dimensional vector corresponding to the j-th data item of the i-th data unit of the first acquired user is used as the growth seed; The growth condition of the region growing algorithm is that the cosine similarity between the C-dimensional vector corresponding to the j-th data item of the i-th data unit of the user and the growing seed is greater than a preset value.
5. The data backup method based on encryption algorithm as claimed in claim 1, characterized in that: The combined clustering algorithm is used to classify all data items of each user, specifically: Obtain the information entropy of the C-dimensional vector corresponding to the content of each data item in the data items of all users; A clustering algorithm is used to cluster all data items of each user to obtain a preset number of clusters; The information entropy mean corresponding to all data items in each cluster is obtained, and all data items in the cluster with the largest information entropy mean are recorded as commodity data items; and all data items in the cluster with the smallest information entropy mean are recorded as personal data items.
6. The data backup method based on encryption algorithm as claimed in claim 1, characterized in that: The confidentiality requirement of each data unit of each user is specifically: Analyze the first leakage probability and category information of the data items in each combination to obtain the second leakage probability and interception probability of each combination of each data unit of each user; The product of the second leakage probability of each combination of each data unit of each user and the interception probability is used as the confidentiality requirement information of each combination of each data unit of each user; The maximum value of the confidentiality requirement information in each data unit is taken as the confidentiality requirement of each data unit.
7. A data backup method based on encryption algorithm as claimed in claim 6, characterized in that: The formula form of the second leakage probability is specifically: ; The formula form of the intercept probability is specifically: ;in, represents the second leakage probability of the bth combination data item of the ith data unit of the eth user; represents the number of commodity data items in the bth combination data item of the ith data unit of the eth user, represents the number of individual data items in the bth combination of data items in the ith data unit of the eth user; u represents the number of all data items contained in a user; represents the first leakage probability of all commodity data items in the bth combination data item of the i-th data unit of the e-th user; The first leakage probability of all individual data items in the b-th combination data item of the i-th data unit of the e-th user; represents the normalization function; represents the interception probability of the bth combination of data items of the i-th data unit of the e-th user; The preset attack type required for each combination of data items representing each data unit of each user.
8. The data backup method based on encryption algorithm as claimed in claim 1, characterized in that: The difference between any two data units belonging to different users is obtained as follows: The cumulative sum of cosine similarities between C-dimensional vectors corresponding to each data item of any two data units belonging to different users is calculated, and the negative correlation mapping results of all the cumulative sums of the any two data units are used as the difference between the any two data units.
9. The data backup method based on encryption algorithm as claimed in claim 1, characterized in that: The process of obtaining the privacy budget of each data unit of each user is specifically as follows: The difference is used as the distance between data units in the clustering algorithm, the data units of all users are clustered, the mean confidentiality requirement of all data units in each cluster is calculated, and the product of the normalized result of the mean confidentiality requirement and the preset parameter value is used as the privacy budget of all data units in each cluster.
10. A data backup system based on an encryption algorithm, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Security protection strategy optimization method and system for energy big data
CN117857019A
Power grid data privacy protection and security processing method and system based on encryption protection
CN118153071A