Privacy protection system based on big data security and privacy calculation

By designing a multi-module privacy protection system in the field of big data security and privacy computing, using data desensitization, masking and fusion technology, the problem of data availability reduction in existing systems when protecting data privacy is solved, and efficient and secure data analysis and calculation are achieved.

CN120162818AInactive Publication Date: 2025-06-17GUANGZHOU VOCATIONAL COLLEGE OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510205385.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

While protecting data privacy, existing privacy protection systems may lead to reduced data availability and difficulty in supporting complex data analysis. Encrypted data cannot directly participate in computing, and there is a risk of data leakage and performance losses.

Method used

A privacy protection system based on big data security and privacy computing is designed, including data acquisition, data processing, protection analysis and management analysis modules. Through technologies such as data desensitization, data masking and data fusion, data privacy protection is achieved while maintaining data availability and accuracy.

Benefits of technology

It realizes the availability and accuracy of data while protecting data privacy, supports complex data analysis and calculations, reduces the risk of data breaches and performance losses, and improves data security and privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162818A_ABST
    Figure CN120162818A_ABST
Patent Text Reader

Abstract

The invention discloses a privacy protection system based on big data security and privacy calculation, which relates to the technical field of data processing and comprises a data acquisition module, a data processing module, a protection analysis module and a management analysis module. According to the method, the privacy of the desensitized data is improved while the original numerical range and precision of the desensitized data are maintained in a data desensitization mode of random disturbance, and the desensitized data can still support complex analysis and calculation due to introduction of random disturbance; and the data mask form based on differential privacy ensures that the masked data can still support effective data analysis while protecting privacy, so that the method can improve the privacy of the data, the data fusion form of homomorphic encryption and differential privacy can realize efficient and safe data fusion on encrypted data, and the data fusion efficiency is improved. The fused data has privacy and meets the requirements of complex analysis and calculation, the defects of a traditional data privacy protection system can be overcome, and the method is more intelligent and practical.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a privacy protection system based on big data security and privacy computing. Background Art

[0002] With the rapid development of big data technology, data has become an important resource in modern society. Therefore, data security and data privacy protection have also become key issues. Big data technology is based on data collection, data storage, data processing, and data analysis. There may be risks of data privacy leakage in these processes. In order to avoid data privacy leakage, privacy protection technology based on big data security and privacy computing is required.

[0003] Although existing privacy protection systems can protect data privacy to a certain extent, there may be a situation where data availability is reduced. Furthermore, it may be difficult to support complex data analysis, and encrypted data may not be directly involved in calculations. Therefore, it may be necessary to decrypt the data before calculation, which poses a risk of data leakage. Moreover, there may be performance losses during data decryption and data encryption operations, resulting in poor functionality and practicability of existing privacy protection systems. Summary of the Invention

[0004] The purpose of the present invention is to provide a privacy protection system based on big data security and privacy computing, which solves the problems presented in the above background art.

[0005] To achieve the above purpose, the present invention provides the following privacy protection system based on big data security and privacy computing, including:

[0006] Data acquisition module: The data acquisition module acquires the original data to be processed and stores the original data to be processed in a database.

[0007] Data processing module: The data processing module is used to input the original data to be processed. The data processing module cleans and sorts the input data, and then stores the processed data in the database.

[0008] Protection analysis module: The protection analysis module extracts the original data matrix and the original data vector from the database, and inputs the element SSVM in the original data matrix corresponding to the original data matrix and the j-th original data point DMSL corresponding to the original data vector into the protection analysis module. The protection analysis module outputs the element in the desensitized data matrix, the j-th masked data point, and the k-th fused data point. p,i and the j-th original data point DMSL corresponding to the original data vector j into the protection analysis module. The protection analysis module outputs the element in the desensitized data matrix, the j-th masked data point, and the k-th fused data point.

[0009] Management Analysis Module: The management analysis module performs statistical analysis on the elements in the de - sensitized data matrix, the j - th masked data point, and the k - th fused data point to evaluate the impact of data privacy protection on task performance. Then, according to the evaluation results, it adjusts the de - sensitized, masked, and fused data. Finally, it continuously monitors and maintains the entire privacy protection process to ensure data security and privacy.

[0010] Optionally, the protection analysis module includes a data de - sensitization analysis sub - module, a data masking analysis sub - module, and a data fusion analysis sub - module.

[0011] Optionally, the calculation formula of the data de - sensitization analysis sub - module is as follows:

[0012]

[0013] Where:

[0014] SSVO p,q Refers to the element in the de - sensitized data matrix, located in the p - th row and q - th column;

[0015] SSVM p,i Refers to the element in the original data matrix, located in the p - th row and i - th column;

[0016] SSVR i,q Refers to the element in the random perturbation factor matrix, located in the i - th row and q - th column;

[0017] SSVN refers to the scaling perturbation factor,... refers to rounding down, d refers to the number of decimal places of the data, SSVM refers to the original data matrix, SSVR refers to the random perturbation matrix, SSVO refers to the de - sensitized data matrix, SSVR i,q +0.5 refers to rounding the result of random perturbation to the nearest integer, Mod 10 d Refers to the modulo operation, 10 d Indicates a decimal number with a modulus of 10;

[0018] The processing process of the data de - sensitization analysis sub - module is as follows:

[0019] Input the element SSVM in the original data matrix corresponding to the original data matrix p,i into the data de - sensitization analysis sub - module. Then, use a random number generator to generate a random perturbation factor matrix with the same size as the original data matrix, and input the element SSVM in the random perturbation factor matrix corresponding to the random perturbation factor matrix p,i into the data de - sensitization analysis sub - module, and output the de - sensitized data matrix SSVO and the corresponding element SSVO in the de - sensitized data matrix p,q .

[0020] Optionally, the calculation formula of the data masking analysis sub-module is as follows:

[0021]

[0022] Where:

[0023] DMSA j Refers to the data point after the j-th mask, DMSL j Refers to the j-th original data point, DMSL refers to the original data vector, DMSA refers to the data vector after masking, sgn(...) refers to the sign function, DMSU j Refers to a random number vector uniformly distributed in the range [0, 1), DQ refers to the sensitivity of the data, DWE refers to the differential privacy parameter of the privacy protection level, DWR refers to the differential privacy parameter of the failure probability;

[0024] Refers to first determining the sign of the Laplace noise through the sign function, and then based on the differential privacy parameter DWE of the privacy protection level, the sensitivity DQ of the data, and SSVR i,q , and finally performing scale adjustment through the natural logarithm ln(2 / DWR), and the result obtained in this way is the Laplace noise added to the original data point;

[0025] The processing process of the data masking analysis sub-module is as follows:

[0026] Input the element SSVR in the random perturbation factor matrix i,q Into the data masking analysis sub-module, then obtain the sensitivity DQ of the data based on the maximum difference between any two adjacent data points in the dataset, and generate a uniformly distributed random number vector DMSU of the same length as the original data vector DMSL using a random number generator j , and input the sensitivity DQ of the data and the random number vector DMSU j Into the data masking analysis sub-module, and output the data vector DMSA after masking and the corresponding j-th data point DMSA after masking j .

[0027] Optionally, the calculation formula of the data fusion analysis sub-module is as follows:

[0028]

[0029] YYD pq = |SSVO p,q - SSERT|;

[0030] Where:

[0031] FNKQ kRefers to the k-th fused data point, FNKQ refers to the fused encrypted data, I refers to the total number of data points participating in the fusion, i refers to the index used to traverse all data points participating in the fusion, ENF pk Refers to the homomorphic encryption function and uses the public key pk; ENF pk (DMSA j ) Refers to the result of homomorphically encrypting the j-th masked data point DMSA j using the public key pk, FNT refers to the fusion weight, YYD pq Refers to the correlation parameter of the masked and desensitized data point, n refers to a large prime number used for modular arithmetic, SSERT refers to the mean value of all elements in the desensitized data matrix SSVO;

[0032] The processing process of the data fusion analysis sub-module is as follows:

[0033] The elements SSVO in the desensitized data matrix p,q and the j-th masked data point DMSA j are input into the data fusion analysis sub-module, and the masked data is encrypted using the homomorphic encryption scheme to obtain the encrypted data ENF pk (DMSA j ), in the encrypted state, based on the output of this data fusion analysis sub-module, the fused encrypted data FNKQ is obtained, and the corresponding decryption algorithm is used to decrypt the fused encrypted data FNKQ to obtain the k-th fused data point FNKQ k .

[0034] Optionally, the data acquisition module is used to receive and acquire the original data to be processed, which is the basis for establishing the database and data privacy protection.

[0035] Optionally, the data cleaning in the data processing module is to remove duplicate data, remove invalid data and delete abnormal data, and verify the data.

[0036] Optionally, the sorting in the data processing module is to perform formatting and unification on the original data to be processed so that it meets the requirements of subsequent processing and analysis.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] 1. The data desensitization analysis sub-module of the present invention outputs the desensitized data matrix and the corresponding elements in the desensitized data matrix. This sub-module can obtain a desensitized data matrix, where each element is the result of the combination of the original data and the random perturbation factor, which can protect the data privacy. The desensitization process of the data through the random perturbation factor can protect the data privacy, while maintaining the availability and numerical range of the data. This sub-module makes the desensitized data different from the original data through random perturbation to protect the data privacy. The desensitized data still retains most of the information of the original data and can be used for subsequent analysis and calculations. This method better retains the availability and accuracy of the data while protecting the data privacy, and at the same time retains the availability and accuracy to provide desensitized data support for subsequent analysis and calculations, improves data security and reduces the risk of data leakage. Protecting data privacy through random perturbation while trying to retain the availability and accuracy of the data. The desensitized data matrix SSVO still retains most of the information of the original data and can be used for subsequent analysis and calculations.

[0039] 2. The data masking analysis sub-module of the present invention outputs the masked data vector DMSA and the corresponding j-th masked data point. This sub-module generates the j-th masked data point by adding Laplace noise to the j-th original data point through the sign function and the uniformly distributed random number. By adding noise, the difference between two adjacent data points is blurred, thus protecting the data privacy. The differential privacy protection is achieved by adding Laplace noise, and the addition of noise is random. The magnitude of the noise is related to the data sensitivity and the differential privacy parameter, and the differential privacy protection has a strict mathematical basis and can resist various attack methods. The calculation of the masked data vector DMSA and the j-th masked data point realizes the differential privacy protection, which protects the data privacy and provides privacy protection support for data publication and sharing. This sub-module achieves the balance between data privacy and availability through differential privacy protection and enables the masked data to still be used for complex analysis and calculations. The masked data maintains the availability and accuracy of the data while protecting the data privacy and can be used for subsequent analysis and calculations.

[0040] III. The data fusion analysis sub-module of the present invention outputs the fused encrypted data and the k-th fused data point. This sub-module realizes the fusion and analysis of data while protecting data privacy, bringing new possibilities to the fields of big data security and privacy computing. This sub-module performs operations on encrypted data based on homomorphic encryption, thus supporting relatively complex data analysis and calculations. Since fusion under encrypted data can protect the privacy of data, the calculation efficiency can be improved and the data leakage risk can be reduced by reducing the number of decryptions and avoiding the leakage of intermediate results. This method can support complex data analysis and calculations while protecting data privacy, improving the calculation efficiency and security. This sub-module fuses and calculates data in the encrypted state based on homomorphic encryption technology, avoiding the risk of the decryption process and data leakage. The fused encrypted data is generated in the encrypted state and can be directly used for subsequent analysis and calculations without decryption, which reduces the risk of data leakage and performance loss, and improves the calculation efficiency and security.

[0041] IV. The present invention performs cyclic iteration on the scaling perturbation factor in the data desensitization analysis sub-module based on the j-th masked data point and the j-th original data point, and continuously optimizes the scaling factor in the data privacy protection process through the iterative method to improve the effect of data privacy protection. The iterative process uses the calculation results of the data masking analysis sub-module as input to ensure the accuracy of the iterative process. The present invention combines the technical advantages of big data security and privacy computing to achieve the high efficiency and accuracy of data privacy protection. By iteratively optimizing the scaling factor, the risk of data privacy leakage can be more effectively controlled. This iterative method improves the effect of data privacy protection, reduces the risk of data leakage, and enhances the flexibility and adaptability of data privacy protection, and can cope with different data privacy protection requirements. Compared with traditional data privacy protection methods, this method realizes the dynamic adjustment and optimization of data privacy protection by iteratively optimizing the scaling factor. By iteratively optimizing the scaling factor, continuous improvement and optimization of data privacy protection can be achieved. This iterative form promotes the integration and development of big data security and privacy computing technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is the method step flow chart of the privacy protection system based on big data security and privacy computing;

[0043] Figure 2 is the overall structure schematic diagram of the privacy protection system based on big data security and privacy computing;

[0044] Figure 3 is the structure schematic diagram of the protection analysis module of the privacy protection system based on big data security and privacy computing. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] Regarding this privacy protection system based on big data security and privacy computing, it is different from the existing privacy protection systems based on big data security and privacy computing. Although the existing privacy protection systems can protect data privacy to a certain extent, the data availability may be reduced. As a result, it is difficult to support complex analysis, and the encrypted data may not be able to directly participate in the calculation. It may be necessary to decrypt it first and then perform the calculation, which poses a risk of data leakage and performance loss. This process has a risk of data leakage, and the decryption and encryption operations will result in performance loss.

[0047] However, the modules of this privacy protection system ensure that the desensitized data has improved privacy while maintaining the original numerical range and accuracy through the data desensitization method of random perturbation. At the same time, due to the introduction of random perturbation, the desensitized data can still support complex analysis and calculation. This method, based on the data masking form of differential privacy, ensures that the masked data can support effective data analysis and mining while protecting privacy, enabling this method to improve the privacy and availability of data. In addition, through the data fusion form of homomorphic encryption and differential privacy, efficient and secure data fusion is achieved on encrypted data, and the fused data not only has privacy but also meets the requirements of complex analysis and calculation, solving the deficiencies of traditional data privacy protection systems and making the present invention more intelligent and practical.

[0048] For the embodiments, please refer to Figures 1 to 3 , this embodiment provides a privacy protection system based on big data security and privacy computing, including:

[0049] A data acquisition module, configured to acquire the original data to be processed and store the original data to be processed in a database.

[0050] A data processing module, configured to input the original data to be processed. The data processing module cleans and sorts the input data and then stores the processed data in the database.

[0051] A protection analysis module, extracts the original data matrix and the original data vector from the database, and the element SSVM in the original data matrix corresponding to the original data matrix p,i and the j-th original data point DMSL corresponding to the original data vector jInput into the protection analysis module, and the protection analysis module outputs the elements in the de-identified data matrix, the j-th masked data point, and the k-th fused data point.

[0052] The management analysis module conducts statistical analysis on the elements in the de-identified data matrix, the j-th masked data point, and the k-th fused data point to evaluate the impact of data privacy protection on task performance. Then, based on the evaluation results, it adjusts the de-identified, masked, and fused data. Finally, it continuously monitors and maintains the entire privacy protection process to ensure the security and privacy of the data.

[0053] The protection analysis module includes a data de-identification analysis sub-module, a data masking analysis sub-module, and a data fusion analysis sub-module.

[0054] In this embodiment, the data de-identification analysis sub-module introduces a random perturbation factor SSVR i,q , making the de-identified data different from the original data, and the difference is random. While protecting data privacy, it better retains the availability and accuracy of the data, improving data security. The data masking analysis sub-module achieves differential privacy protection by adding Laplace noise, and the addition of noise is random and controllable, providing flexible control over the degree of privacy protection, being able to resist various attack methods, and being applicable to different types of data and scenarios. The data fusion analysis sub-module uses homomorphic encryption technology to fuse data in the encrypted state, protecting data privacy, supporting complex data analysis and calculations, improving computational efficiency and security, and reducing the risk of data leakage. Through random perturbation, differential privacy protection, and homomorphic encryption technology, it comprehensively protects data privacy. While protecting data privacy, it tries to retain the availability and accuracy of the data, providing strong support for subsequent analysis and calculations. Using homomorphic encryption technology, it supports complex analysis and calculations on encrypted data, meeting the needs of big data processing. By reducing the number of decryptions and avoiding the leakage of intermediate results, it reduces the risk of data leakage and performance loss.

[0055] Please refer to Figures 1 to 3 , the processing process of the data de-identification analysis sub-module is as follows:

[0056]

[0057] Where:

[0058] SSVO p,q refers to the element in the de-identified data matrix, located in the p-th row and q-th column.

[0059] SSVM p,i refers to the element in the original data matrix, located in the p-th row and i-th column.

[0060] SSVRi,q Refers to the element in the random perturbation factor matrix, located in the i-th row and q-th column, SSVR i,q is a random number uniformly distributed within the range [0, 1).

[0061] SSVN refers to the scaling perturbation factor, ensuring that the desensitized data is within a reasonable numerical range.

[0062] … refers to floor function, d refers to the number of decimal places of the data, SSVM refers to the original data matrix, SSVR refers to the random perturbation matrix, SSVO refers to the desensitized data matrix, SSVR i,q +0.5 refers to rounding the result after random perturbation to the nearest integer, Mod 10 d refers to the modulo operation, 10 d represents a decimal number with a modulus of 10.

[0063] SSVM p,i ×(SSVR i,q +0.5) refers to adding 0.5 to SSVR i,q The purpose of adding 0.5 is to ensure that the values after random perturbation are more evenly distributed within the range of 0 to 1.

[0064] Input the element SSVM in the original data matrix corresponding to the original data matrix into the data desensitization analysis sub-module, then use a random number generator to generate a random perturbation factor matrix with the same size as the original data matrix, and input the element SSVM in the random perturbation factor matrix corresponding to the random perturbation factor matrix into the data desensitization analysis sub-module, and output the desensitized data matrix SSVO and the corresponding element SSVO in the desensitized data matrix p,i Input the element SSVM in the original data matrix corresponding to the original data matrix into the data desensitization analysis sub-module, then use a random number generator to generate a random perturbation factor matrix with the same size as the original data matrix, and input the element SSVM in the random perturbation factor matrix corresponding to the random perturbation factor matrix into the data desensitization analysis sub-module, and output the desensitized data matrix SSVO and the corresponding element SSVO in the desensitized data matrix p,i Input into the data desensitization analysis sub-module, and output the desensitized data matrix SSVO and the corresponding element SSVO in the desensitized data matrix p,q 。

[0065] In this embodiment, it should be noted first that SSVO p,q The index of represents the specific position in the desensitized data matrix, where p represents the row index and q represents the column index. These two indexes together determine the unique position of SSVO p,q in the data matrix, SSVM p,i The index of represents the element position in the original data matrix, where p also represents the row index and i represents the column index. In some cases, in order to maintain data integrity or specific relationships, we may need to select multiple columns, that is, different i values, from the same row of the original data, that is, the same p value, for desensitization processing. This is why the column index of SSVM p,i uses i instead of q, SSVR i,qThe index represents the element position in the random perturbation factor matrix, where i represents the row index and q represents the column index. These two indexes jointly determine the SSVR i,q , the unique position in the random perturbation factor matrix. In the desensitization algorithm, SSVR i,q is introduced to increase the randomness of the data, thereby protecting privacy. Since the random perturbation factor needs to perform a one-to-one correspondence or specific combination operation with the elements in the original data matrix, SSVR i,q 's row and column indexes need to be matched or aligned with those of SSVM p,i in some form. This is why in some cases, the column index of SSVRi,q is the same as that of SSVO p,q , that is, q, while the row index is related to the column index of SSVM p,i . In complex mathematical expressions or algorithms, using unified indexes may lead to ambiguity. For example, if we use the same index to represent different rows or columns, confusion may occur when processing multi-dimensional data. By assigning unique indexes to each row and column, we can avoid this ambiguity and ensure the correctness of the algorithm. When the column index q of SSVO p,q is the same as the column index q of SSVR i,q , this usually means that we are desensitizing a certain column in the original data matrix and using the corresponding random perturbation factor for operation. When the column index i of SSVM p,i is related to the row index i of SSVR i,q , this usually means that we are selecting elements from a certain row of the original data matrix and performing a specific combination operation with a certain row in the random perturbation factor matrix. This combination operation may involve multiple columns, that is, multiple i values, depending on the requirements of the desensitization algorithm and the characteristics of the data.

[0066] The operation of this sub-module yields a desensitized data matrix, where each element is the result of combining the original data with the random perturbation factor, used to protect the privacy of the data. By implementing data desensitization through the random perturbation factor, the privacy of the data can be protected while maintaining the availability and numerical range of the data. This method is applicable to scenarios such as statistical analysis or machine learning that require data processing. This sub-module introduces randomness, making the desensitized data different from the original data, thereby protecting the privacy of the data. The desensitized data still retains most of the information of the original data and can be used for subsequent analysis and calculation. Due to the introduction of randomness, even if the desensitized data is leaked, it is difficult for attackers to infer the specific values of the original data from it. The desensitized data matrix SSVO and the elements in the desensitized data matrix SSVO p,qThe calculation, while protecting data privacy through random perturbation, maintains data availability and introduces a random perturbation factor, making the desensitized data different from the original data. Compared with traditional desensitization methods, this method better preserves data availability and accuracy while protecting data privacy. It protects data privacy while retaining data availability and accuracy, providing desensitized data support for subsequent analysis and calculation, improving data security, and reducing the risk of data leakage. As an element of the desensitized data matrix, it is used for subsequent data processing and analysis. While protecting data privacy, it maintains data availability and accuracy. While protecting data privacy through random perturbation, it tries to retain data availability and accuracy. The desensitized data matrix SSVO still retains most of the information of the original data and can be used for subsequent analysis and calculation.

[0067] Please refer to Figures 1 to 3 , and the processing process of the data masking analysis sub-module is as follows:

[0068]

[0069] Among them:

[0070] DMSA j refers to the j-th masked data point.

[0071] DMSL j refers to the j-th original data point, that is, the data that needs to be protected for privacy.

[0072] DMSL refers to the original data vector, DMSA refers to the masked data vector, sgn(...) refers to the sign function, DMSU j refers to a random number vector uniformly distributed in the range [0, 1).

[0073] DQ refers to the sensitivity of the data, that is, the maximum difference between any two adjacent data points in the dataset established from the data in the database.

[0074] DWE refers to the differential privacy parameter for the degree of privacy protection, and DWR refers to the differential privacy parameter for the failure probability.

[0075] Ln(2 / DWR) refers to the scaling factor related to DWR, which is used to adjust the size of the noise to meet the requirements of differential privacy.

[0076] refers to first determining the sign of the Laplace noise through the sign function, and then based on the differential privacy parameter DWE for the degree of privacy protection, the data sensitivity DQ and SSVR i,q , and finally performing scale adjustment through the natural logarithm ln(2 / DWR). The result obtained in this way is the Laplace noise added to the original data point.

[0077] The processing procedure of the data masking analysis sub-module is as follows:

[0078] Input the element SSVR in the random perturbation factor matrix i,q into the data masking analysis sub-module, and then obtain the data sensitivity DQ based on the maximum difference between any two adjacent data points in the dataset, and generate a uniformly distributed random number vector DMSU with the same length as the original data vector DMSL using a random number generator j , and input the data sensitivity DQ and the random number vector DMSU j into the data masking analysis sub-module, and output the masked data vector DMSA and the corresponding j-th masked data point DMSA j .

[0079] In this embodiment, the operation of this sub-module can obtain a masked data vector, where each element is the result of combining the original data with a random number drawn from the Laplace distribution, used to implement differential privacy protection. By implementing data masking through differential privacy technology, the risk of privacy leakage can be strictly controlled, and data analysis can be allowed under the premise of protecting privacy. This method is applicable to data processing scenarios that need to strictly comply with privacy regulations or policies. This sub-module adds Laplace noise to the j-th original data point DMSL j and generates the j-th masked data point DMSA through the sign function and uniformly distributed random numbers j , so that the difference between any two adjacent data points is blurred by adding noise, thereby protecting the privacy of the data. By adjusting DWE and DWR, the degree of privacy protection and the failure probability can be flexibly controlled. Furthermore, privacy protection has a strict mathematical basis and can resist various attack methods. Differential privacy protection is achieved by adding Laplace noise, and the addition of noise is random, and the magnitude of the noise is related to the data sensitivity and differential privacy parameters. And differential privacy protection has a strict mathematical basis, can resist various attack methods, and provides flexible control of the degree of privacy protection. The calculation of the masked data vector DMSA and the corresponding j-th masked data point DMSA j can achieve differential privacy protection, protect the privacy of the data, provide privacy protection support for data publishing and sharing, provide flexible control of the degree of privacy protection, and meet the needs of different types of data and scenarios. As the masked data point, it is used for data publishing and sharing, while protecting the data privacy, providing flexible control of the degree of privacy protection. This sub-module achieves the balance between data privacy and availability through differential privacy protection, so that the masked data can still be used for complex analysis and calculation. The masked data DMSA jWhile protecting data privacy, the availability and accuracy of the data are maintained, which can be used for subsequent analysis and calculations.

[0080] Please refer to Figures 1 to 3 , the processing procedure of the data fusion analysis sub-module is as follows:

[0081]

[0082] YYD pq = |SSVO p,q - SSERT|.

[0083] Where:

[0084] FNKQ k refers to the k-th fused data point, FNKQ refers to the fused encrypted data, I refers to the total number of data points participating in the fusion, i refers to the index used to traverse all data points participating in the fusion, ENF pk refers to the homomorphic encryption function and uses the public key pk. ENF pk (DMSA j ) refers to the result of homomorphic encryption of the j-th masked data point DMSA j using the public key pk.

[0085] FNT refers to the fusion weight, which is used to adjust the importance of different data points in the fusion process.

[0086] YYD pq refers to the correlation parameter of the masked and desensitized data point, n refers to a large prime number used for modular arithmetic, and SSERT refers to the mean value of all elements in the desensitized data matrix SSVO.

[0087] The processing procedure of the data fusion analysis sub-module is as follows:

[0088] Input the element SSVO in the desensitized data matrix p,q and the j-th masked data point DMSA j into the data fusion analysis sub-module, encrypt the masked data using the homomorphic encryption scheme to obtain the encrypted data ENF pk (DMSA j ). In the encrypted state, based on the output of this data fusion analysis sub-module, the fused encrypted data FNKQ is obtained, and the corresponding decryption algorithm is used to decrypt the fused encrypted data FNKQ to obtain the k-th fused data point FNKQ k .

[0089] In this embodiment, this sub-module can achieve data fusion and analysis while protecting data privacy, bringing new possibilities to the field of big data security and privacy computing. The homomorphic encryption of this sub-module allows operations to be performed on encrypted data, thus supporting complex data analysis and calculations. Since the data is fused in the encrypted state, the privacy of the data can be protected. By reducing the number of decryptions and avoiding the leakage of intermediate results, the computing efficiency can be improved and the risk of data leakage can be reduced. This sub-module supports data fusion on encrypted data, protects data privacy, and uses homomorphic encryption technology to fuse data in the encrypted state. Compared with traditional data fusion methods, this method supports complex data analysis and calculations while protecting data privacy, improving computing efficiency and security. Existing privacy protection methods cannot directly participate in calculations on encrypted data and need to be decrypted first before calculation, which poses risks of data leakage and performance loss. However, this sub-module uses homomorphic encryption technology to fuse and calculate data in the encrypted state, avoiding the risk of the decryption process and data leakage. The fused encrypted data FNKQ is generated in the encrypted state and can be directly used for subsequent analysis and calculations without decryption, which reduces the risk of data leakage and performance loss and improves computing efficiency and security.

[0090] It should be noted that based on the data point DMSA after the j-th mask j and the j-th original data point DMSL j perform cyclic iteration on the scaling perturbation factor SSVN in the data desensitization analysis sub-module, so that the element SSVO in the desensitized data matrix p,q and the data point DMSA after the j-th mask j are continuously optimized and improved. The specific processing process is as follows:

[0091] First:

[0092] Secondly: Set the iteration termination conditions:

[0093] Termination condition 1: The number of iterations is 50 times.

[0094] Termination condition 2: |SSVN new -SSVN old | < 0.001.

[0095] Where:

[0096] SSVN new refers to the scaling perturbation factor after iteration, SSVN old refers to the scaling perturbation factor before iteration, α refers to the learning rate used to control the iteration step size, m refers to the total number of data points, and MAADF refers to the maximum difference amount used to normalize the difference value.

[0097] In this embodiment, through an iterative approach, the scaling factor SSVN in the data privacy protection process is continuously optimized, thereby improving the effect of data privacy protection. During the iterative process, the calculation result DMSA of the data masking analysis sub-module is utilized. j As the input, it ensures the accuracy and effectiveness of the iterative process. The present invention combines the technical advantages of big data security and privacy computing to achieve the high efficiency and accuracy of data privacy protection. By iteratively optimizing the scaling factor SSVN, the risk of data privacy leakage can be more effectively controlled. Utilizing DMSA j As the iterative input, it can achieve fine-grained control of data privacy protection. This iterative method improves the effect of data privacy protection, reduces the risk of data leakage, and enhances the flexibility and adaptability of data privacy protection, enabling it to meet different data privacy protection requirements. Compared with traditional data privacy protection methods, this method realizes the dynamic adjustment and optimization of data privacy protection by iteratively optimizing the scaling factor SSVN. DMSA j is the calculation result of the data masking analysis sub-module, which reflects a certain state or feature in the data privacy protection process. The scaling factor SSVN is an important parameter in the data desensitization analysis sub-module, which determines the degree and effect of data privacy protection. By using DMSA j as the input to iteratively optimize the scaling factor SSVN, the dynamic adjustment and optimization of data privacy protection can be achieved. DMSA j is the basis and basis for the optimization of the scaling factor SSVN. By continuously iterating DMSA j , the value of the scaling factor SSVN can be gradually optimized, thereby achieving the best control of data privacy protection. By iteratively optimizing DMSA j and the scaling factor SSVN, the continuous improvement and optimization of data privacy protection can be achieved. This iterative form promotes the integration and development of big data security and privacy computing technologies, providing new ideas and methods for data privacy protection.

[0098] In the specific implementation process, multiple sub-modules in this method are used to form a privacy protection system architecture based on big data security and privacy computing.

[0099] By inputting the element SSVM in the original data matrix corresponding to the original data matrix into the data desensitization analysis sub-module, the data desensitization analysis sub-module outputs the desensitized data matrix SSVO and the corresponding element SSVO in the desensitized data matrix. p,i input into the data desensitization analysis sub-module, the data desensitization analysis sub-module outputs the desensitized data matrix SSVO and the corresponding element SSVO in the desensitized data matrix. p,q, this sub-module calculates a desensitized data matrix, where each element is the result of combining the original data with a random perturbation factor, which can protect the privacy of the data. The desensitization process of the data is achieved through the random perturbation factor, which can protect the privacy of the data while maintaining the availability and numerical range of the data. This sub-module makes the desensitized data different from the original data through random perturbation to protect the privacy of the data. The desensitized data still retains most of the information of the original data and can be used for subsequent analysis and calculations. This method better retains the availability and accuracy of the data while protecting the privacy of the data, and at the same time retains the availability and accuracy of the data to provide desensitized data support for subsequent analysis and calculations, improves the security of the data, and reduces the risk of data leakage. While protecting the privacy of the data through random perturbation, the availability and accuracy of the data are retained as much as possible. The desensitized data matrix SSVO still retains most of the information of the original data and can be used for subsequent analysis and calculations.

[0100] By inputting the element SSVR in the random perturbation factor matrix i,q into the data masking analysis sub-module, the data masking analysis sub-module outputs the masked data vector DMSA and the corresponding j-th masked data point DMSA j , this sub-module j adds Laplace noise to the j-th original data point DMSL j , and generates the j-th masked data point DMSA through the sign function and uniformly distributed random numbers. By adding noise, the difference between any two adjacent data points is blurred, thereby protecting the privacy of the data. Differential privacy protection is achieved by adding Laplace noise, and the addition of noise is random. The magnitude of the noise is related to the sensitivity of the data and the differential privacy parameter, and differential privacy protection has a strict mathematical basis and can resist various attack methods. The calculation of the masked data vector DMSA and the j-th masked data point DMSA j realizes differential privacy protection, protects the privacy of the data, and provides privacy protection support for data publishing and sharing. This sub-module achieves the balance between data privacy and availability through differential privacy protection, making the masked data still available for complex analysis and calculations. The masked data DMSA j while protecting the privacy of the data, maintains the availability and accuracy of the data and can be used for subsequent analysis and calculations.

[0101] By inputting the element SSVO in the desensitized data matrix p,q and the j-th masked data point DMSA j into the data fusion analysis sub-module, the data fusion analysis sub-module outputs the fused encrypted data FNKQ, and uses the corresponding decryption algorithm to decrypt the fused encrypted data FNKQ to obtain the k-th fused data point FNKQk , this sub-module can achieve data fusion and analysis while protecting data privacy, bringing new possibilities to the fields of big data security and privacy computing. The homomorphic encryption of this sub-module performs operations on encrypted data to support complex data analysis and calculations. Since the data is fused in the encrypted state, it can protect the privacy of the data. By reducing the number of decryptions and avoiding the leakage of intermediate results, it can improve the computing efficiency and reduce the risk of data leakage. This method supports complex data analysis and calculations while protecting data privacy, improving the computing efficiency and security. This sub-module uses homomorphic encryption technology to fuse and calculate data in the encrypted state, avoiding the risk of decryption process and data leakage. The fused encrypted data FNKQ is generated in the encrypted state and can be directly used for subsequent analysis and calculations without decryption, which reduces the risk of data leakage and performance loss and improves the computing efficiency and security.

[0102] Based on the j-th masked data point DMSA j and the j-th original data point DMSL j Perform cyclic iteration on the scaling perturbation factor SSVN in the data desensitization analysis sub-module. By iterating, continuously optimize the scaling factor SSVN in the data privacy protection process, improving the effect of data privacy protection. In the iteration process, use the calculation result DMSA of the data masking analysis sub-module j as the input to ensure the accuracy and effectiveness of the iteration process. The present invention combines the technical advantages of big data security and privacy computing to achieve the high efficiency and accuracy of data privacy protection. By iteratively optimizing the scaling factor SSVN, the risk of data privacy leakage can be more effectively controlled. Using DMSA j as the iteration input can achieve fine-grained control of data privacy protection. This iterative method improves the effect of data privacy protection, reduces the risk of data leakage, and enhances the flexibility and adaptability of data privacy protection, and can meet different data privacy protection requirements. Compared with traditional data privacy protection methods, this method realizes the dynamic adjustment and optimization of data privacy protection by iteratively optimizing the scaling factor SSVN. DMSA j is the calculation result of the data masking analysis sub-module, which reflects a certain state or feature in the data privacy protection process. By iteratively optimizing DMSA j and the scaling factor SSVN, continuous improvement and optimization of data privacy protection can be achieved. This iterative form promotes the integration and development of big data security and privacy computing technologies.

[0103] Furthermore, it enables the overall multiple sub-modules to cooperate with each other pairwise in calculations, and can also perform overall loops and iterations, making the overall system have the effect of automatic optimization and update, and thus having better self-adaptability.

[0104] Furthermore, please refer toFigure 1 , Figure 2 and Figure 3 , the data acquisition module is used to receive and acquire the original data to be processed, which is the basis for establishing the database and data privacy protection. Data cleaning in the data processing module is to remove duplicate data, invalid data, and abnormal data, and perform data verification. Sorting in the data processing module is to perform formatting and unification on the original data to be processed, so that it meets the requirements of subsequent processing and analysis.

[0105] In this embodiment, the processing of data by the data processing module is prior art and can be implemented through data processing software. By the data processing module, the input data is sorted and cleaned, facilitating subsequent data analysis and protection, improving the accuracy of the data. The data is cleaned and sorted once in advance to eliminate invalid and abnormal data in the data, ensuring the validity of the data.

[0106] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A privacy protection system based on big data security and privacy computing, characterized by: include: Data acquisition module: used to acquire the original data to be processed and input the original data to be processed into a database for storage; Data processing module: used to input the raw data to be processed, clean and organize the input data, and then store the processed data in the database; Protection analysis module: used to extract the original data matrix and original data vector from the database, and convert the original data matrix to the element SSVM in the original data matrix p,i The jth original data point DMSL corresponding to the original data vector j Input into the protection analysis module, the protection analysis module outputs the elements in the desensitized data matrix, the jth masked data point and the kth fused data point; Management and analysis module: used to perform statistical analysis on the elements in the desensitized data matrix, the jth masked data point, and the kth fused data point to evaluate the impact of data privacy protection on task performance. Then, based on the evaluation results, the desensitized, masked, and fused data are adjusted. Finally, the entire privacy protection process is continuously monitored and maintained to ensure data security and privacy.

2. The privacy protection system based on big data security and privacy computing according to claim 1 is characterized by: The protection analysis module includes a data desensitization analysis submodule, a data mask analysis submodule and a data fusion analysis submodule.

3. The privacy protection system based on big data security and privacy computing according to claim 2 is characterized by: The calculation formula of the data desensitization analysis submodule is as follows: in: SSVO p,q Refers to the element in the desensitized data matrix, located in the pth row and qth column; SSVM p,i Refers to the element in the original data matrix, located at the pth row and the ith column; SSVR i,q Refers to the element in the random perturbation factor matrix, located in the i-th row and q-th column; SSVN refers to the scaling perturbation factor, … refers to rounding down, d refers to the number of decimal places in the data, SSVM refers to the original data matrix, SSVR refers to the random perturbation matrix, SSVO refers to the desensitized data matrix, SSVR i,q +0.5 means the result after random perturbation is rounded to the nearest integer, Mod 10 d Refers to modular operation, 10 d Represents a decimal number modulo 10; The processing process of the data desensitization analysis submodule is as follows: The elements in the original data matrix corresponding to the original data matrix are SSVM p,i Input it into the data desensitization analysis submodule, and then use the random number generator to generate a random perturbation factor matrix with the same size as the original data matrix, and convert the elements in the random perturbation factor matrix corresponding to the random perturbation factor matrix into SSVM p,i Input to the data desensitization analysis submodule, output the desensitized data matrix SSVO and the corresponding element SSVO in the desensitized data matrix p,q .

4. The privacy protection system based on big data security and privacy computing according to claim 3 is characterized by: The calculation formula of the data mask analysis submodule is as follows: in: DMSA j Refers to the jth masked data point, DMSL j refers to the jth original data point, DMSL refers to the original data vector, DMSA refers to the masked data vector, sgn(……) refers to the symbolic function, DMSU j refers to a random number vector uniformly distributed in the range [0,1), DQ refers to the sensitivity of the data, DWE refers to the differential privacy parameter of the degree of privacy protection, and DWR refers to the differential privacy parameter of the failure probability; It refers to first determining the sign of the Laplace noise through the sign function, and then based on the differential privacy parameter DWE of the privacy protection degree, the sensitivity DQ of the data and SSVR i,q , and finally rescaling is performed by the natural logarithm ln(2 / DWR), so the result is Laplace noise added to the original data points; The processing process of the data mask analysis submodule is as follows: The elements SSVR in the random perturbation factor matrix i,q Input to the data mask analysis submodule, then the sensitivity DQ of the data is obtained based on the maximum difference between any two adjacent data points in the data set, and a uniformly distributed random number vector DMSU with the same length as the original data vector DMSL is generated using a random number generator j , the data sensitivity DQ and random number vector DMSU j Input to the data mask analysis submodule, output the masked data vector DMSA and the corresponding j-th masked data point DMSA j .

5. The privacy protection system based on big data security and privacy computing according to claim 4 is characterized by: The calculation formula of the data fusion analysis submodule is as follows: YYD pq =|SSVO p,q -SSERT|; in: FWf k refers to the kth fused data point, FNKQ refers to the encrypted data after fusion, I refers to the total number of data points involved in fusion, i refers to the index, which is used to traverse all data points involved in fusion, ENF pk Refers to the homomorphic encryption function and uses the public key pk; ENF pk (DMSA j ) refers to the DMSA of the j-th masked data point using the public key pk j The result of homomorphic encryption, FNT refers to the fusion weight, YYD pq refers to the associated parameters of the masked data points, n refers to a large prime number used for modular operations, and SSERT refers to the mean of all elements in the masked data matrix SSVO; The processing process of the data fusion analysis submodule is as follows: The elements SSVO in the desensitized data matrix p,q and the j-th masked data point DMSA j Input to the data fusion analysis submodule, use the homomorphic encryption scheme to encrypt the masked data, and obtain the encrypted data ENF pk (DMSA j ), in the encrypted state, based on the output of the fused encrypted data FNKQ by this data fusion analysis submodule, the fused encrypted data FNKQ is decrypted using the corresponding decryption algorithm to obtain the kth fused data point FNKQ k .

6. The privacy protection system based on big data security and privacy computing according to claim 1 is characterized by: The data acquisition module is used to receive and acquire the original data to be processed, and is the basis for establishing a database and protecting data privacy.

7. The privacy protection system based on big data security and privacy computing according to claim 1 is characterized by: The data cleaning in the data processing module is to remove duplicate data, remove invalid data and delete abnormal data, and verify the data.

8. The privacy protection system based on big data security and privacy computing according to claim 7 is characterized by: The arrangement in the data processing module is to format and unify the raw data to be processed so that it meets the requirements of subsequent processing and analysis.

Citation Information

Cited By

  • Data set construction method based on multi-source fusion and hierarchical security

    CN121705740A