Privacy data security exchange control method based on privacy calculation
By preprocessing, de-identifying, and encrypting data using privacy-preserving computation methods, combined with smart contract decryption and data volume statistics, the problems of low efficiency and privacy leakage in the data exchange process are solved, achieving secure and efficient data exchange.
Patent Information
- Application Number
- CN202511113756.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-11
AI Technical Summary
The lack of analysis and verification in the data exchange process in existing technologies leads to low efficiency in obtaining qualified data to be exchanged, poses a risk of privacy leakage, and results in low exchange efficiency.
By using privacy-based computation, the privacy data obtained from the sending end is preprocessed, classified according to sensitivity level and de-identified, and then encrypted after de-identification. The data is then decrypted through a smart contract, and the statistical data is used to determine the difference rate and null value rate. The field masking ratio, data statistics, or cleaning interval are adjusted to optimize the data interaction process.
It improves the efficiency of obtaining complete privacy data, ensures the security and accuracy of the data exchange process, reduces the risk of privacy data leakage, and optimizes the pass rate of the data exchange process.
Smart Images

Figure CN120614213B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information security technology, and in particular to a privacy data security exchange control method based on privacy computing. Background Technology
[0002] With the acceleration of digital transformation in society, the demand for cross-departmental data sharing has surged. However, traditional data exchange models pose a risk of privacy breaches, primarily stemming from the conflict between data sharing and protection objectives. Typical privacy data often contains proprietary information belonging to individual companies or departments. Such data requires privacy processing before being exchanged with other departments or companies. However, security vulnerabilities exist during this exchange process, resulting in low efficiency. Therefore, improving the efficiency of privacy data exchange is a pressing issue.
[0003] Chinese Patent Publication No. CN115664799B discloses a data exchange method and system for information technology security. The technical solution includes acquiring data to be exchanged from a first client and dividing it into first private data and second private data; encrypting the first private data using a first encryption rule to obtain first encrypted data; encrypting the second private data using a second encryption rule to obtain second encrypted data; and sending the first and second encrypted data to a second client to achieve data exchange. While this technical solution achieves both privacy and computational efficiency during data interaction, it only addresses how to encrypt the data to be exchanged. Because the process of acquiring the data to be exchanged lacks analysis and verification, the acquired data may be substandard during data exchange, resulting in low efficiency in acquiring qualified data. Summary of the Invention
[0004] To address this issue, the present invention provides a privacy-preserving data security exchange control method based on privacy computing, which solves the problem in the prior art where the lack of analysis, verification, and adjustment of data during the interaction process leads to low efficiency in obtaining qualified data to be exchanged.
[0005] To achieve the above objectives, the present invention provides a privacy-based data security exchange control method, comprising:
[0006] The process involves preprocessing several pieces of private data that need to be exchanged at the sending end and dividing them according to their sensitivity level to obtain several pieces of sensitive data.
[0007] Based on the sensitivity level, a corresponding data desensitization strategy is determined, and several sensitive data are processed to obtain the corresponding desensitized data;
[0008] Based on the data desensitization strategy, a corresponding encryption strategy is determined, and several desensitized data are processed to obtain corresponding encrypted data.
[0009] The acquired encrypted data and the corresponding public key are transmitted to the receiving end;
[0010] The smart contract is triggered based on the permissions of the receiving end to determine whether to use the private key corresponding to the public key for decryption to obtain the target data;
[0011] The amount of privacy data and the corresponding amount of target data are statistically analyzed and processed to obtain the difference rate and the null value rate.
[0012] The data interaction process is judged to be qualified based on the difference rate, and the process is judged to be qualified based on the result and the null value rate, or the reason for non-compliance is determined based on the difference rate, and a corresponding instruction is generated based on the reason. The instruction is to determine the field masking ratio in the data desensitization process, or the data statistics in privacy data and corresponding target data, or the data cleaning interval in the preprocessing process.
[0013] Adjust the field masking ratio, the data statistics, or the data cleaning interval based on the instructions.
[0014] Furthermore, the process for determining whether the data interaction process is qualified includes:
[0015] The judgment is made based on the comparison result between the difference rate and the preset difference rate, and the data interaction process is judged as qualified based on the judgment result combined with the comparison result between the null value rate offset and the critical null value rate offset.
[0016] When the data interaction process is deemed unqualified, the cause is determined based on the difference between the difference rate and the preset difference rate.
[0017] The null value rate offset is the difference between the null value rate of the target data and the null value rate of the privacy data.
[0018] Furthermore, the process of determining whether the data interaction process is qualified based on the comparison result between the null value rate offset and the critical null value rate offset includes:
[0019] Based on the comparison result between the null value rate offset and the critical null value rate offset, it is determined whether to reduce the field masking ratio.
[0020] Furthermore, the process of determining whether to reduce the field occlusion ratio includes:
[0021] When it is determined that the null value rate offset is greater than the critical null value rate offset, the field masking ratio is reduced.
[0022] Calculate the difference between the empty value rate offset and the critical empty value rate offset, and record it as the offset difference;
[0023] Based on the comparison result between the offset difference and the preset offset difference, a corresponding instruction is generated to reduce the field occlusion ratio. The reduction in the field occlusion ratio is positively correlated with the offset difference.
[0024] Furthermore, the process of determining whether to reduce the field occlusion ratio also includes:
[0025] When it is determined that the null value rate offset is greater than the critical null value rate offset, the field masking ratio is reduced.
[0026] The sensitivity level of the sensitive data is determined, and corresponding instructions are generated based on the sensitivity level to reduce the field masking ratio. The reduction in the field masking ratio is negatively correlated with the sensitivity level.
[0027] Furthermore, the process of determining the cause based on the difference between the difference rate and the preset difference rate includes:
[0028] Calculate the difference between the difference rate and the preset difference rate and record it as the difference deviation value;
[0029] The reasons for the failure of the data interaction process are determined based on the comparison results between the aforementioned difference deviation value and the preset difference deviation value.
[0030] Based on the aforementioned reasons, corresponding instructions are generated to determine whether to increase the data statistics or adjust the data cleaning interval.
[0031] Furthermore, the preset difference deviation value includes a first preset difference deviation value. When it is determined that the difference deviation value is less than or equal to the first preset difference deviation value, the data statistics are increased.
[0032] Calculate the difference between the first preset difference deviation value and the difference deviation value, and record it as the deviation difference value;
[0033] Based on the comparison result between the deviation difference and the preset deviation difference, a corresponding instruction is generated to increase the data statistics. The increase in the data statistics is negatively correlated with the deviation difference.
[0034] Furthermore, the preset difference deviation value also includes a second preset difference deviation value. When it is determined that the difference deviation value is greater than the first preset difference deviation value and less than or equal to the second preset difference deviation value, the data cleaning interval is adjusted.
[0035] The number of target data whose difference rate is greater than the preset difference rate during the current transmission process is counted and recorded as the number of abnormal data. The number of abnormal data during the historical transmission process is also counted. Variance is calculated based on the number of abnormal data to obtain the abnormal data variance.
[0036] Based on the comparison results between the variance of the outlier data and the variance of the critical outlier data, the data cleaning interval is determined to be expanded or narrowed.
[0037] Furthermore, when it is determined that the variance of the abnormal data is less than or equal to the variance of the critical abnormal data, the data cleaning interval is expanded.
[0038] Calculate the ratio of the number of abnormal data to the total number of target data and record it as the abnormality percentage;
[0039] Based on the comparison result between the anomaly ratio and the preset anomaly ratio, a corresponding instruction is generated to expand the data cleaning interval. The expansion of the data cleaning interval is positively correlated with the anomaly ratio.
[0040] Furthermore, when it is determined that the variance of the abnormal data is greater than the critical variance of the abnormal data, the data cleaning interval is narrowed.
[0041] Calculate the difference between the variance of the outlier data and the variance of the critical outlier data, and record it as the variance difference value;
[0042] Based on the comparison result between the variance difference and the preset variance difference, a corresponding instruction is generated to narrow the data cleaning interval. The reduction of the data cleaning interval is positively correlated with the variance difference.
[0043] Compared with existing technologies, the privacy data security exchange control method based on privacy computing of the present invention has the following advantages: This method obtains the privacy data to be exchanged at the sending end and processes it according to its sensitivity level to obtain sensitive data; it then de-identifies the sensitive data to obtain de-identified data, and finally encrypts the de-identified data to obtain encrypted data. The encrypted data and the corresponding public key are transmitted to the receiving end, where the receiving end triggers a smart contract according to its permissions to determine the use of the private key corresponding to the public key for decryption to obtain the target data. The data volume corresponding to the privacy data and the target data are statistically analyzed and processed to obtain the difference rate and the null value rate. The difference rate is used to determine whether the data interaction process is qualified, and the result combined with the null value rate further determines the state of the data interaction process, thereby identifying the corresponding cause. An instruction is generated based on the cause, and the field masking ratio, data statistics, or data cleaning interval in the data interaction process is adjusted based on the instruction, thereby optimizing the data interaction process and improving the efficiency of obtaining qualified privacy data.
[0044] Furthermore, when determining the reduction of the field masking ratio based on the comparison result of the null value rate offset and the preset null value rate offset, the present invention also determines the reduction range of the field masking ratio based on the comparison result of the offset difference and the preset offset difference, so as to achieve precise adjustment of the field masking ratio when data is desensitized, thereby improving the efficiency of obtaining complete privacy data.
[0045] Furthermore, the present invention also determines the reduction range of the field masking ratio based on the sensitivity level classification of sensitive data, and adjusts the field masking ratio in the desensitization process according to the classification to achieve precise masking, thereby improving the efficiency of obtaining complete privacy data.
[0046] Furthermore, the present invention can also determine the reasons for the failure of the data interaction process based on the comparison results of the difference deviation value and the preset difference deviation value, and then determine the corresponding processing based on the reasons, including: increasing the data statistics or dynamically adjusting the data cleaning interval, thereby improving the pass rate of the data interaction process in a targeted manner, so as to improve the efficiency of obtaining complete privacy data.
[0047] Furthermore, when it is determined that an increase in data statistics is needed, the present invention can compare the deviation difference with a preset deviation difference to determine the increase range of data statistics, thereby accurately increasing the data statistics range to reduce the difference rate, thereby improving the pass rate of the data interaction process and improving the efficiency of obtaining complete privacy data.
[0048] Furthermore, when determining that the data cleaning interval needs to be dynamically adjusted, this invention can determine whether to expand or shrink the data cleaning interval based on the comparison result of the variance of outlier data and the variance of critical outlier data. When it is determined that the data cleaning interval should be expanded, the expansion range of the data cleaning interval is determined based on the comparison result of the proportion of outliers and the preset proportion of outliers, so as to achieve precise adjustment of the data cleaning interval in the data preprocessing process and thus ensure data consistency. When it is determined that the data cleaning interval should be shrunk, the reduction range of the data cleaning interval is determined based on the comparison result of the variance difference and the preset variance difference, so as to achieve precise adjustment of the data cleaning interval in the data preprocessing process and thus avoid process processing. Both of these methods of adjusting the data cleaning interval can improve the pass rate of the data interaction process and improve the efficiency of obtaining complete privacy data. Attached Figure Description
[0049] Figure 1 This is a schematic diagram of the privacy data security exchange control system based on privacy computing in this invention;
[0050] Figure 2 This is a flowchart illustrating the privacy-based data security exchange control method based on privacy computing in this invention.
[0051] Figure 3This is a logic diagram for determining whether the data interaction process is qualified based on the difference rate and the corresponding processing in this invention;
[0052] Figure 4 This is a logic diagram illustrating the determination of the reasons for data interaction process failures based on the difference deviation value and the corresponding processing in this invention. Detailed Implementation
[0053] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0054] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0055] It should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the term "connection" should be interpreted broadly. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0056] Privacy computing refers to a set of technologies that enable data analysis and computation while protecting the data itself from being leaked. Its core goal is to achieve "usable but invisible" data. Privacy computing uses encryption technology to ensure that data is not leaked during the computation process, while allowing specific computational operations to be performed in an encrypted state, and finally obtaining the correct result by decryption.
[0057] Please see Figure 1The diagram shows a module schematic of the privacy-preserving data security exchange control system based on privacy computing in this embodiment. The system includes a sending end, a receiving end, a data acquisition module, a data desensitization module, a data encryption module, a data transmission module, a data decryption module, a statistics module, an analysis module, and a control module. The data acquisition module is connected to the sending end and is used to acquire several privacy data items to be exchanged from the sending end for preprocessing, and to process the privacy data items according to their sensitivity levels to obtain corresponding sensitive data. The data desensitization module is connected to the data acquisition module and is used to determine a corresponding data desensitization strategy based on the sensitivity level, and to process the sensitive data items according to the data desensitization strategy to obtain corresponding desensitized data. The data encryption module is connected to the data desensitization module and is used to determine a corresponding encryption strategy according to the data desensitization strategy, and to process the desensitized data items according to the encryption strategy to obtain corresponding encrypted data. The data transmission module is connected to both the data encryption module and the receiving end, and is used to transmit the encrypted data items and corresponding public keys to the receiving end. The data decryption module is connected to the receiving end and is used to trigger a smart contract based on the receiving end's permissions to determine the activation of the private key corresponding to the public key, and to perform decryption processing based on the private key to obtain target data. Each target data item corresponds to a single piece of privacy data sent by the sending end. The statistics module is connected to both the data acquisition module and the data decryption module, and is used to respectively count the amount of privacy data and the corresponding amount of target data, and process them to obtain the difference rate and null value rate. The analysis module is connected to the statistics module and is used to determine whether the process of obtaining target data is qualified based on the difference rate, and to determine whether the process is qualified based on the determination result combined with the null value rate, or to determine the reason for non-compliance based on the difference rate, and to generate corresponding instructions based on the reasons. The instructions may be to determine the field masking ratio in the data desensitization process, or the data statistics in the privacy data and the corresponding target data, or the data cleaning interval in the preprocessing process. The control module is connected to the analysis module, the data acquisition module, the data desensitization module, and the statistics module, and is used to adjust the field masking ratio, or the data statistics, or the data cleaning interval based on the instructions.
[0058] Please see Figure 2 The diagram shown is a flowchart illustrating the privacy data security exchange control method based on privacy computing in this embodiment. The process includes at least the following steps:
[0059] S1: Obtain several private data that need to be exchanged from the sending end, preprocess them, and divide them according to the sensitivity level to obtain several sensitive data;
[0060] S2: Determine the corresponding data desensitization strategy based on the sensitivity level and process several sensitive data to obtain the corresponding desensitized data;
[0061] S3: Determine the corresponding encryption strategy based on the data desensitization strategy and process several desensitized data to obtain the corresponding encrypted data;
[0062] S4: Transmit the acquired encrypted data and the corresponding public key to the receiving end;
[0063] S5: Trigger the smart contract based on the receiver's permissions to determine whether to use the private key corresponding to the public key for decryption to obtain the target data;
[0064] S6: Calculate the amount of privacy data and the corresponding amount of target data, and process them to obtain the difference rate;
[0065] S7: Determine whether the process of obtaining target data is qualified based on the difference rate, and determine whether the process is qualified based on the judgment result combined with the null value rate, or determine the reason for non-compliance based on the difference rate, and generate corresponding instructions based on the reasons. The instructions are to determine the field masking ratio in the data desensitization process, or the data statistics in privacy data and corresponding target data, or the data cleaning interval in the preprocessing process.
[0066] S8: Adjust field masking ratio, data statistics, or data cleaning interval based on instructions.
[0067] Specifically, in this embodiment, the encryption strategy includes employing a homomorphic encryption algorithm to generate a public-private key pair. The public key is used for data encryption, and the private key is used for result decryption. After sensitive data is input, the public key is called to execute the encryption function and output ciphertext. In the encrypted state, the corresponding calculation can be directly executed, and the ciphertext calculation result is decrypted using the private key to verify the homomorphism of the calculation process. Different homomorphic encryption algorithms are used for sensitive data with different sensitivity levels. For example, highly sensitive data uses fully homomorphic encryption (FHE) or an improved version of the national standard SM9, while medium and low sensitive data uses Paillier additive homomorphic encryption or a modified version of SM2 elliptic curve encryption, and so on. Data anonymization strategies include strong anonymization rules and basic anonymization rules. These strategies directly process sensitive data through techniques such as masking, replacement, or transformation, reducing its identifiability. Sensitivity levels are categorized into highly sensitive fields, generally sensitive fields, and non-sensitive fields. Differentiated strategies are developed based on the data's sensitivity level. For example, highly sensitive fields (such as financial data) require strong anonymization rules (e.g., full masking), while generally sensitive fields (such as occupational categories) and non-sensitive fields (such as public reports or administrative division codes) may only require basic processing (e.g., partial masking). Smart contracts automatically verify data integrity and trigger encrypted transmission processes based on preset conditions. End-to-end encryption can be completed without third-party intervention. The receiving user can obtain the corresponding private key from the smart contract and match it with the corresponding public key for decryption, thus realizing the process of "authorization verification → key release → data decryption." The difference rate is calculated as follows: (|Number of records at the target end - Number of records at the source end| / Number of records at the source end) × 100%, where the number of records at the target end is the amount of data in a single target data set, and the number of records at the source end is the amount of data in a single privacy data set. The number of records counted when calculating the difference rate can be part or all of the total data set. The null value rate is calculated as follows: (Number of null values in a field / Total number of records) × 100%, with a target end null value rate corresponding to the receiving end and a source end null value rate corresponding to the sending end.
[0068] Please see Figure 3 The diagram illustrates the logic for determining the pass / fail status of the data interaction process based on the difference rate in this embodiment, along with the corresponding processing. The process for determining whether the data interaction process is pass / fail includes: determining the pass / fail status based on a comparison between the difference rate and a preset difference rate; and determining the pass / fail status based on the determination result combined with a comparison between the null value rate offset and the critical null value rate offset. If the data interaction process is determined to be pass / fail, the reason is determined based on the difference between the difference rate and the preset difference rate. The null value rate offset is the difference between the null value rate of the target data and the null value rate of the privacy data.
[0069] Specifically, in this embodiment, historical data collected in the past is analyzed, and subsequent preset or critical parameter values are determined by combining statistical methods and application scenarios. To more reliably determine the data interaction process and accurately adjust the corresponding parameters during the interaction, the preset difference rate Q0 can be divided into a first preset difference rate Q1 and a second preset difference rate Q2, using a hierarchical design to perform gradient analysis on the determination process. When the application scenario is determined to be secure interaction involving privacy data, Q1 = 0.9% and Q2 = 1.1%. The comparison process based on the difference rate Q with Q1 and Q2 is as follows:
[0070] If Q is less than or equal to Q1, it indicates that the settings of Q1 and Q2 are accurate and that no other situations occur by default in this scheme. This means that the difference between the privacy data sent by the sender and the corresponding target data received by the receiver is within an acceptable range and has not reached a significant level, so the data interaction process can be judged as qualified. If Q is greater than Q1 and less than or equal to Q2, it indicates that there is an observable deviation in the data interaction process, but it has not yet reached a seriously unqualified level. At this time, further judgment can be made based on the judgment result, combined with the null value rate offset and the critical null value rate offset. Based on the judgment result, the corresponding processing is determined to optimize the data interaction process. In this case, by introducing new parameters, the reliance on a single indicator can be avoided, thereby improving the accuracy of the judgment process. If Q is greater than Q2, it is assumed that no other situations will occur. From the perspective of judging solely by comparing Q with Q0, it indicates that the difference between the privacy data and the corresponding target data is relatively too large and the difference is obvious. Therefore, the data interaction process can be directly judged as unqualified. At this time, the reason for the unqualification can be determined by the difference between Q and Q0, that is, the difference between Q and Q2.
[0071] Furthermore, the process of determining whether the data interaction process is qualified based on the comparison result of the null value rate offset and the critical null value rate offset includes: determining whether to reduce the field occlusion ratio based on the comparison result of the null value rate offset and the critical null value rate offset.
[0072] Specifically, in this embodiment, the null value offset E is an absolute value used to measure the degree of change in the proportion of null values in the target data relative to the privacy data. E = |Target data null value rate − Privacy data null value rate|. It should be noted that the critical null value offset E0 is different for data with different sensitivity levels. For data with a sensitivity level of general sensitive data fields, E0 can be set to 2.3%. The comparison process between the null value offset E and E0 is as follows:
[0073] If E is less than or equal to E0, it indicates that no abnormal interference or excessive anonymization was introduced during the data acquisition process, which conforms to the core principle of privacy computing: "data is usable but not visible." In this case, E does not affect the process of acquiring the target data, and E is not a key factor affecting the data interaction process. If Q is greater than Q1 and less than or equal to Q2, it can be determined that the current data interaction process is unqualified, and the cause can be determined based on the difference between Q and Q2. If E is greater than E0, it indicates that there are problems such as data tampering, excessive anonymization, or violations of the computing protocol, which leads to E being too large. In this case, a large number of abnormal missing values cause Q to be greater than Q1 and less than or equal to Q2. In this case, the field masking ratio during the data anonymization process can be reduced, where the field masking ratio is the ratio of the number of masked fields to the total number of fields. The above judgments are all based solely on comparing E and E0 to determine the cause, assuming that no other situations will occur.
[0074] Furthermore, the process of determining whether to reduce the field occlusion ratio includes: when it is determined that the null value rate offset is greater than the critical null value rate offset, determining to reduce the field occlusion ratio; calculating the difference between the null value rate offset and the critical null value rate offset and recording it as the offset difference; generating a corresponding instruction to reduce the field occlusion ratio based on the comparison result between the offset difference and the preset offset difference, wherein the reduction in the field occlusion ratio is positively correlated with the offset difference.
[0075] Specifically, in this embodiment, when it is determined that E is greater than E0, the preset offset difference Y0 can be divided into a first preset offset difference Y1 and a second preset offset difference Y2. The offset difference Y is compared with Y1 and Y2 to accurately determine the reduction in the field masking ratio. Y1 = 0.4% and Y2 = 0.8%. The comparison process based on the offset difference Y with Y1 and Y2 is as follows: taking a data sensitivity level of general sensitive data field as an example, if Y is less than or equal to Y1, then a corresponding first adjustment of the masking ratio is determined. The data masking process is set to reduce the original field masking ratio by 5%. If Y is greater than Y1 and less than or equal to Y2, a second adjustment instruction to reduce the masking ratio is generated, reducing it by 10%. If Y is greater than Y2, a third adjustment instruction to reduce the masking ratio is generated, reducing it by 18%. Alternatively, a manual review mechanism can be triggered to avoid compliance risks caused by automatic adjustments. After the field masking ratio is reduced, all masked fields are rounded up. It should be noted that the reduction in the field masking ratio can also be set to other compliant values. For example, when Y is less than or equal to Y1, the reduction can be set to 6%. The original field masking ratio may have different initial values depending on the application scenario.
[0076] Furthermore, the process of determining whether to reduce the field masking ratio also includes: when it is determined that the null value rate offset is greater than the critical null value rate offset, determining to reduce the field masking ratio; determining the sensitivity level of the sensitive data, generating a corresponding instruction based on the sensitivity level to reduce the field masking ratio, wherein the reduction in the field masking ratio is negatively correlated with the sensitivity level.
[0077] Specifically, in this embodiment, the field masking ratio can be adjusted hierarchically according to the sensitivity level. If the data sensitivity level is non-sensitive, a fourth instruction to adjust the masking ratio is generated, and the data desensitization process is reduced by 20% based on the original field masking ratio. If the data sensitivity level is generally sensitive, a fifth instruction to adjust the masking ratio is generated, and the data desensitization process is reduced by 13% based on the original field masking ratio. If the data sensitivity level is highly sensitive, a sixth instruction to adjust the masking ratio is generated, and the data desensitization process is reduced by 4% based on the original field masking ratio, or full masking is maintained. It should be noted that the reduction in the field masking ratio can also be set to other values that meet the requirements. For example, when the data sensitivity level is determined to be non-sensitive, the original field masking ratio is reduced by 18%, and the masked fields are rounded up after the reduction in the field masking ratio. It is understood that adjusting the field masking ratio according to the hierarchy is not the same as adjusting the field masking ratio according to the comparison of Y with Y1 and Y2, and the two are not contradictory.
[0078] Please see Figure 4 The diagram shown illustrates the logic for determining the reasons for data interaction process failures based on the difference deviation value in this embodiment, along with the corresponding processing steps. The process of determining the cause based on the difference between the difference rate and the preset difference rate includes: calculating the difference between the difference rate and the preset difference rate and recording it as the difference deviation value; determining the cause of the data interaction process failure based on the comparison result between the difference deviation value and the preset difference deviation value; and generating corresponding instructions based on the cause to determine whether to increase the data statistics or adjust the data cleaning interval.
[0079] Specifically, in this embodiment, the analysis of the cause of non-compliance is based solely on comparing the deviation value D with the preset deviation value D0. It is assumed that other situations will not occur. D0 can be divided into a first preset deviation value D1 and a second preset deviation value D2. By comparing D with D1 and D2, the cause of non-compliance can be accurately determined. D1 = 0.3% and D2 = 0.6%. The comparison process based on D with D1 and D2 is as follows:
[0080] If D is less than or equal to D1, it indicates that the data sampling method does not meet statistical requirements, resulting in insufficient sample representativeness. It is necessary to increase the amount of privacy data and the corresponding target data statistics, i.e., increase the data statistics, to reprocess and analyze and determine the difference rate. It should be noted that the data statistics at this time are partial statistics of the total data volume. If D is greater than D1 and less than or equal to D2, it indicates that the data cleaning rules are not strictly enforced, with unprocessed outliers or format errors. It is necessary to dynamically adjust the data cleaning interval to improve cleaning efficiency and redetermine the difference rate. If D is greater than D2, it means that the current difference rate Q is much greater than the second preset difference rate Q2. In this case, a manual review mechanism should be triggered directly to avoid compliance risks caused by automatic adjustments.
[0081] Furthermore, the preset difference deviation value includes a first preset difference deviation value. When it is determined that the difference deviation value is less than or equal to the first preset difference deviation value, it is determined to increase the data statistics. The difference between the first preset difference deviation value and the difference deviation value is calculated and recorded as the deviation difference value. Based on the comparison result between the deviation difference value and the preset deviation difference value, a corresponding instruction is generated to increase the data statistics value. The increase in the data statistics value is negatively correlated with the deviation difference value.
[0082] Specifically, in this embodiment, when D is less than or equal to D1, the larger the deviation difference N, the greater the difference between D and D1, i.e., the smaller D is, and therefore the smaller the difference rate Q is. In this case, the increase in the data statistic needs to be smaller. The preset deviation difference N0 can be divided into a first preset deviation difference N1 and a second preset deviation difference N2. By comparing N with N1 and N2, the increase in the data statistic can be accurately determined. N1 = 0.1%, N2 = 0.15%. The comparison process based on N with N1 and N2 is as follows: When the data sensitivity level is a general sensitive field, if N is less than or equal to N1, an instruction to generate the corresponding first adjustment statistic is determined, controlling the statistical process to increase the data statistic by 80% based on the original data statistic; if N is greater than N1 and less than or equal to N2, an instruction to generate the corresponding second adjustment statistic is determined, controlling the statistical process to increase the data statistic by 70% based on the original data statistic; if N is greater than N2, an instruction to generate the corresponding third adjustment statistic is determined, controlling the statistical process to increase the data statistic by 55% based on the original data statistic. It should be noted that the increase in the data statistics can also be set to other values that meet the requirements. For example, when N is greater than N2, the data statistics can be increased by 60% based on the original data statistics.
[0083] Furthermore, the preset difference deviation value also includes a second preset difference deviation value. When the difference deviation value is determined to be greater than the first preset difference deviation value and less than or equal to the second preset difference deviation value, the data cleaning interval is adjusted. The number of target data with a difference rate greater than the preset difference rate during the current transmission process is counted and recorded as the number of abnormal data. The number of abnormal data during the historical transmission process is also counted. Variance is calculated based on the number of abnormal data to obtain the abnormal data variance. The data cleaning interval is expanded or reduced based on the comparison result between the abnormal data variance and the critical abnormal data variance.
[0084] Specifically, in this embodiment, when N is greater than N1 and less than or equal to N2, it is necessary to re-determine whether to expand or shrink the data cleaning interval. This can be done by statistically analyzing the current number of abnormal data points and the number of several historical abnormal data points, and then calculating the variance. A critical abnormal data variance M0 = 5.05 can be set. The comparison process between the abnormal data variances M and M0 is as follows: if M is less than or equal to M0, it indicates that the number of abnormal data points fluctuates relatively little and the data distribution is relatively concentrated, suggesting that abnormal data appears in multiple fields or time periods, requiring an expansion of the data cleaning interval to ensure data consistency; if M is greater than M0, it indicates that the number of abnormal data points fluctuates significantly and the data distribution is relatively dispersed, suggesting that only specific target data is abnormal, requiring a reduction in the data cleaning interval to avoid over-processing.
[0085] Furthermore, when it is determined that the variance of the abnormal data is less than or equal to the critical variance of the abnormal data, the data cleaning interval is expanded; the ratio of the number of abnormal data to the total number of target data is calculated and recorded as the abnormality ratio; based on the comparison result of the abnormality ratio and the preset abnormality ratio, a corresponding instruction is generated to expand the data cleaning interval, and the expansion of the data cleaning interval is positively correlated with the abnormality ratio.
[0086] Specifically, in this embodiment, the preset anomaly percentage P0 can be divided into a first preset anomaly percentage P1 and a second preset anomaly percentage P2. The expansion range of the data cleaning interval is accurately determined by comparing the anomaly percentage P with P1 and P2. For example, P1 = 5% and P2 = 10%. The comparison process between P and P1 and P2 is as follows: if P is less than or equal to P1, a corresponding instruction to adjust the cleaning interval is generated, expanding the preprocessing interval by 10% based on the original data cleaning interval; if P is greater than P1 and less than or equal to P2, a corresponding instruction to adjust the cleaning interval is generated, expanding the preprocessing interval by 20% based on the original data cleaning interval; if P is greater than P2, a corresponding instruction to adjust the cleaning interval is generated, expanding the preprocessing interval by 40% based on the original data cleaning interval, or triggering a full review mechanism to re-determine the cause of the anomaly. It should be noted that the expansion range of the data cleaning interval can also be set to other acceptable values. For example, when P is greater than P2, the original data cleaning interval can be expanded by 50%.
[0087] Furthermore, when it is determined that the variance of the abnormal data is greater than the critical variance of the abnormal data, the data cleaning interval is narrowed; the difference between the variance of the abnormal data and the critical variance of the abnormal data is calculated and recorded as the variance difference value; based on the comparison result of the variance difference value and the preset variance difference value, a corresponding instruction is generated to narrow the data cleaning interval, and the narrowing of the data cleaning interval is positively correlated with the variance difference value.
[0088] Specifically, in this embodiment, the preset variance difference value K0 can be divided into a first preset variance difference value K1 and a second preset variance difference value K2. The reduction range of the data cleaning interval is accurately determined by comparing the variance difference value K with K1 and K2. K1=0.55 and K2=0.85. The comparison process between K and K1 and K2 is as follows: if K is less than or equal to K1, then the corresponding instruction for the fourth adjustment cleaning interval is generated, controlling the preprocessing process to reduce the data cleaning interval by 8% based on the original data cleaning interval; if K is greater than K1 and less than or equal to K2, then the corresponding instruction for the fifth adjustment cleaning interval is generated, controlling the preprocessing process to reduce the data cleaning interval by 15% based on the original data cleaning interval; if K is greater than K2, then the corresponding instruction for the sixth adjustment cleaning interval is generated, controlling the preprocessing process to reduce the data cleaning interval by 30% based on the original data cleaning interval, or triggering a full review mechanism to re-determine the cause of the anomaly. It should be noted that the reduction range of the data cleaning interval can also be set to other values that meet the requirements. For example, when K is less than or equal to K1, the original data cleaning interval can be reduced by 10%.
[0089] Furthermore, the process of dividing the sensitive data according to the sensitivity level to obtain a number of sensitive data includes: obtaining the sensitivity score corresponding to the privacy data, determining the sensitivity level based on the comparison result of the sensitivity score and the sensitivity threshold, and processing it to obtain the corresponding sensitive data; wherein, the types of sensitive data include high-sensitivity data, medium-sensitivity data and low-sensitivity data, and the sensitivity score is jointly determined by the type of sensitive data, the strength of the correlation between data, and the leakage impact coefficient.
[0090] Specifically, in this embodiment, the sensitivity score ranges from 0 to 100, with a first sensitivity threshold of 65 and a second sensitivity threshold of 80. When the sensitivity score is less than or equal to the first sensitivity threshold, the privacy data is determined to be a non-sensitive field, and the corresponding sensitive data type is determined to be low-sensitivity data. When the sensitivity score is greater than the first sensitivity threshold and less than or equal to the second sensitivity threshold, the privacy data is determined to be a generally sensitive field, and the corresponding sensitive data type is determined to be medium-sensitivity data. When the sensitivity score is greater than the second sensitivity threshold, the privacy data is determined to be a highly sensitive field, and the corresponding sensitive data type is determined to be highly sensitive data. This achieves the classification of privacy data levels, determining different de-identification and encryption strategies based on different sensitivity levels, while simultaneously considering precise protection and cost control. Operation and maintenance are also convenient, simplifying key management. Furthermore, user permissions differ according to different de-identification and encryption strategies; for example, only authorized users may be allowed to decrypt and obtain highly sensitive data, thereby enhancing the security of privacy data.
[0091] Furthermore, the process of obtaining the corresponding desensitized data includes: determining the data desensitization strategy based on the sensitivity level of the sensitive data, including determining that the data desensitization strategy is irreversible based on the highly sensitive data, and determining that the data desensitization strategy is reversible based on the moderately sensitive data and the low-sensitivity data.
[0092] Specifically, in this embodiment, irreversible desensitization of highly sensitive data can completely block the risk of data leakage to ensure data security, while reversible desensitization of medium and low sensitive data can achieve a balance between security and availability and reduce processing costs; the hierarchical strategy makes it difficult for attackers to reverse-engineer highly sensitive data through medium and low sensitive data.
[0093] Furthermore, the process of obtaining the corresponding encrypted data includes: determining the encryption strategy based on the data desensitization strategy of the desensitized data, including determining to use fully homomorphic encryption on the desensitized data according to the irreversible desensitization, and determining to use additive homomorphic encryption on the desensitized data according to the reversible desensitization.
[0094] Specifically, in this embodiment, the importance of data can be determined based on the data anonymization strategy. Different encryption algorithms are selected according to the importance level to ensure that more important data is protected more securely. Since privacy data cannot be restored after irreversible anonymization, fully homomorphic encryption is added on top of this to ensure that even if decryption is performed, only the anonymized data can be obtained, thus protecting it.
[0095] It is understood that no specific limitation is made to any preset parameter or critical parameter in the embodiments of the present invention, and the above values are not limited thereto. Those skilled in the art can make corresponding adjustments to the preset parameters or critical parameters according to actual needs, analysis of historical data, or equipment usage.
[0096] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0097] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A privacy-preserving data security exchange control method based on privacy computing, characterized in that, include: The process involves preprocessing several pieces of private data that need to be exchanged at the sending end and dividing them according to their sensitivity level to obtain several pieces of sensitive data. Based on the sensitivity level, a corresponding data desensitization strategy is determined, and several sensitive data are processed to obtain the corresponding desensitized data; Based on the data desensitization strategy, a corresponding encryption strategy is determined, and several desensitized data are processed to obtain corresponding encrypted data. The acquired encrypted data and the corresponding public key are transmitted to the receiving end; The smart contract is triggered based on the permissions of the receiving end to determine whether to use the private key corresponding to the public key for decryption to obtain the target data; The amount of privacy data and the corresponding amount of target data are statistically analyzed and processed to obtain the difference rate and the null value rate. The data interaction process is judged to be qualified based on the difference rate, and the process is judged to be qualified based on the result and the null value rate, or the reason for non-compliance is determined based on the difference rate, and a corresponding instruction is generated based on the reason. The instruction is to determine the field masking ratio in the data desensitization process, or the data statistics in privacy data and corresponding target data, or the data cleaning interval in the preprocessing process. Adjust the field masking ratio, the data statistics, or the data cleaning interval based on the instructions.
2. The privacy-based data security exchange control method according to claim 1, characterized in that, The process of determining whether the data interaction process is qualified includes: The judgment is made based on the comparison result between the difference rate and the preset difference rate, and the data interaction process is judged as qualified based on the judgment result combined with the comparison result between the null value rate offset and the critical null value rate offset. When the data interaction process is deemed unqualified, the cause is determined based on the difference between the difference rate and the preset difference rate. The null value rate offset is the difference between the null value rate of the target data and the null value rate of the privacy data.
3. The privacy-based data security exchange control method according to claim 2, characterized in that, The process of determining whether the data interaction process is qualified based on the comparison result between the aforementioned null value rate offset and the critical null value rate offset includes: Based on the comparison result between the null value rate offset and the critical null value rate offset, it is determined whether to reduce the field masking ratio.
4. The privacy-based data security exchange control method according to claim 3, characterized in that, The process of determining whether to reduce the field occlusion ratio includes: When it is determined that the null value rate offset is greater than the critical null value rate offset, the field masking ratio is reduced. Calculate the difference between the empty value rate offset and the critical empty value rate offset, and record it as the offset difference; Based on the comparison result between the offset difference and the preset offset difference, a corresponding instruction is generated to reduce the field occlusion ratio. The reduction in the field occlusion ratio is positively correlated with the offset difference.
5. The privacy-based data security exchange control method according to claim 3, characterized in that, The process of determining whether to reduce the field occlusion ratio also includes: When it is determined that the null value rate offset is greater than the critical null value rate offset, the field masking ratio is reduced. The sensitivity level of the sensitive data is determined, and corresponding instructions are generated based on the sensitivity level to reduce the field masking ratio. The reduction in the field masking ratio is negatively correlated with the sensitivity level.
6. The privacy-based data security exchange control method according to claim 2, characterized in that, The process of determining the cause based on the difference between the difference rate and the preset difference rate includes: Calculate the difference between the difference rate and the preset difference rate and record it as the difference deviation value; The reasons for the failure of the data interaction process are determined based on the comparison results between the aforementioned difference deviation value and the preset difference deviation value. Based on the aforementioned reasons, corresponding instructions are generated to determine whether to increase the data statistics or adjust the data cleaning interval.
7. The privacy-based data security exchange control method according to claim 6, characterized in that, The preset difference deviation value includes a first preset difference deviation value. When it is determined that the difference deviation value is less than or equal to the first preset difference deviation value, the data statistics are increased. Calculate the difference between the first preset difference deviation value and the difference deviation value, and record it as the deviation difference value; Based on the comparison result between the deviation difference and the preset deviation difference, a corresponding instruction is generated to increase the data statistics. The increase in the data statistics is negatively correlated with the deviation difference.
8. The privacy-based data security exchange control method according to claim 7, characterized in that, The preset difference deviation value also includes a second preset difference deviation value. When it is determined that the difference deviation value is greater than the first preset difference deviation value and less than or equal to the second preset difference deviation value, the data cleaning interval is adjusted. The number of target data whose difference rate is greater than the preset difference rate during the current transmission process is counted and recorded as the number of abnormal data. The number of abnormal data during the historical transmission process is also counted. Variance is calculated based on the number of abnormal data to obtain the abnormal data variance. Based on the comparison results between the variance of the outlier data and the variance of the critical outlier data, the data cleaning interval is determined to be expanded or narrowed.
9. The privacy-based data security exchange control method according to claim 8, characterized in that, When it is determined that the variance of the abnormal data is less than or equal to the variance of the critical abnormal data, the data cleaning interval is expanded. Calculate the ratio of the number of abnormal data to the total number of target data and record it as the abnormality percentage; Based on the comparison result between the anomaly ratio and the preset anomaly ratio, a corresponding instruction is generated to expand the data cleaning interval. The expansion of the data cleaning interval is positively correlated with the anomaly ratio.
10. The privacy-based data security exchange control method according to claim 8, characterized in that, When it is determined that the variance of the abnormal data is greater than the critical variance of the abnormal data, the data cleaning interval is narrowed. Calculate the difference between the variance of the outlier data and the variance of the critical outlier data, and record it as the variance difference value; Based on the comparison result between the variance difference and the preset variance difference, a corresponding instruction is generated to narrow the data cleaning interval. The reduction of the data cleaning interval is positively correlated with the variance difference.
Citation Information
Patent Citations
A data exchange method and system for information technology security
CN115664799B
Financial privacy data trusted computing method and system based on multiple data sources
CN116881973A
Smart city big data acquisition system
CN118297438A