Cross-industry universal data sharing privacy protection method and system

Through homomorphic encryption and secure multi-party computing technology, combined with the transformation of homomorphic encryption algorithms based on enterprise credibility scores, the problem of insufficient privacy protection in cross-industry data sharing is solved, and secure sharing and joint modeling of cross-industry data are achieved. The proportion of pseudo-data is dynamically adjusted, and privacy protection and resource utilization are balanced to ensure data security and credibility.

CN120470627BActive Publication Date: 2025-09-19LINGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510948151.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-09-19
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

Existing technologies find it difficult to balance data sharing needs and privacy protection requirements in cross-industry data sharing, and are unable to effectively cope with the diversity and complexity of data in different industries, affecting the efficiency and effectiveness of data sharing, and unable to meet the needs of cross-industry collaborative use of data for innovation.

Method used

Adopting homomorphic encryption and secure multi-party computing technologies, the homomorphic encryption algorithm is transformed through enterprise credibility scoring, combined with the secret sharing technology of secure multi-party computing to achieve data segmentation and encryption, and dynamically adjust based on enterprise credibility scoring. The encryption strength is bound to the decryption conditions to ensure secure data sharing.

Benefits of technology

It achieves secure sharing and joint modeling of cross-industry data while protecting data privacy, dynamically adjusts the proportion of pseudo-data, balances privacy protection and resource utilization efficiency, prevents data leakage risks, and ensures data security and credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470627B_ABST
    Figure CN120470627B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data sharing and privacy protection technology, and specifically to a cross-industry universal data sharing and privacy protection method and system. The system uses homomorphic encryption and secure multi-party computing technologies to achieve secure sharing and joint modeling of data between enterprises across multiple industries through data request and identity authentication, generation of hybrid data sets, modification of homomorphic encryption algorithms, data segmentation and ciphertext distribution, and secure joint modeling. While ensuring data privacy, the system fully leverages the value of cross-industry data, facilitates integrated innovation in industries such as healthcare and finance, meets industry data sharing needs, and promotes collaborative industrial development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data sharing privacy protection, and more specifically, to a cross-industry universal data sharing privacy protection method and system. Background Art

[0002] In today's data-driven era, cross-industry universal data sharing has become a key driver of industrial innovation. This is particularly true in the convergence of finance and healthcare, where universal data sharing can create enormous value. However, in this convergence, how to protect the privacy of data owners while incentivizing them to proactively share data has become a core issue hindering collaborative industrial innovation.

[0003] Existing data sharing technologies have many shortcomings in cross-industry scenarios. For example, they struggle to balance data sharing needs with privacy protection requirements, are unable to effectively address the diversity and complexity of data across different industries, and can affect the efficiency and effectiveness of data sharing while ensuring data security. This makes it difficult to meet the needs of cross-industry collaborative data innovation. Therefore, a universal, cross-industry data sharing privacy protection method and system is urgently needed to address the inadequate privacy protection in cross-industry data sharing in existing technologies and enable secure and efficient data sharing and joint modeling. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the purpose of the present invention is to provide a cross-industry general data sharing privacy protection method and system. By comprehensively using homomorphic encryption and secure multi-party computing technology, it can achieve secure sharing and joint modeling of data between enterprises across industries, and fully realize the value of cross-industry data while ensuring data privacy.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A common data sharing privacy protection method across industries, the process includes:

[0007] Step 1: Data request and identity authentication.

[0008] Data Request: Based on business needs, Company A submits a data request to the cross-industry universal data sharing and privacy protection system (hereinafter referred to as the System) to obtain the user data stored in the blockchain by Company B;

[0009] Identity authentication: The system authenticates the identity of Company A. Once the authentication is passed, the system sends the required data list in the data request application to Company B.

[0010] Enterprise credibility score calculation: Calculate the enterprise credibility score based on the audit process;

[0011] Step 2: Generate mixed data set.

[0012] Adjust the proportion of pseudo data in the generated hybrid dataset based on the calculated privacy, relevance, and historical performance of enterprise data usage;

[0013] Step 3: Transform the homomorphic encryption algorithm.

[0014] In homomorphic encryption algorithms, the enterprise credibility score is directly embedded in the encryption process and participates in the calculation as a key parameter, directly linking encryption strength, decryption conditions, and other aspects with the credibility score. The enterprise credibility score directly affects the specific parameter values ​​of homomorphic encryption, thereby changing the encrypted data form and decryption process, achieving a deep connection between the degree of data security protection and the enterprise credibility.

[0015] Step 4: Data segmentation and ciphertext distribution.

[0016] Utilize secret sharing technology from secure multi-party computing to split data into multiple data shards. Use a modified homomorphic encryption algorithm to encrypt the shared data shards and distribute them to each enterprise node requesting the data.

[0017] Step 5: Security joint modeling.

[0018] After obtaining data shards, enterprises combine their own encrypted data and perform secure joint modeling based on the secure multi-party computing security aggregation protocol.

[0019] A cross-industry universal data sharing privacy protection system, including:

[0020] Data request and authentication module: used to receive data request applications from across industries;

[0021] Hybrid dataset generation module: Calculates the privacy level, relevance value, and enterprise data usage history of the requested data, and generates pseudo data with a corresponding proportional amount of data to form a hybrid dataset;

[0022] Data encryption module: Utilizes secret sharing technology of secure multi-party computing to split data into multiple independent data shards, uses a modified homomorphic encryption algorithm to homomorphically encrypt the data shards, and distributes the encrypted data shards to each enterprise node;

[0023] Secure joint modeling module: Based on the secure aggregation protocol of secure multi-party computing, it performs distributed data computing and model training.

[0024] Furthermore, the enterprise credibility score calculation and scoring rules are as follows:

[0025] Score the enterprise identity and qualifications, data authorization and access scope, and historical access behavior respectively, and use the formula , calculate the enterprise credibility score, where, is the number of audit items, For the The weight of the item review, For the The score of the audit item.

[0026] Furthermore, the data privacy is calculated as follows:

[0027] It is divided into three parts: core evaluation dimension decomposition, weight setting method, and privacy calculation. In the core evaluation dimension decomposition part, a three-level evaluation indicator system is constructed from the aspects of data attributes, usage scenarios, and compliance requirements. In the weight setting method, the first and second level indicators are used to build a judgment matrix by determining the relative importance of the indicators, and the third level indicators are scored by experts. Based on the obtained first and second level indicator weights and the third level indicator scores, the formula is used to calculate the privacy level. , calculate the privacy of the requested data, where, represents the weight of the first-level indicator, represents the weight of the secondary indicator, Indicates the three-level indicator score, Indicates the total number of three-level indicators, The number of secondary indicators corresponding to each primary indicator.

[0028] Furthermore, the correlation value is calculated as follows:

[0029] Construct a graph of the requested data by data type. Nodes represent business data entities, and edges represent relationships between data. Calculate the degree centrality of each type of data, and then calculate the correlation value.

[0030] Furthermore, the evaluation method for the historical use of enterprise data is as follows:

[0031] The assessment is conducted from four dimensions: compliance, security, rationality, and transparency. The initial score for each dimension is calculated using historical data to form an initial score vector. A dynamic weight matrix is ​​constructed, introducing the risk factor matrix and historical performance matrix. The dot product operation is performed on the initial score vector and the dynamic weight matrix to obtain a comprehensive score.

[0032] Furthermore, the homomorphic encryption algorithm is modified as follows:

[0033] The enterprise credibility score is used to adjust the degree of the polynomial ring, plaintext modulus, and ciphertext modulus in the key generation phase. The hash value of the enterprise credibility score is appended to the ciphertext header in the encryption phase. The consistency between the hash value of the plaintext header and the ciphertext is checked in the decryption phase.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] 1. Dynamic encryption strengthens privacy protection: The homomorphic encryption algorithm is modified by enterprise credibility scoring. Parameters such as the number of polynomial rings, plaintext modulus, and ciphertext modulus are dynamically adjusted based on the score. This deeply binds encryption strength to enterprise credibility, achieving differentiated privacy protection and effectively preventing data leakage risks.

[0036] 2. Dynamically adjust the proportion of pseudo data in mixed datasets by comprehensively considering data privacy, relevance, and the company's data usage history. This accurately protects highly private and highly relevant data while avoiding the storage and transmission overhead associated with excessive pseudo data generation, achieving a balance between data privacy protection and resource efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A flowchart of a privacy-preserving approach for common data sharing across industries;

[0038] Figure 2 A flowchart for the transformation of homomorphic encryption algorithms in cross-industry general data sharing privacy protection methods;

[0039] Figure 3 The module block diagram of a cross-industry general data sharing privacy protection system. DETAILED DESCRIPTION

[0040] Example 1, refer to Figure 1 The cross-industry universal data sharing privacy protection method of this embodiment specifically includes the following steps (taking the financial industry and the medical industry as an example to share data and build a comprehensive credit assessment model for patients):

[0041] Step 1: Data request and identity authentication.

[0042] Data Request: Based on business needs, a financial enterprise (including multiple financial enterprises) submits a data request to the cross-industry universal data sharing and privacy protection system (hereinafter referred to as the System) to obtain user data stored on the blockchain by the medical enterprise. The data request includes the applicant's identity and digital signature (to prove the authenticity of the request), the hash value of the target data (to clearly identify the data object requested for access), and a list of the required data (including data attributes, usage scenarios, and compliance and risk descriptions).

[0043] Identity authentication: After receiving the application, the system's data security gateway starts a three-level identity authentication mechanism. First, the digital certificate of the financial enterprise is compared with the evidence information on the blockchain to confirm that the identity has not been tampered with, and to verify whether the business qualifications of the financial enterprise match the application scenario; secondly, the original authorization record of the user data on the blockchain is queried to confirm whether the data is allowed to be shared, and to check whether the access scope applied by the financial enterprise is within the preset permissions of the smart contract; finally, based on the blockchain's tamper-proof historical access log, all data access records of the financial enterprise within a period of time (in this embodiment, the past three years) are retrieved through the smart contract, focusing on checking whether there are high-risk behaviors such as over-range access, illegal use of data, and administrative penalty records; after authentication, the system sends the required data list in the data request application to the medical enterprise;

[0044] Enterprise Credibility Score Calculation: Based on the above review process, the enterprise credibility score is calculated to measure the security level of the data requester and serve as a parameter for subsequent data encryption. The scoring rules are as follows:

[0045] Enterprise identity and qualification verification: If the digital certificate is genuine and the business qualifications fully match, 100 points will be awarded; if there is some discrepancy in the certificate, 20-50 points will be deducted; if the certificate is forged or the qualifications do not match at all, 0 points will be awarded;

[0046] Data authorization and access scope verification: Data access behavior fully complies with the original authorization record, and the access scope is strictly limited to the smart contract permissions. No overstepping of boundaries will result in a score of 100. If there is minor overstepping of permissions (not exceeding the preset threshold), 10 to 30 points will be deducted based on the number of overstepping permissions and the importance of the data involved. If there is serious overstepping of permissions or unauthorized access, 0 points will be awarded.

[0047] Review of historical access behavior: If there are no high-risk data access behaviors in the past three years, including no excessive access, illegal use of data, or administrative penalties, 100 points will be awarded; if there is one excessive access, 10 points will be deducted; if there is one illegal use of data, 30 points will be deducted; if there is one administrative penalty record, 50 points will be deducted; deductions will be calculated cumulatively until all points are deducted;

[0048] Calculating a business's credibility score (Range 0~100), the formula is as follows:

[0049] ;

[0050] in, is the number of audit items. In this embodiment, ; For the The weight of the item review, in this embodiment ; For the Scores for each audit item;

[0051] Step 2: Generate mixed data set.

[0052] S31. Calculate data privacy.

[0053] The system calculates the privacy level of the requested data through three steps: core evaluation dimension decomposition, weight setting method, and privacy calculation. The core evaluation dimension decomposition step considers that data privacy is affected by multiple factors, and constructs a three-level evaluation indicator system based on data attributes, usage scenarios, and compliance requirements, as shown in Table 1:

[0054] Table 1

[0055]

[0056] The weight setting method needs to be adjusted according to the differences in each industry and business. For example, the financial industry focuses more on transaction data, while the medical industry focuses more on the weight of health data. It is also necessary to ensure that the weight of combined data is higher than that of a single data. At the same time, sensitive data that is clearly defined by laws and regulations is given a higher weight. The specific weight calculation method is as follows:

[0057] Taking the first-level indicators as an example, the relative importance of the indicators is determined by expert scoring (1-9 points, 1 = equally important, 9 = extremely important), and the judgment matrix is ​​constructed as shown in Table 2:

[0058] Table 2

[0059]

[0060] As shown in the table above, experts believe that in medical companies, the importance of data attributes for usage scenarios is 3, and the importance of data attributes for compliance and risk is 2. The eigenvalue method is used to solve the eigenvector corresponding to the maximum eigenvalue, and after normalization, the first-level indicator weight is obtained. The specific calculation process is: calculate the product of the elements in each row of the judgment matrix (for example, the data attribute row: ; Usage scenario line: ; Compliance and Risk Management: ), calculate the nth root of the product of each row (such as data attributes: , usage scenarios: ; Compliance and Risk: ), where n is the matrix order, normalize the above results (such as data attribute weights: ; Use scene weight ; Compliance and Risk Weights The weight index of the secondary indicators is calculated in the same way as the primary indicators; the tertiary indicators are scored by experts based on their importance in the actual usage scenario. In this embodiment, the final three-level evaluation indicator system and its weights are shown in Table 3:

[0061] Table 3

[0062]

[0063] Privacy calculation: Based on the obtained weights of the first and second level indicators and the scores of the third level indicators, the privacy of the requested data is calculated. The specific calculation formula is as follows:

[0064] ;

[0065] in, represents the weight of the first-level indicator, represents the weight of the secondary indicator, Indicates the three-level indicator score, Indicates the total number of three-level indicators, The number of secondary indicators corresponding to each primary indicator;

[0066] S32. Pseudo data generation method.

[0067] When generating pseudo data for different types of data, an adaptation method should be adopted based on the characteristics of the data: for numerical data, the mean, variance and other characteristics of the original data are fitted through statistical distributions (such as normal and uniform distributions), or the data distribution is learned using generative models (such as generative adversarial networks and variational autoencoders) and then sampled to ensure that the numerical range and correlation are consistent with the real data; categorical data can be randomly sampled or remapped according to the category frequency, retaining the proportion of each category and the state transition probability; text data uses language models to learn word vector distribution and semantic structure to generate grammatically correct and semantically similar pseudo text, or this can be achieved through synonym replacement and sentence structure reorganization; time series data needs to maintain temporal correlation and periodicity, and the original sequence can be transformed by translation, scaling and other transformations to obtain pseudo data; image data uses generative adversarial networks to learn pixel distribution and feature patterns to generate visually similar but unrealistic images, or pseudo samples are constructed through data enhancement methods such as cropping, rotation, and adding noise. All methods are centered on retaining the key statistical characteristics and distribution laws of the original data to avoid introducing bias that affects the training of the final joint modeling model;

[0068] S33. Generate a mixed data set.

[0069] Data association analysis: Build an undirected graph of the requested data by data type (such as name, phone number, home address, etc.). Nodes represent business data entities, and edges represent relationships between data. Calculate the degree centrality of each type of data as follows:

[0070] Assume that the graph has nodes, nodes The degree of (how many nodes are connected to it) , then the degree centrality calculation formula is:

[0071] ;

[0072] Calculate the The correlation value of each node is as follows:

[0073] ;

[0074] Pseudo data generation: According to the privacy level (the main factor affecting the proportion of pseudo data) and the relevance value (the secondary factor affecting the proportion of pseudo data), the proportion of pseudo data in the authorized access data set is adjusted to make it more difficult to identify and separate the real data in the mixed data set. In this embodiment, the proportion of pseudo data in the data with a privacy level or relevance value greater than 60 is set to be no less than 60%. Equal parts (in this embodiment ), the corresponding pseudo data ratio is shown in Table 4:

[0075] Table 4

[0076]

[0077] Assuming that data is evenly distributed in the correlation value interval and privacy interval, the amount of data in each interval accounts for the total data volume The ratio is When the correlation value is not considered in detail (the default correlation is the highest), the total amount of pseudo data generated is for: ,in, Indicates the proportion of pseudo data corresponding to each privacy level under the default highest relevance (81-100 in this embodiment); when considering the relevance value, the total amount of pseudo data generated for: 2.289N, of which, Indicates the proportion of pseudo data corresponding to all combinations of association values ​​and privacy in the table; further, the proportion of pseudo data generated is saved for: Compared with single-dimensional factors, the dual-factor approach of correlation value and privacy can not only more accurately protect potentially private data in the industry, but also effectively control the storage and transmission overhead caused by generating pseudo-data, thus balancing data privacy protection and resource utilization.

[0078] Furthermore, a multi-dimensional evaluation system for the historical performance of enterprise data usage is introduced to evaluate the performance of enterprises in four dimensions: compliance (C), security (S), rationality (R), and transparency (T). Each dimension is calculated using historical data to obtain an initial score, forming an initial score vector. : ;

[0079] In the compliance dimension, 5 points will be deducted for each violation on a quarterly basis; the violation rectification completion rate is reversely scored according to the proportion of overdue tasks, and 3 points will be deducted for every 10% overdue; in the security dimension, the security response time is divided into four levels according to industry standards, based on the average security response time within the quarter, no points will be deducted for the golden response time (such as less than 1 hour), and 5 points will be deducted for each level lower. The vulnerability repair rate is calculated monthly, and 3 points will be deducted for every 5% below the target value. Major security incidents will trigger triple deductions; in the rationality dimension, the data usage is evaluated through the correlation with the core business, the input-output ratio of resource consumption and output value. The degree of relevance of data usage to core business will be deducted from the basic score in the order of whether the data fully serves the entire process of core business, whether the data directly supports important links of core business, whether the data has indirect connection with core business, whether the data usage only involves edge scenarios of core business, and whether the data usage has no substantial connection with core business. The input-output ratio will be calculated quarterly, and 4 points will be deducted for every 10% drop from the industry average. In the transparency assessment dimension, in the log integrity assessment, 3 points will be deducted for each missing key operation record, and additional points will be deducted for consecutive missing records. 5 points will be deducted for each case of discontinuous log timestamps and data tampering.

[0080] Construct a dynamic weight matrix and introduce a risk coefficient matrix and historical performance matrix , as the basis for dynamic adjustment of weights: the risk coefficient matrix reflects the potential risks of each dimension, which is determined based on industry risk statistics, regulatory trigger records, etc. For example, if a certain industry has frequently experienced data leakage incidents recently, the risk coefficient of the security dimension Increase; the historical performance matrix is ​​based on the company's evaluation data in the past 1-3 years, and calculates the fluctuation coefficient (standard deviation) of each dimension score. The greater the fluctuation, the more unstable the performance of the dimension is, and the more attention should be paid to the weight; the final weight is calculated as the dynamic weight matrix, the formula is as follows:

[0081] ;

[0082] Among them, is the adjustment factor, which is used to balance the impact of risk coefficient and historical performance. In this embodiment, ;

[0083] The initial score vector With dynamic weight matrix Perform dot product operation to get the comprehensive score :

[0084] ;

[0085] The comprehensive score is divided into (Excellent enterprise, excellent data management capabilities), (Good enterprise, good data management capabilities), (medium-sized enterprise, medium-level data management), (Poor companies have more problems with data management and greater risk of privacy leakage) (Risk-based enterprises, with extremely high data management risks). Based on the historical performance of enterprise data usage, excellent enterprises will reduce the amount of false data generated by 20%, good enterprises will reduce the amount of false data generated by 10%, average enterprises will increase the amount of false data generated by 10%, and poor enterprises will increase the amount of false data generated by 20%. Risk-based enterprises will terminate the data sharing process and re-evaluate their historical performance of data usage after they submit appeal materials or their data management level is improved.

[0086] Step 3: Transform the homomorphic encryption algorithm.

[0087] In the homomorphic encryption algorithm, the enterprise credibility score is directly embedded in the encryption process and participates in the calculation as a key parameter, so that the encryption strength, decryption conditions, etc. are directly linked to the credibility score. The level of the enterprise credibility score directly affects the specific parameter values ​​of homomorphic encryption, thereby changing the encrypted data form and decryption process, and achieving a deep binding between the degree of data security protection and the enterprise credibility, such as Figure 2 The specific process is as follows:

[0088] Key generation phase: Assume that the original homomorphic encryption algorithm adopts the BFV scheme (the core principles of other schemes are the same), and the degree of its polynomial ring is , the plaintext modulus is , the ciphertext modulus is The adjusted polynomial ring degree , so that the polynomial dimension is positively correlated with the score; the plaintext modulus , ciphertext modulus , through linear mapping to make the enterprise credibility score Maintain a positive correlation with the size of the modulus, ensuring that security is positively correlated with the enterprise credibility score; after completing the above parameter adjustments, based on the modified parameters, the BFV key generation mechanism is adopted. Now, in a specific polynomial ring, a polynomial is randomly generated. This polynomial is the private key , then use the public parameters and private key To calculate the public key pk, the formula is as follows:

[0089] ;

[0090] in, is a "noise polynomial" with certain random rows.

[0091] Encryption phase: encryption function Expand to ,in Is the public key, used to encrypt plaintext , only those who hold the corresponding private key Only users with ciphertext can decrypt it. Hash value (|| represents string concatenation) to ensure the binding relationship between the ciphertext and the enterprise's credibility. The final ciphertext format is ;

[0092] Decryption phase: After receiving the ciphertext, the receiver first extracts , and recalculate the current enterprise rating hash value , then use your own private key Decryption to obtain plaintext ,verify If the verification fails, it means that the enterprise rating may have been tampered with during the encrypted transmission process or there is a problem with the decryption key, and the system terminates the decryption process.

[0093] Step 4: Data segmentation and ciphertext distribution.

[0094] Data segmentation: Using secret sharing technology from secure multi-party computing, data is segmented into multiple data shards. Each data shard does not contain complete data information, and the original ciphertext can only be restored when a certain number of data shards are combined together.

[0095] Data encryption: The system uses a modified homomorphic encryption algorithm to process the shared data shards (including pseudo-data), converting the original plaintext data into ciphertext data. Homomorphic encryption ensures that the data remains in ciphertext form during subsequent processing. Even if the data is leaked, it cannot be illegally decrypted and the real information cannot be obtained. Furthermore, the data form and decryption method of the ciphertext are deeply tied to the credibility of the enterprise. The lower the credibility of the enterprise, the more complex the data form obtained, and the higher the decryption cost.

[0096] Ciphertext distribution: Distribute encrypted data in shards to each enterprise node that applies for data;

[0097] Step 5: Security joint modeling.

[0098] After acquiring data shards, financial institutions combine them with their own encrypted patient health data and perform secure joint modeling based on the MPC secure aggregation protocol. Financial institutions perform local computations on the patient health data shards and received medical data shards, such as calculating correlation parameters between patient health indicators and financial transactions. Using the secure aggregation protocol, these computational results are then securely merged and sent to the system, gradually building a complete, integrated credit assessment model. Once the model is complete, it is uniformly distributed to participating data sharing companies, ensuring that no company obtains the assessment model in advance and avoids unfair competition. Throughout the modeling process, companies never exchange raw data or intermediate calculation results in plain text, ensuring data privacy and security.

[0099] Based on the above method, the present invention also provides a cross-industry universal data sharing privacy protection system, such as Figure 3 As shown, the system includes:

[0100] Data Request and Authentication Module: This module receives data requests from across industries, performs multi-factor identity authentication on the requester, and strictly reviews the requester's business qualifications, whether they are within the smart contract's permissions, and historical behavior in access logs.

[0101] Hybrid Dataset Generation Module: For data requests that pass identity authentication, the module calculates the privacy level, relevance value, and enterprise data usage history of the requested data, and generates pseudo data of a corresponding proportion based on the privacy level to form a hybrid dataset.

[0102] Data encryption module: Utilizes secret sharing technology of secure multi-party computing to split data into multiple independent data shards, uses a modified homomorphic encryption algorithm to homomorphically encrypt the data shards, and distributes the encrypted data shards to each enterprise node;

[0103] Secure Joint Modeling Module: Based on a secure aggregation protocol for secure multi-party computation, this module supports distributed data computation and model training across different industries without disclosing their individual data shards. Each enterprise node will perform computations locally on its own data shards, and then encrypt and aggregate the computation results through the secure aggregation protocol.

[0104] Through the detailed introduction of the above embodiments, the cross-industry general data sharing privacy protection method and system of the present invention, by comprehensively applying homomorphic encryption and secure multi-party computing technology, and by transforming the homomorphic encryption algorithm, makes the encryption strength and decryption conditions strongly correlated with the enterprise credibility score; comprehensively considering the data privacy, correlation value and enterprise data usage history, dynamically adjusts the proportion of pseudo data in the mixed data set, which not only accurately protects data with high privacy and high correlation value, but also avoids the storage and transmission overhead caused by excessive generation of pseudo data.

[0105] The above formulas are all dimensionless and numerically calculated, and the preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0106] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0107] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0108] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0109] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0110] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0111] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0112] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A cross-industry universal data sharing privacy protection method, characterized by: The method flow is as follows: Step 1: Data Request and Identity Authentication: Based on the data request application issued by the data requester, the data requester's identity is authenticated and the data requester's corporate credibility score is calculated; Step 2: Generate a hybrid dataset: Based on the data request, calculate the privacy level, relevance value, and enterprise data usage history of the requested data. Based on the calculation results, generate a hybrid dataset containing pseudo data. Different proportions of pseudo data are generated for data with different privacy levels and relevance values, and the proportion of pseudo data is further adjusted based on the enterprise data usage history. Step 3: Modify the homomorphic encryption algorithm: Use the enterprise credibility score to modify the homomorphic encryption algorithm. The specific process is as follows: The enterprise credibility score is used to adjust the degree of the polynomial ring, plaintext modulus, and ciphertext modulus during the key generation phase. During the encryption phase, a hash value of the enterprise credibility score is appended to the ciphertext header. During the decryption phase, the consistency between the hash value of the plaintext header and the ciphertext is verified. Step 4: Data Segmentation and Ciphertext Distribution: Use secure multi-party computing technology to segment the mixed data set, encrypt the segmented data using the modified homomorphic encryption algorithm, and send it to the data requester; Step 5: Secure joint modeling: Perform secure joint modeling based on the secure aggregation protocol of secure multi-party computing.

2. The cross-industry universal data sharing privacy protection method according to claim 1 is characterized in that: In step 1, the enterprise credibility score of the data requester is calculated. The specific process is as follows: Score the enterprise identity and qualifications, data authorization and access scope, and historical access behavior respectively, and use the formula , calculate the enterprise credibility score, where, is the number of audit items, For the The weight of the item review, For the The score of the audit item.

3. The cross-industry universal data sharing privacy protection method according to claim 1 is characterized in that: The method for generating pseudo data in step 2 is as follows: If the data being applied for is numerical data, the pseudo data is generated by fitting the characteristics of the original data through statistical distribution; If the data applied for is categorized data, random sampling or remapping is performed based on the category frequency to generate the pseudo data; If the data being applied for is text data, the language model is used to learn the word vector distribution and semantic structure, and generate grammatically correct and semantically similar pseudo text as the pseudo data; If the requested data is time series data, the original sequence is translated and scaled to obtain pseudo data; If the requested data is image data, a generative adversarial network is used to generate pseudo samples to serve as the pseudo data.

4. The cross-industry universal data sharing privacy protection method according to claim 1 is characterized in that: In step 2, the privacy level of the requested data is calculated as follows: Based on the data's own attributes, usage scenarios, and compliance requirements, the core evaluation dimensions are broken down to build a three-level evaluation indicator system; The weights of the first-level and second-level indicators are set; the third-level indicators are scored by experts; Using the formula , calculate the privacy of the requested data, where, represents the weight of the first-level indicator, represents the weight of the secondary indicator, Indicates the three-level indicator score, Indicates the total number of three-level indicators, The number of secondary indicators corresponding to each primary indicator.

5. The cross-industry universal data sharing privacy protection method according to claim 1 is characterized in that: In step 2, the relevance value of the requested data is calculated as follows: Construct a graph of the requested data by data type. Nodes represent business data entities, and edges represent relationships between data. Calculate the degree centrality of each type of data, and then calculate the correlation value.

6. The cross-industry universal data sharing privacy protection method according to claim 1 is characterized in that: In step 2, the historical usage performance of enterprise data is calculated as follows: The assessment is conducted from four dimensions: compliance, security, rationality, and transparency. The initial score for each dimension is calculated using historical data to form an initial score vector. A dynamic weight matrix is ​​constructed, introducing the risk factor matrix and historical performance matrix. The dot product operation is performed on the initial score vector and the dynamic weight matrix to obtain a comprehensive score.

7. A cross-industry universal data sharing privacy protection system, characterized by: It includes data request and authentication module, hybrid data set generation module, data encryption module, and security joint modeling module; Data request and authentication module: used to receive data request applications from across industries, authenticate the identity of the data requester based on the data request application, and calculate the enterprise credibility score of the data requester; A hybrid data set generation module is used to calculate the privacy level, relevance value and enterprise data usage history of the requested data according to the data request application, and generate pseudo data of a corresponding proportion of data volume to form a hybrid data set; Data encryption module: Utilizes secret sharing technology from secure multi-party computing to split mixed data sets into multiple independent data shards. A modified homomorphic encryption algorithm is used to homomorphically encrypt the data shards, and the encrypted data shards are distributed to each enterprise node. The specific process of modifying the homomorphic encryption algorithm is as follows: The enterprise credibility score is used to adjust the degree of the polynomial ring, plaintext modulus, and ciphertext modulus during the key generation phase. During the encryption phase, a hash value of the enterprise credibility score is appended to the ciphertext header. During the decryption phase, the consistency between the hash value of the plaintext header and the ciphertext is verified. Secure joint modeling module: Based on the secure aggregation protocol of secure multi-party computing, it performs distributed data computing and model training.

Citation Information

Patent Citations

  • Federated learning anti-reasoning attack privacy protection method based on double perturbation

    CN115481431A

  • Privacy protection method for multi-mechanism joint data value sharing service

    CN118869260A