Data quality verification method and system in multi-party security computing
By combining polynomial zero-knowledge proof and distributed private key share distribution technology, the contradiction between privacy protection and inefficiency of data quality verification in multi-party secure computing is solved, efficient cross-platform data quality verification is achieved, and the adaptability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510471439.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In multi-party security computing, it is difficult for the existing technology to achieve efficient and reliable data quality verification while protecting data privacy, and the cross-platform compatibility is poor, making it impossible to effectively identify carefully constructed data that does not conform to the true distribution.
Using polynomial zero-knowledge proof and distributed private key share distribution combined with polynomial commitment technology, it supports data quality verification across protocols and multiple frameworks by sorting the private key shares of data holders and generating polynomial ciphertext commitments.
It realizes efficient and reliable data quality verification while protecting data privacy, improves verification accuracy and system universality, and adapts to the deployment needs of different computing environments.
Smart Images

Figure CN120372653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data security technology, and in particular to a data quality verification method and system in multi-party secure computing, and belongs to the technical fields of cryptography, distributed computing, data quality management, and privacy protection. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, multi-party data collaborative computing and analysis has become an important technology application model. Multi-party secure computing allows multiple data holders to jointly complete specific computing tasks without leaking the original data, which is very important for VIP, Communication It is of great significance to industries involving sensitive data such as finance, medical care, and government affairs.
[0003] However, data quality verification faces severe challenges in the process of multi-party secure computing. First, traditional data quality verification methods usually require direct access to the original data, which contradicts the privacy protection goal of multi-party secure computing. Second, since the data is scattered among different participants, a centralized quality verification method cannot be adopted. Third, malicious participants may provide low-quality or untrue data, resulting in distorted calculation results, which is difficult for other participants to discover and verify.
[0004] Homomorphic encryption, trusted execution environment and other methods are often used in existing technologies for data protection, but these methods have limitations in data quality verification: on the one hand, although homomorphic encryption can protect data privacy, it has high computational complexity and is difficult to support complex data quality verification; on the other hand, the trusted execution environment requires specific hardware support, lacks universality, and still has security risks.
[0005] In addition, most of the existing data quality verification methods in multi-party secure computing focus on integrity verification and consistency verification, lacking in-depth verification of data distribution characteristics, and cannot effectively identify carefully constructed data that does not conform to the actual distribution. At the same time, existing methods usually rely on specific computing frameworks, have poor cross-platform compatibility, and are difficult to flexibly deploy in a variety of application environments.
[0006] Therefore, how to achieve efficient and reliable data quality verification while protecting data privacy and supporting cross-platform applications has become a technical problem that needs to be urgently solved in the field of multi-party secure computing. Summary of the invention
[0007] The main purpose of the present invention is to provide a data quality verification method and system in multi-party secure computing, aiming to solve technical problems in the prior art such as the difficulty in balancing data quality verification and privacy protection, low verification efficiency, and poor cross-platform compatibility.
[0008] The present invention proposes a data quality verification method in multi-party secure computing, including:
[0009] It includes the following steps: S1. A distributed private key share distribution method based on polynomial zero-knowledge verifiability. For a given access policy and a set of data holders, first sort the private key shares of the data holders, then generate a polynomial, calculate the ciphertext commitment and coefficient commitment of the polynomial, and outsource and store these commitments to an authoritative institution. Finally, the data holders use a secret multiplication sharing scheme to generate a private key share proof and calculate the zero-knowledge proof of the polynomial; S2. A verifiable random sampling method based on polynomial commitment, by extracting the statistical features of the dataset distribution, encoding its distribution, and binding the encoded value to the polynomial commitment; S3. Support cross-protocol and protocol implementation based on multiple frameworks; S4. A verifiable data quality proof framework method based on polynomial commitment, verifying the data distribution characteristics of the data holder's share and checking whether it conforms to the distribution of the statistical features.
[0010] Preferably, the method of S1 specifically includes: S11. Generate the private key shares of the data holders; S12. A verifiable representation of distributed private key share distribution, generating commitment values; S13. Verifiable data quality proof; S14. Privacy protection.
[0011] Preferably, S11 includes: 1A1). Select an access policy. For a given access policy, use the Shahal secret sharing scheme for private key share distribution to obtain the shares of all n data holders; 1B1). According to the obtained private key shares of the n holders and the specified access structure, generate the key of the ciphertext matrix composed of all shares.
[0012] Preferably, S12 includes: 1A2). A set of data holders; 1B2). Outsource and store the commitment to an authoritative institution.
[0013] Preferably, S13 includes: 1A3). The holder obtains the proof of the private key share; 1B3). Calculate the zero-knowledge proof of the ciphertext.
[0014] Preferably, S14 includes: 1A4). The verifier verifies the proof of the private key share. If the proof is invalid, an exception warning is issued.
[0015] Preferably, the method of S2 specifically includes: S21. Encoding the distribution characteristics of the dataset; S22. Verifiable representation of the dataset distribution characteristics; S23. A verifiable data quality verification method based on polynomial commitment.
[0016] Preferably, the S21 includes: 2A1), selecting the feature X of the data set. For the set of N data samples, taking the value of each feature X in each sample as the feature vector x, and binning all the feature vectors x to obtain the bin set V; 2B1), generating the within-bin feature mean through the bin set V after binning; 2C1), for each bin set Vi, calculating the number of data samples in each bin set, and by comparing with the data set N, obtaining the proportion of the number of data samples in each bin set Vi in the whole, that is, the bin distribution ratio pv.
[0017] Preferably, the S22 includes: 2A2), for the generated bin set V, first sorting the bin set to obtain the order l of each bin set, encrypting the within-bin feature mean according to the order l to obtain a hash value h, and using the encrypted within-bin feature mean set and h as the distributed ciphertext; 2B2), for the data set Di of the data holder, statistically obtaining the proportion distribution of the bin set V of the feature X to obtain the distribution feature set pi of each bin set Vi; 2C2), for the encrypted bin set Vi, calculating its private key share di using the secret sharing scheme, and calculating the verification share according to di and pi; 2D2), for the share of each data holder, calculating the ciphertext m aggregated by multiple shares and its aggregation coefficient commitment value c, using c and m as the ciphertext value of the shared data set, generating a polynomial based on the ciphertext of the polynomial commitment, and finally the holder sending the ciphertext commitment s of its polynomial and the coefficient commitment ci.
[0018] A data quality verification system in multi-party secure computation includes a system server side, a data holder side, and a data training side. The system server side is connected to the data holder side and the data training side through network signals or wireless signals; the system server side is used for uploading and downloading private data and public data, and providing data quality verification technology to the data holder side; the data holder side is used for generating data holder privacy protection shares for private data, data quality and data anomaly detection; the data training side is used for training public data and verifying data quality; wherein, the system server side includes a data holder data set module, a data training side data set module, a system data set module, a multi-security protocol module, and an algorithm model module. The data holder data set module, the data training side data set module, and the system data set module are signal-connected to the multi-security protocol module, and the multi-security protocol module is signal-connected to the algorithm model module.
[0019] The beneficial effects of the present invention are mainly reflected in:
[0020] 1. By innovatively combining polynomial zero-knowledge proof and distributed private key share distribution, the present invention realizes efficient data quality verification while protecting data privacy, and solves the contradiction between privacy protection and quality verification;
[0021] 2. The verifiable random sampling method based on polynomial commitment proposed by the present invention makes the verification of data distribution characteristics more accurate and efficient, significantly improving the reliability of verification.
[0022] 3. The present invention supports the implementation across protocols and multiple frameworks, significantly enhancing the generality and adaptability of the system, and facilitating the deployment and application in different computing environments.
[0023] 4. The present invention constructs a complete data quality proof framework, integrating verification, anomaly detection and handling mechanisms, and improving the robustness and availability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic flow chart of a data quality verification method in a multi-party secure computation of the present invention;
[0025] Figure 2 is a schematic flow chart of a distributed private key share distribution method based on polynomial zero-knowledge verifiability of the present invention;
[0026] Figure 3 is a schematic flow chart of a verifiable random sampling method based on polynomial commitment of the present invention;
[0027] Figure 4 is a schematic diagram of the implementation of the present invention across protocols and based on multiple frameworks;
[0028] Figure 5 is a schematic flow chart of a method for a verifiable data quality proof framework based on polynomial commitment of the present invention;
[0029] Figure 6 is a schematic structural diagram of a data quality verification system in a multi-party secure computation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] Please refer to the appended Figures 1-6 , and the technical solutions of the present invention will be described in detail below in conjunction with the specific embodiments.
[0031] Embodiment 1: A data quality verification method in multi-party secure computation
[0032] Referring to Figure 1 , a data quality verification method in multi-party secure computation proposed by the present invention includes the following steps:
[0033] S1. A distributed private key share distribution method based on polynomial zero - knowledge verifiability. For a given access policy and a set of data holders, first, sort the private key shares of the data holders, then generate a polynomial, calculate the ciphertext commitment and coefficient commitment of the polynomial, and outsource these commitments to a trusted authority for storage. Finally, the data holders use the secret multiplication sharing scheme to generate a private key share proof and calculate the zero - knowledge proof of the polynomial;
[0034] S2. A verifiable random sampling method based on polynomial commitment. By extracting the statistical features of the dataset distribution and encoding its distribution, bind the encoded values to the polynomial commitment;
[0035] S3. Support cross - protocol and protocol implementation based on multiple frameworks;
[0036] S4. A verifiable data quality proof framework method based on polynomial commitment, which verifies the data distribution characteristics of the data holder's share and checks whether it conforms to the statistical feature distribution.
[0037] The following elaborates on each step in detail.
[0038] S1: A distributed private key share distribution method based on polynomial zero - knowledge verifiability
[0039] Refer to Figure 2 , step S1 specifically includes the following sub - steps:
[0040] S11. Generate the private key shares of the data holders;
[0041] S12. A verifiable representation for distributed private key share distribution, generate commitment values;
[0042] S13. Verifiable data quality proof;
[0043] S14. Privacy protection.
[0044] Among them, the specific implementation process of S11 includes:
[0045] 1A1). Select an access policy. For the given access policy, use the Shahal secret sharing scheme for private key share distribution to obtain the shares of all n data holders;
[0046] In a preferred embodiment of the present invention, the Shahal secret sharing scheme adopts a (t, n) threshold scheme, that is, at least t shares are required to reconstruct the secret, where 1 ≤ t ≤ n. The specific process is as follows:
[0047] First, select a large prime number p and a polynomial f(x) = a0 + a1x + a2x 2 +...+a t-1 xt- 1 mod p, where a0 is the secret to be shared, and a1, a2, ..., a t-1 are randomly selected coefficients.
[0048] Then, for each data holder i (1 ≤ i ≤ n), calculate the share s i = f(i) mod p. In this way, only when at least t data holders cooperate can the polynomial f(x) be reconstructed by Lagrange interpolation and the secret a0 be obtained.
[0049] 1B1), Based on the obtained private key shares of the n holders and according to the specified access structure, generate the key of the ciphertext matrix composed of all shares.
[0050] In this embodiment, the access structure defines which subsets of data holders have the right to reconstruct the secret. Based on the obtained n private key shares s1, s2, ..., s n , construct the ciphertext matrix M, where M i,j = E(s i , k j ), E represents the encryption function, and k j is the key corresponding to the jth authorized subset in the access structure. This ensures that only members of the authorized subset can access the corresponding private key share.
[0051] The specific implementation process of S12 includes:
[0052] 1A2), The set of data holders;
[0053] In this embodiment, the set of data holders D = {D1, D2, ..., D n} represents the n entities participating in the secure calculation, and each entity holds part of the data and the corresponding private key share.
[0054] 1B2), Outsource the commitment to a trusted authority for storage.
[0055] To ensure the fairness and reliability of verification, the present invention outsources the generated polynomial ciphertext commitment C f = {c0, c1, ..., c t-1} and the coefficient commitment to a trusted authority for storage. g and h are public generators, and r i is a random number. The trusted authority only stores these commitments and does not access the original data, ensuring data privacy.
[0056] The specific implementation process of S13 includes:
[0057] 1A3), The holder obtains the proof of the private key share;
[0058] For each data holder D i , generate its private key share s i Proof of π i In the present invention, the Schnorr protocol is used to generate zero-knowledge proof:
[0059] 1. Data holder D i Choose a random number r i ,calculate
[0060] 2. Computational Challenges i =H(g,h,s i ,t i ), where H is a secure hash function;
[0061] 3. Calculate the response z i =r i +c i ·s i mod p;
[0062] 4. Private key share proof π i =(t i ,z i );
[0063] 1B3) Calculate the zero-knowledge proof of the ciphertext.
[0064] Based on private key share s i And polynomial f(x), calculate the zero-knowledge proof of the ciphertext ∏ i , prove that the data holder D i The correct private key share is indeed held, and the share is the value of the polynomial f(x) at point i. In order to improve efficiency, this embodiment adopts the Bulletproofs scheme, which has a smaller proof size and higher verification efficiency.
[0065] A specific zero-knowledge proof calculation process is as follows:
[0066] 1. Data holder D i Calculating commitment
[0067] 2. Construct and prove the polynomial f(x)-s i =a1(xi)+a2(xi) 2 +...+a t-1 (xi) t-1 ;
[0068] 3. Generate the coefficient commitment of the polynomial and construct the zero-knowledge proof ∏ i ;
[0069] The specific implementation process of S14 includes:
[0070] 1A4), The verifier verifies the proof of the private key share. If the proof is invalid, an exception warning is issued.
[0071] The verifier receives the proof of the private key share π i from the data holder D i =(t i , z i ), and then executes the following verification steps:
[0072] 1. Calculate the challenge c i =H(g, h, s i , t i );
[0073] 2. Verify
[0074] If the verification is successful, the proof is accepted; otherwise, the proof is considered invalid and an exception warning is issued. In practice, to improve security, a verification threshold can be set. For example, an exception warning is issued only when the verification fails three times in a row, which can reduce the false alarm rate caused by temporary factors such as network fluctuations.
[0075] S2: Verifiable Random Sampling Method Based on Polynomial Commitment
[0076] Referring to Figure 3 , step S2 specifically includes the following sub-steps:
[0077] S21. Encoding the distribution characteristics of the data set;
[0078] S22. Representing the verifiable data set distribution characteristics;
[0079] S23. Verifiable Data Quality Verification Method Based on Polynomial Commitment.
[0080] Among them, the specific implementation process of S21 includes:
[0081] 2A1), Select the feature X of the data set. For the set of N data samples, take the value of each feature X in each sample as the feature vector x, and bucket all the feature vectors x to obtain the bucket set V;
[0082] In the preferred embodiment of the present invention, the feature X can be a numerical feature (such as age, income) or a categorical feature (such as gender, occupation). The bucketing strategy varies according to the feature type:
[0083] For numerical features, equal-width bucketing (dividing the numerical range into k equal intervals) or equal-frequency bucketing (making the number of samples in each bucket approximately equal) can be used
[0084] For categorical features, each category corresponds to a bucket.
[0085] For example, for the age feature, the bucketing intervals can be set as [0 - 18, 19 - 35, 36 - 50, 51 - 65, 66+], and data samples are assigned to the corresponding buckets according to the age values. The number of buckets k is usually determined based on the data distribution and the requirements for verification accuracy. Practice shows that a value of k between 5 and 10 usually achieves a good balance.
[0086] 2B1), Generate the mean value of the features within the bucket from the bucket set V after bucketing;
[0087] For each bucket V i (1 ≤ i ≤ k), calculate its mean value of the features:
[0088]
[0089] where |V i | represents the number of samples in bucket V i , and x represents the feature value of the samples within the bucket. For categorical features, the encoded numerical values can be used to calculate the mean, or the mode can be directly used instead of the mean.
[0090] 2C1), For each bucket set Vi, calculate the number of data samples in each bucket set. By comparing with the data set N, obtain the proportion of the number of data samples in each bucket set Vi in the whole, that is, the bucket distribution ratio pv.
[0091] The calculation formula for the bucket distribution ratio is:
[0092]
[0093] where |V i | is the number of samples in bucket V i , and N is the total number of samples. These distribution ratios form the distribution feature fingerprint of the data and are an important basis for data quality verification.
[0094] The specific implementation process of S22 includes:
[0095] 2A2), For the generated bucket set V, first sort the bucket set to obtain the order l of each bucket set. According to the order l, encrypt the mean value of the features within the bucket to obtain a hash value h. The encrypted set of mean values of the features within the bucket and h are used as distributed ciphertexts;
[0096] The sorting of the bucket set can be based on the boundary values of the buckets (for numerical features) or the alphabetical order of the categories (for categorical features). After sorting, each bucket V i obtains a unique serial number l i , and then the mean value of the features within the bucket μi Perform encryption:
[0097] E i = Enc(k, μ i || l i ),
[0098] where Enc is the encryption function, k is the key, and || represents the concatenation operation. Finally, calculate the hash value of all encrypted values:
[0099] h = H(E1 || E2 ||... || E k ),
[0100] where H is a secure hash function, such as SHA-256. The distributed ciphertext consists of {E1, E2,..., E k , h}.
[0101] 2B2) For the data set Di of the data holder, count the proportion distribution of the bucket set V of the statistical feature X to obtain the distribution feature set pi of each bucket set Vi;
[0102] For each data holder D i , calculate the distribution ratio of its data in each bucket:
[0103] p i = {p i,1 , p i,2 ,..., p i,k},
[0104] where |V i,j | is the number of samples of the data holder D i in the bucket V j , and N i is the total number of samples owned by D i .
[0105] 2C2) For the encrypted bucket set Vi, use the secret sharing scheme to calculate its private key share di, and calculate the verification share according to di and pi;
[0106] For each encrypted bucket set V i , adopt the aforementioned Shamir secret sharing scheme to generate the private key share d i . Then, calculate the verification share based on the private key share and the distribution feature:
[0107] v i = H(d i || p i ). G,
[0108] Where H is a hash function that maps to an elliptic curve point, G is the base point on the elliptic curve, and || represents the concatenation operation.
[0109] 2D2) For the shares of each data holder, calculate the ciphertext m aggregated from multiple shares and its aggregated coefficient commitment value c. Use c and m as the ciphertext values of the shared dataset, generate a polynomial based on the polynomial commitment of the ciphertext, and finally the holder sends the ciphertext commitment s of its polynomial and the coefficient commitment ci.
[0110] The aggregation calculation of multiple verification shares is as follows:
[0111]
[0112] Where α i is the aggregation coefficient, satisfying ∑ i = 1 n α i = 1. The aggregated coefficient commitment value is:
[0113] c = H(α1||α2||...||α n ),
[0114] Based on the aggregated ciphertext m, construct the polynomial P(x) = m + ∑ j = 1 t β j ·x j , where β j are random coefficients.
[0115] Finally, calculate the ciphertext commitment s = Commit(P) of the polynomial and the coefficient commitment c i = H(β1||β2||...||β t ), and send them to the verifier.
[0116] The specific implementation process of S23 is relatively complex and involves the interactive verification between data holders. Here, its core steps are given:
[0117] 1. Holder A distributes data to Holder B for training;
[0118] 2. Holder B generates the features of multiple data and sends them to Holder A;
[0119] 3. Holder A conducts training and verifies the data quality;
[0120] 4. Sample the feature distribution from the trained dataset;
[0121] 5. Calculate the distribution features of each bucket set;
[0122] 6. Conduct verifiable data quality detection;
[0123] 7. Generate commitment values and share private key proofs;
[0124] 8. Verify the correctness of the holder;
[0125] 9. Extract polynomial commitments and calculate the distribution characteristics of the encoding;
[0126] 10. Compare with the stored share distribution characteristics;
[0127] 11. Output data quality assurance or anomaly warnings;
[0128] S3: Support cross - protocol and protocol implementation based on multiple frameworks
[0129] Refer to Figure 4 , the present invention supports multiple implementation frameworks, mainly including PyTorch implementation and TensorFlow implementation. Taking the PyTorch implementation as an example, the specific steps include:
[0130] S21. Generate data holder private key shares;
[0131] S22. Sort the private key shares;
[0132] S23. Generate ciphertext commitments and coefficient commitments;
[0133] S24. Generate share proofs and zero - knowledge proofs;
[0134] S25. Divide the dataset and calculate the feature distribution;
[0135] S26. Encode the features and generate ciphertext commitments;
[0136] S27. Send information such as ciphertext commitments to data requesters.
[0137] This multi - framework support design enables the present invention to be flexibly deployed in different computing environments, meet the needs of different users, and improve the versatility and practical value of the system.
[0138] S4: Verifiable data quality proof framework method based on polynomial commitments
[0139] Refer to Figure 5 , the data quality verification process of the present invention mainly includes:
[0140] S41. Extract the ciphertext and coefficients in the polynomial commitment and calculate the corresponding distribution characteristics;
[0141] S42. Compare the calculated features with the features of the data holder to verify whether they are consistent.
[0142] The verification process adopts zero-knowledge proof technology to ensure that data privacy is not leaked. When the features are consistent and the zero-knowledge proof is valid, the data quality is considered normal; otherwise, the data quality is considered problematic and the corresponding exception handling mechanism is triggered.
[0143] In practical applications, a fault tolerance threshold can be set. For example, when the feature deviation does not exceed 5%, the data quality is still considered acceptable. This design takes into account the reasonable fluctuations that may exist in the actual data, improving the practicality and robustness of the system.
[0144] Embodiment 2: A data quality verification system in multi-party secure computing
[0145] Refer to Figure 6 This invention also proposes a data quality verification system in multi-party secure computing corresponding to the above method. It is characterized by including a system server side 1, a data holder 2, and a data training party 3. The system server side 1 is connected to the data holder 2 and the data training party 3 through network signals or wireless signals; the system server side 1 is used to upload and download private data and public data, and provide data quality verification technology to the data holder 2; the data holder 2 is used to generate data holder privacy protection shares for private data, detect data quality and data anomalies; the data training party 3 is used to train public data and verify data quality; among them, the system server side 1 includes a data holder dataset module 11, a data training party dataset module 12, a system dataset module 13, a multi-security protocol module 14, and an algorithm model module 15. The data holder dataset module 11, the data training party dataset module 12, and the system dataset module 13 are signal-connected to the multi-security protocol module 14, and the multi-security protocol module 14 is signal-connected to the algorithm model module 15.
[0146] In specific implementation, the functions of each module are as follows:
[0147] 1. Data holder dataset module 11: Responsible for verifying and uploading the private dataset provided by the data holder 2 to ensure that the data format and structure meet the system requirements.
[0148] 2. Data training party dataset module 12: Responsible for training the dataset provided by the data training party 3 and verifying the quality of the trained dataset to detect possible quality problems.
[0149] 3. System dataset module 13: Responsible for managing the public dataset generated by the system, including functions such as data storage, access control, and update maintenance.
[0150] 4. Multi - security protocol module 14: It realizes the core functions of protecting the privacy of the private data provided by the data holder 2 and the data training party 3, including key technologies such as distributed private key share distribution, polynomial commitment generation and verification in the above - mentioned methods.
[0151] 5. Algorithm model module 15: It is responsible for training the data uploaded by the system server side 1, generating a model, and performing quality detection on the trained data set to evaluate the impact of data quality on model performance.
[0152] The system operation process can be summarized as follows: The data holder 2 and the data training party 3 upload data to the system server side 1. The multi - security protocol module 14 performs privacy protection processing on the data. The algorithm model module 15 conducts training and quality verification based on the processed data, and finally feeds back the verification results to the data holder 2 and the data training party 3. During the whole process, the original data will not be leaked, and at the same time, the system can effectively detect and handle data quality problems.
[0153] This system supports flexible deployment modes and can run on cloud servers, edge devices or hybrid environments. In scenarios with particularly high security requirements, blockchain technology can also be introduced to record the verification results and key operations on the blockchain, further enhancing the transparency and immutability of the system.
[0154] To more intuitively illustrate the implementation process of the present invention, a specific application scenario example is given below.
[0155] Suppose three institutions A, B, and C are data holders and hope to jointly train a risk assessment model without sharing the original customer data. Each institution is worried that the other institutions may provide data of poor quality, affecting the accuracy of the model. At this time, the method and system of the present invention can be applied for data quality verification.
[0156] First, each of the three institutions performs bucketing on its data features (such as customer age, income, credit record, etc.) and calculates the mean and distribution ratio of the features within the buckets. Taking the age feature as an example, 5 buckets can be set: [18 - 25], [26 - 35], [36 - 45], [46 - 60], [61+], and calculate the sample ratio and age mean within each bucket.
[0157] Then, each institution generates private key shares and distributes them to other institutions using the Shamir secret sharing scheme. For example, adopting a (2,3) threshold scheme, any two institutions can cooperate to reconstruct the secret. Each institution calculates the polynomial ciphertext commitment and coefficient commitment and stores them in a commonly recognized authoritative institution (such as an industry regulatory agency).
[0158] Next, each institution generates a zero - knowledge proof of its private key share and calculates a polynomial commitment of the data distribution characteristics. In the verification phase, institution A can verify whether the data provided by institutions B and C conforms to the declared distribution characteristics without accessing the original data.
[0159] The method of the present invention can effectively detect the following data quality problems:
[0160] 1. Distribution shift: The data distribution is significantly inconsistent with the declared distribution;
[0161] 2. Excessive outlier ratio: The distribution ratio of some buckets is abnormal;
[0162] 3. Data forgery: The feature mean does not match the actual situation;
[0163] In this way, the three institutions can ensure the quality of the data participating in the calculation while protecting the privacy of customers, improving the reliability and effectiveness of the final model.
[0164] The data quality verification method and system in multi - party secure computing proposed by the present invention, through the innovative combination of polynomial zero - knowledge proof, distributed private key share distribution, and polynomial commitment technology, realizes efficient data quality verification while protecting data privacy. This method supports multiple implementation frameworks, adapts to different computing environments, and has good generality and scalability. Compared with the prior art, the present invention has significantly improved in terms of privacy protection strength, verification reliability, and application flexibility, providing important technical support for the practical application of multi - party secure computing.
[0165] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data quality verification method in multi-party secure computing, characterized in that It includes the following steps: S1. A distributed private key share distribution method based on polynomial zero-knowledge verification. For a given access policy and a set of data holders, first sort the private key shares of the data holders, then generate a polynomial, calculate the ciphertext commitment and coefficient commitment of the polynomial, and outsource these commitments to a trusted authority for storage. Finally, the data holders use a secret multiplication sharing scheme to generate a proof of the private key share and calculate a zero-knowledge proof of the polynomial; S2. A verifiable random sampling method based on polynomial commitment, by extracting the statistical features of the dataset distribution, encoding its distribution, and binding the encoded values to polynomial commitments; S3. Support for cross-protocol and protocol implementation based on multiple frameworks; S4. A verifiable data quality proof framework method based on polynomial commitment, which verifies the data distribution characteristics of the data holder's share and checks whether it conforms to the distribution of the statistical features.
2. The method for verifying data quality in multi-party secure computing according to claim 1, wherein The method of S1 specifically includes: S11. Generate the private key share of the data holder; S12. Verifiable representation of distributed private key share distribution, generate commitment values; S13. Verifiable data quality proof; S14. Privacy protection.
3. The method for validating data quality in multi-party secure computing according to claim 2, wherein S11 includes: 1A1). Select an access policy. For the given access policy, use the Shahal secret sharing scheme for private key share distribution to obtain the shares of all n data holders; 1B1). According to the obtained private key shares of the n holders and the specified access structure, generate the key of the ciphertext matrix composed of all shares.
4. The data quality verification method in multi-party secure computing according to claim 2, wherein S12 includes: 1A2). The set of data holders; 1B2). Outsource the commitment to a trusted authority for storage.
5. The method for validating data quality in multi-party secure computing according to claim 2, wherein S13 includes: 1A3). The holder obtains a proof of the private key share; 1B3). Calculate the zero-knowledge proof of the ciphertext.
6. The data quality verification method in multi-party secure computing according to claim 2, wherein S14 includes: 1A4). The verifier verifies the proof of the private key share. If the proof is invalid, an exception warning is issued.
7. The method for validating data quality in multi-party secure computing according to claim 1, wherein The method of S2 specifically includes: S21. Encoding of the distribution characteristics of the dataset; S22. Verifiable representation of the dataset distribution characteristics; S23. Verifiable data quality verification method based on polynomial commitment.
8. The method for verifying data quality in multi-party secure computing according to claim 7, wherein S21 includes: 2A1). Select the feature X of the dataset. For a set of N data samples, take the value of each feature X in each sample as the feature vector x, bucket all the feature vectors x to obtain the bucket set V; 2B1). Generate the within-bucket feature mean through the bucket set V after bucketing; 2C1). For each bucket set Vi, calculate the number of data samples in each bucket set, and by comparing with the dataset N, obtain the proportion of the number of data samples in each bucket set Vi in the overall proportion, that is, the bucket distribution proportion pv.
9. The data quality verification method in multi-party secure computing according to claim 7, wherein The S22 includes: 2A2), for the generated bucket set V, first sort the bucket set to obtain the order l of each bucket set, encrypt the mean value of the features within the bucket according to the order l to obtain a hash value h, and use the encrypted mean value set of the features within the bucket and h as the distributed ciphertext; 2B2), for the data set Di of the data holder, count the proportion distribution of the bucket set V of the feature X to obtain the distribution feature set pi of each bucket set Vi; 2C2), for the encrypted bucket set Vi, use the secret sharing scheme to calculate its private key share di, and calculate the verification share according to di and pi; 2D2), for the share of each data holder, calculate the ciphertext m aggregated by multiple shares and its aggregation coefficient commitment value c, use c and m as the ciphertext value of the shared data set, generate a polynomial of the ciphertext based on the polynomial commitment, and finally the holder sends the ciphertext commitment s of its polynomial and the coefficient commitment ci.
10. A data quality verification system in multi-party secure computing, including a system server side, a data holder, and a data training party. The system server side is connected to the data holder and the data training party through network signals or wireless signals. The system server side is used to upload and download private data and public data, and provide data quality verification technology to the data holder. The data holder is used to generate data holder privacy protection shares for private data, detect data quality and data anomalies. The data training party is used to train public data and verify data quality. Among them, The system server side includes a data holder data set module, a data training party data set module, a system data set module, a multi-security protocol module, and an algorithm model module. The data holder data set module, the data training party data set module, and the system data set module are signal-connected to the multi-security protocol module, and the multi-security protocol module is signal-connected to the algorithm model module.
Citation Information
Patent Citations
Secure multi-party computing data system, method and equipment and data processing terminal
CN114172655A
Enterprise credit data privacy protection method and device based on zero-knowledge proof technology
CN119363357A
Federal learning service method and system oriented to data circulation
CN119670147A
Apparatus for secure multiparty computations for machine-learning
US20230269090A1
Method for implementing threshold signature, computer device, and storage medium
WO2025043917A1