Large-scale private data union set method

By constructing a verifiable commitment structure through multi-view encoding and cryptographic commitments, and combining secure computation and contribution assessment, the problem of consistent user identity identification in secure multi-party computation is solved, achieving fair collaboration of high-quality data and sustainable trust.

CN121834896APending Publication Date: 2026-04-10HANGZHOU CHUANGXIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, secure multi-party computation protocols cannot effectively identify the consistency of user identities, resulting in low-quality and erroneous associations during the data union process. This undermines data accuracy and trust, and in turn triggers adverse selection in the incentive system, leading to unfair data contributions and increased compliance risks.

Method used

A verifiable commitment structure is constructed using multi-view encoding and cryptographic commitment. Through multiple rounds of privacy set intersection and identity authenticity verification, combined with secure computation and contribution evaluation, weights and incentive mechanisms are dynamically allocated to ensure a fair evaluation and incentive mechanism for data quality and contribution.

Benefits of technology

It enables trusted collaboration and fair value distribution of high-quality data, ensures data accuracy and the sustainability of trust, avoids pollution by low-quality data and incentivizes adverse selection, and improves the accuracy and compliance of the joint data pool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834896A_ABST
    Figure CN121834896A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale private data union method, belongs to the technical field of data union, and solves the problems of identity alignment failure, stealth data stripping, excitation reversion and the like caused by the fact that an MPC technology cannot verify the authenticity and contribution degree of an identifier during multi-party computing cross-platform data union. And the problems of joint data asset degradation, business model collapse and alliance trust disruption are solved. Comprising the following steps: S1, an initialization and consensus establishment step: all participants jointly determine a set of global operation parameters comprising a plurality of threshold values and weight vectors through a secure multi-party negotiation protocol; according to the method, a verifiable identity alignment protocol and a fine-grained contribution evaluation and dynamic excitation compatible economic model based on security calculation are introduced into an MPC privacy protection framework by fusing cryptography and game theory mechanisms, so that the problems of identity authenticity verification deficiency, data contribution quantity misalignment, excitation structure distortion and the like are systematically solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data union technology, and more particularly to a method for large-scale privacy data union. Background Technology

[0002] Large-scale privacy data union is the process of securely fusing massive amounts of sensitive data from multiple sources under the principle of data usability without visibility. Privacy data refers to legally protected data such as medical records and financial transactions containing personally identifiable information. Data union merges multi-source data to form a complete new dataset, primarily releasing the value of aggregated data while protecting privacy. It has significant implications for many fields: merging hospital medical records can aid disease research, joint analysis between banks can strengthen fraud prevention, data fusion in government departments can optimize public services, and it can also provide diverse data for AI training.

[0003] In an alliance comprised of e-commerce platforms, ride-hailing service providers, and video platforms, the three parties aimed to construct a joint user profile by combining privacy data through secure multi-party computation (MPC) technology. However, due to the inherent inability of the MPC protocol to perform semantic verification of input data, it was unable to identify and correct issues such as inconsistencies in user identities, invalid identifiers, or deliberate "watering down" of the hash identifiers submitted by each party. This first undermined the identity alignment foundation of the data combination. This technical defect, combined with the black-box nature of the MPC process, made it impossible for any participant to measure the true data contribution of other parties to the joint profile. This led to an implicit power struggle under asymmetric data contributions and the silent erosion of data sovereignty. Furthermore, this opaque and unfair value exchange environment, lacking a verifiable contribution feedback mechanism, ultimately triggered adverse selection in the incentive system, prompting participants to rationally choose to submit low-quality identifiers with a quantity rather than quality strategy to "free-ride."

[0004] The aforementioned issues have led to the continuous contamination of the joint data pool by low-quality and incorrectly associated "stitched users," causing the accuracy of the generated profiles to collapse, resulting in the failure of precise services. Even the revenue-sharing business model that relies on transparent contribution metrics cannot operate. Trust within the alliance has been completely eroded due to an unauditable chain of suspicion. At the same time, the entire system faces hidden compliance risks because it cannot guarantee the accuracy of data association.

[0005] Therefore, a large-scale privacy data union method is proposed to solve or alleviate the above problems. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for large-scale privacy data union.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: A method for large-scale privacy-preserving data union includes the following steps: The S1 initialization and consensus establishment steps involve all participating parties jointly determining a set of global operating parameters, including multiple thresholds and weight vectors, through a secure multi-party negotiation protocol. In the S2 identifier preprocessing and verifiable commitment steps, each participant performs multi-view encoding on its local user identifier set and generates corresponding cryptographic commitments to construct a verifiable commitment structure. The S3 enhanced privacy set intersection and identity verification steps involve each party performing multiple rounds of privacy set intersection based on the encoded identifiers, and verifying the consistency of users' behavior and metadata across platforms within the intersection, calculating and filtering out the valid set of users whose identity authenticity meets the standards. The S4 data quality and feature contribution security assessment steps involve each participating party conducting a security assessment of the data in the effective user set from multiple dimensions, including completeness, accuracy, timeliness, and cross-platform consistency, calculating a platform-level data quality score, and securely assessing the global importance of each data feature. The S5 contribution matrix construction and dynamic weight allocation steps are based on identity authenticity, data quality, feature importance and synergy effect. The contribution components of each participant in multiple dimensions are calculated, the contribution matrix is ​​constructed and the dimension weights are determined through optimization learning. Finally, the comprehensive contribution of each participant is calculated and the sample and feature weights in joint modeling are allocated accordingly. The S6 joint modeling and performance verification steps employ a weighted federated learning framework to train the joint model based on the assigned weights, securely calculate the marginal contribution of each participant to the model performance, and monitor the consistency of model predictions to trigger potential realignment mechanisms. The S7 dynamic incentive mechanism and system optimization steps dynamically allocate benefits based on comprehensive contribution, marginal contribution, and consistency performance through a nonlinear benefit function and penalty mechanism, update the multidimensional reputation scores of each participant, and dynamically optimize global system parameters based on online learning feedback loop.

[0008] Preferably, the S1 initialization and consensus establishment step involves each participating party jointly determining a set of global operating parameters, including multiple thresholds and weight vectors, through a secure multi-party negotiation protocol. Specifically, this includes the following steps: Each participant first submits the encryption form of its proposed initial parameters, which include an identifier authenticity threshold, a data quality threshold, and multiple weight vectors. Through a secure multi-party computation protocol, the encrypted parameters are calculated by weighting the amount of data from each party and performing multiple rounds of iterative weighted average calculation until the variance of the parameter value is lower than a predetermined convergence threshold, thereby reaching a consensus. Based on the consensus results, a cryptographic infrastructure is jointly established through a distributed key generation protocol, including generating homomorphic encryption key pairs and commitment scheme parameters for subsequent computation, wherein the encryption private key is securely divided and held by each party respectively; Multiple secure salt values ​​are generated together and used as common inputs for subsequent hash calculations to ensure the consistency of the encoding.

[0009] Preferably, the S2 identifier preprocessing and verifiable commitment step, in which each participant performs multi-view encoding on its local user identifier set and generates corresponding cryptographic commitments to construct a verifiable commitment structure, specifically includes the following steps: Each participant calculates three independent security hashes for each of its user identifiers: a base layer hash based on the global salt and the original identifier, used for core alignment; a behavior layer hash based on different salts, platform identifiers, and user behavior feature summaries, used for consistency verification; and a metadata layer hash based on metadata features. A quantitative activity score is calculated for each user, which is obtained by applying a time-decay-weighted average to the user's historical activity records; A user-level cryptographic commitment is generated for each user, which binds the user's base layer hash to their activity score; simultaneously, a behavioral commitment is generated based on the user's behavioral feature vector. All user commitments are aggregated into a batch commitment, and a Merkle tree is constructed. The root hash value is then publicly disclosed as proof of the immutability of the initial state of the data.

[0010] Preferably, the S3 enhanced privacy set intersection and identity verification step involves each party performing multiple rounds of privacy set intersection based on the encoded identifier, and verifying the consistency of users' behavior and metadata across platforms within the intersection. The step then calculates and filters out the valid user set whose identity verification meets the criteria. Specifically, this includes the following steps: Three rounds of privacy set intersection are performed: the first round uses base layer hashing to perform multi-party intersection and obtain an initial intersection; the second round uses behavior layer hashing to cross-validate the initial intersection and filter out users who have behavioral evidence on at least two platforms; the third round evaluates the temporal consistency of the behavioral patterns of the remaining users across platforms and filters out inconsistent users. For users who pass the intersection, a comprehensive cross-platform identity authenticity score is securely calculated. This score is composed of the following weighted sums: the cross-validation index of activity scores across platforms, the consistency measure of temporal patterns in behavioral sequences, and the cross-platform similarity measure of metadata features. Based on the distribution of all users' authenticity scores in the current round, the authenticity filtering threshold is dynamically adjusted, and users with scores higher than the adjusted threshold are included in the final set of valid users. Generate a zero-knowledge proof for the entire selection process of the valid user set, proving that the selection process complies with the protocol rules and does not disclose original sensitive information.

[0011] Preferably, the S4 data quality and feature contribution security assessment step involves each participating party conducting a security assessment of the data in the effective user set from multiple dimensions, including completeness, accuracy, timeliness, and cross-platform consistency, calculating a platform-level data quality score, and securely assessing the global importance of each data feature. Specifically, this includes the following steps: For each user in the valid user set, each participant securely computes four quality metrics at the feature level: a completeness metric based on the proportion of non-null values, an accuracy metric obtained through leave-one-out cross-validation, a timeliness metric based on data freshness, and a consistency metric based on data similarity with other platforms. By using predefined prior weights for feature importance, the aforementioned feature-level quality indicators are aggregated into a data quality score for users on this platform. Each participant calculates the importance of each of its features to the prediction target locally, and through a secure multi-party computation protocol, the local feature importance of all participants is weighted and averaged using the platform-level data quality scores of each party as weights to obtain a global feature importance score. Based on the importance of global features, the platform-level data quality scores of each participant are recalculated and updated, forming a feedback loop for quality assessment.

[0012] Preferably, the S5 contribution matrix construction and dynamic weight allocation step, based on identity authenticity, data quality, feature importance, and synergy, calculates the contribution components of each participant in multiple dimensions, constructs a contribution matrix, determines the dimension weights through optimization learning, and finally calculates the comprehensive contribution of each participant, and allocates sample and feature weights in joint modeling accordingly. Specifically, it includes the following steps: For each participant, four contribution components are calculated: identifier contribution, based on the number, quality, and uniqueness of the effective users it provides; feature contribution, based on the global importance and diversity of the features it provides; data quality contribution, based on the absolute value, improvement, and stability of its platform-level quality score; and synergy contribution, based on the additional value gain generated by combining its data with data from other parties. The contribution components of each party in the four dimensions are arranged to form a contribution matrix; Using the actual marginal contribution of each party to the performance of the joint model as the objective, a regularized regression analysis is performed on the contribution matrix to solve for the optimal dimension weights, and the weights are then normalized. The contribution matrix is ​​weighted and summed using normalized weights to obtain the overall contribution of each participant. Based on the overall contribution of each participant and their data quality scores for each user, weights from different platforms are assigned to the user in the joint modeling process; at the same time, weights are assigned to each feature within the platform based on the importance of the global features.

[0013] Preferably, the S6 joint modeling and performance verification step, based on the assigned weights, uses a weighted federated learning framework to train the joint model, securely calculates the marginal contribution of each participant to the model performance, and monitors the consistency of model predictions to trigger a potential realignment mechanism, specifically including the following steps: Each participant uses locally allocated sample weights and feature weights to train the model on local data, and securely uploads model updates using homomorphic encryption technology. A safe weighted average algorithm is used to aggregate the model updates from all parties, where the aggregation weight is positively correlated with the overall contribution of each party and its local model performance; By combining Monte Carlo simulation with secure multi-party computation, the Shapley value of each participant is approximately calculated as a quantitative indicator of their marginal contribution to the performance of the joint model. Continuously monitor the consistency between the local model predictions of each participant and the global model predictions, and calculate the average consistency score and its stability; when the score is below the threshold or the stability is insufficient, automatically trigger the identity re-verification and data re-alignment process for the inconsistent user subset.

[0014] Preferably, the S7 dynamic incentive mechanism and system optimization steps, based on comprehensive contribution, marginal contribution, and consistency performance, dynamically allocate benefits through a nonlinear benefit function and penalty mechanism, update the multidimensional reputation scores of each participant, and dynamically optimize global system parameters based on online learning feedback loops, specifically including the following steps: Construct a total revenue pool, the total amount of which consists of three parts: basic service revenue, incentive revenue linked to the performance of the joint model, and incentive revenue linked to the growth of the effective user base. Design a nonlinear revenue distribution function that maps the overall contribution of participants to the basic revenue. This function adopts a piecewise power function form to incentivize medium contributors and reward high contributors. The base payout is adjusted according to the proportion of each participant's Shapley value, and then a penalty calculated based on their data quality defects, prediction inconsistencies and instabilities is deducted to obtain the final payout. Maintain a multidimensional reputation score vector, update it with an exponentially decaying moving average based on the historical performance of each participant in four aspects: identity authenticity, data quality, contribution and consistency, and use the reputation score for parameter weighting and permission allocation in subsequent rounds; By treating global system parameters as adjustable actions and using the overall performance of the system in terms of total revenue, fairness, and risk as rewards, the system can adapt to environmental changes through online learning and dynamic optimization using a contextual slot machine algorithm.

[0015] The present invention has the following beneficial effects: This invention integrates cryptography and game theory mechanisms, introducing a verifiable identity alignment protocol, a fine-grained contribution assessment based on secure computation, and a dynamic incentive-compatible economic model within the MPC privacy protection framework. This systematically addresses issues such as the lack of identity authenticity verification, inaccurate data contribution measurement, and distorted incentive structure, ensuring the high-quality evolution and sustainable value distribution of joint data assets. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0020] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0021] In the description of this invention, it should be understood that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used to facilitate the description of this invention and to simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0022] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0023] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0024] A method for large-scale privacy-preserving data union, such as Figure 1 As shown, it includes the following steps: The S1 initialization and consensus establishment steps involve all participating parties jointly determining a set of global operating parameters, including multiple thresholds and weight vectors, through a secure multi-party negotiation protocol. Each participant first submits the encryption form of its proposed initial parameters, which include an identifier authenticity threshold, a data quality threshold, and multiple weight vectors. Through a secure multi-party computation protocol, the encrypted parameters are calculated by weighting the amount of data from each party and performing multiple rounds of iterative weighted average calculation until the variance of the parameter value is lower than a predetermined convergence threshold, thereby reaching a consensus. Based on the consensus results, a cryptographic infrastructure is jointly established through a distributed key generation protocol, including generating homomorphic encryption key pairs and commitment scheme parameters for subsequent computation, wherein the encryption private key is securely divided and held by each party respectively; Multiple secure salt values ​​are generated together and used as common inputs for subsequent hash calculations to ensure the consistency of the encoding; In the S2 identifier preprocessing and verifiable commitment steps, each participant performs multi-view encoding on its local user identifier set and generates corresponding cryptographic commitments to construct a verifiable commitment structure. Each participant calculates three independent security hashes for each of its user identifiers: a base layer hash based on the global salt and the original identifier, used for core alignment; a behavior layer hash based on different salts, platform identifiers, and user behavior feature summaries, used for consistency verification; and a metadata layer hash based on metadata features. A quantitative activity score is calculated for each user, which is obtained by applying a time-decay-weighted average to the user's historical activity records; A user-level cryptographic commitment is generated for each user, which binds the user's base layer hash to their activity score; simultaneously, a behavioral commitment is generated based on the user's behavioral feature vector. Aggregate all users' commitments into a batch commitment, construct a Merkle tree, and publicly disclose the root hash value as proof of the immutability of the initial state of the data. The S3 enhanced privacy set intersection and identity verification steps involve each party performing multiple rounds of privacy set intersection based on the encoded identifiers, and verifying the consistency of users' behavior and metadata across platforms within the intersection, calculating and filtering out the valid set of users whose identity authenticity meets the standards. Three rounds of privacy set intersection are performed: the first round uses base layer hashing to perform multi-party intersection and obtain an initial intersection; the second round uses behavior layer hashing to cross-validate the initial intersection and filter out users who have behavioral evidence on at least two platforms; the third round evaluates the temporal consistency of the behavioral patterns of the remaining users across platforms and filters out inconsistent users. For users who pass the intersection, a comprehensive cross-platform identity authenticity score is securely calculated. This score is composed of the following weighted sums: the cross-validation index of activity scores across platforms, the consistency measure of temporal patterns in behavioral sequences, and the cross-platform similarity measure of metadata features. Based on the distribution of all users' authenticity scores in the current round, the authenticity filtering threshold is dynamically adjusted, and users with scores higher than the adjusted threshold are included in the final set of valid users. Generate a zero-knowledge proof for the entire selection process of the valid user set, proving that the selection process complies with the protocol rules and does not disclose original sensitive information; The S4 data quality and feature contribution security assessment steps involve each participating party conducting a security assessment of the data in the effective user set from multiple dimensions, including completeness, accuracy, timeliness, and cross-platform consistency, calculating a platform-level data quality score, and securely assessing the global importance of each data feature. For each user in the valid user set, each participant securely computes four quality metrics at the feature level: a completeness metric based on the proportion of non-null values, an accuracy metric obtained through leave-one-out cross-validation, a timeliness metric based on data freshness, and a consistency metric based on data similarity with other platforms. By using predefined prior weights for feature importance, the aforementioned feature-level quality indicators are aggregated into a data quality score for users on this platform. Each participant calculates the importance of each of its features to the prediction target locally, and through a secure multi-party computation protocol, the local feature importance of all participants is weighted and averaged using the platform-level data quality scores of each party as weights to obtain a global feature importance score. Based on the importance of global features, the platform-level data quality scores of each participant are recalculated and updated to form a feedback loop for quality assessment. The S5 contribution matrix construction and dynamic weight allocation steps are based on identity authenticity, data quality, feature importance and synergy effect. The contribution components of each participant in multiple dimensions are calculated, the contribution matrix is ​​constructed and the dimension weights are determined through optimization learning. Finally, the comprehensive contribution of each participant is calculated and the sample and feature weights in joint modeling are allocated accordingly. For each participant, four contribution components are calculated: identifier contribution, based on the number, quality, and uniqueness of the effective users it provides; feature contribution, based on the global importance and diversity of the features it provides; data quality contribution, based on the absolute value, improvement, and stability of its platform-level quality score; and synergy contribution, based on the additional value gain generated by combining its data with data from other parties. The contribution components of each party in the four dimensions are arranged to form a contribution matrix; Using the actual marginal contribution of each party to the performance of the joint model as the objective, a regularized regression analysis is performed on the contribution matrix to solve for the optimal dimension weights, and the weights are then normalized. The contribution matrix is ​​weighted and summed using normalized weights to obtain the overall contribution of each participant. Based on the overall contribution of each participant and their data quality scores for each user, weights from different platforms are assigned to the user in the joint modeling; at the same time, weights are assigned to each feature within the platform based on the importance of the global features. The S6 joint modeling and performance verification steps employ a weighted federated learning framework to train the joint model based on the assigned weights, securely calculate the marginal contribution of each participant to the model performance, and monitor the consistency of model predictions to trigger potential realignment mechanisms. Each participant uses locally allocated sample weights and feature weights to train the model on local data, and securely uploads model updates using homomorphic encryption technology. A safe weighted average algorithm is used to aggregate the model updates from all parties, where the aggregation weight is positively correlated with the overall contribution of each party and its local model performance; By combining Monte Carlo simulation with secure multi-party computation, the Shapley value of each participant is approximately calculated as a quantitative indicator of their marginal contribution to the performance of the joint model. Continuously monitor the consistency between the local model predictions of each participant and the global model predictions, and calculate the average consistency score and its stability; when the score is below the threshold or the stability is insufficient, automatically trigger the identity re-verification and data re-alignment process for the inconsistent user subset. The S7 dynamic incentive mechanism and system optimization steps dynamically allocate benefits based on comprehensive contribution, marginal contribution, and consistency performance through a nonlinear benefit function and penalty mechanism, update the multidimensional reputation scores of each participant, and dynamically optimize global system parameters based on online learning feedback loop. Construct a total revenue pool, the total amount of which consists of three parts: basic service revenue, incentive revenue linked to the performance of the joint model, and incentive revenue linked to the growth of the effective user base. Design a nonlinear revenue distribution function that maps the overall contribution of participants to the basic revenue. This function adopts a piecewise power function form to incentivize medium contributors and reward high contributors. The base payout is adjusted according to the proportion of each participant's Shapley value, and then a penalty calculated based on their data quality defects, prediction inconsistencies and instabilities is deducted to obtain the final payout. Maintain a multidimensional reputation score vector, update it with an exponentially decaying moving average based on the historical performance of each participant in four aspects: identity authenticity, data quality, contribution and consistency, and use the reputation score for parameter weighting and permission allocation in subsequent rounds; By treating global system parameters as adjustable actions and using the overall performance of the system in terms of total revenue, fairness, and risk as rewards, the system can adapt to environmental changes through online learning and dynamic optimization using a contextual slot machine algorithm.

[0025] The methodology also includes an auditability and dispute resolution step throughout: After key steps such as commitment generation, authenticity verification, and contribution calculation are completed, corresponding zero-knowledge proofs or verifiable computational proofs are generated, and the hash digests of the key data are stored on the blockchain. Establish a dispute triggering mechanism based on statistical anomaly detection, which automatically marks a dispute when a participant's benefit or data statement deviates significantly from its historical pattern or alliance consensus; After a dispute arises, the parties involved submit encrypted evidence to a trusted execution environment for arbitration calculation. The trusted execution environment verifies the evidence within an isolated and secure zone and outputs the ruling result. This process does not disclose the original data to any participating party. Based on the arbitration result, the pre-agreed compensation or penalty agreement shall be executed, and the credit scores of the relevant parties shall be updated; All operations involving multi-party collaborative computation in this method, including parameter negotiation, set intersection, quality assessment, contribution calculation, model aggregation, and marginal contribution assessment, are executed under the protection of a secure multi-party computation protocol and homomorphic encryption technology. This ensures that no participant can obtain the original input data, intermediate computation status, or individual information beyond its aggregation permissions from other participants, except for the outputs agreed upon in the protocol.

[0026] The large-scale privacy data union method proposed in this invention reverses the vicious cycle of collaborative failure and lays a sustainable foundation of trust and value exchange for cross-platform privacy data collaboration.

[0027] This method first introduces a secure multi-party parameter negotiation protocol and a distributed cryptographic infrastructure co-construction during the initialization and consensus-building phase. All participants reach a consensus on key thresholds and weights through multiple rounds of encrypted interaction and weighted average calculation, and jointly generate the secure salt value and encryption key required for subsequent calculations. This avoids the risk of initial unfairness and single point of failure caused by parameter bias or key centralization.

[0028] Following the identifier preprocessing and verifiable commitment stage, the method requires all parties to perform multi-view hashing of user identifiers and generate cryptographic commitments that bind behavioral characteristics, thereby constructing a publicly auditable Merkle tree. This transforms the previously invisible user activity and behavioral patterns into publicly verifiable but unreverse-inferable cryptographic evidence, making the data submission behavior itself auditable. This effectively deters speculative behavior that submits false or low-quality identifiers and provides a solid evidentiary anchor for subsequent authenticity verification.

[0029] Entering the enhanced privacy set intersection and identity verification phase, the multi-round progressive alignment protocol, after basic hash matching, further introduces behavioral consistency verification and cross-platform metadata similarity analysis. Through secure computation, it comprehensively evaluates the identity authenticity score of each alignment candidate and dynamically adjusts the filtering threshold based on the score distribution, ultimately selecting a highly credible and effective user set. This complex step solves the problem of identity alignment failure. It can not only exclude real user fragments using inconsistent identifiers such as different mobile phone numbers and email addresses, but also accurately identify and filter out water-filling attacks such as deliberately injected canceled numbers or short-term rental accounts, thereby ensuring the purity and consistency of the joint user pool and curbing the generation of ghost users.

[0030] Subsequently, in the data quality and security assessment phase, the method did not stop at verifying identity. Instead, it further conducted a detailed quantitative assessment of the data of valid users in an encrypted state from four dimensions: completeness, accuracy, timeliness, and cross-platform consistency. It also securely aggregated the platform-level quality scores of each participant and the global importance weight of each feature. The significance of this step lies in realizing the objective measurement of the intrinsic value of data in a privacy-preserving environment. This allows high-dimensional features such as high-value luxury consumption records and intercontinental flight information to be accurately identified in terms of their importance, while low-density movie viewing preferences are reasonably assessed for their auxiliary value. This provides a scientific basis for measuring fair contribution and breaks the exploitative dilemma of high-quality data being silently appropriated and assigned the value of low-quality data in the past.

[0031] Based on the high-quality identity and data inputs mentioned above, the method comprehensively deconstructs the inputs of each participant from four dimensions—identifier contribution, feature contribution, data quality contribution, and synergy contribution—in the contribution matrix construction and dynamic weight allocation stages. By constructing a contribution matrix and using the actual marginal contribution of each party to the joint model performance as the target feedback, the optimal weights of each dimension are safely learned, and a fair comprehensive contribution is calculated. Based on this, precise platform and feature weights are assigned to each sample in the joint modeling.

[0032] Transforming contribution measurement from subjective debate into a verifiable and optimizable objective calculation process ensures that the cost of maintaining high-quality identifiers and the value of providing core travel characteristics are accurately reflected and rewarded, eliminating the economic incentives for free-riding and incentivizing reverse behavior.

[0033] In the joint modeling and performance validation phase, the method adopts a federated learning framework that weights contribution and quality to ensure that high-value data receives an impact commensurate with its importance during model training. It also integrates secure multi-party computation and Monte Carlo sampling to approximate the Shapley values ​​of each participant, thereby quantifying the marginal contribution of each party to the final model performance. In addition, it continuously monitors prediction consistency and establishes a realignment trigger mechanism. These steps ensure that the joint model is not only high-performing but also fair and reliable. Data quality issues or malicious behavior from any party will be exposed through a decline in model consistency and trigger a remediation process, thus ensuring that the joint data assets are continuously optimized rather than degraded.

[0034] Ultimately, in the incentive mechanism and dynamic adjustment phase, a complex economic system integrating nonlinear payoff mapping, multidimensional reputation systems, and online parameter optimization is established. This system transforms comprehensive contribution and Shapley value into payoffs through piecewise power functions, imposes economic penalties on behaviors such as poor data quality and inconsistent predictions, and maintains a long-term evolving reputation score for each participant. Reputation directly affects the weight and payoff in future collaborations. The parameters of the entire system can also be dynamically optimized based on historical rewards through a contextual slot machine model to achieve long-term incentive compatibility. The establishment of this entire economic engine makes providing high-quality, real data the dominant strategy for all participants in repeated games. The business revenue sharing model can operate fairly based on transparent and verifiable metrics, and the trust in the alliance can be continuously strengthened rather than disintegrated under the guarantee of algorithms.

[0035] In summary, this approach not only overcomes the shortcomings of traditional MPC in terms of input quality blindness using cryptographic methods, but also solves the incentive problem in collaboration using game theory and economic models.

[0036] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for large scale private data union, the method comprising: The method comprises the following steps: S1 initialization and consensus establishment step, each participant determines a set of global running parameters including multiple thresholds and weight vectors through a secure multi-party negotiation protocol; S2 identifier preprocessing and verifiable commitment step, each participant encodes the local user identifier set and generates the corresponding cryptographic commitment, and constructs a verifiable commitment structure; S3 enhanced privacy set intersection and identity authenticity verification step, each party performs multiple rounds of privacy set intersection based on the encoded identifier, and performs cross-platform behavior and metadata consistency verification on the users in the intersection, and calculates and filters out the effective user set that meets the identity authenticity standard; S4 data quality and feature contribution safety evaluation step, each participant evaluates the data in the effective user set from multiple dimensions of integrity, accuracy, timeliness and cross-platform consistency, calculates the platform-level data quality score, and safely evaluates the global importance of each data feature; S5 contribution matrix construction and dynamic weight allocation step, based on identity authenticity, data quality, feature importance and synergy effect, the contribution of each participant in multiple dimensions is calculated, the contribution matrix is constructed, and the dimension weight is determined through optimization learning, and finally the comprehensive contribution of each participant is calculated, and the sample and feature weight in joint modeling are allocated accordingly; S6 joint modeling and performance verification step, based on the allocated weight, the joint model is trained using the weighted federated learning framework, and the marginal contribution of each participant to the model performance is safely calculated, and the consistency of the model prediction is monitored to trigger the potential realignment mechanism; S7 dynamic incentive mechanism and system optimization step, according to the comprehensive contribution, marginal contribution and consistency performance, the revenue is dynamically allocated through a nonlinear revenue function and a penalty mechanism, the multi-dimensional reputation points of each participant are updated, and the global system parameters are dynamically optimized based on the online learning feedback cycle.

2. The method of claim 1, wherein, The S1 initialization and consensus establishment step, each participant determines a set of global running parameters including multiple thresholds and weight vectors through a secure multi-party negotiation protocol, specifically comprising the following steps: Each participant first submits an encrypted form of their initial parameter proposal, which includes identifier authenticity threshold, data quality threshold and multiple weight vectors; Through a secure multi-party computation protocol, the encrypted parameters are calculated by weighted average in multiple rounds of iteration according to the data volume of each party, until the variance of the parameter value is lower than the predetermined convergence threshold, so as to reach a consensus; Based on the consensus result, a distributed key generation protocol is used to establish a cryptographic infrastructure, including generating a homomorphic encryption key pair and commitment scheme parameters for subsequent calculation, wherein the encryption private key is securely divided into separate possession by each party; A plurality of secure salt values are generated, which are used as public inputs for subsequent hash calculation to ensure consistency of encoding.

3. The method of claim 1, wherein, The S2 identifier preprocessing and verifiable commitment step, each participant encodes the local user identifier set and generates the corresponding cryptographic commitment, and constructs a verifiable commitment structure, specifically comprising the following steps: Each party computes three independent secure hash values for each of its user identifiers, respectively: a base layer hash based on a global salt value and the original identifier, for core alignment; a behavior layer hash based on different salt values, platform identifier and user behavior feature digest, for consistency verification; a metadata layer hash based on metadata features; A quantified activity score is calculated for each user, which is obtained by time-decay weighted average of the user's historical activity records; A user-level cryptographic commitment is generated for each user, which binds the user's base layer hash and activity score; meanwhile, a behavior commitment is generated based on the user behavior feature vector; All user commitments are aggregated into a batch commitment, and a Merkle tree is constructed, with the tree root hash value being publicly disclosed as the tamper-proof proof of the initial state of the party's data.

4. The method of claim 1, wherein, The S3 enhanced privacy set intersection and identity authenticity verification steps, each party performs multiple rounds of privacy set intersection based on the encoded identifiers, and performs consistency verification of cross-platform behavior and metadata on the users in the intersection, calculates and filters out the effective user set that meets the identity authenticity, specifically including the following steps: Perform three rounds of privacy set intersection: the first round uses the base layer hash to perform multi-party set intersection to obtain a preliminary intersection; the second round uses the behavior layer hash to cross-verify the preliminary intersection to filter out users with behavior evidence on at least two platforms; the third round evaluates the time consistency of the behavior patterns of the remaining users across platforms and filters out inconsistent users; For the users passed through the intersection, the cross-platform identity authenticity comprehensive score is securely calculated, which is composed of the following components: the cross-verification index of the activity score of each platform, the time pattern consistency measure of the behavior sequence, and the cross-platform similarity measure of the metadata features; According to the distribution of the authenticity scores of all users in the current round, the authenticity filtering threshold is dynamically adjusted, and the users with scores higher than the adjusted threshold are included in the final effective user set; A zero-knowledge proof is generated for the entire effective user set screening process, proving that the screening process complies with the protocol rules and does not leak original sensitive information.

5. The method of claim 1, wherein, The S4 data quality and feature contribution security evaluation step, each participant securely evaluates the data in the effective user set from multiple dimensions of integrity, accuracy, timeliness and cross-platform consistency, calculates the platform-level data quality score, and securely evaluates the global importance of each data feature, specifically including the following steps: For each user in the effective user set, each participant securely computes four quality indicators at the feature level: an integrity measure based on the proportion of non-empty values, an accuracy measure obtained by leave-one-platform cross-validation, a timeliness measure based on data freshness, and a consistency measure with other platform data; The above feature-level quality indicators are aggregated into the data quality score of the user on the platform using the pre-defined feature importance prior weight. The S3 enhanced privacy set intersection and identity authenticity verification steps, each party performs multiple rounds of privacy set intersection based on the encoded identifiers, and performs consistency verification of cross-platform behavior and metadata on the users in the intersection, calculates and filters out the effective user set that meets the identity authenticity, specifically including the following steps: Perform three rounds of privacy set intersection: the first round uses the base layer hash to perform multi-party set intersection to obtain a preliminary intersection; the second round uses the behavior layer hash to cross-verify the preliminary intersection to filter out users with behavior evidence on at least two platforms; the third round evaluates the time consistency of the behavior patterns of the remaining users across platforms and filters out inconsistent users; For the users passed through the intersection, the cross-platform identity authenticity comprehensive score is securely calculated, which is composed of the following components: the cross-verification index of the activity score of each platform, the time pattern consistency measure of the behavior sequence, and the cross-platform similarity measure of the metadata features; According to the distribution of the authenticity scores of all users in the current round, the authenticity filtering threshold is dynamically adjusted, and the users with scores higher than the adjusted threshold are included in the final effective user set; A zero-knowledge proof is generated for the entire effective user set screening process, proving that the screening process complies with the protocol rules and does not leak original sensitive information. Each participant calculates the importance of each feature to the prediction target locally, and through a secure multi-party computation protocol, the local feature importance of all participants is weighted and averaged to obtain the global feature importance score, with the platform-level data quality score of each party as the weight; Based on the global feature importance, the platform-level data quality score of each participant is recalculated and updated to form a feedback loop for quality assessment.

6. The method of claim 1, wherein, The S5 contribution matrix construction and dynamic weight allocation step calculates the contribution components of each participant in multiple dimensions based on identity authenticity, data quality, feature importance, and synergy effect, constructs a contribution matrix, and determines the dimension weight through optimization learning. Finally, the comprehensive contribution of each participant is calculated, and the sample and feature weights in joint modeling are allocated accordingly, which includes the following steps: Calculate the contribution component of each participant in four dimensions: identifier contribution, based on the number of valid users, quality, and uniqueness provided; Feature contribution, based on the global importance and diversity of the features provided; Data quality contribution, based on the absolute value, improvement, and stability of the platform-level quality score; Synergy effect contribution, based on the additional value gain generated by combining the data of other parties; Arrange the contribution components of each party in four dimensions to form a contribution matrix; Take the actual marginal contribution of each party to the joint model performance as the target, and perform a regression analysis with regularization on the contribution matrix to solve the optimal dimension weight, and normalize the weight; Use the normalized weight to perform weighted summation on the rows of the contribution matrix to obtain the comprehensive contribution of each participant; According to the comprehensive contribution of each participant and its data quality score for each user, allocate the weight from different platforms in joint modeling; At the same time, according to the global feature importance, allocate the weight to each feature within the platform.

7. The method of claim 1, wherein, The S6 joint modeling and performance verification step trains the joint model based on the allocated weight using the weighted federated learning framework, and securely calculates the marginal contribution of each participant to the model performance, while monitoring the consistency of the model prediction to trigger the potential realignment mechanism. It includes the following steps: Each participant uses the locally allocated sample weight and feature weight to train the model on local data and securely uploads the model update through homomorphic encryption technology; Use a secure weighted average algorithm to aggregate the model updates of each party, where the aggregation weight is positively correlated with the comprehensive contribution of each party and its local model performance; Through the combination of Monte Carlo simulation and secure multi-party computation, the Sharpe value of each participant is approximately calculated as a quantitative indicator of its marginal contribution to the joint model performance; Continuously monitor the consistency of the local model prediction and the global model prediction of each participant, and calculate the average consistency score and its stability; When the score is below the threshold or the stability is insufficient, automatically trigger the identity re-verification and data realignment process for the inconsistent user subset.

8. The method of claim 1, wherein, The S7 dynamic incentive mechanism and system optimization step dynamically allocates the revenue through a nonlinear revenue function and a penalty mechanism according to the comprehensive contribution, marginal contribution and consistency performance, updates the multi-dimensional credit points of each participant, and dynamically optimizes the global system parameters based on the online learning feedback cycle, and specifically includes the following steps: A total revenue pool is constructed, and the total amount of the pool is composed of three parts, i.e., a basic service revenue, an incentive revenue linked to the performance of the joint model, and an incentive revenue linked to the growth of the effective user scale; A nonlinear revenue allocation function is designed to map the comprehensive contribution of the participants to the basic revenue, and the function adopts a piecewise power function form to encourage medium contributors and reward high contributors; The basic revenue is adjusted according to the Sharpe value proportion of each participant, and then the penalty items calculated based on the data quality defects, prediction inconsistency and instability are deducted to obtain the final revenue; A multi-dimensional credit point vector is maintained, and the historical performance of each participant in terms of identity authenticity, data quality, contribution and consistency is updated by an exponentially decaying moving average, and the credit points are used for subsequent round parameter weighting and permission allocation; The system global parameters are taken as adjustable actions, the comprehensive performance of the total system revenue, fairness and risk are taken as rewards, the contextual tiger machine algorithm is used for online learning and dynamic optimization, so that the system can adapt to environmental changes.