Cross-domain collaborative privacy protection high-dimensional data construction method and system

By employing a cross-domain collaborative privacy-preserving high-dimensional data construction method, utilizing FM-Sketch and differential privacy techniques, the privacy protection problem of high-dimensional data under multi-data-party cross-domain collaboration is solved. This method achieves joint distribution estimation of cross-party attribute pairs and efficient construction of high-dimensional data, and is applicable to scenarios such as medical diagnosis and financial risk control.

CN121935962APending Publication Date: 2026-04-28XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI AN JIAOTONG UNIV
Filing Date
2026-01-15
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies struggle to protect the privacy of high-dimensional data in cross-domain collaboration scenarios involving multiple data parties, especially when original data is not shared. They cannot accurately estimate the joint distribution of cross-party attributes, and the preservation of statistical features and data utility of high-dimensional data are affected by noise superposition and computational costs.

Method used

A cross-domain collaborative method for constructing privacy-preserving high-dimensional data is adopted. By combining FM-Sketch technology and marginal stepwise construction strategy with differential privacy mechanism, the joint distribution estimation of cross-domain attribute pairs is achieved. The low-dimensional marginal distribution is optimized by greedy algorithm and gradient descent method to generate high-quality high-dimensional dataset.

Benefits of technology

Without sharing the original data, it achieves efficient and accurate estimation of joint distribution of cross-dimensional attributes, alleviates the dimensionality curse of high-dimensional data, reduces noise and computational overhead, ensures data statistical quality, and supports subsequent data analysis tasks such as medical diagnosis and financial risk control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935962A_ABST
    Figure CN121935962A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain collaborative privacy protection high-dimensional data construction method and system, and aims to solve the problems of'data islands', 'dimension chants' and privacy and utility imbalance faced by high-dimensional data construction when multiple data parties do not share original data in a cross-domain collaborative scene. According to the method, accurate estimation of joint distribution of cross-party attributes is realized through an FM-Sketch technology, and cross-party correlation can be captured without sharing original data; a low-dimensional marginal construction-high-dimensional extension strategy is adopted, a key low-dimensional marginal is screened based on InDif correlation measurement and a greedy algorithm, high-dimensional data are constructed step by step, and a dimension curse is relieved; the method comprises the following steps: constructing a full-process differential privacy system, fusing a Gaussian mechanism and zero-set differential privacy (zCDP) in the links of preprocessing, local data processing, cross-square estimation and high-dimensional data construction, and distributing privacy budget according to 2 / 3 power of the size of an attribute domain; and iteratively optimizing the high-dimensional data set through a step-by-step updating (GUM) method, and retaining the distribution characteristics and attribute correlation of the original data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of cross-domain data collaborative processing and privacy protection technology, specifically relating to a cross-domain collaborative method and system for constructing privacy-protected high-dimensional data. Background Technology

[0002] In the era of big data, data has become a key production factor, widely used in fields such as medical diagnosis optimization, financial risk assessment, and government decision support. However, the risk of privacy breaches during data collection and sharing (such as the illegal acquisition of sensitive personal information) hinders the development of cross-institutional data collaboration, necessitating reliable privacy protection technologies. Differential privacy technology, by introducing controllable random noise, limits the impact of individual data on the results, becoming a core means of privacy protection and has been applied in scenarios such as Google and the U.S. Census Bureau. However, traditional differential privacy has limitations in high-dimensional data and distributed scenarios: the "curse of dimensionality" in high-dimensional data leads to noise superposition, reducing the statistical utility of the data; centralized architectures rely on trusted third parties and cannot adapt to the collaborative needs of multiple parties that do not share original data. Multi-data-party cross-domain collaborative processing is suitable for cross-institutional collaborations involving "same user, different attributes" (e.g., banks holding financial data and e-commerce platforms holding consumer data for joint modeling). However, preserving the statistical characteristics of high-dimensional data under this model still faces core challenges: First, cross-party attribute correlation estimation is prone to information loss due to privacy measures (noise, anonymization), affecting the quality of data statistics; second, it is difficult to balance local and cross-party attribute information, and distributed scenarios cannot replicate the accuracy of centralized statistics; third, the large number of high-dimensional attribute combinations increases noise, communication, and computational costs. Existing technologies are difficult to adapt to cross-domain collaboration needs: methods based on probabilistic graphical models, such as PrivBayes and PGM, rely on sparse structures and manual parameter tuning, and cannot automate the processing of high-dimensional data; while PrivSyn can be extended to high-dimensional statistical scenarios through low-dimensional margins, it is only applicable to single data nodes and does not support cross-node attribute collaboration. In summary, there is an urgent need for a high-dimensional data processing technology that integrates cross-domain collaboration characteristics and differential privacy to solve the core problems of multi-data-party collaboration, cross-party correlation preservation, and high-dimensional data utility improvement, so as to meet the practical needs of cross-institutional data privacy sharing and in-depth analysis. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a cross-domain collaborative method and system for constructing privacy-preserving high-dimensional data, aiming to solve the following problems: 1. Achieving accurate estimation of the joint distribution and correlation of cross-party attribute pairs without sharing original data among multiple data parties; 2. Alleviating the "curse of dimensionality" of high-dimensional data by refining and expanding low-dimensional statistical features to reduce noise and computational overhead; 3. Strictly meeting differential privacy requirements by using Gaussian mechanisms, zero-centralized differential privacy (zCDP), and other technologies to ensure privacy and security throughout the entire process of data preprocessing, local statistical analysis, cross-node correlation estimation, and high-dimensional data construction; 4. Improving the quality of data statistics by preserving the distribution characteristics and attribute correlations of the original data, supporting subsequent data analysis tasks such as classification and clustering.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: A cross-domain collaborative method for constructing privacy-preserving high-dimensional data includes the following steps: Step 1, Parameter Setting and Data Preprocessing Stage: The data collector confirms the number of participating data parties. Attribute sets of each data source Privacy Budget GUM iteration count and convergence threshold Data providers and initializing the user dataset , For a global attribute set, perform equal-width or equal-frequency binning on each data attribute to divide the attribute values ​​into... Each bin is divided into intervals; the binned intervals are then numerically or symbolically encoded and injected with Gaussian noise. The encoded dataset is obtained. , Allocated by the privacy budget; Step 2, Local Data Processing Stage: Each data provider independently calculates the InDif relevance metric locally. ,in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; allocating local privacy budgets to powers of 2 / 3 of the attribute domain size. , , For the size of the attribute field, For the total privacy budget; a low-dimensional boundary is selected using a greedy algorithm to construct an attribute association graph. And select eligible clusters to form a marginal set. ; Step 3, Data Upload and Cross-Party Attribute Estimation Stage: Each data party collaboratively generates hash keys through secure multi-party computation. Each data provider adds virtual elements to the user set corresponding to the local attribute values, generates an FM-Sketch based on differential privacy, and records the maximum hash value and the least significant bit; the data collector calculates the cardinality of the union of attribute value sets based on the FM-Sketch of multiple data providers. Combined with the cardinality of single data attribute values , Through the principle of inclusion and exclusion After obtaining the intersection cardinality, the intersection cardinality is standardized to obtain the marginal distribution estimate of the cross-attribute pairs. ; Step 4, Consistency Processing Phase: Apply privacy budget to shared attributes based on the marginal table. With contribution cell count Assign weights Calculate the weighted average probability of attribute values. , ( For the first Each marginal table pair The probability estimate is obtained; the initial probability estimate is projected onto the nearest effective probability distribution; the weighted average and projection operations are performed alternately until the marginal table is reached. or If the distance is less than the convergence threshold, convergence is considered achieved, and iteration stops. Step 5, High-Dimensional Data Construction Stage: Initialize the Iteration Counter Attenuation rate and the unselected boundary set; when Maximum number of iterations At that time, traverse the unselected edges Calculate high-dimensional datasets With the target margin The difference, dynamically calculate and update the step size Reduce step size , This is the current record number. The target number of records; through Update the high-dimensional dataset, if Then terminate the current marginal update. The convergence threshold; join in Clear after traversal. And increasing After iteration, a high-dimensional dataset is output. .

[0005] A further improvement of this invention is that, in step 1, the specific method of data binning is selected based on data distribution and privacy requirements: for continuous attributes, equal-frequency binning is preferentially used to divide them into discrete intervals; for discrete attributes, equal-width binning is used to supplement the interval division; Gaussian noise. of The value is determined based on the privacy budget allocated to each data provider, ensuring that attribute boundary privacy is not leaked during the coding process.

[0006] A further improvement of this invention is that, in step 2, the core calculation and filtering logic for local data processing specifically includes the following steps: Step 2.1 Correlation measurement calculation: based on The measurement method is defined as follows: ;in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; Step 2.2, Optimal Privacy Budget Allocation: Allocate local privacy budget Distribute the attribute domain size to each boundary to minimize the overall noise error, using the formula: Total budget Allocate based on the size of the data source attributes; Step 2.3, Marginal Selection: A greedy algorithm is used to select attribute pairs to minimize the overall error; first, the selected marginal set is initialized. Iteration counting Calculate the initial total error Iterate through unselected 2D edges ,for Privacy budget allocation at various margins Calculate the margin Total error after From the noise error term With residual dependency error term Composition; selection to make Minimum margin If the total error does not decrease, the iteration will terminate; otherwise, it will... join in X And continue iterating until the selected marginal set is obtained. ; Step 2.4, Marginal Combination: First, combine the two-dimensional margins... Convert to corresponding attributes, and construct a graph based on the association relationships between attribute pairs. Then initialize the selected attribute set. Marginal set after combination From the size of the group arrive Traverse the graph sequentially. Each of the following sizes is The group ,like and The number of overlapping attributes does not exceed 2 and field size Then join in and will Add the attribute to , The number of attribute pairs is used to obtain the marginal set that preserves the joint distribution characteristics of high-dimensional attributes. .

[0007] A further improvement of this invention lies in the initial total error. Dependency error for all unselected two-dimensional boundaries sum.

[0008] A further improvement of this invention is that, in step 3, the FM-Sketch generation and marginal calculation of cross-domain attribute estimation specifically includes the following steps: Step 3.1, Hash Key Collaborative Generation: Each data party collaboratively generates a hash key through secure multi-party computation. This ensures the privacy of the hashing process; Step 3.2, Local FM-Sketch Generation: Data side Each attribute Its attribute value Corresponding user attribute set To each Add virtual elements and compute the corresponding FM-Sketch based on differential privacy; calculate the hash value of the values ​​in the user set using a hash function, update the maximum hash value in the FM-Sketch structure using this hash value, and maintain an array to track the least significant bit of the hash value to update the value at the corresponding position, finally obtaining the FM-Sketch record set of the data party: , Size of the attribute domain; Step 3.3, Cross-Side Marginal Distribution Estimation: Cross-side marginal distribution estimation is based on data... Each FM-Sketch Implementation: First, initialize the marginal distribution estimation set of cross-attribute pairs. , then traverse and FM-Sketch: All attributes and their corresponding values , Then calculate the union FM-Sketch: = max( , And calculate the cardinality of the union. Combined with attribute values , cardinality , Through the principle of inclusion and exclusion Obtain the intersection cardinality, and use this intersection cardinality as the cross-attribute pair. , Marginal distribution counts are stored. Finally, for The counts of all attribute pairs are standardized to obtain the marginal distribution estimate of cross-domain attribute pairs. .

[0009] A further improvement of this invention is that, in step 4, the weighted average and projection optimization for consistency processing specifically includes the following steps: Step 4.1, Weighted Average Integration: For shared attributes Shared attributes quilt Each marginal table contains a privacy budget for each marginal table. With contribution cell count Assign weights ,in, It is assigned to the marginal table Privacy budget, It is a marginal table Used for attributes The number of cells; then calculate the attributes. Each possible value Weighted average probability: ;in For the first Each marginal table pair The probability estimate; Step 4.2, Projecting to the efficient distribution: First, perform an initial estimation based on the current distribution, assuming a marginal table. The initial probability estimate is Next, the initial probability estimate will be projected onto the nearest effective probability distribution. This makes it the distance to the initial probability estimate. Minimum; this problem is described as an optimization problem:

[0010] Its constraints are: ; Then, the gradient descent method is used to solve the optimization problem; Step 4.3, Iterative Convergence: To address the consistency issue among multiple marginal tables simultaneously, weighted averaging and projection operations are performed alternately until convergence. In each iteration, a weighted average operation is first performed on all shared attributes, and then each marginal table is projected onto an effective distribution. When the difference between the marginal tables is less than a set threshold, convergence is considered achieved, and the iteration stops.

[0011] A further improvement of this invention is that, in step 5, the update operation rule for the GUM iteration is as follows: first, initialize the iteration counter. And set the attenuation rate At the same time, initialize the selected edge set. It is an empty set; when < Maximum number of iterations At that time, traverse all sets belonging to the marginal set. And not selected the edge : Computing high-dimensional datasets With the target margin Differences And based on the current number of records and according to Calculate the target number of records Dynamically calculate and update step size and reduce step size Based on the number of iterations Choose either "Replace" or "Copy" to proceed. If the norm of a high-dimensional dataset changes after the update, then... Then terminate the current marginal update, and subsequently adjust the step size and... join in Clear after traversal. And increasing Finally, return the high-dimensional dataset after the iteration is complete. .

[0012] A cross-domain collaborative system for constructing privacy-preserving high-dimensional data includes the following steps: Parameter setting and data preprocessing unit: The data collector confirms the number of participating data parties. Attribute sets of each data source Privacy Budget GUM iteration count and convergence threshold Data providers and initializing the user dataset , For a global attribute set, perform equal-width or equal-frequency binning on each data attribute to divide the attribute values ​​into... Each bin is divided into intervals; the binned intervals are then numerically or symbolically encoded and injected with Gaussian noise. The encoded dataset is obtained. , Allocated by the privacy budget; Local data processing unit: Each data provider independently calculates the InDif relevance metric locally. ,in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; allocating local privacy budgets to powers of 2 / 3 of the attribute domain size. , , For the size of the attribute field, For the total privacy budget; a low-dimensional boundary is selected using a greedy algorithm to construct an attribute association graph. And select eligible clusters to form a marginal set. ; Data upload and cross-party attribute estimation unit: Data parties collaboratively generate hash keys through secure multi-party computation. Each data provider adds virtual elements to the user set corresponding to the local attribute values, generates an FM-Sketch based on differential privacy, and records the maximum hash value and the least significant bit; the data collector calculates the cardinality of the union of attribute value sets based on the FM-Sketch of multiple data providers. Combined with the cardinality of single data attribute values , Through the principle of inclusion and exclusion After obtaining the intersection cardinality, the intersection cardinality is standardized to obtain the marginal distribution estimate of the cross-attribute pairs. ; Consistency processing unit: Privacy budget for shared attributes based on marginal table With contribution cell count Assign weights Calculate the weighted average probability of attribute values. , ( For the first Each marginal table pair The probability estimate is obtained; the initial probability estimate is projected onto the nearest effective probability distribution; the weighted average and projection operations are performed alternately until the marginal table is reached. or If the distance is less than the convergence threshold, convergence is considered achieved, and iteration stops. High-dimensional data building unit: Initializing the iteration counter Attenuation rate and the unselected boundary set; when Maximum number of iterations At that time, traverse the unselected edges Calculate high-dimensional datasets With the target margin The difference, dynamically calculate and update the step size Reduce step size , This is the current record number. The target number of records; through Update the high-dimensional dataset, if Then terminate the current marginal update. The convergence threshold; join in Clear after traversal. And increasing After iteration, a high-dimensional dataset is output. .

[0013] A further improvement of this invention lies in that, in the parameter setting and data preprocessing unit, the specific method of data binning is selected according to data distribution and privacy requirements: for continuous attributes, equal-frequency binning is preferentially used to divide them into discrete intervals, and for discrete attributes, equal-width binning is used to supplement the interval division; Gaussian noise. of The value is determined based on the privacy budget allocated to each data provider, ensuring that attribute boundary privacy is not leaked during the coding process.

[0014] A further improvement of this invention is that, in the local data processing unit, the core calculation and filtering logic for local data processing specifically includes the following steps: Step 2.1 Correlation measurement calculation: based on The measurement method is defined as follows: ;in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; Step 2.2, Optimal Privacy Budget Allocation: Allocate local privacy budget Distribute the attribute domain size to each boundary to minimize the overall noise error, using the formula: Total budget Allocate based on the size of the data source attributes; Step 2.3, Marginal Selection: A greedy algorithm is used to select attribute pairs to minimize the overall error; first, the selected marginal set is initialized. Iteration counting Calculate the initial total error Iterate through unselected 2D edges ,for Privacy budget allocation at various margins Calculate the margin Total error after From the noise error term With residual dependency error term Composition; selection to make Minimum margin If the total error does not decrease, the iteration will terminate; otherwise, it will... join in X And continue iterating until the selected marginal set is obtained. ; Step 2.4, Marginal Combination: First, combine the two-dimensional margins... Convert to corresponding attributes, and construct a graph based on the association relationships between attribute pairs. Then initialize the selected attribute set. Marginal set after combination From the size of the group arrive Traverse the graph sequentially. Each of the following sizes is The group ,like and The number of overlapping attributes does not exceed 2 and field size Then join in and will Add the attribute to , The number of attribute pairs is used to obtain the marginal set that preserves the joint distribution characteristics of high-dimensional attributes. .

[0015] Compared with the prior art, the present invention has at least the following beneficial technical effects: This invention provides a cross-domain collaborative method and system for constructing privacy-preserving high-dimensional data. By leveraging FM-Sketch technology and a marginal step-by-step data construction strategy, it overcomes the limitation of "data silos" among multiple data parties: under the premise that each data party does not share original sensitive data, it achieves efficient and accurate estimation of the joint distribution of cross-party attributes; at the same time, through the technical approach of "constructing high-dimensional data with low-dimensional margins," it effectively alleviates the "curse of dimensionality" of high-dimensional data, making the construction of high-dimensional datasets of hundreds of dimensions and above practically feasible for deployment, and providing technical support for cross-party data collaboration in high-dimensional scenarios.

[0016] Furthermore, a comprehensive differential privacy technology system is constructed, combined with an adaptive privacy budget allocation strategy: differential privacy mechanisms are strictly followed throughout all stages of data preprocessing (binning and encoding), local margin selection (greedy algorithm balancing noise and correlation error), cross-square margin estimation (FM-Sketch noise injection), and high-dimensional data construction (GUM iteration). By dynamically allocating the privacy budget according to attribute domain size and attribute correlation, the impact of a single data record on the output is strictly limited, while the limited privacy budget is used efficiently. This achieves a better balance between "privacy leakage risk" and "data utility loss," solving the problem of insufficient utility or protection caused by the coarse allocation of privacy budgets in traditional methods.

[0017] Furthermore, a stepwise update method (GUM) is employed to iteratively optimize the high-dimensional dataset, accurately preserving the distribution characteristics of the original data (such as the marginal probability density of each attribute) and cross-attribute correlations (such as the joint distribution pattern of attribute pairs from different data sources), generating high-quality high-dimensional data. This method not only effectively captures the complex relationships between cross-attributes but also significantly improves the distribution consistency and practicality of high-dimensional data. High-dimensional data can directly support high-precision data analysis tasks such as medical diagnosis and financial risk control.

[0018] Furthermore, experiments were conducted using multiple real-world datasets to verify that the VertiSyn method can generate high-quality high-dimensional data under different privacy budgets and effectively support various data analysis tasks. It has significant advantages over mainstream algorithms in high-dimensional data and complex data distribution scenarios, and can better preserve the distribution characteristics and correlations of the original data. Attached Figure Description

[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating the method of the present invention.

[0021] Figure 2 This is a schematic diagram of the logical architecture of the present invention.

[0022] Figure 3 This is a schematic diagram of the results of l-way TVD, which is the evaluation index in an embodiment of the present invention.

[0023] Figure 4This is a schematic diagram showing the results of an embodiment of the present invention where the evaluation metric is the SVM classification error rate.

[0024] Figure 5 This is a structural block diagram of the system of the present invention. Detailed Implementation

[0025] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0026] In the description of this invention, it should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0027] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0028] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0029] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.

[0030] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0031] Example 1 Reference Figure 1This invention provides a privacy-preserving high-dimensional data construction method for cross-domain collaboration, comprising: addressing the need for high-dimensional data construction in cross-domain collaboration scenarios where multiple data parties do not share original data, achieving accurate estimation of the joint distribution of cross-party attributes through FM-Sketch technology, breaking through the limitation of "data silos"; adopting a strategy of first constructing a low-dimensional marginal distribution and then expanding and constructing high-dimensional statistical data based on this distribution to alleviate the "curse of dimensionality" of high-dimensional data and reduce noise interference and computational overhead; constructing a full-process differential privacy technology system, strictly adhering to the differential privacy mechanism in the data preprocessing, local data processing, cross-party attribute estimation, and high-dimensional data construction stages, and combining a privacy budget adaptive allocation strategy to balance privacy protection and data utility; and iteratively optimizing the high-dimensional dataset through the stepwise update method (GUM) to retain the distribution characteristics and attribute correlations of the original data, applicable to data analysis tasks such as classification and clustering in scenarios such as medical diagnosis and financial risk control.

[0032] The present invention specifically includes the following steps: Step 1: Parameter setting and data preprocessing stage Step 1.1, Parameter Settings: The data collection party (central server) confirms the number of participating data parties. Attribute sets of each data source ( For global attribute sets), privacy budget GUM iteration count and convergence threshold Data sources and initializing the user dataset (The data has been aligned using some kind of record ID).

[0033] Step 1.2, Data Binning: For continuous or discrete attributes of each data source, use equal-width binning or equal-frequency binning (selected according to data distribution and privacy requirements) to divide the attribute values ​​into bins. To mitigate the dimensionality curse, the attribute domain size can be reduced by dividing the "income" attribute (a continuous value) into three intervals: "low," "medium," and "high," based on equal frequency binning.

[0034] Step 1.3, Data Encoding: The binned intervals are numerically or symbolically encoded (e.g., mapping "low, medium, high" to 1, 2, 3), while simultaneously incorporating local noise perturbation (Gaussian noise) during the encoding process. , (Assigned by the privacy budget), protecting bin boundary privacy, resulting in the encoded dataset. .

[0035] Step 2: Local Data Processing Stage: Gaussian noise is added to the local data repository to ensure differential privacy requirements are met. Simultaneously, a greedy algorithm is used for margin selection, carefully choosing a set of low-dimensional margins to effectively describe the key features of the dataset. Each data repository independently performs the following operations locally, without disclosing the original data: Step 2.1, InDif Correlation Measurement Calculation: This invention requires a measurement method capable of measuring the correlation between attributes, and therefore proposes a measurement method called Independent Differerce (InDif), which is defined as: ;in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent.

[0036] Step 2.2, Optimal Privacy Budget Allocation: Allocate local privacy budget (Total Budget) (Allocation based on data domain attribute size) Allocating data to each boundary by raising the attribute domain size to a power of 2 / 3 can minimize the overall noise error. The formula is: .

[0037] Step 2.3, Marginal Selection: A greedy algorithm is used to select attribute pairs to minimize the overall error. First, the selected marginal set is initialized. Iteration counting Calculate the initial total error (Dependency error of all unselected two-dimensional boundaries) (sum). Iterative traversal of unselected 2D edges. ,for Privacy budget allocation at various margins Calculate the margin Total error after (from noise error term) With residual dependency error term Composition); selection to make Minimum margin If the total error does not decrease, the iteration will terminate; otherwise, it will... join in X And continue iterating until the selected marginal set is obtained. .

[0038] Step 2.4, Marginal Combination: First, combine the two-dimensional margins... Convert to corresponding attributes, and construct a graph based on the association relationships between attribute pairs. Then initialize the selected attribute set. Marginal set after combination From the size of the group ( (for the number of attribute pairs) to Traverse the graph sequentially. Each of the following sizes is The group ,like and The number of overlapping attributes does not exceed 2 and field size Then join in and will Add the attribute to Finally, a marginal set that preserves the joint distribution characteristics of high-dimensional attributes is obtained. .

[0039] Step 3: Data Upload and Cross-Side Attribute Estimation Stage: The central server uses Flajolet-Martin (FM) Sketch technology to estimate the joint distribution of cross-side attribute pairs without acquiring the original data.

[0040] Step 3.1, Hash Key Collaborative Generation: Data parties collaboratively generate hash keys through secure multi-party computation (such as the Diffie-Hellman protocol). (The central server is not visible), ensuring the privacy of the hashing process.

[0041] Step 3.2, Local FM-Sketch Generation: Data side Each attribute Its attribute value Corresponding user attribute set To each Add virtual elements and calculate the corresponding FM-Sketch based on differential privacy (DP); calculate the hash value of the values ​​in the user set using a hash function, update the maximum hash value in the FM-Sketch structure using this hash value, and maintain an array to track the least significant bit (LSB) of the hash value to update the value at the corresponding position, finally obtaining the FM-Sketch record set of the data party: ( (This refers to the size of the attribute domain).

[0042] Step 3.3, Cross-Side Marginal Distribution Estimation: Cross-side marginal distribution estimation is based on data... Each FM-Sketch Implementation: First, initialize the marginal distribution estimation set of cross-attribute pairs. , then traverse and FM-Sketch: All attributes and their corresponding values , Then calculate the union FM-Sketch: = max( , And calculate the cardinality of the union. Combined with attribute values , cardinality , By using the principle of inclusion-exclusion Obtain the intersection cardinality, and use this intersection cardinality as the cross-attribute pair. , Marginal distribution counts are stored. Finally, for The counts of all attribute pairs are standardized to obtain the marginal distribution estimate of cross-domain attribute pairs. .

[0043] Step 4: Consistency Processing Stage: In the previous steps, we have obtained the one-dimensional and two-dimensional marginal distributions with added noise. The two-dimensional marginal distribution here also includes the distribution of cross-domain attribute pairs. Next, we need to perform consistency processing on the currently selected low-dimensional marginals, integrate the two-dimensional marginals corresponding to cross-domain attribute pairs and local attribute pairs, and ensure the consistency and integrity of the data.

[0044] Step 4.1, Weighted Average Integration: For shared attributes (quilt Each marginal table contains privacy budgets based on the privacy of each marginal table. With contribution cell count Assign weights ,in, It is assigned to the marginal table Privacy budget, It is a marginal table Used for attributes The number of cells. Then calculate the attributes. Each possible value Weighted average probability: ;in For the first Each marginal table pair The probability estimate.

[0045] Step 4.2, Projecting to the efficient distribution: First, perform an initial estimation based on the current distribution, assuming a marginal table. The initial probability estimate is This may contain invalid values ​​(such as negative probabilities or probabilities whose sum is not equal to 1). Next, the initial probability estimate will be projected onto the nearest effective probability distribution. This makes it the distance to the initial probability estimate. Minimize. This problem can be described as an optimization problem:

[0046] Its constraints are:

[0047] Then, the gradient descent method is used to solve the above optimization problem.

[0048] Step 4.3, Iterative Convergence: To simultaneously address the consistency issue among multiple marginal tables, weighted averaging and projection operations need to be performed alternately until convergence. In each iteration, a weighted average operation is first performed on all shared attributes, followed by a projection operation onto each marginal table to an efficient distribution. When the differences between marginal tables (such as...) Distance or When the distance is less than a certain threshold, convergence is considered achieved, and iteration stops.

[0049] Through these steps, this method can gradually reduce inconsistencies between marginal tables, ultimately generating a more consistent high-dimensional dataset.

[0050] Step 5: High-Dimensional Data Construction Stage: The Gradually Update Method (GUM) is employed. Based on the already consistent low-dimensional marginal distribution (including local and cross-domain attribute pairs), the distribution of the high-dimensional dataset is iteratively adjusted to gradually approximate the target marginal distribution, ultimately generating high-dimensional data that meets differential privacy requirements. First, the iteration counter is initialized. And set the attenuation rate At the same time, initialize the selected edge set. It is an empty set; when < Maximum number of iterations At that time, traverse all sets belonging to the marginal set. And not selected the edge : Computing high-dimensional datasets With the target margin Differences And based on the current number of records and according to Calculate the target number of records Dynamically calculate and update step size and reduce step size Based on the number of iterations Choose either "Replace" or "Copy" to proceed. If the norm of a high-dimensional dataset changes after the update, then... ( If the convergence threshold is reached, the current marginal update is terminated, and then the step size is adjusted and... join in Clear after traversal. And increasing Finally, return the high-dimensional dataset after the iteration is complete. .

[0051] Example 2 like Figure 5 As shown, the present invention provides a cross-domain collaborative privacy-preserving high-dimensional data construction system, comprising the following steps: Parameter setting and data preprocessing unit: The data collector confirms the number of participating data parties. Attribute sets of each data source Privacy Budget GUM iteration count and convergence threshold Data providers and initializing the user dataset , For a global attribute set, perform equal-width or equal-frequency binning on each data attribute to divide the attribute values ​​into... Each bin is divided into intervals; the binned intervals are then numerically or symbolically encoded and injected with Gaussian noise. The encoded dataset is obtained. , Allocated by the privacy budget; Local data processing unit: Each data provider independently calculates the InDif relevance metric locally. ,in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; allocating local privacy budgets to powers of 2 / 3 of the attribute domain size. , , For the size of the attribute field, For the total privacy budget; a low-dimensional boundary is selected using a greedy algorithm to construct an attribute association graph. And select eligible clusters to form a marginal set. ; Data upload and cross-party attribute estimation unit: Data parties collaboratively generate hash keys through secure multi-party computation. Each data provider adds virtual elements to the user set corresponding to the local attribute values, generates an FM-Sketch based on differential privacy, and records the maximum hash value and the least significant bit; the data collector calculates the cardinality of the union of attribute value sets based on the FM-Sketch of multiple data providers. Combined with the cardinality of single data attribute values , Through the principle of inclusion and exclusion After obtaining the intersection cardinality, the intersection cardinality is standardized to obtain the marginal distribution estimate of the cross-attribute pairs. ; Consistency processing unit: Privacy budget for shared attributes based on marginal table With contribution cell count Assign weights Calculate the weighted average probability of attribute values. , ( For the first Each marginal table pair The probability estimate is obtained; the initial probability estimate is projected onto the nearest effective probability distribution; the weighted average and projection operations are performed alternately until the marginal table is reached. or If the distance is less than the convergence threshold, convergence is considered achieved, and iteration stops. High-dimensional data building unit: Initializing the iteration counter Attenuation rate and the unselected boundary set; when Maximum number of iterations At that time, traverse the unselected edges Calculate high-dimensional datasets With the target margin The difference, dynamically calculate and update the step size Reduce step size , This is the current record number. The target number of records; through Update the high-dimensional dataset, if Then terminate the current marginal update. The convergence threshold; join in Clear after traversal. And increasing After iteration, a high-dimensional dataset is output. .

[0052] In the parameter setting and data preprocessing unit of this embodiment, the specific method of data binning is selected according to data distribution and privacy requirements: for continuous attributes, equal-frequency binning is preferentially used to divide them into discrete intervals, and for discrete attributes, equal-width binning is used to supplement the interval division; Gaussian noise. of The value is determined based on the privacy budget allocated to each data provider, ensuring that attribute boundary privacy is not leaked during the coding process.

[0053] In the local data processing unit of this embodiment, the core calculation and filtering logic for local data processing specifically includes the following steps: Step 2.1 Correlation measurement calculation: based on The measurement method is defined as follows: ;in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; Step 2.2, Optimal Privacy Budget Allocation: Allocate local privacy budget Distribute the attribute domain size to each boundary to minimize the overall noise error, using the formula: Total budget Allocate based on the size of the data source attributes; Step 2.3, Marginal Selection: A greedy algorithm is used to select attribute pairs to minimize the overall error; first, the selected marginal set is initialized. Iteration counting Calculate the initial total error Iterate through unselected 2D edges ,for Privacy budget allocation at various margins Calculate the margin Total error after From the noise error term With residual dependency error term Composition; selection to make Minimum margin If the total error does not decrease, the iteration will terminate; otherwise, it will... join in X And continue iterating until the selected marginal set is obtained. ; Step 2.4, Marginal Combination: First, combine the two-dimensional margins... Convert to corresponding attributes, and construct a graph based on the association relationships between attribute pairs. Then initialize the selected attribute set. Marginal set after combination From the size of the group arrive Traverse the graph sequentially. Each of the following sizes is The group ,like and The number of overlapping attributes does not exceed 2 and field size Then join in and will Add the attribute to , The number of attribute pairs is used to obtain the marginal set that preserves the joint distribution characteristics of high-dimensional attributes. .

[0054] In the description of this invention, "cross-domain collaboration" refers to a distributed scenario in which multiple data parties hold different attribute data of the same user group and achieve collaborative statistical processing without sharing the original data; "differential privacy (DP)" refers to ensuring that the probability difference in the algorithm output of adjacent datasets that differ by only one record is constrained by the privacy budget by introducing controllable random noise; "total variation distance (TVD)" refers to an index that measures the distribution difference between the original dataset and the constructed high-dimensional privacy-compatible dataset, and the smaller the value, the higher the distribution similarity.

[0055] To test the impact of different privacy budgets on the performance of our method (VertiSyn), and to compare it with existing cross-domain high-dimensional data processing methods, experiments were conducted using three high-dimensional datasets: Adult, BR2000, and NLTCS. The evaluation metric was l-way TVD (results are shown in the figure). Figure 3 (As shown). Experiments show that as the privacy budget increases, the l-way TVD of each method tends to improve. This method performs better in processing the BR2000 dataset with a large number of attributes and has better distribution preservation performance.

[0056] To test the usability of this method in downstream tasks, the Adult dataset was used as an example, and the evaluation metric was the SVM classification error rate (results are shown below). Figure 4 (As shown in the figure). Experiments show that the classification error rate of our method is lower than that of other comparative methods, and remains stable when the privacy budget changes, demonstrating that the constructed high-dimensional privacy-compatible data has good task adaptability.

[0057] Reference method logic flowchart ( Figure 2 The logical architecture of this invention includes five stages: data preprocessing, local processing, cross-domain attribute estimation, consistency processing, and high-dimensional data construction. Based on this invention, modifications or improvements can be made to its algorithm efficiency, scenario adaptability, etc., all of which fall within the scope of protection claimed by this invention.

[0058] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the scope of the invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0059] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other embodiments that can be understood by those skilled in the art. The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A cross-domain collaborative method for constructing privacy-preserving high-dimensional data, characterized in that, Includes the following steps: Step 1, Parameter Setting and Data Preprocessing Stage: The data collector confirms the number of participating data parties. Attribute sets of each data source Privacy Budget GUM iteration count and convergence threshold Data providers and initializing the user dataset , For a global attribute set, perform equal-width or equal-frequency binning on each data attribute to divide the attribute values ​​into... Each bin is divided into intervals; the binned intervals are then numerically or symbolically encoded and injected with Gaussian noise. The encoded dataset is obtained. , Allocated by the privacy budget; Step 2, Local Data Processing Stage: Each data provider independently calculates the InDif relevance metric locally. ,in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; allocating local privacy budgets to powers of 2 / 3 of the attribute domain size. , , For the size of the attribute field, For the total privacy budget; a low-dimensional boundary is selected using a greedy algorithm to construct an attribute association graph. And select eligible clusters to form a marginal set. ; Step 3, Data Upload and Cross-Party Attribute Estimation Stage: Each data party collaboratively generates hash keys through secure multi-party computation. Each data provider adds virtual elements to the user set corresponding to the local attribute values, generates an FM-Sketch based on differential privacy, and records the maximum hash value and the least significant bit; the data collector calculates the cardinality of the union of attribute value sets based on the FM-Sketch of multiple data providers. Combined with the cardinality of single data attribute values , Through the principle of inclusion and exclusion After obtaining the intersection cardinality, the intersection cardinality is standardized to obtain the marginal distribution estimate of the cross-attribute pairs. ; Step 4, Consistency Processing Phase: Apply privacy budget to shared attributes based on the marginal table. With contribution cell count Assign weights Calculate the weighted average probability of attribute values. , ( For the first Each marginal table pair The probability estimate is obtained; the initial probability estimate is projected onto the nearest effective probability distribution; the weighted average and projection operations are performed alternately until the marginal table is reached. or If the distance is less than the convergence threshold, convergence is considered achieved, and iteration stops. Step 5, High-Dimensional Data Construction Stage: Initialize the Iteration Counter Attenuation rate and the unselected boundary set; when Maximum number of iterations At that time, traverse the unselected edges Calculate high-dimensional datasets With the target margin The difference, dynamically calculate and update the step size Reduce step size , This is the current record number. The target number of records; through Update the high-dimensional dataset, if Then terminate the current marginal update. The convergence threshold; join in Clear after traversal. And increasing After iteration, a high-dimensional dataset is output. .

2. The method for constructing privacy-preserving high-dimensional data through cross-domain collaboration according to claim 1, characterized in that, In step 1, the specific method of data binning is selected based on data distribution and privacy requirements: for continuous attributes, equal-frequency binning is preferred to divide them into discrete intervals; for discrete attributes, equal-width binning is used to supplement the interval division; Gaussian noise. of The value is determined based on the privacy budget allocated to each data provider, ensuring that attribute boundary privacy is not leaked during the coding process.

3. The method for constructing privacy-preserving high-dimensional data through cross-domain collaboration according to claim 2, characterized in that, In step 2, the core calculation and filtering logic for local data processing specifically includes the following steps: Step 2.1 Correlation measurement calculation: based on The measurement method is defined as follows: ;in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; Step 2.2, Optimal Privacy Budget Allocation: Allocate local privacy budget Distribute the attribute domain size to each boundary to minimize the overall noise error, using the formula: Total budget Allocate based on the size of the data source attributes; Step 2.3, Marginal Selection: A greedy algorithm is used to select attribute pairs to minimize the overall error; first, the selected marginal set is initialized. Iteration counting Calculate the initial total error Iterate through unselected 2D edges ,for Privacy budget allocation at various margins Calculate the margin Total error after From the noise error term With residual dependency error term Composition; selection to make Minimum margin If the total error does not decrease, the iteration will terminate; otherwise, it will... join in X And continue iterating until the selected marginal set is obtained. ; Step 2.4, Marginal Combination: First, combine the two-dimensional margins... Convert to corresponding attributes, and construct a graph based on the association relationships between attribute pairs. Then initialize the selected attribute set. Marginal set after combination From the size of the group arrive Traverse the graph sequentially. Each of the following sizes is The group ,like and The number of overlapping attributes does not exceed 2 and field size Then join in and will Add the attribute to , The number of attribute pairs is used to obtain the marginal set that preserves the joint distribution characteristics of high-dimensional attributes. .

4. The method for constructing privacy-preserving high-dimensional data through cross-domain collaboration according to claim 3, characterized in that, Initial total error Dependency error for all unselected two-dimensional boundaries sum.

5. The method for constructing privacy-preserving high-dimensional data through cross-domain collaboration according to claim 3, characterized in that, Step 3, the FM-Sketch generation and marginal calculation of cross-sectoral attribute estimation specifically includes the following steps: Step 3.1, Hash Key Collaborative Generation: Each data party collaboratively generates a hash key through secure multi-party computation. This ensures the privacy of the hashing process; Step 3.2, Local FM-Sketch Generation: Data side Each attribute Its attribute value Corresponding user attribute set To each Add virtual elements and compute the corresponding FM-Sketch based on differential privacy; calculate the hash value of the values ​​in the user set using a hash function, update the maximum hash value in the FM-Sketch structure using this hash value, and maintain an array to track the least significant bit of the hash value to update the value at the corresponding position, finally obtaining the data party's FM-Sketch record set: , Size of the attribute domain; Step 3.3, Cross-Side Marginal Distribution Estimation: Cross-side marginal distribution estimation is based on data... Each FM-Sketch Implementation: First, initialize the marginal distribution estimation set of cross-attribute pairs. , then traverse and FM-Sketch: All attributes and their corresponding values , Then calculate the union FM-Sketch: = max( , And calculate the cardinality of the union. Combined with attribute values , cardinality , Through the principle of inclusion and exclusion Obtain the intersection cardinality, and use this intersection cardinality as the cross-attribute pair. , Marginal distribution counts are stored. Finally, for The counts of all attribute pairs are standardized to obtain the marginal distribution estimate of cross-domain attribute pairs. .

6. The method for constructing privacy-preserving high-dimensional data through cross-domain collaboration according to claim 5, characterized in that, Step 4, the weighted average and projection optimization for consistency processing specifically includes the following steps: Step 4.1, Weighted Average Integration: For shared attributes Shared attributes quilt Each marginal table contains a privacy budget for each marginal table. With contribution cell count Assign weights ,in, It is assigned to the marginal table Privacy budget, It is a marginal table Used for attributes The number of cells; then calculate the attributes. Each possible value Weighted average probability: ;in For the first Each marginal table pair The probability estimate; Step 4.2, Projecting to the efficient distribution: First, perform an initial estimation based on the current distribution, assuming a marginal table. The initial probability estimate is Next, the initial probability estimate will be projected onto the nearest effective probability distribution. This makes it the distance to the initial probability estimate. Minimum; this problem is described as an optimization problem: Its constraints are: ; Then, the gradient descent method is used to solve the optimization problem; Step 4.3, Iterative Convergence: To address the consistency issue among multiple marginal tables simultaneously, weighted averaging and projection operations are performed alternately until convergence. In each iteration, a weighted average operation is first performed on all shared attributes, and then each marginal table is projected onto an effective distribution. When the difference between the marginal tables is less than a set threshold, convergence is considered achieved, and the iteration stops.

7. The method for constructing privacy-preserving high-dimensional data through cross-domain collaboration according to claim 6, characterized in that, In step 5, the update operation rule for GUM iteration is as follows: First, initialize the iteration counter. And set the attenuation rate At the same time, initialize the selected edge set. It is an empty set; when < Maximum number of iterations At that time, traverse all sets belonging to the marginal set. And not selected the edge : Computing high-dimensional datasets With the target margin Differences And based on the current number of records and according to Calculate the target number of records Dynamically calculate and update step size and reduce step size Based on the number of iterations Choose either "Replace" or "Copy" to proceed. If the norm of a high-dimensional dataset changes after the update, then... Then terminate the current marginal update, and subsequently adjust the step size and... join in Clear after traversal. And increasing Finally, return the high-dimensional dataset after the iteration is complete. .

8. A cross-domain collaborative system for constructing privacy-preserving high-dimensional data, characterized in that, Includes the following steps: Parameter setting and data preprocessing unit: The data collector confirms the number of participating data parties. Attribute sets of each data source Privacy Budget GUM iteration count and convergence threshold Data providers and initializing the user dataset , For a global attribute set, perform equal-width or equal-frequency binning on each data attribute to divide the attribute values ​​into... Each bin is divided into intervals; the binned intervals are then numerically or symbolically encoded and injected with Gaussian noise. The encoded dataset is obtained. , Allocated by the privacy budget; Local data processing unit: Each data provider independently calculates the InDif relevance metric locally. ,in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; allocating local privacy budgets to powers of 2 / 3 of the attribute domain size. , , For the size of the attribute field, For the total privacy budget; a low-dimensional boundary is selected using a greedy algorithm to construct an attribute association graph. And select eligible clusters to form a marginal set. ; Data upload and cross-party attribute estimation unit: Data parties collaboratively generate hash keys through secure multi-party computation. Each data provider adds virtual elements to the user set corresponding to the local attribute values, generates an FM-Sketch based on differential privacy, and records the maximum hash value and the least significant bit; the data collector calculates the cardinality of the union of attribute value sets based on the FM-Sketch of multiple data providers. Combined with the cardinality of single data attribute values , Through the principle of inclusion and exclusion After obtaining the intersection cardinality, the intersection cardinality is standardized to obtain the marginal distribution estimate of the cross-attribute pairs. ; Consistency processing unit: Privacy budget for shared attributes based on marginal table With contribution cell count Assign weights Calculate the weighted average probability of attribute values. , ( For the first Each marginal table pair The probability estimate is obtained; the initial probability estimate is projected onto the nearest effective probability distribution; the weighted average and projection operations are performed alternately until the marginal table is reached. or If the distance is less than the convergence threshold, convergence is considered achieved, and iteration stops. High-dimensional data building unit: Initializing the iteration counter Attenuation rate and the unselected boundary set; when Maximum number of iterations At that time, traverse the unselected edges Calculate high-dimensional datasets With the target margin The difference, dynamically calculate and update the step size Reduce step size , This is the current record number. The target number of records; through Update the high-dimensional dataset, if Then terminate the current marginal update. The convergence threshold; join in Clear after traversal. And increasing After iteration, a high-dimensional dataset is output. .

9. A cross-domain collaborative privacy-preserving high-dimensional data construction system according to claim 8, characterized in that, In the parameter settings and data preprocessing unit, the specific method of data binning is selected based on data distribution and privacy requirements: for continuous attributes, equal-frequency binning is preferred to divide them into discrete intervals; for discrete attributes, equal-width binning is used to supplement the interval division; Gaussian noise. of The value is determined based on the privacy budget allocated to each data provider, ensuring that attribute boundary privacy is not leaked during the coding process.

10. A cross-domain collaborative privacy-preserving high-dimensional data construction system according to claim 9, characterized in that, In the local data processing unit, the core calculation and filtering logic for local data processing specifically includes the following steps: Step 2.1 Correlation measurement calculation: based on The measurement method is defined as follows: ;in, For attributes Two-dimensional marginal distribution, Assumption Marginal distribution when independent; Step 2.2, Optimal Privacy Budget Allocation: Allocate local privacy budget Distribute the attribute domain size to each boundary to minimize the overall noise error, using the formula: Total budget Allocate based on the size of the data source attributes; Step 2.3, Marginal Selection: A greedy algorithm is used to select attribute pairs to minimize the overall error; first, the selected marginal set is initialized. Iteration counting Calculate the initial total error Iterate through unselected 2D edges ,for Privacy budget allocation at various margins Calculate the margin Total error after From the noise error term With residual dependency error term Composition; selection to make Minimum margin If the total error does not decrease, the iteration will terminate; otherwise, it will... join in X And continue iterating until the selected marginal set is obtained. ; Step 2.4, Marginal Combination: First, combine the two-dimensional margins... Convert to corresponding attributes, and construct a graph based on the association relationships between attribute pairs. Then initialize the selected attribute set. Marginal set after combination From the size of the group arrive Traverse the graph sequentially. Each of the following sizes is The group ,like and The number of overlapping attributes does not exceed 2 and field size Then join in and will Add the attribute to , The number of attribute pairs is used to obtain the marginal set that preserves the joint distribution characteristics of high-dimensional attributes. .