Federal learning-based privacy data calculation method and system
By injecting dynamic differential privacy noise and encryption technology into federated learning, combining zero-knowledge proof and dual-channel encrypted transmission, the problem of difficult to balance privacy protection and model accuracy in cross-domain data computing is solved, and efficient and secure data computing and privacy protection are achieved.
Patent Information
- Application Number
- CN202510391711.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art fails to fully consider privacy protection when processing cross-domain data, resulting in an increase in the risk of privacy leakage and is difficult to achieve efficient and secure neighbor search and binning interval calculations, and cannot balance privacy protection and model accuracy.
The privacy data calculation method based on federated learning is adopted, and the local data is preprocessed and dense feature alignment is aligned, the federated mode is selected, dynamic differential privacy noise is injected, and the gradient is encrypted, so as to achieve the aggregation and verification of security gradients, dynamically adjust the privacy budget and noise intensity, and ensure data security by using dual-channel encrypted transmission and zero-knowledge proof.
It effectively improves the accuracy and completeness of data, realizes secure computing of cross-domain data, prevents original data leakage, balances privacy protection and model accuracy, and ensures the security and consistency of data during the calculation process.
Smart Images

Figure CN120162828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of privacy data computing, and particularly to a privacy data computing method and system based on federated learning. Background Art
[0002] In the prior art, when standardizing numerical features, filling missing values, and processing high-cardinality features, conventional methods do not fully consider cross-domain data privacy. The ordinary method of filling missing values with the mean does not utilize a secure collaboration mechanism and cannot ensure filling accuracy while protecting data privacy. In scenarios of cross-institutional data integration, for example, when financial institutions and healthcare institutions cooperate in data analysis, simple data processing methods are prone to increasing the risk of privacy leakage. Traditional methods are difficult to achieve efficient and secure neighbor search and bin interval calculation. Conventional practices cannot guarantee the security of data during the calculation process through secure multi-party computing (MPC) technology, resulting in sacrificing privacy to obtain calculation results in practical applications. Due to excessive protection of privacy, multi-source data cannot be effectively utilized to improve data quality. Existing feature alignment technologies are difficult to ensure the security of feature consistency verification in multi-party collaborative scenarios. Traditional federated mode selection methods often rely on simple empirical judgments and cannot perform intelligent and accurate mode selection based on the actual overlap of features and sample IDs among participating parties. In terms of gradient processing, it is difficult for the prior art to balance privacy protection and model accuracy. Existing federated transfer learning technologies are relatively single in dealing with the feature distribution differences between the source domain and the target domain and lack the ability of dynamic adjustment. Traditional privacy protection mechanisms cannot dynamically adjust privacy budgets and noise intensities according to changes in data sensitivity; therefore, there is a need to provide a privacy data computing method and system based on federated learning. Summary of the Invention
[0003] The purpose of the present invention is to provide a privacy data computing method and system based on federated learning. To solve the above-mentioned prior art problems, the present invention is realized through the following technical solutions:
[0004] In the first aspect, a privacy data computing method based on federated learning includes the following steps:
[0005] Preprocess local data, perform encrypted feature alignment on the preprocessed data, select a federated mode, inject dynamic differential privacy noise based on the local model gradient, and encrypt the gradient with injected noise to obtain a secure gradient;
[0006] Aggregate the obtained secure gradients, verify the data consistency, establish a federated transfer learning component to minimize the feature distribution differences between the source domain and the target domain, generate a final prediction label through multi-party collaborative decision-making to establish a secure boosting tree model, and perform zero-knowledge audit and evidence storage;
[0007] Dynamically enhance privacy protection, classify the sensitivity of local data, verify and adjust the intensity of differential privacy noise, optimize privacy security using dual-channel encrypted transmission, and strengthen the federated transfer learning component.
[0008] In a second aspect, embodiments of the present invention provide a privacy data computing system based on federated learning, including the following steps:
[0009] Data storage module: Support accessing local structured data from multiple data sources;
[0010] Data preprocessing module: Preprocess local data, perform encrypted feature alignment on the preprocessed data, and select the federated mode;
[0011] Gradient encryption module: Inject differential privacy noise dynamically based on the local model gradient, and encrypt the gradient with injected noise to obtain a secure gradient;
[0012] Federated learning component construction module: Aggregate the obtained secure gradients, verify the data consistency, establish a federated transfer learning component to minimize the feature distribution difference between the source domain and the target domain, and generate the final prediction label by establishing a secure boosting tree model through multi-party collaborative decision-making; Zero-knowledge audit and evidence storage module: Select an appropriate zero-knowledge proof algorithm according to different business scenarios and security requirements, and select an appropriate blockchain platform for deployment;
[0013] Privacy protection dynamic enhancement module: Dynamically enhance privacy protection, classify the sensitivity of local data, verify and adjust the intensity of differential privacy noise, optimize privacy security using dual-channel encrypted transmission, and strengthen the federated transfer learning component.
[0014] Advantages of the present invention:
[0015] 1. Standardize numerical features, fill in missing values with KNN interpolation, and perform binning on high-cardinality features to improve data accuracy and integrity. Use secure multi-party computation (MPC) to achieve cross-domain neighbor search and collaborative binning interval calculation, preventing the leakage of original data. Utilize the consortium chain architecture combined with zero-knowledge proof (ZKP) to achieve feature alignment among multiple parties, verify feature consistency, use independent channels according to different business scenarios to avoid cross-leakage of data; connect to the Ethereum private chain through the IBC protocol, support the alignment of heterogeneous data sources, apply zero-knowledge proof (ZKP), select an appropriate federated mode, inject dynamic differential privacy noise into the local model gradient, balance privacy protection and model accuracy by controlling the privacy budget, and aggregate secure gradients; verify data consistency to ensure the reliability and stability of data during model training, minimize the feature distribution difference between the source domain and the target domain, generate zero-knowledge proofs and store them on the blockchain for evidence preservation, ensuring the authenticity and traceability of audit information, facilitating audits by regulatory agencies and authorized parties without revealing sensitive data;
[0016] 2. Dynamically adjust the privacy budget according to the data sensitivity level, increase the noise injection intensity for highly sensitive data, and the central node verifies the loss of model accuracy after noise injection. If the loss exceeds the threshold, the privacy budget value is proportionally reduced and retraining is performed. Use the homomorphic encryption algorithm to encrypt the gradients after noise injection to prevent data leakage. By comparing the consistency of the results of the two channels by the central node, effectively detect whether the data has been tampered with, providing an additional security protection mechanism for data transmission and processing. Perform secondary hashing on non-shared features and enhance them with salt values. The smart contract negotiates the global salt value and automatically updates it after each round of alignment. Call the pre-compiled ZKP verification function through the smart contract to ensure that the hash calculation conforms to the preset rules, verify that the submitted salt-enhanced hash value is consistent with the local original data, and the salt value or timestamp has not been tampered with. After verification, the result is uploaded to the chain, otherwise, trigger the node reputation penalty mechanism to enhance the trust between nodes and the credibility of data. Dynamically adjust the migration weight according to the similarity between the source domain and the target domain, enabling federated transfer learning to better adapt to the relationship between the source domain and the target domain in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 is the flowchart of the steps of the privacy data calculation method based on federated learning provided in Embodiment 1 of the present invention;
[0019] Figure 2 It is a schematic structural diagram of a privacy data calculation system based on federated learning provided in Embodiment 2 of the present invention. Detailed implementation manners
[0020] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] Embodiment 1
[0022] As Figure 1 shown, the privacy data calculation method based on federated learning provided in the embodiment of the present invention specifically includes the following steps:
[0023] Step 1: Preprocess the local data, perform encrypted feature alignment on the preprocessed data, perform joint modeling based on the preprocessed data, select the federated mode, inject dynamic differential privacy noise based on the local model gradient, and encrypt the injected noise gradient to obtain a secure gradient;
[0024] In Step 1:
[0025] Specifically, the specific process of preprocessing the local data and performing encrypted feature alignment on the preprocessed data is as follows:
[0026] Obtain the original data table D = {x1, x2, x3,..., x n} of the local data. The original data includes structured data such as but not limited to bank user transaction records and hospital electronic medical records;
[0027] Perform preprocessing on the obtained original data table to obtain a preprocessed data feature set;
[0028] It should be noted that the preprocessed data feature set includes but is not limited to: a numerical preprocessing feature set, a missing value preprocessing data feature set, and a high-cardinality preprocessing data feature set
[0029] Specifically, perform standardization processing on numerical features to obtain a numerical preprocessing feature set, for example: transaction amount and age;
[0030] Fill in the missing values of the original data table using KNN interpolation to obtain a missing value preprocessing data feature set, and perform binning processing on the high-cardinality features of the original data table to obtain a high-cardinality preprocessing data feature set;
[0031] It should be noted that KNN interpolation represents a method for filling missing values based on neighborhood similarity. By securely collaborating, the k nearest neighbors most similar to the missing sample are found, and the feature values of these neighbors are used to infer the missing values. Bin processing represents a key step in discretizing continuous features into multiple intervals, improving the robustness of the model, enhancing interpretability, and ensuring cross-domain data distribution alignment.
[0032] Specifically, based on the cosine similarity between the original data, k nearest neighbors most similar to the missing sample are found in the local data. In the federated scenario, cross-domain neighbor search is achieved through secure multi-party computation (MPC) to obtain the feature set of preprocessed missing value data.
[0033] Using quantile binning for processing to ensure that the feature distributions of different participants are within the same intervals. Secure multi-party computation (MPC) is used to collaboratively compute the bin intervals in the encrypted domain to obtain the feature set of preprocessed high-cardinality data. Exemplarily, for the income type feature x i Using quantile binning for processing the income type feature x i Divided into a intervals, bin interval = Quantile(x i , bins = a);
[0034] It should be noted that secure multi-party computation (MPC) represents the core technology for realizing cross-domain data collaborative computing. Its core goal is to enable multiple participants to jointly complete complex computing tasks without revealing the original data.
[0035] Based on the obtained feature set of preprocessed data, adopting the Hyperledger Fabric consortium blockchain architecture, multi-party feature alignment is achieved through smart contracts, and zero-knowledge proof (ZKP) is combined to verify feature consistency.
[0036] It should be noted that the Hyperledger Fabric consortium blockchain architecture represents an enterprise-level distributed ledger platform designed specifically for consortium blockchains, with high modularity and scalability, suitable for various complex business scenarios.
[0037] Specifically, asymmetric encryption hashing is performed on the feature name and value. Through the formula H(F j ) = SHA_3_256(F j || salt p ) to obtain the feature hash value H(F j ), where F j represents the data feature, and salt p represents the unique random salt value of participant p to prevent rainbow table attacks.
[0038] Construct data nodes based on the calculated feature hash values, and the local data holder stores the preprocessed feature hash value H(F j ). The local data holders include, but are not limited to: financial institutions and medical institutions; generate transaction sorting and establish sorting nodes, supporting multi-channel isolation, including but not limited to: financial channels and medical channels;
[0039] Deploy feature alignment logic and zero-knowledge proof ZKP verification rules to establish smart contracts;
[0040] It should be noted that zero-knowledge proof ZKP represents a cryptographic technique that allows a prover to prove to a verifier that a certain statement is true without revealing any additional information other than the fact that the statement is true;
[0041] Specifically, use the zero-knowledge proof ZKP verification rules for verification: zk-SNARKs (Groth16) is applicable to high-frequency and low-latency scenarios, such as financial transactions, with a generated proof volume of approximately 200 bytes and a verification time < 30ms;
[0042] zk-STARKs is resistant to quantum attacks and is used in high-security scenarios, such as medical scenarios, with a proof volume of 100KB and a verification time < 200ms;
[0043] Based on the established data nodes, sorting nodes and smart contracts, establish a Hyperledger Fabric consortium chain architecture;
[0044] Based on the calculated feature hash values, compare the feature hash values with a preset sensitive hash threshold. If the feature hash value is greater than or equal to the preset sensitive hash threshold, the data node corresponding to the feature hash value is marked as a sensitive data node, and the sensitive feature hash pair corresponding to the sensitive data node is visible to the authorized node;
[0045] Use independent channels based on different business scenarios to avoid data cross-leakage;
[0046] Connect to the Ethereum private chain through the IBC protocol to support heterogeneous data source alignment;
[0047] It should be noted that the Ethereum private chain represents a blockchain network based on Ethereum blockchain technology that focuses on protecting user privacy;
[0048] Second specifically, perform joint modeling based on the preprocessed data. The specific process of selecting the federated mode is as follows:
[0049] If there is a lot of feature overlap and little sample ID overlap among the participants, then select the vertical federated mode and use the private set intersection PSI technology to determine the common user IDs;
[0050] If there is a large overlap in sample IDs and a small overlap in features among the participating parties, the horizontal federated mode is selected, and the obtained aligned feature set is directly used;
[0051] Specifically, the specific process of injecting dynamic differential privacy noise based on the local model gradient and encrypting the gradient with injected noise to obtain the secure gradient is as follows:
[0052] Inject Laplace noise into the local model gradient g i where Δf is the gradient sensitivity, representing the maximum gradient change, and Δf is calculated by the formula where g and g i and g i ' are local model gradients calculated at different times or for different samples, ∈ is the privacy budget, controlling the noise intensity, and its value range is between (0, 1];
[0053] Exemplarily, assuming N = 101×103 = 10403, and the gradient g i after injecting noise is 0.6, then through the formula [[g i = 0.6×10403 mod 10403 2 the encrypted secure gradient [[g i is obtained;
[0054] Step 2: Aggregate the obtained secure gradients and verify the data consistency, establish a federated transfer learning component to adapt to minimize the feature distribution difference between the source domain and the target domain, generate the final prediction label through multi-party collaborative decision-making to establish a secure boosting tree model, and conduct zero-knowledge audit and evidence preservation;
[0055] In Step 2:
[0056] Specifically, the specific process of aggregating the obtained secure gradients and verifying the data consistency is as follows:
[0057] The weights w i of the participating parties are dynamically allocated according to the data quality, and the weights w of the participating parties are obtained through the formula i , where YB i represents the sample size, ZL i represents the data quality coefficient, which is comprehensively evaluated by the data integrity W zx and the missing rate Q sl indicators;
[0058] The central node calculates the aggregated gradient, and the aggregated gradient G is obtained through the formula , where w i represents the weights of the participating parties, and g iIt represents the local model gradient, and k represents the total number of nodes;
[0059] The improved Practical Byzantine Fault Tolerance PBFT protocol is adopted, and the number of nodes for verifying data consistency is greater than or equal to times the total number of nodes k;
[0060] Through the verification formula the data consistency Y is obtained ZX , where λ is the preset gradient fluctuation threshold, with a value of 0.1, and k represents the total number of nodes;
[0061] Exemplarily, the federated learning system includes 3 participating parties, i.e., k = 3, and the parameters of the participating parties are shown in Table 1:
[0062] Participant <![CDATA[Sample size YB i > <![CDATA[Data integrity W zx > <![CDATA[Missing rate Q sl > <![CDATA[Local model gradient g i > Node A 500 0.95 0.05 0.32 Node B 300 0.8 0.2 0.28 Node C 200 0.6 0.4 0.35
[0063] Table 1 Statistical Table of Participating Party Data
[0064] Based on the parameter data of the participating parties in Table 1, calculate the data quality coefficient ZL of the participating parties i :
[0065] Calculate the weight w i :
[0066]
[0067] Calculate the aggregated gradient G:
[0068]
[0069]
[0070] Calculate the threshold: λ * |G| = 0.1 × |0.313| = 0.0313;
[0071] Calculate the gradient deviation of each node: Gradient deviation of node A: |0.32 - 0.313| = 0.007, Gradient deviation of node A: |0.28 - 0.313| = 0.033, Gradient deviation of node A: |0.35 - 0.313| = 0.037;
[0072] The deviation of node B, 0.033 > 0.0313, and the deviation of node C, 0.037 > 0.0313, both exceed the threshold;
[0073] According to the PBFT rules, nodes need to reach an agreement. Node A and node B pass the verification, node C is excluded, and the aggregated gradient is recalculated, only using the local model gradients of node A and node B;
[0074] Specifically, the specific process of establishing a federated transfer learning component to minimize the feature distribution difference between the source domain and the target domain, generating the final prediction label through multi-party collaborative decision-making to establish a secure boosting tree model, and performing zero-knowledge audit and evidence storage is as follows:
[0075] Adopt a method that combines the maximum mean discrepancy (MMD) and the KL divergence to minimize the feature distribution difference between the source domain D S and the target domain D T . The objective function L adapt = MMD(D S , D T ) + KL(P S , P T ), where MMD(D S , D T ) represents minimizing the distance between the source domain D S and the target domain D T in the reproducing kernel Hilbert space (RKHS). Here, φ(x) represents the feature mapping function, and KL is the KL divergence;
[0076] It should be noted that the reproducing kernel Hilbert space (RKHS) represents a Hilbert space composed of functions. A Hilbert space is a complete inner product space. The core property of the reproducing kernel Hilbert space is that there exists a kernel function K such that the value of any function f in the space at point x can be expressed in the form of an inner product;
[0077] Each participating party generates a zero-knowledge proof for its own operations and data usage during the entire process;
[0078] Exemplarily, the participating party needs to prove that it complies with the protocol regulations during the gradient calculation, model training, and data provision processes, without leaking sensitive information. The zero-knowledge proof is generated using the ZK-SNARK proof system;
[0079] Store the generated zero-knowledge proof and related audit information on the blockchain for evidence storage. For example, the operation time and the identity of the participating party. The non-tamperable feature of the blockchain ensures the authenticity and traceability of the audit information. The regulatory agency and other authorized parties audit the behavior of the participating party by verifying the zero-knowledge proof without obtaining specific sensitive data;
[0080] The technical solution of the embodiment of the present invention is as follows: perform standardization on numerical features, KNN interpolation filling for missing values, and binning processing on high-cardinality features to improve the accuracy and integrity of data. During the preprocessing process, use secure multi-party computing (MPC) to implement cross-domain neighbor search and binning interval collaborative calculation, ensuring the security of cross-domain data during the calculation process and preventing the leakage of original data. Utilize the Hyperledger Fabric consortium chain architecture combined with zero-knowledge proof (ZKP), and through asymmetric encryption hashing of feature names and values, construct data nodes, sorting nodes, and smart contracts to achieve feature alignment among multiple parties, verify feature consistency, and ensure data accuracy and consistency. Use independent channels according to different business scenarios to avoid data cross-leakage; connect to the Ethereum private chain through the IBC protocol to support heterogeneous data source alignment. The application of zero-knowledge proof (ZKP) allows proving the authenticity and compliance of data without revealing sensitive information. According to the overlap situation of features and sample IDs among participating parties, select an appropriate federated mode, inject dynamic differential privacy noise into the local model gradient, balance privacy protection and model accuracy by controlling the privacy budget, aggregate secure gradients, and dynamically allocate participant weights according to data quality to improve the effect of model training; use an improved practical Byzantine fault tolerance (PBFT) protocol to verify data consistency, ensuring the reliability and stability of data during the model training process. Use a method combining maximum mean discrepancy (MMD) and KL divergence to minimize the feature distribution difference between the source domain and the target domain, which helps improve the model's migration ability between different data domains. Each participant generates a zero-knowledge proof and stores it on the blockchain for evidence retention. Utilize the immutable feature of the blockchain to ensure the authenticity and traceability of audit information, facilitating auditing by regulatory agencies and authorized parties without revealing sensitive data;
[0081] Embodiment 2
[0082] As Figure 1 shown, the privacy data calculation method based on federated learning provided by the embodiment of the present invention specifically includes the following steps:
[0083] Step 3: Dynamically enhance privacy protection, classify the sensitivity of local data, verify and adjust the intensity of dynamic differential privacy noise, optimize privacy security using dual-channel encryption transmission, and strengthen the federated transfer learning component;
[0084] In Step 3:
[0085] Specifically, the specific process of dynamically enhancing privacy protection and classifying the sensitivity of local data is as follows:
[0086] Dynamically adjust the privacy budget according to the data sensitivity level S ∈ {1, 2, 3} (level 1 is the lowest, level 3 is the highest), through the formula where ∈base is the basic privacy budget value. For example, since medical data involves patient privacy, the sensitivity level is usually set to 3. If ∈ base is set to 0.9, then at this time Accordingly, the injected noise intensity will increase to better protect data privacy;
[0087] The central node verifies the model accuracy loss after noise injection. If the loss exceeds the loss threshold δ, then the ∈ value is reduced proportionally and retrained;
[0088] Use the Paillier homomorphic encryption algorithm to encrypt the gradient g i after noise injection; through the formula [[g i = g i * N mod N 2 is obtained, where N = p * q (p and q are large prime numbers, which are public key parameters);
[0089] The central node compares the consistency of the dual-channel results. If the deviation exceeds the deviation threshold, it is determined as a tampering attack;
[0090] It should be noted that the Paillier homomorphic encryption algorithm has additive homomorphicity, that is, [[g i + g j = [[g i * [[g j , enabling some calculation operations to be performed on the encrypted gradient without decryption, ensuring the security of data during the calculation process;
[0091] Second specifically, verify and adjust the dynamic differential privacy noise intensity. The specific process of optimizing privacy security using dual-channel encrypted transmission is as follows:
[0092] Perform secondary hashing on non-shared features to prevent replay attacks
[0093] Perform salt value enhancement on the feature name and value, and perform secondary hashing on non-shared features. Through the formula: H'(F j ) = SHA_3_256(F j || salt p || timestamp) to obtain the salt value enhanced hash value H'(F j ), negotiate the global salt value salt global through the smart contract, and automatically update it after each round of alignment;
[0094] The participating party submits the salt value enhanced hash value H'(F j ) to the Fabric channel, triggering the smart contract to execute the alignment logic;
[0095] The participant generates a proof π, stating that: the submitted salt-enhanced hash value H'(F j ) is consistent with the local original data, and the salt value or timestamp has not been tampered with;
[0096] Define a constraint system using the Circom language to ensure that the hash calculation complies with the preset rules;
[0097] It should be noted that the Circom language represents a domain-specific language DSL, which is specifically designed and developed for zero-knowledge proof ZKP circuits;
[0098] The smart contract calls the pre-compiled ZKP verification function Verify(π, H'(F j )) After successful verification, the alignment result and the proof hash are uploaded to the chain, otherwise the node reputation penalty mechanism is triggered;
[0099] Thirdly, specifically, the specific process of strengthening the federated transfer learning component is as follows:
[0100] Based on the obtained objective function L adapt , according to the similarity sim(D S , D T ) between the source domain and the target domain, dynamically adjust the transfer weight β, through the formula L adapt = MMD(D S , D T ) + β * KL(P S , P T ), where β represents a preset balance coefficient, and max(KL) represents the maximum KL divergence;
[0101] It should be noted that through the above method of dynamically adjusting the transfer weight, the differences in data characteristics between the source domain and the target domain in different scenarios can be flexibly adapted. When the similarity between the source domain and the target domain is relatively high, the transfer weight is appropriately increased, so that the target domain can make full use of the rich knowledge and information in the source domain, accelerate the convergence speed of the model, and improve the generalization performance of the model in the target domain; while when the similarity between the two is relatively low, the transfer weight is reduced to avoid the interference of the noisy knowledge in the source domain on the learning of the target domain model, and ensure that the model focuses on the characteristics and laws of the target domain data itself;
[0102] Based on the dynamic enhancement of privacy protection, build a comprehensive and dynamic privacy protection and federated transfer learning collaborative system.
[0103] The technical solution of the embodiment of the present invention is as follows: Dynamically adjust the privacy budget according to the data sensitivity level, increase the noise injection intensity for highly sensitive data, while protecting data privacy, avoid excessive protection resulting in too large a loss of model accuracy. The central node verifies the loss of model accuracy after noise injection. If the loss exceeds the threshold, the privacy budget value is reduced proportionally and retraining is performed to ensure that the model can still maintain high performance under the premise of privacy protection. The Paillier homomorphic encryption algorithm is used to encrypt the gradients after injecting noise. Using its additive homomorphic property, calculations can be performed on encrypted data, ensuring the security of data during the calculation process and preventing data leakage. By comparing the consistency of the dual-channel results by the central node, it effectively detects whether the data has been tampered with, providing an additional security protection mechanism for data transmission and processing. The non-shared features are secondarily hashed and enhanced in combination with a salt value to prevent replay attacks. The smart contract negotiates the global salt value and automatically updates it after each round of alignment. The Circom language is used to define the constraint system, and the pre-compiled ZKP verification function is called through the smart contract to ensure that the hash calculation conforms to the preset rules, verifying that the submitted salt value-enhanced hash value is consistent with the local original data, and the salt value or timestamp has not been tampered with. After verification, the result is uploaded to the chain, otherwise the node reputation penalty mechanism is triggered to enhance the trust between nodes and the credibility of data. Dynamically adjust the migration weight according to the similarity between the source domain and the target domain, enabling federated transfer learning to better adapt to the relationship between the source domain and the target domain in different scenarios;
[0104] Embodiment 3
[0105] As Figure 2 shown, the privacy data calculation system based on federated learning provided by the embodiment of the present invention specifically includes the following modules:
[0106] Data storage module: Supports accessing local structured data from multiple data sources, including but not limited to: bank user transaction records and hospital electronic medical records. It has a data format conversion function, which can uniformly convert data in different formats into a format that can be processed within the system. A distributed file system or object storage is used to store the original data to ensure the high availability and scalability of the data;
[0107] Data preprocessing module: Preprocesses the local data and performs encrypted feature alignment on the preprocessed data, and selects the federated mode
[0108] Gradient encryption module: Based on the local model gradient, injects differential privacy noise and encrypts the gradient after injecting noise to obtain a secure gradient;
[0109] Federated learning component construction module: Aggregates the obtained secure gradients and verifies the consistency of the data, establishes a federated transfer learning component to adapt to minimizing the feature distribution difference between the source domain and the target domain, and generates the final prediction label by establishing a secure boosting tree model through multi-party collaborative decision-making;
[0110] Zero - knowledge audit and evidence - storage module: Select appropriate zero - knowledge proof algorithms according to different business scenarios and security requirements, and select appropriate blockchain platforms for deployment;
[0111] Privacy - protection dynamic enhancement module: Dynamically enhance privacy protection, classify the sensitivity of local data, verify and adjust the intensity of dynamic differential privacy noise, optimize privacy security using dual - channel encrypted transmission, and strengthen the federated transfer learning component.
[0112] The above has described an embodiment of the present invention in detail, but the content described is only the preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention; the above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation and historical experience and can be adjusted according to the actual situation; the above is only the preferred embodiment of the present invention and does not limit the present invention. Any equivalent changes and improvements made according to the scope of the present invention application should still fall within the scope covered by the patent of the present invention.
Claims
1. A method for calculating private data based on federated learning, characterized in that: The following steps are involved: Preprocess the local data and perform dense feature alignment on the preprocessed data. Select the federation mode, inject dynamic differential privacy noise based on the local model gradient, and encrypt the gradient of the injected noise to obtain a secure gradient. Aggregate the obtained security gradients and verify the consistency of the data. Establish a federated transfer learning component to minimize the difference in feature distribution between the source domain and the target domain. Establish a security boosting tree model through multi-party collaborative decision-making to generate the final prediction label, and perform zero-knowledge audit and storage. Dynamically enhance privacy protection, grade the sensitivity of local data, verify and adjust the dynamic differential privacy noise intensity, optimize privacy security using dual-channel encrypted transmission, and strengthen federated transfer learning components.
2. The method for calculating private data based on federated learning according to claim 1, characterized in that: The specific process of preprocessing local data is as follows: Get the original data table D of local data = {x1,x2,x3,...,x n }, standardize the numerical features to obtain the numerical preprocessing feature set, use KNN interpolation to fill the missing values of the original data table to obtain the missing value preprocessing data feature set, and bin the high cardinality features of the original data table to obtain the high cardinality preprocessing data feature set; Based on the obtained pre-processed data feature set, the Hyperledger Fabric consortium chain architecture is adopted to achieve multi-party feature alignment through smart contracts, and the feature consistency is verified in combination with zero-knowledge proof ZKP.
3. The method for calculating private data based on federated learning according to claim 1, characterized in that: The specific process of performing dense feature alignment is as follows: Perform asymmetric encryption hashing on the feature name and value, and obtain the feature hash value through the formula; The data node is constructed based on the calculated feature hash value, and the local data holder stores the preprocessed feature hash value; Deploy feature alignment logic and zero-knowledge proof ZKP verification rules to establish smart contracts; Use zero-knowledge proof ZKP verification rules for verification; Use independent channels based on different business scenarios to avoid cross-data leakage; Connect to the Ethereum privacy chain through the IBC protocol to support alignment of heterogeneous data sources.
4. The method for calculating private data based on federated learning according to claim 1, characterized in that: The specific process of obtaining the safety gradient is as follows: In the local model gradient g i Inject Laplace noise into Where Δf is the gradient sensitivity, Δf is calculated by the formula Calculated, where g i and g i ' is the local model gradient calculated at different times or different samples, ∈ is the privacy budget, which controls the noise intensity and has a value range of (0,1].
5. The method for calculating private data based on federated learning according to claim 1, characterized in that: The specific process of the security gradient aggregation and verification of data consistency is as follows: The weight of the participants is dynamically allocated according to the data quality, and the weight of the participants is obtained by the formula w i ; The central node calculates the aggregate gradient and obtains the aggregate gradient through the formula; By verifying the formula Get data consistency Y ZX , where λ is the preset gradient fluctuation threshold, which is set to 0.
1.
6. The method for calculating private data based on federated learning according to claim 1, characterized in that: The process of minimizing the difference in feature distribution between the source domain and the target domain is: The maximum mean difference (MMD) and KL divergence are combined to minimize the source domain D S With the target domain D T The characteristic distribution difference of the objective function L adapt =MMD(D S ,D T )+KL(P S ,P T ), where MMD(D S ,D T ) represents the minimization of the source domain D S With the target domain D T The distance in the reproducing kernel Hilbert space RKHS, where φ(x) represents the feature mapping function, KL(P S ,P T ) is the KL divergence.
7. The method for calculating private data based on federated learning according to claim 1, characterized in that: The process of dynamically enhancing privacy protection is as follows: The gradient g after injecting noise is encrypted using homomorphic encryption algorithm. i Encryption; through the formula [[g i ]]=g i *N mod N 2 It is obtained that n=p*q, p and q are large prime numbers and are public key parameters; the central node compares the consistency of the dual-channel results, and if the deviation exceeds the deviation threshold, it is determined to be a tampering attack.
8. The method for calculating private data based on federated learning according to claim 1, characterized in that: The process of verifying and adjusting the dynamic differential privacy noise intensity is as follows: The feature names and values are salted and the non-shared features are hashed twice. The formula is: H'(F j )=SHA_3_256(F j ||salt p ||timestamp) to obtain the salt value to enhance the hash value H'(F j ), negotiate the global salt value salt through smart contracts global , automatically updated after each round of alignment; The smart contract calls the precompiled ZKP verification function Verify(π,H'(F j )), after verification, the alignment result and the proof hash are uploaded to the chain, otherwise the node reputation penalty mechanism is triggered.
9. The method for calculating private data based on federated learning according to claim 1, characterized in that: The process of strengthening the federated transfer learning component is: The migration weight β is dynamically adjusted according to the similarity between the source domain and the target domain, and the formula L adapt =MMD(D S ,D T )+β*KL(P S ,P T ) is used to obtain the enhanced objective function, where β represents the preset balance coefficient.
10. A private data computing system based on federated learning, characterized in that: include: Data storage module: supports access to local structured data from multiple data sources; Data preprocessing module: preprocess local data and perform dense feature alignment on the preprocessed data, and select the federation mode; Gradient encryption module: Injects dynamic differential privacy noise based on the local model gradient, and encrypts the gradient injected with noise to obtain a secure gradient; Federated learning component building module: Aggregate the obtained security gradients and verify the consistency of the data, establish a federated transfer learning component to minimize the feature distribution difference between the source domain and the target domain, and establish a security boosting tree model through multi-party collaborative decision-making to generate the final prediction label; Zero-knowledge audit and evidence storage module: select the appropriate zero-knowledge proof algorithm and the appropriate blockchain platform for deployment according to different business scenarios and security requirements; Privacy protection dynamic enhancement module: Dynamically enhances privacy protection, grades the sensitivity of local data, verifies and adjusts the intensity of dynamic differential privacy noise, optimizes privacy security using dual-channel encrypted transmission, and strengthens the federated transfer learning component.
Citation Information
Patent Citations
Physical medical data fusion privacy protection method based on cloud and mist architecture longitudinal federal learning
CN116595584A
Method and device for realizing data privacy protection processing based on federal model training, processor and computer readable storage medium thereof
CN118036067A
Multi-strategy federated learning method suitable for defending DLG attack
CN119005248A
Medical data privacy protection method and system based on artificial intelligence
CN119577841A
Cited By
E-commerce platform multi-source data security fusion method and system based on federated learning
CN120930169A
Cross-device biological characteristic multi-mimicry learning system
CN120977023A
Federal transfer learning driven multi-scene communication parameter optimization system and method
CN121037870A
Public facility abnormal data processing and analysis method based on space-time diagram neural network
CN121580247A
A public facility abnormal data processing and analysis method based on a space-time graph neural network
CN121580247B