Privacy data computing method and system based on federated learning
By employing a federated learning approach, local data is preprocessed and dense-state feature aligned, dynamic differential privacy noise is injected and gradients are encrypted, secure gradients are aggregated and data consistency is verified, a federated transfer learning component is established, and dual-channel encrypted transmission and zero-knowledge audit evidence storage are used to solve the privacy leakage and data consistency problems in cross-institutional data integration, achieving secure and efficient data computation and model training.
Patent Information
- Application Number
- CN202510391711.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Existing technologies fail to effectively protect data privacy in cross-institutional data integration scenarios, leading to an increased risk of privacy leakage. Furthermore, they are difficult to achieve efficient and secure neighbor search and binning interval computation. Feature alignment technology lacks security in multi-party collaborative scenarios. It is difficult to balance privacy protection and model accuracy during gradient processing. Federated transfer learning lacks dynamic adjustment capabilities. Traditional privacy protection mechanisms cannot dynamically adjust privacy budgets and noise intensity.
By employing a federated learning approach, local data is preprocessed and its dense features aligned. A federated mode is selected, dynamic differential privacy noise is injected and gradients are encrypted. Secure gradients are aggregated and data consistency is verified. A federated transfer learning component is established, utilizing dual-channel encrypted transmission and zero-knowledge audit evidence storage to dynamically adjust the strength of privacy protection. Secure multi-party computation and consortium blockchain architecture are adopted to ensure data security and consistency.
It achieves secure collaborative computation of cross-domain data, prevents leakage of original data, and improves data accuracy and integrity. It realizes cross-domain neighbor search and bin-based inter-domain collaborative computation through secure multi-party computation to prevent leakage of original data. It uses consortium blockchain architecture and zero-knowledge proof to verify feature consistency, ensuring data accuracy and consistency. It dynamically adjusts the privacy budget to balance privacy protection and model accuracy, and generates zero-knowledge proofs for storage to ensure the authenticity and traceability of audit information.
Smart Images

Figure CN120162828B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of privacy data computing, in particular to a privacy data computing method and system based on federated learning. BACKGROUND
[0002] In the prior art, when numerically valued features are standardized, missing values are filled, and high-base features are processed, the conventional method does not fully consider cross-domain data privacy. The ordinary mean filling method for missing values does not use a secure collaboration mechanism and cannot ensure filling accuracy while protecting data privacy. In the cross-institutional data integration scenario, for example, when financial institutions and medical and health institutions cooperate to perform data analysis, simple data processing methods are prone to increase the risk of privacy leakage. Traditional methods are difficult to achieve efficient and secure neighbor search and bin interval calculation. The conventional method cannot guarantee the security of data in the computing process through secure multi-party computation (MPC) technology, resulting in the sacrifice of privacy to obtain the calculation result in actual application. Excessive protection of privacy cannot effectively utilize multi-source data to improve data quality. The existing feature alignment technology is difficult to guarantee the security of feature consistency verification in a multi-party collaboration scenario. The traditional federated mode selection method often relies on simple experience judgment and cannot intelligently and accurately select the mode according to the actual overlap of features and sample IDs between participants. In the gradient processing aspect, the existing technology is difficult to balance privacy protection and model accuracy. The existing federated transfer learning technology has a relatively single method and lacks dynamic adjustment capability when dealing with feature distribution differences between source domains and target domains. The traditional privacy protection mechanism cannot dynamically adjust the privacy budget and noise intensity according to the change of data sensitivity. Therefore, it is necessary to provide a privacy data computing method and system based on federated learning. SUMMARY
[0003] The purpose of the present application is to provide a privacy data computing method and system based on federated learning. To solve the above-mentioned problems in the prior art, the present application is implemented by the following technical solutions:
[0004] In a first aspect, the privacy data computing method based on federated learning includes the following steps:
[0005] The local data is preprocessed, and the preprocessed data is subjected to cryptic feature alignment. A federated mode is selected. Dynamic differential privacy noise is injected based on the local model gradient. The gradient with the injected noise is encrypted to obtain a secure gradient.
[0006] The obtained secure gradient is aggregated and the consistency of the data is verified. A federated transfer learning component is established to adapt to the minimization of the feature distribution difference between the source domain and the target domain. A secure promotion tree model is established through multi-party collaborative decision-making to generate a final prediction label, and zero-knowledge audit evidence is stored.
[0007] The privacy protection is dynamically enhanced, the sensitivity of local data is classified, the intensity of dynamic differential privacy noise is verified and adjusted, double-channel encryption transmission is utilized to optimize privacy security, and the federated transfer learning component is strengthened.
[0008] In the second aspect, the embodiments of the present application provide a privacy data computing system based on federated learning, comprising the following steps:
[0009] The data storage module supports accessing local structured data from various data sources.
[0010] The data preprocessing module preprocesses the local data and aligns the features of the preprocessed data in a secret state, and selects a federated mode.
[0011] The gradient encryption module injects dynamic differential privacy noise based on the local model gradient, and encrypts the gradient with the injected noise to obtain a secure gradient.
[0012] The federated learning component construction module aggregates the obtained secure gradient and verifies the consistency of the data, establishes a federated transfer learning component to adapt to the minimum feature distribution difference between the source domain and the target domain, establishes a security promotion tree model through multi-party collaborative decision to generate a final prediction label; the zero-knowledge audit evidence module selects a suitable zero-knowledge proof algorithm according to different business scenarios and security requirements, and selects a suitable blockchain platform for deployment.
[0013] The privacy protection dynamic enhancement module dynamically enhances the privacy protection, classifies the sensitivity of local data, verifies and adjusts the intensity of dynamic differential privacy noise, optimizes privacy security by using double-channel encryption transmission, and strengthens the federated transfer learning component.
[0014] The present application has the following advantages:
[0015] 1. Numerical value type feature standardization, missing value KNN interpolation filling and high base number feature binning processing operation, improve the accuracy and integrity of the data, use secure multi-party computation MPC to realize cross-domain neighbor search and binning interval collaborative calculation, prevent original data leakage, use alliance chain architecture combined with zero-knowledge proof ZKP to realize feature alignment between multiple parties, verify feature consistency, use independent channels according to different business scenarios to avoid data cross leakage; connect the Ethereum privacy chain through the IBC protocol, support heterogeneous data source alignment, the application of zero-knowledge proof ZKP, select the appropriate federal mode, inject dynamic differential privacy noise into the local model gradient, balance privacy protection and model accuracy by controlling the privacy budget, aggregate the secure gradient; verify data consistency to ensure the reliability and stability of the data in the model training process, minimize the feature distribution difference between the source domain and the target domain, generate zero-knowledge proof and store it on the blockchain for evidence, ensure the authenticity and traceability of the audit information, facilitate the audit of regulatory agencies and authorized parties, while not leaking sensitive data;
[0016] 2. Dynamically adjust the privacy budget according to the data sensitivity level, increase the noise injection intensity for high sensitivity data, and the center node verifies the model accuracy loss after noise injection, if the loss exceeds the threshold, reduce the privacy budget value by a certain proportion and retrain, use homomorphic encryption algorithm to encrypt the gradient after noise injection, prevent data leakage, compare the consistency of the double-channel results through the center node, effectively detect whether the data has been tampered with, provide an additional security protection mechanism for data transmission and processing, perform secondary hashing on non-common features and enhance them with salt values, the smart contract negotiates the global salt value and automatically updates it after each alignment, call the precompiled ZKP verification function through the smart contract, ensure that the hash calculation meets the preset rules, verify that the submitted salt value enhanced hash value is consistent with the local original data, the salt value or timestamp has not been tampered with, if the verification is passed, the result is chained, otherwise, trigger the node reputation punishment mechanism, enhance the trust between nodes and the credibility of data, dynamically adjust the migration weight according to the similarity of the source domain and the target domain, which can make the federal migration learning better adapt to the relationship between the source domain and the target domain in different scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creating any inventive labor.
[0018] Figure 1 is the step flow chart of the privacy data calculation method based on federal learning provided by embodiment 1 of the present application;
[0019] Figure 2 is a structural schematic diagram of the privacy data calculation system based on federated learning provided by Embodiment 2 of the present application. DETAILED DESCRIPTION
[0020] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.
[0021] Embodiment 1
[0022] As shown in Figure 1 , the privacy data calculation method based on federated learning provided by the embodiments of the present application specifically includes the following steps:
[0023] Step one: pre-processing the local data and performing secure feature alignment on the pre-processed data, jointly modeling based on the pre-processed data, selecting a federated mode, injecting dynamic differential privacy noise based on the local model gradient, and encrypting the gradient with injected noise to obtain a secure gradient;
[0024] In step one:
[0025] First, the specific process of pre-processing the local data and performing secure feature alignment on the pre-processed data is as follows:
[0026] Obtain the original data table D = {x1, x2, x3,..., x n} of the local data, and the original data includes but is not limited to structured data such as bank user transaction records and hospital electronic medical records;
[0027] Pre-process the obtained original data table to obtain a pre-processed data feature set;
[0028] It should be noted that the pre-processed data feature set includes but is not limited to: a numerical pre-processed feature set, a missing value pre-processed data feature set, and a high base pre-processed data feature set
[0029] Specifically, the numerical features are standardized to obtain the numerical pre-processed feature set, such as transaction amount and age.
[0030] The missing values of the original data table are filled by KNN interpolation to obtain the missing value pre-processed data feature set, and the high base features of the original data table are binned to obtain the high base pre-processed data feature set.
[0031] It should be noted that KNN interpolation represents a missing value filling method based on neighborhood similarity, and the most similar k neighbors of the missing sample are found through secure collaboration, and the feature values of these neighbors are used to infer the missing values; the binning process represents a key step of discretizing continuous features into multiple intervals, which improves the robustness of the model, enhances the interpretability and ensures the alignment of cross-domain data distribution;
[0032] Specifically, based on the cosine similarity between the original data, the most similar k neighbors of the missing sample are found in the local data, and in the federal scenario, cross-domain neighbor search is realized through secure multi-party computation MPC, and the missing value preprocessing data feature set is obtained;
[0033] The quantile binning is used for processing to ensure that the feature distribution of different participants is in the same interval, and the secure multi-party computation MPC is used to collaboratively calculate the binning interval in the encrypted domain to obtain a high-base preprocessing data feature set. For example, the income type feature x i is processed using quantile binning to ensure that the income type feature x i is divided into a intervals, and the binning interval is equal to Quantile(x i , bins=a).
[0034] It should be noted that secure multi-party computation MPC represents a core technology for realizing cross-domain data collaborative computation, and its core goal is to enable multiple participants to jointly complete complex computation tasks without revealing the original data;
[0035] Based on the obtained preprocessing data feature set, Hyperledger Fabric consortium chain architecture is adopted, multi-party feature alignment is realized through smart contract, and feature consistency is verified through zero-knowledge proof ZKP;
[0036] It should be noted that Hyperledger Fabric consortium chain architecture represents an enterprise-level distributed ledger platform designed for consortium chains, which has high modularity and scalability, and is suitable for various complex business scenarios;
[0037] Specifically, the feature name and value are asymmetrically encrypted and hashed, and the feature hash value H(F j ) is obtained through the formula H(F j ) = SHA_3_256(F p || salt j ), where F i represents the data feature, and salt i represents a random salt value unique to participant p to prevent rainbow table attacks;
[0038] Based on the calculated feature hash value, a data node is constructed, and the local data holder stores the preprocessed feature hash value H(F j ), and the local data holder includes but is not limited to financial institutions and medical institutions; a transaction ordering and block ordering node is generated to support multi-channel isolation, including but not limited to financial channels and medical channels;
[0039] Deploy feature alignment logic and zero-knowledge proof ZKP verification rules to establish a smart contract;
[0040] It should be noted that zero-knowledge proof ZKP represents a kind of cryptography technology, which allows the prover to prove to the verifier that a statement is true without revealing any additional information other than the statement is true;
[0041] Specifically, the verification is performed using the zero-knowledge proof ZKP verification rule: zk-SNARKs(Groth16) is suitable for high-frequency low-latency scenarios, such as financial transactions, with a proof volume of about 200 bytes and a verification time of <30ms;
[0042] zk-STARKs is resistant to quantum attacks and is used in high-security scenarios such as medical scenarios, with a proof volume of 100KB and a verification time of <200ms;
[0043] Based on the established data node, ordering node and smart contract, a Hyperledger Fabric consortium chain architecture is established;
[0044] Based on the calculated feature hash value, the feature hash value is compared with the preset sensitive hash threshold, and if the feature hash value is greater than or equal to the preset sensitive hash threshold, the feature hash value corresponding to the data node is marked as a sensitive data node, and the sensitive feature hash corresponding to the sensitive data node is visible to the authorized node;
[0045] Based on different business scenarios, independent channels are used to avoid data cross-leakage;
[0046] Connect to the Ethereum privacy chain through the IBC protocol to support heterogeneous data source alignment;
[0047] It should be noted that the Ethereum privacy chain represents a kind of blockchain network based on Ethereum blockchain technology, which focuses on protecting user privacy;
[0048] Secondly, based on the preprocessed data, a joint modeling is performed, and the specific process of selecting a federal mode is as follows:
[0049] If the features overlap between the participants are more and the sample IDs overlap are less, a longitudinal federal mode is selected, and a private set intersection PSI technology is used to determine the common user ID;
[0050] If the sample IDs of the participants overlap a lot and the features overlap a little, a horizontal federated mode is selected, and the obtained aligned feature set is directly used;
[0051] Thirdly, the specific process of injecting dynamic differential privacy noise based on the local model gradient and encrypting the gradient with the injected noise to obtain a secure gradient is as follows:
[0052] In the local model gradient g i Laplace noise is injected where Δf is the gradient sensitivity, representing the maximum gradient change, and Δf is calculated by the formula where g i and g i are local model gradients calculated at different times or different samples, ∈ is the privacy budget, controlling the noise intensity, and the value range is between (0, 1];
[0053] For example, assuming N = 101 x 103 = 10403, the gradient g i after injecting noise = 0.6, then the encrypted secure gradient [[g i ]] = 0.6 x 10403 mod 10403 2 is obtained by the formula i ;
[0054] Step two: aggregate the obtained secure gradient and verify the consistency of the data, establish a federated transfer learning component to minimize the feature distribution difference between the source domain and the target domain, generate the final prediction label through the multi-party collaborative decision to establish a security promotion tree model, and perform zero-knowledge audit evidence;
[0055] In step two:
[0056] Firstly, the specific process of aggregating the obtained secure gradient and verifying the consistency of the data is as follows:
[0057] The participant weight w i is dynamically allocated according to the data quality, and the participant weight w i is obtained by the formula where YB i represents the sample size, ZL i represents the data quality coefficient, and is obtained by comprehensive evaluation of the data integrity W zx and the missing rate Q sl ;
[0058] The center node calculates the aggregated gradient, and the aggregated gradient G is obtained by the formula where w i represents the participant weight, and g irepresents the local model gradient, k represents the total number of nodes;
[0059] The improved practical Byzantine fault tolerance PBFT protocol is adopted, and the number of nodes verifying data consistency is greater than or equal to times the total number of nodes k;
[0060] The data consistency Y is obtained by verifying the formula ZX wherein λ is a preset gradient fluctuation threshold, and the value is 0.1, and k represents the total number of nodes;
[0061] For example, the federated learning system contains 3 participants, that is, k=3, and the participant parameters are shown in Table 1:
[0062] Participants Sample size YB i ]] Data integrity W zx ]]> Deletion rate Q sl ]] Local model gradient g i ]] Node A 500 0.95 0.05 0.32 Node B 300 0.8 0.2 0.28 Node C 200 0.6 0.4 0.35
[0063] Table 1 Participant data statistics table
[0064] Based on the parameter data of the participants in Table 1, the data quality coefficient ZL of the participants is calculated i :
[0065] The weight w is calculated i :
[0066]
[0067] The aggregated gradient G is calculated:
[0068]
[0069]
[0070] The threshold value is calculated: λ*|G|=0.1×|0.313|=0.0313;
[0071] The gradient deviation of each node is calculated: the gradient deviation of node A is |0.32-0.313|=0.007, the gradient deviation of node A is |0.28-0.313|=0.033, and the gradient deviation of node A is |0.35-0.313|=0.037.
[0072] The deviation of node B is 0.033>0.0313, the deviation of node C is 0.037>0.0313, and both exceed the threshold value;
[0073] According to the PBFT rule, consensus needs to be reached by nodes, node A and node B pass the verification, node C is excluded, the aggregated gradient is recalculated, and only the local model gradient of node A and node B is used;
[0074] Secondly, the federated transfer learning component is established to minimize the feature distribution difference between the source domain and the target domain, and through multi-party collaborative decision, a security promotion tree model is established to generate the final prediction label, and the specific process of zero-knowledge audit evidence is as follows:
[0075] The maximum mean difference MMD and the KL divergence are combined to minimize the feature distribution difference between the source domain D S and the target domain D T , and the objective function L adapt = MMD(D S , D T ) + KL(P S , P T ), wherein, MMD(D S , D T ) represents the distance between the source domain D S and the target domain D T in the reproducing kernel Hilbert space RKHS, wherein φ(x) represents the feature mapping function, and KL is the KL divergence.
[0076] It should be noted that the reproducing kernel Hilbert space RKHS represents a Hilbert space composed of functions, and the Hilbert space is a complete inner product space. The core feature of the reproducing kernel Hilbert space is that there is a kernel function K, so that the value of any function f in the space at point x can be represented in the form of inner product.
[0077] Each participant generates a zero-knowledge proof for the operation and data usage in the entire process;
[0078] For example, the participants need to prove that they comply with the protocol in the gradient calculation, model training and data provision process, and there is no leakage of sensitive information. The zero-knowledge proof is generated by the ZK-SNARK proof system;
[0079] The generated zero-knowledge proof and related audit information are stored on the blockchain for evidence storage, such as operation time and participant identity. The non-tamperable feature of the blockchain ensures the authenticity and traceability of the audit information. The regulatory authorities and other authorized parties can audit the behavior of the participants by verifying the zero-knowledge proof without obtaining specific sensitive data;
[0080] The technical scheme of the embodiment of the present application is: numerical type feature standardization, missing value KNN interpolation filling and high base number feature binning processing operation, which improves the accuracy and integrity of the data. In the preprocessing process, the cross-domain neighbor search and binning interval collaborative calculation are realized by using secure multi-party computation MPC, the security of the cross-domain data in the calculation process is ensured, the original data leakage is prevented, the data nodes, ordering nodes and smart contracts are constructed by asymmetric encryption hashing of the feature name and value, the feature alignment between multiple parties is realized, the feature consistency is verified, the accuracy and consistency of the data are ensured, independent channels are used according to different business scenarios to avoid data cross leakage; the IBC protocol is connected with the Ethereum privacy chain to support heterogeneous data source alignment, the application of zero-knowledge proof ZKP allows the authenticity and compliance of the data to be proved without leaking sensitive information, according to the overlapping situation of the features and sample IDs between the participants, a suitable federal mode is selected, dynamic differential privacy noise is injected into the local model gradient, the privacy protection and model accuracy are balanced by controlling the privacy budget, the secure gradient is aggregated, and the participant weight is dynamically allocated according to the data quality to improve the effect of model training; the improved practical Byzantine fault tolerance PBFT protocol is used to verify the data consistency, ensuring the reliability and stability of the data in the model training process, the maximum mean difference MMD and KL divergence are combined to minimize the feature distribution difference between the source domain and the target domain, which helps to improve the migration ability of the model between different data domains, each participant generates zero-knowledge proof and stores it on the blockchain for evidence, the tamper-proof property of the blockchain is used to ensure the authenticity and traceability of the audit information, which is convenient for the authorized parties to audit, and the sensitive data is not leaked;
[0081] Embodiment 2
[0082] As Figure 1 shown, the privacy data calculation method based on federated learning provided by the embodiment of the present application specifically includes the following steps:
[0083] Step three: dynamically enhancing privacy protection, classifying the sensitivity of local data, and verifying and adjusting the intensity of dynamic differential privacy noise, using double-channel encryption transmission to optimize privacy security, and strengthening the federated transfer learning component;
[0084] In step three:
[0085] Firstly, the specific process of classifying the sensitivity of local data in the dynamic enhancement of privacy protection is as follows:
[0086] According to the data sensitivity level S∈{1,2,3} (1 is the lowest and 3 is the highest), the privacy budget is dynamically adjusted by the formula where ∈base Based on the basic privacy budget value, for example, medical data, due to its involvement with patient privacy, is typically set to a sensitivity level of 3. If ∈ base If we set it to 0.9, then at this time... Correspondingly, the intensity of the injected noise will be increased to better protect data privacy;
[0087] The central node verifies the model accuracy loss after noise injection. If the loss exceeds the loss threshold δ, the ∈ value is reduced proportionally and the model is retrained.
[0088] The Paillier homomorphic encryption algorithm is used to analyze the gradient g after injecting noise. i Encryption is performed; using the formula [[g i ]]=g i *Nmod N 2 We obtain N = p * q (p and q are large prime numbers, which are public key parameters);
[0089] The central node compares the consistency of the results from the two channels. If the deviation exceeds the deviation threshold, it is determined to be a tampering attack.
[0090] It should be noted that the Paillier homomorphic encryption algorithm has additive homomorphism, that is, [[g i +g j ]]=[[g i ]]*[[g j This allows computational operations to be performed on the encryption gradient without decryption, ensuring data security during the computation process.
[0091] Secondly, and specifically, the process of verifying and adjusting the intensity of dynamic differential privacy noise, and optimizing privacy and security using dual-channel encrypted transmission is as follows:
[0092] Perform secondary hashing on non-common features to prevent replay attacks.
[0093] Salt enhancement is applied to feature names and values, and secondary hashing is performed on non-public features, using the formula: H'(F j )=SHA_3_256(F j ||salt p ||timestamp) to obtain the salt-enhanced hash value H'(F) j The global salt value is negotiated through a smart contract. global It updates automatically after each round of alignment;
[0094] The participants will salt-enhance the hash value H'(F) j Submitting to the Fabric channel triggers the smart contract to execute alignment logic;
[0095] The participant generates a proof pi, which states that the submitted salt value enhanced hash value H'(F j ) is consistent with the local original data, and the salt value or timestamp has not been tampered with;
[0096] A constraint system is defined using the Circom language to ensure that the hash calculation meets the preset rules;
[0097] It should be noted that the Circom language represents a domain-specific language DSL designed and developed specifically for zero-knowledge proof ZKP circuits;
[0098] The smart contract calls the pre-compiled ZKP verification function Verify(pi, H'(F j )), and after verification, the result is aligned with the proof hash on the chain, otherwise the node reputation punishment mechanism is triggered;
[0099] Thirdly, the specific process of strengthening the federal transfer learning component is:
[0100] Based on the obtained target function L adapt , the similarity sim(D S , D T ) between the source domain and the target domain is dynamically adjusted to adjust the transfer weight β, through the formula L adapt = MMD(D S , D T ) + β * KL(P S , P T ), wherein β represents the preset balance coefficient, and max(KL) represents the maximum KL divergence;
[0101] It should be noted that by dynamically adjusting the transfer weight in the above manner, the difference between the source domain and the target domain data features in different scenarios is flexibly adapted. When the similarity between the source domain and the target domain is high, the transfer weight is appropriately increased, so that the target domain can fully utilize the rich knowledge and information of the source domain, accelerate the convergence speed of the model, and improve the generalization performance of the model in the target domain. When the similarity between the two is low, the transfer weight is reduced to avoid the interference of the noise knowledge of the source domain on the learning of the target domain model, and ensure that the model focuses on the features and rules of the target domain data itself;
[0102] Based on the dynamic enhancement of privacy protection, a comprehensive and dynamic privacy protection and federal transfer learning collaborative system is constructed.
[0103] The technical solution of this invention is as follows: The privacy budget is dynamically adjusted based on the data sensitivity level. For highly sensitive data, the noise injection intensity is increased. While ensuring data privacy, this avoids excessive model accuracy loss due to over-protection. The central node verifies the model accuracy loss after noise injection. If the loss exceeds a threshold, the privacy budget value is reduced proportionally and the model is retrained. This ensures that the model maintains high performance under privacy protection. The Paillier homomorphic encryption algorithm is used to encrypt the gradient after noise injection. Utilizing its additive homomorphism, calculations can be performed on encrypted data, ensuring data security during the calculation process and preventing data leakage. The central node compares the consistency of the dual-channel results to effectively detect whether the data has been tampered with. This attack provides additional security mechanisms for data transmission and processing. It performs secondary hashing on non-public features and combines it with salt enhancement to prevent replay attacks. The smart contract negotiates the global salt value and automatically updates it after each round of alignment. It uses the Circom language to define a constraint system and calls the pre-compiled ZKP verification function through the smart contract to ensure that the hash calculation conforms to the preset rules. It verifies that the submitted salt-enhanced hash value is consistent with the local original data and that the salt value or timestamp has not been tampered with. After verification, the result is uploaded to the chain; otherwise, a node reputation penalty mechanism is triggered to enhance the trust between nodes and the credibility of the data. It dynamically adjusts the migration weight based on the similarity between the source and target domains, enabling federated transfer learning to better adapt to the relationship between the source and target domains in different scenarios.
[0104] Example 3
[0105] like Figure 2 As shown in the embodiments of the present invention, the privacy-preserving data computing system based on federated learning specifically includes the following modules:
[0106] Data storage module: Supports access to local structured data from multiple data sources, including but not limited to: bank user transaction records and hospital electronic medical records. It has data format conversion function, which can uniformly convert data of different formats into a format that the system can process. It uses a distributed file system or object storage to store the original data, ensuring high availability and scalability of the data.
[0107] Data preprocessing module: preprocesses local data and performs dense-state feature alignment on the preprocessed data, selecting federated mode.
[0108] Gradient encryption module: Based on the local model gradient, dynamic differential privacy noise is injected, and the gradient of the injected noise is encrypted to obtain a secure gradient;
[0109] Federated learning component building module: Aggregates the obtained safety gradients and verifies the consistency of the data, establishes a federated transfer learning component to adapt and minimize the feature distribution difference between the source domain and the target domain, and establishes a safety boosting tree model through multi-party collaborative decision-making to generate the final predicted label;
[0110] Zero-knowledge audit evidence module: according to different business scenarios and security requirements, select appropriate zero-knowledge proof algorithm, select appropriate blockchain platform for deployment;
[0111] Privacy protection dynamic enhancement module: dynamically enhance privacy protection, classify the sensitivity of local data, and verify and adjust the dynamic differential privacy noise intensity, optimize privacy security by using double-channel encryption transmission, and strengthen the federal transfer learning component.
[0112] The above describes one embodiment of the present application in detail, but the content is only the preferred embodiment of the present application, and cannot be considered to limit the scope of the present application; the above formula is de-dimensioned to calculate its numerical value, the formula is obtained by collecting a large amount of data to simulate a formula of the most recent real situation, and the preset parameters in the formula are set by the person skilled in the art according to the actual situation and historical experience, which can be adjusted according to the actual situation; the above is only the preferred embodiment of the present application, and cannot be used to limit the present application, any equivalent changes and improvements made according to the scope of the present application should still belong to the patent coverage range of the present application.
Claims
1. A federated learning based privacy data computation method, characterized in that, Includes the following steps: The local data is preprocessed and the preprocessed data is aligned with dense state features. The federated mode is selected, dynamic differential privacy noise is injected based on the gradient of the local model, and the gradient of the injected noise is encrypted to obtain a secure gradient. The specific process for performing dense-state feature alignment is as follows: Perform asymmetric cryptographic hashing on the feature name and value, and obtain the feature hash value using a formula; Data nodes are constructed based on the calculated feature hash values, and the local data holder stores the preprocessed feature hash values. Deploy feature alignment logic and zero-knowledge proof (ZKP) verification rules to establish smart contracts; Verification is performed using the zero-knowledge proof ZKP verification rules; Use independent channels for different business scenarios to avoid cross-data leakage; Connect to the Ethereum privacy chain via the IBC protocol to support heterogeneous data source alignment; The obtained security gradients are aggregated and the consistency of the data is verified. A federated transfer learning component is established to adapt and minimize the feature distribution difference between the source domain and the target domain. A security boosting tree model is established through multi-party collaborative decision-making to generate the final predicted label, and zero-knowledge audit evidence is performed. The privacy protection is dynamically enhanced by classifying the sensitivity of local data, verifying and adjusting the intensity of dynamic differential privacy noise, optimizing privacy security through dual-channel encrypted transmission, and strengthening the federated transfer learning component. The process of dynamically enhancing privacy protection is as follows: The homomorphic encryption algorithm is used to encrypt the gradient after injecting noise ; and the formula is obtained, wherein , and are large prime numbers, and is a public key parameter; the central node compares the consistency of the double-channel results, and if the deviation exceeds a deviation threshold, it is determined as a tampering attack; The process of verifying and adjusting the dynamic differential privacy noise intensity is as follows: Salt enhancement is applied to feature names and values, and secondary hashing is performed on non-common features using the formula. Obtain salt-enhanced hash value The global salt value is negotiated through smart contracts. It updates automatically after each round of alignment, among which, This represents data characteristics. This refers to the participating parties. Unique random salt value to prevent rainbow table attacks; The smart contract calls a pre-compiled ZKP verification function wherein, The proof generated for the participant is verified, and after verification, the alignment result is chained with the proof hash, otherwise, the node reputation punishment mechanism is triggered.
2. The federated learning based private data computation method according to claim 1, wherein, The specific process for preprocessing local data is as follows: The original data table of the local data is acquired The numerical value type features are standardized to obtain a numerical preprocessed feature set, KNN interpolation is used to fill in the missing values of the original data table to obtain a missing value preprocessed data feature set, and high-base number features of the original data table are subjected to binning processing to obtain a high-base number preprocessed data feature set. Based on the obtained preprocessed data feature set, the following is adopted: The consortium blockchain architecture uses smart contracts to align features among multiple parties and combines zero-knowledge proofs (ZKP) to verify feature consistency.
3. The privacy data computation method based on federated learning according to claim 1, characterized in that, The specific process for obtaining the security gradient is as follows: Local model gradient Injecting Laplace noise ,in, For gradient sensitivity, Through formula The calculation shows that, and These are the local model gradients calculated at different times or for different samples. To maintain privacy, noise intensity is controlled, with values ranging from [value range missing]. between.
4. The privacy data computation method based on federated learning according to claim 1, characterized in that, The specific process of aggregating the security gradient and verifying data consistency is as follows: Participant weights are dynamically allocated based on data quality, and are obtained through a formula. ; The central node calculates the aggregate gradient, which is obtained through a formula. By verifying the formula Achieving data consistency ,in, This represents the total number of nodes. The preset gradient fluctuation threshold is set to 0.
1.
5. The privacy data computation method based on federated learning according to claim 1, characterized in that, The process of minimizing the difference in feature distribution between the source domain and the target domain is as follows: A method combining maximum mean difference (MMD) and KL divergence is used to minimize the source domain. With the target domain The difference in characteristic distribution, objective function ,in, This represents minimizing the source domain. With the target domain Distance in the regenerating nucleus Hilbert space RKHS It is the KL divergence.
6. The privacy data computation method based on federated learning according to claim 1, characterized in that, The process of enhancing the federated transfer learning component is as follows: The migration weights are dynamically adjusted based on the similarity between the source and target domains. Through formula The enhanced objective function is obtained, where, This represents minimizing the source domain. With the target domain Distance in the regenerating nucleus Hilbert space RKHS It is the KL divergence. This represents the preset balance coefficient.
7. A privacy-preserving data computation system based on federated learning, the system being used to execute the privacy-preserving data computation method according to any one of claims 1-6, characterized in that, include: Data storage module: Supports accessing locally structured data from multiple data sources; Data preprocessing module: preprocesses local data and aligns the preprocessed data with dense features, selecting federated mode; Gradient encryption module: Based on the local model gradient, dynamic differential privacy noise is injected, and the gradient of the injected noise is encrypted to obtain a secure gradient; Federated learning component building module: Aggregates the obtained safety gradients and verifies the consistency of the data, establishes a federated transfer learning component to adapt and minimize the feature distribution difference between the source domain and the target domain, and establishes a safety boosting tree model through multi-party collaborative decision-making to generate the final predicted label; Zero-knowledge audit and evidence storage module: Select appropriate zero-knowledge proof algorithms and suitable blockchain platforms for deployment based on different business scenarios and security requirements; The privacy protection dynamic enhancement module dynamically enhances privacy protection by classifying the sensitivity of local data, verifying and adjusting the intensity of dynamic differential privacy noise, optimizing privacy security through dual-channel encrypted transmission, and strengthening the federated transfer learning component.
Citation Information
Patent Citations
Physical medical data fusion privacy protection method based on cloud and mist architecture longitudinal federal learning
CN116595584A
Method and device for realizing data privacy protection processing based on federal model training, processor and computer readable storage medium thereof
CN118036067A