Privacy Data Analysis Method and System Based on Collaborative Learning and Dynamic Encryption
By adopting a method based on collaborative learning and dynamic encryption in data analysis, using homomorphic encryption and quantum key distribution technology to dynamically adjust the privacy protection level, the contradiction between privacy protection and data analysis efficiency in the existing technology is solved, and efficient and accurate privacy data analysis is achieved.
Patent Information
- Application Number
- CN202510112140.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-24
AI Technical Summary
It is difficult for the existing technology to maintain high data analysis efficiency and accuracy while protecting data privacy. Especially in large-scale and heterogeneous data environments, static privacy protection policies are difficult to adapt to dynamically changing privacy needs.
The privacy data analysis method based on collaborative learning and dynamic encryption is adopted, and the initial encryption is realized through homomorphic encryption algorithms. The dynamic encryption key is generated using quantum key distribution technology, and the privacy protection level is dynamically adjusted in combination with the federated learning architecture and secure computing protocol.
It achieves the protection of data privacy while maintaining high data analysis efficiency and accuracy, adapting to the needs of large-scale and heterogeneous data environments, and balancing the relationship between privacy protection and data utilization by dynamically adjusting the privacy protection level.
Smart Images

Figure CN119557909B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to privacy technologies, and in particular to a privacy data analysis method and system based on collaborative learning and dynamic encryption. Background Art
[0002] With the advent of the big data era, data analysis is increasingly widely used in various fields. However, the problem of data privacy protection has also become prominent and has become an important issue that needs to be solved urgently. Traditional data analysis methods often require centralized collection and processing of a large amount of raw data, which not only increases the risk of data leakage but also faces challenges in terms of laws and regulations.
[0003] Most existing privacy protection schemes adopt static security policies and are difficult to adapt to different application scenarios and dynamically changing privacy requirements. In practical applications, the sensitivity of data may change with factors such as time and environment, and a fixed privacy protection level may lead to overprotection or underprotection.
[0004] Therefore, there is an urgent need for a new method that can protect data privacy while maintaining high data analysis efficiency and accuracy. This method should be able to adapt to large-scale and heterogeneous data environments and dynamically adjust the privacy protection level according to actual needs to balance the relationship between privacy protection and data utility. Summary of the Invention
[0005] Embodiments of the present invention provide a privacy data analysis method and system based on collaborative learning and dynamic encryption, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention,
[0007] A privacy data analysis method based on collaborative learning and dynamic encryption is provided, including:
[0008] Receiving the raw data of multiple participants, the raw data of each participant is in an initial encrypted state, and the initial encrypted state is implemented by a homomorphic encryption algorithm; assigning a unique participant identifier to each participant, and the participant identifier includes a timestamp and a random hash value; generating a corresponding dynamic encryption key based on the participant identifier, and the dynamic encryption key is generated by a quantum key distribution technology; using the dynamic encryption key to perform secondary encryption on the raw data of each participant to obtain secondary encrypted data, and the secondary encryption adopts a hierarchical encryption strategy, and different encryption algorithms with different strengths are used for different sensitive levels of data;
[0009] Input the secondary encrypted data into a preset collaborative learning model. The collaborative learning model includes multiple sub-models, each sub-model corresponding to a participating party, and the sub-models adopt a federated learning architecture; use a secure computing protocol under homomorphic encryption to decrypt the collaborative learning model in combination with the dynamic encryption keys of each participating party to obtain a privacy-preserving analysis model that can be used by each participating party.
[0010] When performing data analysis, use the privacy-preserving analysis model to process the input data to be analyzed, and ensure the correctness of the analysis process through verifiable computing technology; based on a predefined access control policy, only return the analysis results that the corresponding participating party has the right to access, and the analysis results are processed by k-anonymization to further enhance privacy protection; regularly execute the model update and retraining process to adapt to the dynamic changes in the data distribution.
[0011] In an alternative embodiment,
[0012] Generate a corresponding dynamic encryption key based on the participating party identifier. The dynamic encryption key is generated using quantum key distribution technology, including:
[0013] Use a periodically poled lithium niobate crystal to generate entangled photon pairs through the spontaneous parametric down-conversion process. The temperature of the periodically poled lithium niobate crystal is controlled at 163.5 ± 0.1 °C, and the wavelength of the entangled photon pairs is 1550 nm; send one photon of the entangled photon pairs to the sender and the other photon to the receiver, where the photon transmission uses a photonic crystal fiber with an optical loss rate not higher than 0.2 dB / km; set a quantum repeater every 50 km on the transmission path of the photonic crystal fiber. The quantum repeater uses a rare-earth doped YSO crystal as a quantum storage medium with a storage time not less than 1 ms.
[0014] The sender and the receiver respectively use superconducting nanowire single-photon detectors to perform quantum state measurements on the received photons. The quantum efficiency of the superconducting nanowire single-photon detector at a wavelength of 1550 nm is not less than 93%, and the dark count rate is not higher than 1 Hz.
[0015] Based on the quantum state measurement results, generate a quantum key through post-processing steps. The post-processing steps include basis comparison, parameter estimation, information reconciliation, and privacy amplification. Among them, information reconciliation uses a rate-adaptive low-density parity-check code, and privacy amplification uses a privacy matrix with a dynamically adjustable size.
[0016] In an alternative embodiment,
[0017] Use the dynamic encryption key to perform secondary encryption on the original data of each party to obtain secondary encrypted data. The secondary encryption adopts a hierarchical encryption strategy, and different encryption algorithms with different strengths are used for data with different sensitivity levels, including:
[0018] Use a data sensitivity evaluation model based on deep learning to evaluate the sensitivity of the data to be encrypted. The data sensitivity evaluation model based on deep learning includes 12 basic layers, 2 fully connected layers, 1 Dropout layer, and 1 output layer. The dropout rate of the Dropout layer is 0.1, and the output layer uses the Sigmoid activation function;
[0019] Based on the quantum key, use a key derivation function to generate three sub-keys, which are respectively used for data encryption at high, medium, and low sensitivity levels;
[0020] For high-sensitivity data with a sensitivity score greater than 0.8, use the Kyber1024 algorithm for key encapsulation and use AES-256-GCM for authenticated encryption; for medium-sensitivity data with a sensitivity score between 0.4 and 0.8, use the NTRU-HRSS-701 parameter set for key generation and encapsulation and use ChaCha20-Poly1305 for authenticated encryption; for low-sensitivity data with a sensitivity score less than 0.4, directly use the derived sub-key and use XChaCha20-Poly1305 for authenticated encryption.
[0021] In an alternative embodiment,
[0022] Before inputting the secondary encrypted data into a preset collaborative learning model, the method further includes training the collaborative learning model:
[0023] During the training process of the collaborative learning model, each sub-model only processes the secondary encrypted data of its own party and generates intermediate results, where the intermediate results include model parameter gradients and local loss function values; use a secure multi-party computation protocol to collect the intermediate results generated by each sub-model and aggregate the intermediate results to form a global model. The aggregation process uses differential privacy technology to prevent reverse derivation of the original data; use the global model to perform analysis and processing under privacy protection on newly input data to be analyzed, where the analysis and processing include feature extraction, classification, regression, and clustering; during the training process of the collaborative learning model, periodically update the dynamic encryption keys of each party based on a Bloom filter and a time decay function, and re-encrypt the original data of each party using the updated dynamic encryption keys;
[0024] Monitor the training process of the collaborative learning model. When it is detected that the model performance reaches a preset threshold or a potential model poisoning attack is detected, trigger a model verification program;
[0025] The model verification program includes: verifying the model integrity using zero - knowledge proof technology, evaluating the model robustness by adversarial sample testing, and ensuring the model fairness through multi - party consistency checking; dynamically adjusting the collaborative learning strategy according to the model verification results, including updating the weights of participants, adjusting the learning rate, and modifying the model structure; when the model verification passes and the performance meets the standard, terminating the training process and outputting the final collaborative learning model.
[0026] In an alternative embodiment,
[0027] During the training process of the collaborative learning model, each sub - model only processes the double - encrypted data of its own participant and generates intermediate results, where the intermediate results include model parameter gradients and local loss function values; using a secure multi - party computation protocol to collect the intermediate results generated by each sub - model and aggregating the intermediate results to form a global model, and the aggregation process uses differential privacy technology to prevent reverse derivation of the original data; using the global model to perform analysis and processing under privacy protection on newly input data to be analyzed, and the analysis and processing include feature extraction, classification, regression, and clustering, including:
[0028] Encrypt the original data of the participant using the CKKS approximate homomorphic encryption scheme, where the polynomial modulus degree of the CKKS approximate homomorphic encryption scheme is 2^15, the plaintext modulus is 2^40, the ciphertext modulus is approximately 2^440, and the security level is greater than 128 bits; in the local sub - model of each participant, perform batch processing encoding on the encrypted data, pack multiple real numbers into a polynomial, and use polynomial approximation of Taylor expansion to calculate the non - linear activation function; perform forward propagation and backward propagation calculations on the encrypted data to generate encrypted model parameter gradients and encrypted local loss function values;
[0029] Generate a public key and a private key using the (T,n) - threshold Paillier encryption scheme, where T is the reconstruction threshold and n is the number of participants, and use the Shamir secret sharing scheme to split the private key into n shares and distribute them to each participant; each participant uses the public key to encrypt its model parameter gradients and local loss function values, and add calibrated Gaussian noise before encryption, and the noise variance σ² of the calibrated Gaussian noise is calculated according to the preset privacy budget ε and sensitivity Δf;
[0030] Select one participant as the aggregator, and other participants send the encrypted intermediate results to the aggregator, and the aggregator calculates the sum of the encrypted gradients and encrypted losses; at least n participants cooperate to perform partial decryption on the encrypted total gradients and total losses using their respective private key shares, and the aggregator collects the partial decryption results and completes the final decryption to obtain the global gradients and global losses;
[0031] Update the global model parameters using the global gradient, and use the global model to perform analysis and processing under privacy protection on newly input data to be analyzed; in the analysis and processing, use the SimHash algorithm to extract features, map high-dimensional feature vectors to a low-dimensional binary space; use the exponential mechanism to select the best splitting feature to construct an ensemble of privacy-preserving decision trees; adopt inner product functional encryption to achieve privacy-preserving linear regression; introduce Laplace noise in each iteration of the K-means clustering algorithm to protect the cluster center coordinates.
[0032] In an alternative embodiment,
[0033] During the training process of the collaborative learning model, periodically update the dynamic encryption keys of each party based on the Bloom filter and the time decay function, and re-encrypt the original data of each party using the updated dynamic encryption keys, including:
[0034] Set up a Bloom filter, the size m and the number k of hash functions of which are calculated according to the expected number of elements N and the false positive probability p; use the exponential decay function w(t) = exp(-λt) to calculate the key update weight, where λ is the decay rate and t is the key usage time;
[0035] For each encryption operation, calculate its feature hash and insert it into the Bloom filter, and periodically check the filling rate of the Bloom filter; when the filling rate exceeds the preset filling threshold, trigger key update, generate a new key, and use the exponential decay function to calculate the mixing weight, mix the new and old keys to obtain the final key, and re-encrypt the original data of each party using the updated dynamic encryption key.
[0036] In the second aspect of the embodiments of the present invention,
[0037] Provide a privacy data analysis system based on collaborative learning and dynamic encryption, including:
[0038] A first unit for receiving the original data of multiple parties, the original data of each party being in an initial encrypted state, the initial encrypted state being implemented using a homomorphic encryption algorithm; assigning a unique party identifier to each party, the party identifier including a timestamp and a random hash value; generating a corresponding dynamic encryption key based on the party identifier, the dynamic encryption key being generated using quantum key distribution technology; using the dynamic encryption key to perform secondary encryption on the original data of each party to obtain secondary encrypted data, the secondary encryption adopting a hierarchical encryption strategy and using different strength encryption algorithms for different sensitive levels of the data;
[0039] A second unit for inputting the secondary encrypted data into a preset collaborative learning model, where the collaborative learning model includes multiple sub-models, each sub-model corresponding to a participating party, and the sub-models adopt a federated learning architecture; using a secure computing protocol under homomorphic encryption, decrypting the collaborative learning model in combination with the dynamic encryption keys of each participating party to obtain a privacy-preserving analysis model available for each participating party;
[0040] A third unit for, when performing data analysis, processing the input data to be analyzed using the privacy-preserving analysis model and ensuring the correctness of the analysis process through verifiable computing technology; based on a predefined access control policy, only returning the analysis results that the corresponding participating party has the right to access, where the analysis results are processed by k-anonymization to further enhance privacy protection; regularly executing a model update and retraining process to adapt to the dynamic changes in the data distribution.
[0041] In the third aspect of the embodiments of the present invention,
[0042] There is provided an electronic device, including:
[0043] A processor;
[0044] A memory for storing instructions executable by the processor;
[0045] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0046] In the fourth aspect of the embodiments of the present invention,
[0047] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0048] The quantum key generated by this application has high security and randomness and can be used for subsequent dynamic encryption processes. This method for generating dynamic encryption keys based on quantum key distribution technology combines the basic principles of quantum physics and advanced information processing technology, providing a solid security foundation for privacy data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flowchart of the privacy data analysis method based on collaborative learning and dynamic encryption according to the embodiments of the present invention;
[0050] Figure 2 It is a structural diagram of the privacy data analysis system based on collaborative learning and dynamic encryption according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0053] Figure 1 The following is a schematic flowchart of the privacy data analysis method based on collaborative learning and dynamic encryption according to the embodiments of the present invention, as Figure 1 shown. The method includes:
[0054] S101. Receive the original data of multiple participating parties. The original data of each participating party is in an initial encrypted state, and the initial encrypted state is implemented using a homomorphic encryption algorithm. Assign a unique participating party identifier to each participating party. The participating party identifier includes a timestamp and a random hash value. Generate a corresponding dynamic encryption key based on the participating party identifier. The dynamic encryption key is generated using quantum key distribution technology. Use the dynamic encryption key to perform secondary encryption on the original data of each participating party to obtain secondary encrypted data. The secondary encryption adopts a hierarchical encryption strategy, and different encryption algorithms with different strengths are used for data with different sensitivity levels.
[0055] S102. Input the secondary encrypted data into a preset collaborative learning model. The collaborative learning model includes multiple sub-models, and each sub-model corresponds to a participating party. The sub-model adopts a federated learning architecture. Use a secure computing protocol under homomorphic encryption to decrypt the collaborative learning model in combination with the dynamic encryption keys of each participating party to obtain a privacy protection analysis model that can be used by each participating party.
[0056] S103. When performing data analysis, use the privacy protection analysis model to process the input data to be analyzed, and ensure the correctness of the analysis process through verifiable computing technology. Based on a predefined access control policy, only return the analysis results that the corresponding participating party has the right to access to the corresponding participating party. The analysis results are processed by k-anonymization to further enhance privacy protection. Regularly execute the model update and retraining process to adapt to the dynamic changes in the data distribution.
[0057] In an alternative embodiment,
[0058] Generate a corresponding dynamic encryption key based on the participant identifier, and the dynamic encryption key is generated using quantum key distribution technology, including:
[0059] Use a periodically poled lithium niobate crystal to generate entangled photon pairs through the process of spontaneous parametric down-conversion. The temperature of the periodically poled lithium niobate crystal is controlled at 163.5 ± 0.1 °C, and the wavelength of the entangled photon pairs is 1550 nm. Send one photon of the entangled photon pairs to the sender and the other photon to the receiver, where the photon transmission uses a photonic crystal fiber with an optical loss rate not higher than 0.2 dB / km. Set a quantum repeater every 50 km on the transmission path of the photonic crystal fiber. The quantum repeater uses a rare-earth doped YSO crystal as a quantum storage medium with a storage time not less than 1 ms.
[0060] The sender and the receiver respectively use superconducting nanowire single-photon detectors to perform quantum state measurements on the received photons. The quantum efficiency of the superconducting nanowire single-photon detector at a wavelength of 1550 nm is not less than 93%, and the dark count rate is not higher than 1 Hz.
[0061] Based on the quantum state measurement results, generate a quantum key through a post-processing step. The post-processing step includes basis comparison, parameter estimation, information reconciliation, and privacy amplification. Among them, information reconciliation uses a rate-adaptive low-density parity-check code, and privacy amplification uses a privacy matrix with a dynamically adjustable size.
[0062] Exemplarily, first, prepare a periodically poled lithium niobate crystal as a source for generating entangled photon pairs. Precisely control the crystal temperature at 163.5 ± 0.1 °C. This temperature range is crucial for maintaining the nonlinear optical properties of the crystal. Under this condition, generate entangled photon pairs with a wavelength of 1550 nm through the process of spontaneous parametric down-conversion. The 1550 nm wavelength is selected because it has low losses in fiber optic transmission and is suitable for long-distance quantum communication.
[0063] Next, separate the generated entangled photon pairs, send one photon to the sender, and the other to the receiver. The photon transmission uses a specially designed photonic crystal fiber with an optical loss rate not exceeding 0.2 dB / km. This low-loss characteristic can significantly improve the transmission distance and quality of quantum signals. For example, at a transmission distance of 100 km, the signal intensity only attenuates by about 20%, which is much better than traditional fibers.
[0064] Considering the attenuation problem of quantum signals during long-distance transmission, a quantum repeater is set every 50 km along the transmission path of the photonic crystal fiber. These repeaters use rare-earth doped YSO crystals as quantum storage media and can achieve a storage time of no less than 1 ms. This means that quantum information can be temporarily stored in the repeater for at least 1 millisecond, providing an adequate time window for subsequent quantum state operations and transmissions.
[0065] When photons reach the sender and the receiver, superconducting nanowire single-photon detectors are used by both parties for quantum state measurement. The selected detectors have a quantum efficiency of at least 93% at a wavelength of 1550 nm, and at the same time, the dark count rate is lower than 1 Hz. Such high-efficiency and low-noise detectors can accurately capture weak quantum signals, ensuring the reliability of the measurement results. For example, within 1 second, the detector can effectively identify more than 93,000 photons, while the number of false detection signals generated is no more than 1.
[0066] Based on the results of quantum state measurement, the sender and the receiver perform a series of post-processing steps to generate the final quantum key. First is the basis comparison, where both parties publicly exchange their choices of measurement bases and retain the measurement results with consistent bases. Then parameter estimation is carried out to evaluate the error rate of the quantum channel and the possible amount of information leakage.
[0067] Next is the information reconciliation stage, where rate-adaptive low-density parity-check codes are used to correct errors in the measurement results. This coding method can dynamically adjust the coding rate according to the actual channel conditions, minimizing information leakage while ensuring the error correction effect. For example, when the channel error rate is 5%, a coding rate of 0.8 may be adopted; while when the error rate rises to 10%, the coding rate may be reduced to 0.6 to provide stronger error correction capabilities.
[0068] The final privacy amplification step uses a privacy matrix with a dynamically adjustable size. The dimension of this matrix can change dynamically according to the current security requirements and computing resources. For example, under high security requirements, a large matrix of 10000×5000 may be used; while in the case of resource constraints, it may be reduced to a scale of 1000×500. By adjusting the matrix size, a balance can be achieved between security and efficiency.
[0069] After this series of steps, the finally generated quantum key has a high level of security and randomness and can be used for subsequent dynamic encryption processes. This method of generating dynamic encryption keys based on quantum key distribution technology combines the basic principles of quantum physics and advanced information processing technologies, providing a solid security foundation for privacy data analysis.
[0070] In an alternative implementation,
[0071] Use the dynamic encryption key to perform secondary encryption on the original data of each participant to obtain secondary encrypted data. The secondary encryption adopts a hierarchical encryption strategy, and different strength encryption algorithms are used for data with different sensitivity levels, including:
[0072] Use a data sensitivity assessment model based on deep learning to assess the sensitivity of the data to be encrypted. The data sensitivity assessment model based on deep learning includes 12 basic layers, 2 fully connected layers, 1 Dropout layer, and 1 output layer. The dropout rate of the Dropout layer is 0.1, and the output layer uses the Sigmoid activation function;
[0073] Based on the quantum key, use a key derivation function to generate three sub-keys, which are respectively used for data encryption at high, medium, and low sensitivity levels;
[0074] For high-sensitivity data with a sensitivity score greater than 0.8, use the Kyber1024 algorithm for key encapsulation and use AES-256-GCM for authenticated encryption; for medium-sensitivity data with a sensitivity score between 0.4 and 0.8, use the NTRU-HRSS-701 parameter set for key generation and encapsulation and use ChaCha20-Poly1305 for authenticated encryption; for low-sensitivity data with a sensitivity score less than 0.4, directly use the derived sub-key and use XChaCha20-Poly1305 for authenticated encryption.
[0075] Exemplarily, first, assess the sensitivity of the data to be encrypted. This step uses a data sensitivity assessment model based on deep learning, and the model architecture includes 12 basic layers, 2 fully connected layers, 1 Dropout layer, and 1 output layer. The basic layer is responsible for feature extraction, which can be a convolutional layer or a recurrent layer, selected according to the data type. The fully connected layer is used to map the extracted features to the sensitivity score. The dropout rate of the Dropout layer is set to 0.1, which helps prevent overfitting. The output layer uses the Sigmoid activation function to map the final result to a value between 0 and 1, representing the sensitivity score of the data.
[0076] For example, for a piece of data containing name, age, address, and medical records, the model may give a sensitivity score of 0.95, indicating that this is highly sensitive information. While for data containing only age and city, the model may give a score of 0.3, indicating lower sensitivity.
[0077] Next, based on the previously generated quantum key, three sub-keys are generated using a key derivation function. This step ensures that even if one sub-key is cracked, the other sub-keys remain secure. For example, HKDF (HMAC-based Key Derivation Function) can be used as the key derivation function to derive three 256-bit sub-keys from a 256-bit quantum key, which are respectively used for data encryption at high, medium, and low sensitivity levels.
[0078] Subsequently, according to the sensitivity score of the data, the corresponding encryption algorithm is selected. For high-sensitivity data with a sensitivity score greater than 0.8, the Kyber1024 algorithm is used for key encapsulation. Kyber1024 is a post-quantum cryptography algorithm that can resist attacks from quantum computers. The encapsulated key is used for the AES-256-GCM algorithm for authenticated encryption. This combination provides extremely high security and is suitable for protecting the most sensitive data.
[0079] For medium-sensitivity data with a sensitivity score between 0.4 and 0.8, the NTRU-HRSS-701 parameter set is used for key generation and encapsulation. NTRU is also a post-quantum cryptography algorithm that provides strong security. The encapsulated key is used for authenticated encryption with the ChaCha20-Poly1305 algorithm. This combination achieves a good balance between security and performance and is suitable for handling medium-sensitivity data.
[0080] For low-sensitivity data with a sensitivity score less than 0.4, the derived sub-key is directly used, and XChaCha20-Poly1305 is used for authenticated encryption. XChaCha20 is an extended version of ChaCha20 that uses a longer nonce, enhancing security. This method provides sufficient protection for low-sensitivity data while maintaining high encryption efficiency.
[0081] In practical applications, for example, for a data entry containing name, ID number, and bank account information, the sensitivity score is 0.92. The system will use Kyber1024 to encapsulate a temporary key and then use AES-256-GCM to encrypt this data. The encryption process not only ensures the confidentiality of the data but also provides integrity protection through the GCM mode.
[0082] For a data entry containing age, occupation, and work city, the sensitivity score is 0.6. The system will use NTRU-HRSS-701 to generate and encapsulate the key and then use ChaCha20-Poly1305 to encrypt the data. This method ensures security while having a faster encryption speed compared to the encryption method for high-sensitivity data.
[0083] For data entries that contain only public information such as company names and founding years, the sensitivity score is 0.2. The system directly uses the derived sub-key and XChaCha20-Poly1305 for encryption. This method provides fast encryption speed and is suitable for processing a large amount of low-sensitivity data.
[0084] Through this hierarchical encryption strategy, the system can flexibly select the encryption strength according to the sensitivity of the data, optimizing the use of computing resources while ensuring security. This method not only improves the overall encryption efficiency but also ensures that the most sensitive data is protected with the strongest protection.
[0085] In an alternative embodiment,
[0086] Before inputting the secondarily encrypted data into a preset collaborative learning model, the method further includes training the collaborative learning model:
[0087] During the training process of the collaborative learning model, each sub-model only processes the secondarily encrypted data of its own participating party and generates intermediate results, which include model parameter gradients and local loss function values; uses a secure multi-party computation protocol to collect the intermediate results generated by each sub-model and aggregates the intermediate results to form a global model, and the aggregation process uses differential privacy technology to prevent reverse derivation of the original data; uses the global model to perform analysis and processing under privacy protection on newly input data to be analyzed, and the analysis and processing include feature extraction, classification, regression, and clustering; during the training process of the collaborative learning model, periodically updates the dynamic encryption keys of each participating party based on a Bloom filter and a time decay function, and re-encrypts the original data of each participating party using the updated dynamic encryption keys;
[0088] Monitor the training process of the collaborative learning model, and trigger a model verification program when it is detected that the model performance reaches a preset threshold or a potential model poisoning attack is detected;
[0089] The model verification program includes: using zero-knowledge proof technology to verify model integrity, using adversarial sample testing to evaluate model robustness, and ensuring model fairness through multi-party consistency checking; dynamically adjusting the collaborative learning strategy according to the model verification results, including updating the participating party weights, adjusting the learning rate, and modifying the model structure; when the model verification passes and the performance meets the standard, terminate the training process and output the final collaborative learning model.
[0090] Exemplarily, before inputting the secondarily encrypted data into a preset collaborative learning model, it is necessary to train the model. This training process involves multiple key steps, aiming to ensure the performance, security, and privacy protection of the model.
[0091] First, during the training process of the collaborative learning model, the sub-models of each participating party only process the secondarily encrypted data of their own party. For example, for a medical data analysis task, the sub-model of Hospital A may process 1000 encrypted patient records, while the sub-model of Hospital B processes 800. Each sub-model generates intermediate results after processing, including model parameter gradients and local loss function values. These intermediate results do not directly expose the original data but reflect the progress of model learning.
[0092] Next, a secure multi-party computation protocol is used to collect the intermediate results generated by each sub-model. This step ensures the security of the intermediate results during transmission and aggregation. For example, homomorphic encryption technology can be used to enable each party to perform calculations in the encrypted state without decrypting the intermediate results.
[0093] Then, the collected intermediate results are aggregated to form a global model. During the aggregation process, differential privacy technology is applied to prevent reverse derivation of the original data. Specifically, calibrated noise can be added to the gradients before aggregation. For example, for gradients with a sensitivity of 1, Gaussian noise with a mean of 0 and a standard deviation of 0.1 may be added, which not only protects privacy but also does not significantly affect the model performance.
[0094] Using the formed global model, privacy-preserving analysis and processing can be performed on newly input data to be analyzed. This includes tasks such as feature extraction, classification, regression, and clustering. For example, in a medical diagnosis scenario, the model may extract key features from patient data, perform disease classification, or predict treatment effects.
[0095] During the training process of the collaborative learning model, it is also necessary to regularly update the dynamic encryption keys of each participating party. This step is implemented based on a Bloom filter and a time decay function. The Bloom filter is used to quickly check whether the key needs to be updated, and the time decay function determines the update frequency. For example, the key can be set to be updated every 24 hours, or when a certain position in the Bloom filter is accessed more than 1000 times, an update is triggered. The updated dynamic encryption key is used to re-encrypt the original data of each participating party, further enhancing data security.
[0096] Throughout the training process, it is necessary to continuously monitor the performance of the collaborative learning model. When it is detected that the model performance reaches a preset threshold (for example, the accuracy exceeds 95%) or a potential model poisoning attack is found (such as the model accuracy suddenly drops by more than 10%), the model verification program is triggered.
[0097] The model verification process includes multiple steps. First, zero-knowledge proof technology is used to verify the model integrity. This can ensure that the model has not been tampered with and that all participating parties have followed the protocol. For example, the zk-SNARK protocol can be used to allow each participating party to prove that they have correctly executed the model update without revealing the specific update content.
[0098] Second, adversarial sample testing is adopted to evaluate the model robustness. This involves generating carefully designed inputs that attempt to deceive the model. For example, in an image classification task, small perturbations may be added to normal images to check whether the model can still classify correctly. If the model maintains correct predictions on more than 95% of the adversarial samples, it can be considered to have good robustness.
[0099] In addition, multi-party consistency checking is carried out to ensure the model fairness. This may include checking whether the performance of the model is consistent across different demographic groups. For example, in a loan approval model, ensuring that different applicants are treated fairly with a difference of no more than 5%.
[0100] According to the results of model verification, the collaborative learning strategy is dynamically adjusted. This may include updating the weights of the participating parties (such as reducing the weights of underperforming participating parties), adjusting the learning rate (such as reducing the learning rate from 0.01 to 0.001 when the model performance stagnates), or modifying the model structure (such as adding or reducing convolutional layers in a convolutional neural network).
[0101] When the model verification passes and the performance meets the standard, the training process is terminated and the final collaborative learning model is output. For example, when the accuracy of the model on the validation set exceeds 98% and all security and fairness checks have been passed, it can be considered that the training is completed.
[0102] This complete training process not only ensures the performance of the collaborative learning model but also maintains data privacy and model security throughout the process. Through periodic key updates, continuous monitoring and verification, and dynamic policy adjustment, this method can address various potential security threats and performance challenges, and finally produce a high-quality and trustworthy collaborative learning model.
[0103] In an alternative implementation,
[0104] During the training process of the collaborative learning model, each sub-model only processes the double-encrypted data of its own participating party and generates intermediate results, where the intermediate results include model parameter gradients and local loss function values; a secure multi-party computation protocol is used to collect the intermediate results generated by each sub-model and aggregate the intermediate results to form a global model, and differential privacy technology is adopted in the aggregation process to prevent reverse derivation of the original data; the global model is used to perform analysis and processing on the newly input data to be analyzed under privacy protection, and the analysis and processing include feature extraction, classification, regression, and clustering including:
[0105] Use the CKKS approximate homomorphic encryption scheme to encrypt the original data of the participants. The polynomial modulus degree of the CKKS approximate homomorphic encryption scheme is 2^15, the plaintext modulus is 2^40, the ciphertext modulus is approximately 2^440, and the security level is greater than 128 bits; in the local sub-model of each participant, batch process the encrypted data, pack multiple real numbers into a polynomial, and use the polynomial approximation of Taylor expansion to calculate the non-linear activation function; perform forward propagation and backward propagation calculations on the encrypted data to generate the encrypted model parameter gradients and the encrypted local loss function values.
[0106] Use the (T,n) threshold Paillier encryption scheme to generate the public key and the private key, where T is the reconstruction threshold and n is the number of participants, and use the Shamir secret sharing scheme to split the private key into n shares and distribute them to each participant; each participant uses the public key to encrypt its model parameter gradients and local loss function values, and add calibrated Gaussian noise before encryption. The noise variance σ² of the calibrated Gaussian noise is calculated according to the preset privacy budget ε and sensitivity Δf.
[0107] Select a participant as the aggregator, and other participants send the encrypted intermediate results to the aggregator. The aggregator calculates the sum of the encrypted gradients and the encrypted losses; at least n participants cooperate to perform partial decryption on the encrypted total gradients and total losses using their respective private key shares. The aggregator collects the partial decryption results and completes the final decryption to obtain the global gradients and the global losses.
[0108] Use the global gradients to update the global model parameters, and use the global model to perform analysis processing on the newly input data to be analyzed under privacy protection; in the analysis processing, use the SimHash algorithm for feature extraction to map the high-dimensional feature vectors to the low-dimensional binary space; use the exponential mechanism to select the best split features to construct a privacy protection decision tree ensemble; use inner product functional encryption to achieve privacy protection linear regression; introduce Laplace noise in each iteration of the K-means clustering algorithm to protect the clustering center coordinates.
[0109] Exemplarily, during the training process of the collaborative learning model, the sub-models of each participant only process the double-encrypted data of their own parties and generate intermediate results. This process involves multiple complex technical steps, aiming to ensure data privacy and model security.
[0110] First, use the CKKS approximate homomorphic encryption scheme to encrypt the original data of the participants. The polynomial modulus degree of the CKKS scheme is set to 2^15, the plaintext modulus is 2^40, and the ciphertext modulus is approximately 2^440. The selection of these parameters ensures a security level greater than 128 bits. For example, for a medical data set containing a patient's age, blood pressure, and blood sugar level, each value will be converted into an encrypted polynomial.
[0111] In each participant's local sub-model, batch encoding is performed on the encrypted data. This step packs multiple real numbers into a polynomial, greatly improving the computational efficiency. For example, the age data of 10 patients may be packed into a polynomial. For non-linear activation functions such as ReLU or Sigmoid, polynomial approximation using Taylor expansion is used for calculation. This allows approximate non-linear operations to be performed in the encrypted domain.
[0112] Next, forward propagation and backpropagation calculations are performed on the encrypted data. In forward propagation, the encrypted input data passes through the layers of the model to generate prediction results. In backpropagation, the gradients of the loss function with respect to the parameters of each layer are calculated. All these operations are performed in the encrypted state, and finally, the encrypted model parameter gradients and the encrypted local loss function values are generated.
[0113] To securely aggregate these encrypted intermediate results, the (T,n) threshold Paillier encryption scheme is introduced. Here, T is the reconstruction threshold and n is the number of participants. For example, in the case of 5 participants, T may be set to 3, meaning that at least 3 participants need to cooperate to decrypt the data. The private key is split into n shares using the Shamir secret sharing scheme and distributed to each participant. Each participant gets a private key fragment and cannot decrypt the data alone.
[0114] Each participant uses the public key to encrypt its model parameter gradients and local loss function values. Before encryption, calibrated Gaussian noise is added to achieve differential privacy. The variance of the noise is calculated based on the preset privacy budget and sensitivity. For example, if the privacy budget ε is set to 0.1 and the sensitivity is 1, Gaussian noise with a standard deviation of approximately 10 may be added.
[0115] In the aggregation phase, one participant is selected as the aggregator. The other participants send their encrypted intermediate results to the aggregator. The aggregator calculates the sum of the encrypted gradients and encrypted losses, but due to the nature of homomorphic encryption, the aggregator cannot see the actual values.
[0116] At least T parties (3 in the above example) cooperate to partially decrypt the encrypted total gradient and total loss using their respective private key shares. The aggregator collects these partial decryption results and completes the final decryption to obtain the global gradient and global loss. This process ensures that the aggregated result can only be obtained when a sufficient number of parties cooperate.
[0117] Update the global model parameters using the decrypted global gradient. This updated global model is used for privacy-preserving analysis and processing of newly input data to be analyzed.
[0118] In the analysis and processing stage, first use the SimHash algorithm for feature extraction. SimHash can map high-dimensional feature vectors to a low-dimensional binary space, greatly reducing storage and computational requirements. For example, a 1000-dimensional patient feature vector may be compressed into a 64-bit binary string.
[0119] For classification tasks, use the exponential mechanism to select the best splitting feature to construct a privacy-preserving decision tree ensemble. The exponential mechanism randomly selects the splitting feature based on the information gain of the feature and the privacy budget, ensuring both model performance and data privacy.
[0120] In regression tasks, use inner product functional encryption to achieve privacy-preserving linear regression. This allows the inner product of the feature vector and the weight vector to be calculated in the encrypted state, enabling prediction without exposing the original data.
[0121] For clustering tasks, introduce Laplace noise in each iteration of the K-means algorithm to protect the coordinates of the cluster centers. For example, for two-dimensional data, Laplace noise with a mean of 0 and a scale parameter of 1 may be added to the x and y coordinates of each cluster center respectively.
[0122] This entire process forms a closed loop: from the encryption of the original data, to the training of the local model, then to secure aggregation, and finally to the update and application of the global model. Each step builds on the previous step and provides the necessary input for the next step. By comprehensively applying technologies such as homomorphic encryption, secret sharing, and differential privacy, this method achieves efficient collaborative learning and data analysis while protecting data privacy.
[0123] In an alternative embodiment,
[0124] During the training process of the collaborative learning model, periodically update the dynamic encryption keys of each party based on the Bloom filter and the time decay function, and re-encrypt the original data of each party, including:
[0125] Set up a Bloom filter, whose size m and the number of hash functions k are calculated based on the expected number of elements N and the false positive probability p; use the exponential decay function w(t) = exp(-λt) to calculate the key update weight, where λ is the decay rate and t is the key usage time;
[0126] For each encryption operation, calculate its characteristic hash and insert it into the Bloom filter, and periodically check the filling rate of the Bloom filter; when the filling rate exceeds the preset filling threshold, trigger key update, generate a new key, and use the exponential decay function to calculate the mixing weight, mix the new and old keys to obtain the final key, and re-encrypt the original data of each party using the updated dynamic encryption key.
[0127] Exemplarily, during the training process of a collaborative learning model, in order to ensure data security and model performance, it is necessary to periodically update the dynamic encryption keys of each party. This process is implemented based on a Bloom filter and a time decay function, and specifically involves multiple technical steps, forming a closed-loop key management mechanism.
[0128] First, set up a Bloom filter. A Bloom filter is a space-efficient probabilistic data structure used to determine whether an element belongs to a set. In this solution, the Bloom filter is used to track the usage of encryption operations. The size m of the Bloom filter and the number of hash functions k need to be calculated based on the expected number of elements N and the acceptable false positive probability p. For example, assume there are 100,000 encryption operations expected within an hour and the desired false positive probability is no more than 0.1%, then the size of the Bloom filter can be calculated to be approximately 1,437,759 bits and 10 hash functions are needed.
[0129] Next, introduce a time decay function to calculate the key update weight. Here, the exponential decay function is used, where λ is the decay rate and t is the key usage time. For example, if λ = 0.1 is set, then when the key has been used for 10 time units, its weight will drop to approximately 36.8% of the initial value. This function ensures that newly used keys have a higher weight, while the weight of keys that have not been used for a long time will gradually decrease.
[0130] During each encryption operation, the system calculates the characteristic hash of the operation and inserts it into the Bloom filter. The characteristic hash can include information such as the encryption timestamp, data type, user ID, etc. For example, for an encryption operation on patient records by user A at 10:00:00 on October 1, 2023, a hash value "2023100110000_userA_patientRecord" may be generated.
[0131] The system periodically checks the fill rate of the Bloom filter. The fill rate refers to the proportion of bits in the Bloom filter that are set to 1. When this fill rate exceeds a preset fill threshold, the key update process is triggered. For example, the fill threshold can be set to 75%, meaning that when more than 75% of the bits in the Bloom filter are set to 1, the key needs to be updated.
[0132] After triggering the key update, the system generates a new key. This new key may be generated by a cryptographically secure random number generator. For example, a 256-bit AES key may be generated.
[0133] Then, the mixing weight is calculated using the exponential decay function defined previously. This mixing weight determines the proportion of the new key and the old key in the final key. For example, if the old key has been used for 5 time units, its weight may be exp(-0.1*5) ≈ 0.61, and the weight of the new key is 1 - 0.61 = 0.39.
[0134] Based on these weights, the system mixes the new key and the old key to obtain the final updated dynamic encryption key. The mixing process may involve performing a bit-level XOR operation on the old key and the new key according to the weights.
[0135] Finally, the original data of each party is re-encrypted using this updated dynamic encryption key. This process may take a relatively long time, depending on the size of the data volume. For example, for a dataset containing 1 million patient records, re-encryption may take several hours.
[0136] This process forms a complete closed loop: from the setting of the Bloom filter, to the tracking of encryption operations, to the triggering of key updates, and finally to the generation and application of new keys. Each step is closely connected to the previous and subsequent steps, jointly constituting a dynamic and secure key management mechanism.
[0137] In this way, the system can adapt to changes in data usage patterns while maintaining high security. Frequently used data will trigger more frequent key updates, while less frequently used data may use relatively stable keys. This not only improves the security of the system but also optimizes the use of computing resources because it is not necessary to update keys for all data at the same frequency.
[0138] Overall, this dynamic key update mechanism based on the Bloom filter and the time decay function provides a flexible, efficient, and secure data protection solution for the collaborative learning model, effectively balancing security, performance, and resource utilization.
[0139] Figure 2This is a schematic structural diagram of the privacy data analysis system based on collaborative learning and dynamic encryption according to an embodiment of the present invention. As Figure 2 shown, the system includes:
[0140] A first unit, configured to receive the original data of multiple participating parties. The original data of each participating party is in an initial encrypted state, and the initial encrypted state is implemented using a homomorphic encryption algorithm; assign a unique participating party identifier to each participating party, where the participating party identifier includes a timestamp and a random hash value; generate a corresponding dynamic encryption key based on the participating party identifier, and the dynamic encryption key is generated using quantum key distribution technology; use the dynamic encryption key to perform secondary encryption on the original data of each participating party to obtain secondary encrypted data, and the secondary encryption adopts a hierarchical encryption strategy, and different encryption algorithms with different strengths are used for different sensitive levels of data;
[0141] A second unit, configured to input the secondary encrypted data into a preset collaborative learning model. The collaborative learning model includes multiple sub-models, and each sub-model corresponds to a participating party, and the sub-model adopts a federated learning architecture; use a secure computing protocol under homomorphic encryption to decrypt the collaborative learning model in combination with the dynamic encryption keys of each participating party to obtain a privacy protection analysis model that can be used by each participating party;
[0142] A third unit, configured to, when performing data analysis, use the privacy protection analysis model to process the input data to be analyzed, and ensure the correctness of the analysis process through verifiable computing technology; based on a predefined access control policy, only return the analysis results that the corresponding participating party has the right to access to the corresponding participating party, and the analysis results are processed by k-anonymization to further enhance privacy protection; regularly execute the model update and retraining process to adapt to the dynamic changes of the data distribution.
[0143] In a third aspect of the embodiments of the present invention,
[0144] An electronic device is provided, including:
[0145] A processor;
[0146] A memory for storing instructions executable by the processor;
[0147] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0148] In a fourth aspect of the embodiments of the present invention,
[0149] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0150] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0151] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A privacy data analysis method based on collaborative learning and dynamic encryption, characterized in that: include: Receive original data from multiple participants, where the original data of each participant is in an initial encryption state, and the initial encryption state is implemented using a homomorphic encryption algorithm; assign a unique participant identifier to each participant, and the participant identifier includes a timestamp and a random hash value; generate a corresponding dynamic encryption key based on the participant identifier, and the dynamic encryption key is generated using quantum key distribution technology; use the dynamic encryption key to perform secondary encryption on the original data of each participant to obtain secondary encrypted data, and the secondary encryption adopts a layered encryption strategy, and uses encryption algorithms of different strengths for different sensitivity levels of data; Inputting the secondary encrypted data into a preset collaborative learning model, wherein the collaborative learning model includes a plurality of sub-models, each sub-model corresponds to a participant, and the sub-model adopts a federated learning architecture; Using a secure computing protocol under homomorphic encryption and combining the dynamic encryption keys of each participant to decrypt the collaborative learning model, a privacy-preserving analysis model that can be used by each participant is obtained; When performing data analysis, the input data to be analyzed is processed using the privacy protection analysis model, and the correctness of the analysis process is ensured through verifiable computing technology; based on the predefined access control policy, only the analysis results that the corresponding participants are entitled to access are returned, and the analysis results are k-anonymized to further enhance privacy protection; Regularly perform model update and retraining processes to adapt to dynamic changes in data distribution; Generating a corresponding dynamic encryption key based on the participant identifier, wherein the dynamic encryption key is generated using quantum key distribution technology, includes: An entangled photon pair is generated by a periodically poled lithium niobate crystal through a spontaneous parametric down-conversion process, wherein the temperature of the periodically poled lithium niobate crystal is controlled at 163.5±0.1°C, and the wavelength of the entangled photon pair is 1550nm; one photon in the entangled photon pair is sent to a sender, and the other photon is sent to a receiver, wherein the photon transmission adopts a photonic crystal fiber with an optical loss rate not higher than 0.2dB / km; a quantum repeater is arranged every 50km on the transmission path of the photonic crystal fiber, wherein the quantum repeater adopts a rare earth-doped YSO crystal as a quantum storage medium, and the storage time is not less than 1ms; The sender and the receiver respectively use a superconducting nanowire single-photon detector to measure the quantum state of the received photons, wherein the quantum efficiency of the superconducting nanowire single-photon detector at a wavelength of 1550nm is not less than 93%, and the dark count rate is not higher than 1Hz; Based on the quantum state measurement result, a quantum key is generated through a post-processing step, which includes basis comparison, parameter estimation, information coordination and privacy amplification, wherein information coordination adopts a rate-adaptive low-density parity-check code, and privacy amplification adopts a privacy matrix with dynamically adjustable size.
2. The method according to claim 1, characterized in that The original data of each participant is re-encrypted using the dynamic encryption key to obtain re-encrypted data. The re-encryption adopts a layered encryption strategy, and encryption algorithms of different strengths are adopted for different sensitivity levels of data, including: A data sensitivity assessment model based on deep learning is used to perform sensitivity assessment on the encrypted data, wherein the data sensitivity assessment model based on deep learning includes 12 base layers, 2 fully connected layers, 1 Dropout layer and 1 output layer, wherein the dropout rate of the Dropout layer is 0.1, and the output layer uses a Sigmoid activation function; Based on the quantum key, three sub-keys are generated using a key derivation function, which are used for data encryption at three sensitivity levels: high, medium and low; For high-sensitivity data with a sensitivity score greater than 0.8, the Kyber1024 algorithm is used for key encapsulation, and AES-256-GCM is used for authenticated encryption; for medium-sensitivity data with a sensitivity score between 0.4 and 0.8, the NTRU-HRSS-701 parameter set is used for key generation and encapsulation, and ChaCha20-Poly1305 is used for authenticated encryption; for low-sensitivity data with a sensitivity score less than 0.4, the derived subkey is directly used, and XChaCha20-Poly1305 is used for authenticated encryption.
3. The method according to claim 1, characterized in that Before inputting the secondary encrypted data into the preset collaborative learning model, the method further includes training the collaborative learning model: During the training process of the collaborative learning model, each sub-model only processes the secondary encrypted data of the participant and generates an intermediate result, which includes the model parameter gradient and the local loss function value; a secure multi-party computing protocol is used to collect the intermediate results generated by each sub-model, and the intermediate results are aggregated to form a global model, and the aggregation process adopts differential privacy technology to prevent the reverse deduction of the original data; the global model is used to perform analysis and processing on the newly input data to be analyzed under privacy protection, and the analysis and processing includes feature extraction, classification, regression and clustering; during the training process of the collaborative learning model, the dynamic encryption key of each participant is periodically updated based on the Bloom filter and the time decay function, and the original data of each participant is re-encrypted using the updated dynamic encryption key; Monitoring the training process of the collaborative learning model, and triggering a model verification procedure when it is detected that the model performance reaches a preset threshold or a potential model poisoning attack is detected; The model verification procedure includes: verifying the integrity of the model using zero-knowledge proof technology, evaluating the robustness of the model using adversarial sample testing, and ensuring the fairness of the model through multi-party consistency checks; According to the model verification results, the collaborative learning strategy is dynamically adjusted, including updating the weights of participants, adjusting the learning rate, and modifying the model structure; when the model verification passes and the performance meets the requirements, the training process is terminated and the final collaborative learning model is output.
4. The method according to claim 3, characterized in that During the training process of the collaborative learning model, each sub-model processes only the secondary encrypted data of the participant and generates intermediate results, which include model parameter gradients and local loss function values; a secure multi-party computing protocol is used to collect the intermediate results generated by each sub-model, and the intermediate results are aggregated to form a global model. The aggregation process uses differential privacy technology to prevent reverse derivation of original data; The global model is used to perform analysis and processing of the newly input data to be analyzed under privacy protection, wherein the analysis and processing includes feature extraction, classification, regression and clustering, including: The original data of the participants is encrypted using the CKKS approximate homomorphic encryption scheme, wherein the polynomial modulus of the CKKS approximate homomorphic encryption scheme is 2^15, the plaintext modulus is 2^40, the ciphertext modulus is 2^440, and the security level is greater than 128 bits; in the local sub-model of each participant, the encrypted data is batch encoded, multiple real numbers are packaged into a polynomial, and the nonlinear activation function is approximately calculated using the Taylor expansion polynomial; forward propagation and back propagation calculations are performed on the encrypted data to generate encrypted model parameter gradients and encrypted local loss function values; Generate a public key and a private key using the (T,n) threshold Paillier encryption scheme, where T is the reconstruction threshold and n is the number of participants. Use the Shamir secret sharing scheme to split the private key into n parts and distribute them to each participant. Each participant uses the public key to encrypt its model parameter gradient and local loss function value, and adds calibrated Gaussian noise before encryption. The noise variance σ² of the calibrated Gaussian noise is calculated based on the preset privacy budget ε and sensitivity Δf. One participant is selected as an aggregator, and the other participants send the encrypted intermediate results to the aggregator, and the aggregator calculates the sum of the encrypted gradient and the encrypted loss; at least n participants cooperate to partially decrypt the encrypted total gradient and total loss using their respective private key shares, and the aggregator collects the partial decryption results and completes the final decryption to obtain the global gradient and global loss; The global model parameters are updated using the global gradient, and the global model is used to perform analysis and processing of newly input data to be analyzed under privacy protection; in the analysis and processing, the SimHash algorithm is used to extract features and map high-dimensional feature vectors to low-dimensional binary space; an exponential mechanism is used to select the best splitting features to construct a privacy-preserving decision tree ensemble; inner product function encryption is used to realize privacy-preserving linear regression; and Laplace noise is introduced in each iteration of the K-means clustering algorithm to protect the coordinates of the cluster center.
5. The method according to claim 4, characterized in that During the training process of the collaborative learning model, the dynamic encryption key of each participant is periodically updated based on the Bloom filter and the time decay function, and the original data of each participant is re-encrypted using the updated dynamic encryption key, including: Set a Bloom filter whose size m and number of hash functions k are calculated based on the expected number of elements N and the false positive probability p; use the exponential decay function w(t) = exp(-λt) to calculate the key update weight, where λ is the decay rate and t is the key usage time; For each encryption operation, its characteristic hash is calculated and inserted into the Bloom filter, and the filling rate of the Bloom filter is periodically checked; when the filling rate exceeds a preset filling threshold, a key update is triggered, a new key is generated, and the exponential decay function is used to calculate the mixing weight, the new and old keys are mixed to obtain the final key, and the updated dynamic encryption key is used to re-encrypt the original data of each participant.
6. A privacy data analysis system based on collaborative learning and dynamic encryption, used to implement the method described in any one of claims 1 to 5, characterized in that: include: The first unit is used to receive original data of multiple participants, where the original data of each participant is in an initial encryption state, and the initial encryption state is implemented by a homomorphic encryption algorithm; a unique participant identifier is assigned to each participant, and the participant identifier includes a timestamp and a random hash value; a corresponding dynamic encryption key is generated based on the participant identifier, and the dynamic encryption key is generated by a quantum key distribution technology; the original data of each participant is secondary encrypted by using the dynamic encryption key to obtain secondary encrypted data, and the secondary encryption adopts a layered encryption strategy, and encryption algorithms of different strengths are used for different sensitivity levels of data; A second unit is used to input the secondary encrypted data into a preset collaborative learning model, wherein the collaborative learning model includes a plurality of sub-models, each sub-model corresponds to a participant, and the sub-model adopts a federated learning architecture; Using a secure computing protocol under homomorphic encryption and combining the dynamic encryption keys of each participant to decrypt the collaborative learning model, a privacy-preserving analysis model that can be used by each participant is obtained; The third unit is used to process the input data to be analyzed by using the privacy protection analysis model when performing data analysis, and ensure the correctness of the analysis process by verifiable computing technology; based on the predefined access control policy, only the analysis results that the corresponding participants are entitled to access are returned, and the analysis results are k-anonymized to further enhance privacy protection; Perform model update and retraining processes regularly to adapt to dynamic changes in data distribution.
7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 5.
8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Data security gateway method and system based on edge privacy calculation
CN118921161A
Privacy protection method and system based on big data security and privacy calculation
CN119128960A