Genomics-oriented anti-quantum robust parameter aggregation federated learning method and device

By performing quantization and TFHE encryption processing within TEE and encrypted aggregation processing on the central server, the problems of privacy leakage, computational accuracy and insufficient anti-malignant attack capabilities in single-cell RNA sequencing data analysis are solved, and efficient, secure and robust model updates are achieved.

CN120146158AActive Publication Date: 2025-06-13HANGZHOU NUOWEI INFORMATION TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510631368.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Traditional federated learning has problems such as privacy leakage, insufficient computational accuracy and insufficient anti-malignant attack capabilities in single-cell RNA sequencing data analysis.

Method used

Using TFHE algorithm and enhanced robust aggregation algorithm, the client quantizes and encrypts the local model parameters within TEE, and the central server performs robust aggregation in the encryption domain to enhance the model's quantum robustness.

Benefits of technology

By keeping data and model updated in an encrypted state, the risk of privacy leakage is eliminated, computing accuracy and anti-malicious attack capabilities are enhanced, and zero exposure guarantees for sensitive data are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146158A_ABST
    Figure CN120146158A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a genomics-oriented anti-quantum robust parameter aggregation federated learning method and device. The method comprises the following steps: a client performs quantization processing and TFHE encryption processing on local model parameters in a TEE; the client exports a ciphertext of an encryption model parameter from the TEE and uploads the ciphertext to a central server; the central server carries out robust aggregation processing on the encryption model parameters in an encryption domain to obtain optimization model parameters and sends the optimization model parameters to the client; and the client performs decryption and inverse quantization processing on the optimization model parameters and then updates the local model. According to the technical scheme provided by the embodiment of the invention, the TEE for local calculation and the TFHE for parameter transmission and aggregation are combined, so that the original data and model updating can be kept encrypted in the whole process, and the risk of data leakage caused by inference attack is eliminated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of secure computing, and in particular, to a quantum-resistant robust parameter aggregation federated learning method and device for genomics. Background Art

[0002] Federated Learning (FL) is a distributed machine learning paradigm that allows multiple institutions to collaboratively train a shared model without exchanging raw data. This approach is particularly important in areas where strict data privacy protection is required, such as the field of genomics, where single-cell RNA sequencing (scRNA-seq) data contains sensitive genetic information. In traditional federated learning, each participating institution trains a local model on its private dataset and only shares model parameters (such as weights or gradients) with a central server. The central server aggregates these parameters to update the global model and then redistributes the updated model to the institutions for further training. This iterative process maintains data locality and reduces the risk of direct data exposure.

[0003] Despite the above advantages of federated learning, when applied to single-cell RNA sequencing (scRNA-seq) data analysis, traditional federated learning still has the following problems: Transmitting model parameters in plaintext poses a significant privacy threat. An attacker may use these parameters to reconstruct parts of the original data or infer sensitive information of individuals through inference attacks (such as membership inference attacks). This vulnerability undermines the privacy guarantees required for genomic data processing.

[0004] Existing encryption methods generally allow computations on encrypted data but are limited to approximate arithmetic. These schemes cannot support the exact comparison operations required by advanced aggregation algorithms, limiting their ability to ensure robust and accurate model updates in a secure manner.

[0005] Traditional federated learning systems usually assume that participants are benign. However, in real-world scenarios, malicious nodes (such as Byzantine adversaries) may submit corrupted updates to disrupt the training process or degrade the model performance. Standard aggregation methods lack mechanisms to detect or mitigate such attacks, and when a large number of nodes are compromised, the system reliability will decrease. Summary of the Invention

[0006] Based on the above situation of the prior art, the purpose of the embodiments of the present invention is to provide a quantum-resistant robust parameter aggregation federated learning method and device for genomics, which solves the problems of privacy, computational accuracy, and anti-adversarial interference existing in traditional federated learning through the TFHE algorithm and an enhanced robust aggregation algorithm.

[0007] To achieve the above object, according to the first aspect of the present invention, a quantum-resistant robust parameter aggregation federated learning method for genomics is provided, including the steps: The client quantizes and TFHE encrypts the local model parameters within the TEE; The client exports the ciphertext of the encrypted model parameters from the TEE and uploads it to the central server; The central server performs robust aggregation processing on the encrypted model parameters in the encrypted domain, obtains the optimized model parameters, and sends them to the client; After the client decrypts and dequantizes the optimized model parameters, it updates the local model; Among them, the local model parameters are obtained by the client training the local model within the TEE based on the local private dataset; the robust aggregation processing optimizes and updates the global model parameters based on the similarity between the encrypted model parameters of each client.

[0008] Further, the method further includes the steps: The client obtains the change value of each updated optimized model parameter relative to the model parameter in the previous round, and uploads the change value and the ciphertext of the encrypted model parameter in this round to the central server; The central server determines the importance of each parameter based on the change value, and truncates the encrypted model parameters of each client according to the importance.

[0009] Further, the local private dataset includes original single-cell RNA sequencing data; the method further includes the steps: Preprocess the original single-cell RNA sequencing data in the local private dataset.

[0010] Further, the preprocessing includes clearing zero-expressed genes in the data, converting the original single-cell RNA sequencing data to CPM, selecting highly expressed genes according to the converted data, and performing PCA dimensionality reduction processing on the selected genes.

[0011] Further, the robust aggregation processing includes: Calculating the Euclidean distance between the encrypted model parameter vectors of each client pairwise; Based on the calculated Euclidean distance, adjusting the encrypted model parameters participating in the global model; Based on the result of the adjustment, optimizing and updating the global model parameters.

[0012] Further, based on the calculated Euclidean distance, adjusting the encrypted model parameters participating in the global model includes: For each client, calculate the sum of the Euclidean distances between the encryption model parameters of this client and those of other clients, and use it as the distance set of this client; Among the distance sets of all clients, remove the client model parameters corresponding to the largest distance set.

[0013] Furthermore, optimize and update the global model parameters based on the adjustment result, including: After removing the client model parameters corresponding to the largest distance set, calculate the average value of the encrypted model parameter vectors of the remaining clients; Use this average value as the optimized update parameter of the global model.

[0014] Furthermore, the central server performs robust aggregation processing on the ciphertext by executing the homomorphic operation in the compiled TFHE calculation circuit.

[0015] Furthermore, the TFHE calculation circuit is pre-generated on any client according to the local model parameter vector of this client.

[0016] According to another aspect of the present invention, there is provided a quantum-resistant robust parameter aggregation federated learning device for genomics, including: Client parameter encryption module, used to perform quantization processing and TFHE encryption processing on local model parameters within the TEE; Client parameter upload module, used to export the ciphertext of the encrypted model parameters from the TEE and upload it to the central server; Central server aggregation module, used to perform robust aggregation processing on the encrypted model parameters in the encrypted domain, obtain optimized model parameters and send them to the client; Client parameter update module, used to decrypt and dequantize the optimized model parameters, and then update the local model; Wherein, the local model parameters are obtained by the client training the local model within the TEE based on the local private data set; the robust aggregation processing optimizes and updates the global model parameters based on the similarity between the encrypted model parameters of each client.

[0017] In summary, the embodiments of the present invention provide a quantum-resistant robust parameter aggregation federated learning method and apparatus based on genomics. The method includes the steps of: the client performs quantization processing and TFHE encryption processing on local model parameters within the TEE; the client exports the ciphertext of the encrypted model parameters from the TEE and uploads it to the central server; the central server performs robust aggregation processing on the encrypted model parameters in the encrypted domain to obtain optimized model parameters and sends them to the client; after the client decrypts and dequantizes the optimized model parameters, it updates the local model. The technical solution of the embodiments of the present invention can keep the original data and model updates encrypted throughout the process by combining the TEE for local computing and the TFHE method for parameter transmission and aggregation, eliminating the risk of data leakage caused by inference attacks and providing a zero-exposure guarantee for sensitive single-cell RNA sequencing (scRNA-seq) data; the enhanced Krum algorithm can effectively reduce the impact of malicious nodes by selecting updates that are consistent with most updates, and can maintain the integrity of the model even in the presence of adversarial interference. This robustness is crucial for real-world deployments where not all participants can be assumed to be trustworthy; the use of dynamic truncation reduces the computational overhead of TFHE operations while retaining sufficient accuracy required for model training and aggregation. Local training based on the TEE further optimizes performance by avoiding encryption during computationally intensive training phases, balancing security and efficiency. The technical solution provided by the embodiments of the present invention supports accurate cell type classification and other single-cell RNA sequencing (scRNA-seq) tasks across distributed datasets, enabling each client to collaborate without sacrificing data privacy or model quality and being compatible with various machine learning models. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of the quantum-resistant robust parameter aggregation federated learning method based on genomics provided by the embodiments of the present invention; Figure 2 is a flowchart of the quantum-resistant robust parameter aggregation federated learning method provided by another embodiment of the present invention; Figure 3 is a flowchart of the quantum-resistant robust parameter aggregation federated learning method provided by another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0020] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present invention should have the ordinary meaning understood by those of ordinary skill in the field to which the present invention pertains. The terms "first", "second" and similar words used in one or more embodiments of the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0021] For single-cell RNA sequencing (scRNA-seq) data, federated learning promotes cross-institutional collaboration by leveraging diverse datasets to enhance cell type classification and other analyses while complying with privacy regulations. Typically, model parameters are transmitted in plaintext for efficient computation and aggregation. Existing implementations often rely on basic aggregation methods (such as the averaging method) and may introduce encryption schemes to enhance security. However, this basic approach, while suitable for distributed training, has vulnerabilities in highly sensitive data scenarios, is difficult to resist data poisoning attacks, etc., limiting its application, and may also have issues such as privacy risks, encryption limitations, and vulnerability to malicious attacks.

[0022] Embodiments of the present invention provide a quantum-resistant robust parameter aggregation federated learning method for genomics. Figure 1 The flowchart of the federated learning method in an embodiment of the present invention is shown in Figure 1 As shown, the method includes the following steps: S11. The client performs quantization processing and TFHE encryption processing on the local model parameters within the TEE. The local model parameters are obtained by the client training a local model within the TEE based on the local private dataset. In this embodiment of the present invention, the local private dataset includes raw single-cell RNA sequencing data and also includes the step of preprocessing the raw single-cell RNA sequencing data of the local private dataset.

[0023] S12. After exporting the ciphertext of the encrypted model parameters from the TEE, the client uploads it to the central server.

[0024] S13. The central server performs robust aggregation processing on the encrypted model parameters in the encrypted domain to obtain optimized model parameters and sends them to the client. The robust aggregation processing optimizes and updates the global model parameters based on the similarity between the encrypted model parameters of each client.

[0025] After the client decrypts and dequantizes the optimized model parameters, it updates the local model.

[0026] An embodiment of the present invention also provides a quantum-resistant robust parameter aggregation federated learning device for genomics, which includes: A client parameter encryption module for quantizing and TFHE encrypting local model parameters within the TEE; A client parameter upload module for exporting the ciphertext of the encrypted model parameters from the TEE and uploading it to the central server; A central server aggregation module for robustly aggregating the encrypted model parameters in the encrypted domain to obtain optimized model parameters and sending them to the client; A client parameter update module for decrypting and dequantizing the optimized model parameters and then updating the local model; Among them, the local model parameters are obtained by the client training the local model within the TEE based on the local private dataset; the robust aggregation process optimizes and updates the global model parameters based on the similarity between the encrypted model parameters of each client.

[0027] The quantum-resistant robust parameter aggregation federated learning method and system involved in the above embodiments of the present invention complete the process of federated learning through the interaction of multiple clients and a central server. The technical solutions of the above embodiments will be described below from the perspectives of the client and the central server respectively.

[0028] An embodiment of the present invention also provides a quantum-resistant robust parameter aggregation federated learning method for genomics, which is applied to the client. Figure 2 The flowchart of the federated learning method in the embodiment of the present invention is shown in Figure 2 As shown, the method includes the following steps: S102. Train a local model within the TEE based on the local private dataset to obtain local model parameters. In the embodiments of the present invention, the local model training is carried out within a Trusted Execution Environment (TEE for short). The TEE is a hardware-isolated enclave that can ensure the confidentiality and integrity of the local model training process. Devices supporting TEE (such as devices supporting Intel SGX, AMD SEV, or equivalent technologies) can be deployed on each client participating in the federated learning. The TEE encrypts the memory and isolates the computing process from the host operating system, preventing unauthorized access even if the device is compromised. The key management for subsequent encryption steps can also be securely processed within the TEE. Within the TEE, preprocess the original single-cell RNA sequencing (scRNA-seq) data of the local private dataset, including removing zero-expression genes from the data, converting the original single-cell RNA sequencing (scRNA-seq) data to CPM (Counts Per Million), selecting highly expressed genes based on the converted data, and performing PCA dimensionality reduction on the selected genes. Single-cell data usually has a large number of zero values, which may be caused by technical noise or true biological non-expression. Removing zero-expression genes generally means filtering out genes that are expressed as zero in most cells (for example, retaining genes expressed in at least 10% of the cells), thereby reducing noise and lowering the data dimension. By normalizing the original data to CPM, the differences in sequencing depth can be eliminated, making the expression levels between different cells comparable. Highly expressed or highly variable genes can be selected based on the mean or variance of gene expression levels. Perform PCA dimensionality reduction on the selected genes to extract the main variation directions, which can reduce redundant information and accelerate the subsequent analysis process. After data preprocessing, train the local model using plaintext data. The local model is, for example, a neural network or a support vector machine for cell type classification.

[0029] S104. Quantize the local model parameters and perform TFHE processing within the TEE. TFHE (Tiny Fully Homomorphic Encryption) is an improved method of homomorphic encryption. After the model training is completed, the local model parameters are extracted from the model within the TEE. At this time, the local model parameters are in plaintext floating-point form. To make subsequent TFHE and homomorphic calculations more efficient and feasible, while retaining sufficient precision, the extracted floating-point local model parameters are quantized within the TEE. The quantization process can map floating-point numbers to a discrete integer range with a fixed bit width (e.g., 8 bits, 16 bits). In the embodiments of the present invention, the quantization process includes pre-quantization processing and post-quantization processing. Among them, the pre-quantization processing is used to determine the range of the model parameters. The relevant parameters of each client participating in federated learning can be collected in advance, including the maximum value and the minimum value of the parameters, and the range of the model parameters is determined through the maximum value and the minimum value. The post-quantization processing is used to map the floating-point model parameters to quantized integers, and the quantized model parameters are represented in fixed-point or integer form.

[0030] According to some optional embodiments, the method further includes the steps of: S1041. Obtain the change value of each updated optimized model parameter relative to the model parameter in the previous round, and upload the change value and the ciphertext of the encrypted model parameter in this round to the central server, so that the central server determines the importance of each parameter based on the change value, and truncates the encrypted model parameters of each client according to the importance. In this embodiment of the present invention, a dynamic truncation method is used for processing. For the local model parameters of each client, calculate its importance score. Multiple methods can be used for scoring, such as calculating the gradient norm of the parameter, the absolute value size of the parameter, etc. In an embodiment of the present invention, the average value of the absolute value of the parameter change is used as the importance scoring index. Specifically, in each round of local training of the client, for a local model, a number of local model parameters will be generated. For each model parameter , where i represents the client number and j represents the model parameter index, calculate the absolute value of the change value of the model parameter in the current round and the previous round , and then calculate the average value of the absolute values of all change values of the model parameter in each round , as the index of the importance score of the model parameter in this round, and upload the average value to the central server in encrypted or plaintext form, so that the central server determines the importance of each parameter based on the change value. The larger the value of the average value, the higher the importance of the corresponding model parameter.

[0031] The quantized model parameters are encrypted using the TFHE algorithm within the TEE. In the embodiments of the present invention, a TFHE library that supports exact integer arithmetic and comparison is preferably used, such as the TFHE library provided by Zama (e.g., Concrete). The client uses the private key to encrypt each quantized parameter (or block of parameters) to generate the corresponding TFHE ciphertext. TFHE supports exact arithmetic and comparison operations on encrypted data, overcoming the limitations of approximate encryption schemes. After training, the local model parameters are quantized into a discrete format within the TEE and then encrypted into TFHE ciphertext. In the embodiments of the present invention, the quantized model parameters are encrypted using the TFHE algorithm within the TEE in the client, solving the problem that the existing encryption methods (such as homomorphic encryption, fully homomorphic encryption not based on TFHE, etc.) cannot perform complex aggregation operations in the ciphertext state. For example, homomorphic encryption cannot perform addition and multiplication operations simultaneously, and fully homomorphic encryption not based on TFHE cannot perform complex calculation operations (such as CKKS).

[0032] S106. After exporting the ciphertext of the encrypted model parameters and the absolute value of the model parameter change value from the TEE, upload them to the central server so that the central server can perform robust aggregation processing on the encrypted model parameters in the encrypted domain. The ciphertext of the encrypted model parameters is securely exported from the TEE and uploaded to the central server through a secure channel. At this time, even if the ciphertext is intercepted, the parameter information of the original local model cannot be directly obtained.

[0033] S108. Receive the optimized model parameters sent by the central server. After the central server completes the encrypted aggregation, it sends the optimized model parameters, that is, the encrypted global model update ciphertext (or the selected client update ciphertext), back to the relevant client through a secure channel. After receiving the ciphertext within the TEE, the client performs the following steps to update its local model: S110. Decrypt and dequantize the optimized model parameters to obtain the updated optimized model parameters. Within the TEE, use the TFHE private key held locally by the client to decrypt the received ciphertext. The decryption operation is performed in the secure environment of the TEE, which can ensure that the private key is not leaked. The result of decryption is the quantized integer parameter. For the quantized integer parameter obtained by decryption, perform a dequantization algorithm within the TEE to convert the quantized integer parameter back to an approximate original floating-point representation. Dequantization is the inverse process of quantization.

[0034] S112. Update the local model using the updated optimized model parameters. Use the floating-point parameters obtained after dequantization to update the local model of the client (e.g., perform operations such as replacing or weighted averaging of model parameters).

[0035] An embodiment of the present invention also provides a quantum-resistant robust parameter aggregation federated learning method for genomics, which is applied to a central server. Figure 3 The flowchart of the federated learning method according to the embodiment of the present invention is shown in Figure 3 As shown, the method includes the following steps: S202. Receive the ciphertext of the encrypted model parameters uploaded by the client. The encrypted model parameters are obtained by the client locally quantifying and performing TFHE processing on the local model parameters. The central server receives the ciphertext of the encrypted model parameters uploaded by all participating federated learning clients and executes an enhanced Krum robust aggregation algorithm in the encrypted domain. The central server does not need to hold the encrypted private key, and all calculations in the central server are homomorphically performed on the ciphertext.

[0036] S204. Perform robust aggregation processing on the encrypted model parameters in the encrypted domain to obtain optimized model parameters. In the embodiment of the present invention, the robust aggregation processing adopts an enhanced Krum aggregation algorithm to optimize and update the global model parameters based on the similarity between the encrypted model parameters of each client. The enhanced Krum aggregation algorithm includes a series of calculation steps, including calculating the Euclidean distance between model update vectors, sorting the distances, selecting the nearest neighbors, calculating the sum of the distances, and finding the client with the smallest sum. These steps involve basic arithmetic operations (subtraction, multiplication, addition) and comparison operations. In the embodiment of the present invention, these calculation logics are defined as a calculation graph or program, and using a compiler tool (such as Concrete) in the ZamaTFHE ecosystem, this calculation graph is compiled into a TFHE calculation circuit suitable for efficient execution on TFHE ciphertext. This compilation process is part of the algorithm definition phase and can be completed offline before system deployment, or completed by a coordinator or other trusted role at the beginning of the federated learning session and provide the compiled circuit to the central server. The compiled circuit defines a series of homomorphic gate operations (such as homomorphic addition gates, homomorphic multiplication gates, etc.) and noise reduction (Bootstrapping) operations that the server needs to execute in sequence. The central server loads the ciphertext of the encrypted model parameters received from the client and the pre-compiled TFHE calculation circuit. The central server uses the loaded ciphertext of the encrypted model parameters as input and starts the execution of the TFHE calculation circuit. The entire calculation process is completely performed in the encrypted domain, and the central server processes the ciphertext by executing the homomorphic operations in the compiled TFHE calculation circuit. The TFHE calculation circuit can be pre-generated on any client: input the local model vector of the client into the TFHE toolkit, input the required logical operations into the TFHE toolkit at the same time, and generate a TFHE calculation circuit capable of calculating functions such as Euclidean distance and comparing sizes in the encrypted state through the TFHE toolkit. The robust aggregation processing executed by the central server through the TFHE calculation circuit includes the following steps: S2041. Calculate the Euclidean distance between the encrypted model parameter vectors of each pair of clients. The encrypted model parameters uploaded by each client form an encrypted model parameter vector. In this step, calculate the Euclidean distance between the encrypted model parameter vectors uploaded by each pair of clients.

[0037] S2042. Based on the calculated Euclidean distance, adjust the encrypted model parameters participating in the global model. For each client, calculate the sum of the Euclidean distances between the encrypted model parameters of this client and those of other clients as the distance set of this client. Among the distance sets of all clients, remove the client model parameters corresponding to the largest distance set.

[0038] S2043. Select the encrypted model parameters with the smallest sum to update the parameters of the global model. After removing the client model parameters corresponding to the largest distance set, calculate the average value of the encrypted model parameter vectors of the remaining clients, and use this average value as the optimized update parameter of the global model.

[0039] According to some optional embodiments, before performing the robust aggregation process, it further includes determining the importance degree of each model parameter based on the change value of the model parameters uploaded by the client, and performing truncation processing on the encrypted model parameters of each client according to this importance degree to improve the consistency and efficiency of data processing. Specifically, integrate the absolute values of the change values of the model parameters uploaded by each client to obtain the global parameter importance score, and simple average method or weighted average method, etc. can be used for integration. In this embodiment, the simple average method is adopted, that is, for each model parameter index j, calculate the average value of the importance scores of this model parameter of all clients: ; where N represents the number of clients, represents the importance score of the j-th parameter of the i-th client. This importance score can be obtained by the central server sorting according to the absolute value of the change value of the model parameters uploaded by the client.

[0040] According to the global parameter importance score, sort all parameters. Ascending or descending order can be used. In this embodiment, ascending order is adopted, that is, the parameters with smaller importance scores are arranged in the front. According to the sorting result, allocate different truncation bit numbers to different parameters. Specifically, a set of selectable truncation bit number sets can be preset in advance, where ; then according to the sorting order of the parameters, assign the elements in the quantization bit number set to each parameter in turn. For example, if there are M parameters and the truncation bit number set has n elements, then the A quantization bit number is allocated to the k-th parameter (k starts counting from 0). The central server receives the truncated parameters sent by each client, and performs truncation according to the data with different truncation bits during the robust aggregation process to obtain the truncated model parameters.

[0041] S206. Send the optimized update parameter to the client, so that after the client decrypts and dequantizes the optimized update parameter, the local model is updated.

[0042] By adopting the method based on the similarity of model vectors in the above steps, the model parameters uploaded by malicious nodes can be identified and removed, thus ensuring the reliability of the federated learning system.

[0043] In the above embodiments, the processes of client processing and central server aggregation constitute a training iteration of federated learning. By repeatedly executing these iteration processes, the global model is gradually optimized until the predetermined convergence condition or the maximum number of iterations is reached. During the whole process, the plaintext forms of the original data and model parameters only exist in the TEE of the client, and the central server always processes encrypted data, thus greatly ensuring the security of the data.

[0044] An embodiment of the present invention also provides a quantum-resistant robust parameter aggregation federated learning device for genomics, which is applied to a client. The device includes: A model training module, configured to train a local model in the TEE based on a local private data set to obtain local model parameters; A parameter processing module, configured to perform quantization processing and TFHE processing on the local model parameters in the TEE; A parameter uploading module, configured to export the ciphertext of the encrypted model parameters from the TEE and upload it to the central server, so that the central server performs robust aggregation processing on the encrypted model parameters in the encrypted domain.

[0045] An embodiment of the present invention also provides a quantum-resistant robust parameter aggregation federated learning device for genomics, which is applied to a central server. The device includes: A parameter receiving module, configured to receive the ciphertext of the encrypted model parameters uploaded by the client, and the encrypted model parameters are obtained by the client performing quantization and TFHE processing on the local model parameters locally; A parameter aggregation module, configured to perform robust aggregation processing on the encrypted model parameters in the encrypted domain to obtain optimized model parameters.

[0046] In the device for quantum-resistant robust parameter aggregation federated learning for genomics provided in each of the above embodiments of the present invention, the specific processes for each module to implement its functions are the same as the steps of the method for quantum-resistant robust parameter aggregation federated learning for genomics provided in each of the above embodiments of the present invention, and the repeated description thereof will be omitted here.

[0047] In addition, an embodiment of the present invention may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the method for quantum-resistant robust parameter aggregation federated learning for genomics in each of the embodiments of the present invention.

[0048] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0049] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0050] In summary, the embodiments of the present invention relate to a quantum-resistant robust parameter aggregation federated learning method and device based on genomics. The method includes the steps: the client performs quantization processing and TFHE encryption processing on local model parameters within the TEE; the client exports the ciphertext of the encrypted model parameters from the TEE and uploads it to the central server; the central server performs robust aggregation processing on the encrypted model parameters in the encrypted domain to obtain optimized model parameters and sends them to the client; the client decrypts and dequantizes the optimized model parameters and then updates the local model. The technical solution of the embodiments of the present invention can keep the original data and model updates encrypted throughout the process by combining the TEE for local computing and the TFHE method for parameter transmission and aggregation, eliminating the risk of data leakage caused by inference attacks and providing a zero-exposure guarantee for sensitive single-cell RNA sequencing (scRNA-seq) data; the enhanced Krum algorithm can effectively reduce the impact of malicious nodes by selecting updates that are consistent with most updates, and can maintain the integrity of the model even in the presence of adversarial interference. This robustness is crucial for realistic deployments where not all participants can be assumed to be trustworthy; the use of dynamic truncation reduces the computational overhead of TFHE operations while retaining sufficient accuracy required for model training and aggregation. Local training based on the TEE further optimizes performance by avoiding encryption during computationally intensive training phases, balancing security and efficiency; the technical solution provided by the embodiments of the present invention supports accurate cell type classification and other single-cell RNA sequencing (scRNA-seq) tasks across distributed datasets, enabling each client to collaborate without sacrificing data privacy or model quality and being compatible with various machine learning models.

[0051] It should be understood that the discussion of any above embodiments is only exemplary and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples; within the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present invention as described above, which are not provided in detail for the sake of brevity. The above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principles of the present invention and do not constitute a limitation to the present invention. Therefore, any modifications, equivalent replacements, improvements, etc. made without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all variations and modification examples that fall within the scope and boundaries of the appended claims or the equivalent forms of such scope and boundaries.

Claims

1. A quantum-resistant robust parameter aggregation federated learning method for genomics, characterized in that: Includes steps: The client quantizes and encrypts the local model parameters in TEE; The client exports the ciphertext of the encrypted model parameters from the TEE and uploads it to the central server; The central server performs robust aggregation processing on the encrypted model parameters in the encryption domain to obtain optimized model parameters and sends them to the client; The client updates the local model after decrypting and dequantizing the optimization model parameters; Among them, the local model parameters are obtained by the client training the local model in TEE based on the local private data set; The robust aggregation process optimizes and updates the global model parameters based on the similarity between the encryption model parameters of each client.

2. The method according to claim 1, characterized in that: The method further comprises the steps of: The client obtains the change value of each updated optimization model parameter relative to the model parameter of the previous round, and uploads the change value and the ciphertext of the encrypted model parameter of this round to the central server; The central server determines the importance of each parameter based on the change value, and performs truncation processing on the encryption model parameters of each client according to the importance.

3. The method according to claim 1, characterized in that The local private data set includes original single-cell RNA sequencing data; the method further includes the steps of: Preprocess the raw single-cell RNA sequencing data of the local private dataset.

4. The method according to claim 3, characterized in that The preprocessing includes removing zero-expression genes in the data, converting the original single-cell RNA sequencing data into CPM, selecting high-expression genes according to the converted data, and performing PCA dimensionality reduction processing on the selected genes.

5. The method according to any one of claims 1 to 4, characterized in that: The robust aggregation process includes: Calculate the Euclidean distance between each client's encryption model parameter vector; Based on the calculated Euclidean distance, the parameters of the encrypted models participating in the global model are adjusted; The global model parameters are optimized and updated based on the adjustment results.

6. The method according to claim 5, characterized in that Based on the calculated Euclidean distance, the encryption model parameters participating in the global model are adjusted, including: For each client, the sum of the Euclidean distances between the client and the encryption model parameters of other clients is calculated as the distance set of the client; Among the distance sets of all clients, the client model parameters corresponding to the largest distance set are eliminated.

7. The method according to claim 6, characterized in that Optimizing and updating the global model parameters based on the adjustment results includes: After removing the client model parameters corresponding to the largest distance set, the average value of the encrypted model parameter vectors of the remaining clients is calculated; The average value is used as the optimization update parameter of the global model.

8. The method according to claim 7, characterized in that The central server performs robust aggregation on the ciphertext by executing homomorphic operations in the compiled TFHE computation circuit.

9. The method according to claim 8, characterized in that The TFHE calculation circuit is pre-generated at any client according to the local model parameter vector of the client.

10. A quantum-resistant robust parameter aggregation federated learning device for genomics, characterized in that: include: The client parameter encryption module is used to quantize and encrypt local model parameters in TEE. A client parameter uploading module, used to export the ciphertext of the encryption model parameters from the TEE and upload it to the central server; A central server aggregation module, used for performing robust aggregation processing on the encrypted model parameters in the encryption domain, obtaining optimized model parameters and sending them to the client; A client parameter updating module, used to update the local model after decrypting and dequantizing the optimization model parameters; Among them, the local model parameters are obtained by the client training the local model in TEE based on the local private data set; The robust aggregation process optimizes and updates the global model parameters based on the similarity between the encryption model parameters of each client.

Citation Information

Patent Citations

  • TEE environment-oriented low-overhead federated learning method and system

    CN116881965A

  • Federal learning aggregation method based on function encryption on lattice

    CN118428487A

  • Quantum security federated learning method based on homomorphic encryption and electronic equipment

    CN118449676A

  • Data multi-level coding method and device based on fully homomorphic encryption, equipment and medium

    CN118536130A

  • Non-intrusive load monitoring method based on hybrid federated learning and cross Transform

    CN119962582A