Quantum security data processing method and system based on federal learning
By combining federated learning with quantum technology, using the key center distribution, decryption key files and the service center for multiple rounds of federated learning parameter iteration, the problems of data privacy and security in federated learning are solved, and efficient and secure data transmission and accurate machine learning models are achieved.
Patent Information
- Application Number
- CN202510279417.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In federated learning, the existing technology is difficult to effectively protect the privacy of user data, especially facing the risks of poisoning attacks, counterattacks and backdoor attacks. At the same time, QKD's sensitive transmission over long distances and environmental interference limits its application.
By combining federated learning with quantum technology, using the key center to distribute the encrypted and decrypted key files, the service center conducts multiple rounds of federated learning parameter iterations, and sets up the encrypted and decrypted key pool in the client and service center to ensure the security of data transmission.
It realizes that while protecting the privacy of user data, it uses client data to perform machine learning, obtains an accurate learning model, improves the security and efficiency of data transmission, and solves the problem of low distance and coding rate of QKD in practical applications.
Smart Images

Figure CN120217410A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information security technology, and in particular, to a quantum-secure data processing method and system based on federated learning. Background Art
[0002] In daily life and work, numerous powerful intelligent devices play an increasingly important role. These intelligent devices such as computers, smartphones, tablets, smartwatches, etc. not only provide us with a lot of convenience but also collect a large amount of data. On the other hand, artificial intelligence technology is also developing rapidly. As one of the most important technologies, deep learning technology requires extremely large amounts of data as a basis. The data on these intelligent devices is undoubtedly very attractive. However, the data in these intelligent devices often involves users' privacy information, and users do not allow this privacy information to leave the intelligent devices and be provided for training by traditional machine learning methods.
[0003] In view of the requirement that data does not leave the local device, Federated Learning (FL) has emerged. Federated learning is a distributed machine learning technology that does not collect users' data but allows this data to remain locally, enabling user devices to train machine learning models in place and upload the trained models to the server. This method well protects the security of local private data. Although federated learning aims to protect user data privacy, it also faces the risk of privacy leakage. The main security threats include poisoning attacks, adversarial attacks, and backdoor attacks. Poisoning attacks affect model training by injecting malicious data; adversarial attacks interfere with the model by generating adversarial samples, while backdoor attacks implant backdoors in the model to make it behave abnormally under specific inputs. These attacks not only affect the accuracy of the model but may also be used for malicious purposes.
[0004] To ensure the security of data transmission, currently, the method of QKD (quantum key distribution) is often adopted. However, the communication distance using the QKD method is affected by the survival situation of photons during the transmission process and is difficult to play a role in long-distance transmission. Moreover, the QKD system is very sensitive to environmental interference, and any minor interference may cause the key generation to fail. All these situations greatly limit the application of QKD in reality.
[0005] In view of this, how to ensure the security of global data in machine learning using federated learning in practical applications is an important issue faced by the current deep learning field. Here, global data security includes the data security of each device participating in machine learning locally and the security of federated learning parameters during transmission. Summary of the Invention
[0006] Objective of the Invention: To solve the related technical problems raised in the background art, the present invention provides a quantum-secure data processing method and system based on federated learning, which combines federated learning with quantum technology to protect the data privacy of each client while achieving the effect of using the data of the client for machine learning to obtain an accurate learning model.
[0007] Technical Solution: A quantum-secure data processing method based on federated learning according to the present invention includes the following steps:
[0008] (1) The key center distributes a set of encryption and decryption key files to the client cluster and the service center respectively; then, the service center sets the number of iterations T of federated learning, and according to the number of iterations T, the service center randomly selects T sub-client clusters from the client cluster; among them, the sub-client cluster is denoted as The number of clients in each sub-client cluster is N;
[0009] (2) The service center selects the sub-client cluster S 1 Execute the parameter iteration of the first round of federated learning, and after completion, select the sub-client cluster S 2 Execute the parameter iteration of the second round of federated learning, and so on, until the sub-client cluster S T Execute the parameter iteration of the T-th round of federated learning, and finally complete the parameter iteration of T rounds of federated learning;
[0010] (3) The service center sends the parameter results of the T-round federated learning parameter iteration to each client in the client cluster, and each client uses the parameter results as the parameters of the local federated learning model to perform predictions on the input data received by the client.
[0011] Furthermore, the encryption and decryption key files in the client cluster and the service center are the same.
[0012] Furthermore, the specific process of step (2) is as follows:
[0013] 1) The key center assigns a symmetric key to each pair of clients in the sub-client cluster S 1 ;
[0014] 2) Then, each client in the sub-client cluster S 1 encodes the local data set into a data set in a quantum state;
[0015] 3) The service center randomly initializes the global model parameters representing an M-dimensional real number space, where M represents the size of the global model parameter; and then encrypts and distributes the global model parameter θ 0 to each client in the sub-client cluster S 1 ;
[0016] 4) Each client in the sub-client cluster S 1 decrypts, and based on the received initialized global model parameters θ 0 trains the local local model to obtain local model optimization update parameters, encrypts and feeds back the local model optimization update parameters to the service center;
[0017] 5) The service center receives and decrypts to obtain the local model optimization update parameters of all clients in the sub-client cluster S 1 performs weighted aggregation on these local model optimization update parameters to obtain the global model optimization update parameters Δθ of the first round 1 ; The service center processes the global model optimization update parameters of the first round to obtain the global model parameters θ of the first round 1 and sends the global model parameters θ 1 to each client in the sub-client cluster S 2 ;
[0018] 6) By analogy, the sub-client cluster S 2 to the sub-client cluster S T adopt the same method as the sub-client cluster S 1 to generate the corresponding local model optimization update parameters, and the service center processes the global model parameters of the corresponding round according to the local model optimization update parameters of each round; Finally, through the parameter iteration of T rounds of federated learning, the final global model parameters θ T are obtained.
[0019] Further, the specific process of the key center allocating symmetric keys to each pair of clients in the sub-client cluster S 1 is as follows:
[0020] A: Renumber all clients in the sub-client cluster S 1 to obtain the first to the Nth clients;
[0021] B: The key center selects the first client from the sub-client cluster S 1 and counts the total number N of clients in the sub-client cluster S 1 and issues N - 1 key files to the first client; The first client numbers the N - 1 received key files, with the numbers 1-2, 1-3,..., 1-N, and feeds back the corresponding relationship between these numbers and the key files to the key center;
[0022] C: The key center selects from the sub-client cluster S 1Select the second client from it, send the key files numbered 1 - 2 to the second client, and then send (N - 2) key files to the second client; the second client modifies the numbers of the received key files numbered 1 - 2 to 2 - 1, and at the same time, numbers the received (N - 2) key files as 2 - 3, 2 - 4, ……, 2 - N, and feeds back the corresponding relationship between these numbers and the key files to the key center;
[0023] D: The key center selects the third client from the sub - client cluster S 1 Select the third client from it, send the key files numbered 1 - 3 and 2 - 3 to the third client, and then send (N - 3) key files to the third client; the third client modifies the number of the received key file numbered 1 - 3 to 3 - 1, modifies the number of the received key file numbered 2 - 3 to 3 - 2, and at the same time, numbers the received (N - 3) key files as 3 - 4, 3 - 5, ……, 3 - N; and feeds back the corresponding relationship between these numbers and the key files to the key center;
[0024] E: And so on, until the key center selects the Nth client from the sub - client cluster S 1 Select the Nth client from it, send the key files numbered 1 - N, 2 - N, ……, (N - 1) - N to the Nth client, and the Nth client modifies the numbers of the received key files to N - 1, N - 2, ……, N - (N - 1); thus far, the key center has completed the key distribution for all clients in the entire sub - client cluster S 1 All clients in it.
[0025] Furthermore, each client in the sub - client cluster S 1 Encoding the local data set into a data set of quantum states means:
[0026] Select the kth client in the sub - client cluster S 1 where k ∈ {1, 2, …, N}, and the kth client encodes its local data set to obtain a data set of quantum states representing the quantum state of each data in the data set, is the corresponding label for each data, n k represents the total number of data in the data set of the kth client, i represents the ith data in the data set; the kth client encodes the data set in the way of where is the data in the local data set of the kth client, represents the d - dimensional real - number space;
[0027]
[0028] where, is the unitary embedding operator, are conjugate, and the data is encoded into quantum bit quantum state data of n bits through the above encoding method.
[0029] Furthermore, the specific process of training the local local model based on the received initialized global model parameter θ 0 is as follows:
[0030] S1: Select the k-th client in the sub-client cluster S 1 , where k ∈ {1, 2,..., N}, and then randomly select a client from the sub-client cluster S 1 as the communication peer. The client of this communication peer is the r-th client, where r ∈ {1, 2,..., N}; the k-th client and the r-th client compare the initialized global model parameters θ 0 decrypted by each other. If they are the same, the training process continues; if they are different, the service center redistributes the global model parameter θ 0 ;
[0031] S2: The k-th client obtains the key file numbered (k - r) from the local, and obtains the key from this key file
[0032] S3: The k-th client obtains the data set D k , and based on this data set D k obtains the basic parameters of the local model by minimizing the loss function and calculates the basic update parameters of the local model in the first round where the loss function represents the difference between the predicted value of the local model and the true data value of the data set;
[0033] S4: The k-th client uses the obtained key to process the basic update parameters of the local model as follows to obtain the optimized update parameters of the local model
[0034]
[0035] where means taking (-1) when k > r; otherwise taking 1; Q q represents the quantization function, q represents the quantization length:
[0036]
[0037] This quantization function quantizes a scalar s to [-(2 q-1 - 1), 2 q-1 Signed integers in the range of [[-1]]; where, represents the set of real numbers, and the value range of s is [-β, β]; sgn() in the quantization function represents the sign function, abs() represents the absolute value, and Round() means mapping the input to the nearest integer value;
[0038] S5: The k-th client encrypts and feeds back the local model optimization update parameters of this first round to the service center.
[0039] Furthermore, the step 5) means:
[0040] The service center receives and decrypts to obtain the local model optimization update parameters of all clients in the sub-client cluster S 1 in, and performs weighted aggregation on the local model optimization update parameters of all clients in the sub-client cluster S 1 to obtain the global model optimization update parameters Δθ of the first round 1 :
[0041]
[0042] Apply the basic attributes to the above formula and simplify the above formula to get:
[0043]
[0044] The service center processes based on the global model optimization update parameters of the first round to obtain the global model parameters θ of the first round 1 , θ 1 = θ 0 + D q (Δθ 1 ); Send the global model parameters θ 1 to each client in the sub-client cluster S 2 ;
[0045] where, D q is the dequantization function corresponding to the quantization function Q q :
[0046]
[0047] The service center performs the following adjustment for this dequantization function: If v > 2 q-1 - 1, then update v to v - 2 q ; Otherwise, v remains unchanged.
[0048] Furthermore, obtaining the final global model parameters θ T through T rounds of parameter iteration in the step 6) means:
[0049] For the t-th round, the service center selects the sub-client cluster S t Perform parameter iteration for the t-th round of federated learning: the sub-client cluster S t Each client in it trains its local model based on the global model parameters θ of the previous round received t-1 to obtain the global model parameters θ of the t-th round t :
[0050] θ t = θ t-1 + D q (Δθ t )
[0051]
[0052] Finally, after T times of parameter iteration, the final global model parameters θ are obtained T :
[0053] θ T = θ T-1 + D q (Δθ T )
[0054]
[0055] The present invention further includes a system based on the above data processing method. The system includes a service center, a key center, and a client cluster, and the client cluster is composed of multiple clients; wherein, the service center is connected to the key center, and the service center and the key center are respectively connected to each client in the client cluster;
[0056] The service center is used to train the global model in the system;
[0057] The key center is used to generate and distribute the keys involved in the communication process for the service center and each client in the client cluster;
[0058] Each client in the client cluster has a local data set, and the local data set is a data set in a quantum state. Each client trains its local model based on the local data set.
[0059] Further, each client in the client cluster includes a first communication area, a first isolation area, and a first privacy area. The first communication area includes a first communication proxy unit. The first isolation area includes a first encryption / decryption unit and a first encryption / decryption key pool that are connected to each other. The first encryption / decryption unit is also connected to the first communication proxy unit. The first privacy area includes a local model training key pool, a local model training unit, and a data storage unit that are connected in sequence. The local model training key pool and the local model training unit are also respectively connected to the first encryption / decryption unit;
[0060] The first communication proxy unit is used for data transmission between the client and the outside world;
[0061] The first encryption / decryption unit is used for encrypting and decrypting data;
[0062] The first encryption / decryption key pool is used to provide keys for encryption and decryption operations;
[0063] The local model training key pool is used to provide keys for the local model training unit;
[0064] The local model training unit is used to train the local model on the local data set;
[0065] The data storage unit is used to store the local data set, and the local data set is a data set in a quantum state;
[0066] The service center includes a second communication area, a second isolation area, and a second privacy area. The second communication area includes a second communication proxy unit. The second isolation area includes a second encryption / decryption unit and a second encryption / decryption key pool that are connected to each other. The second encryption / decryption unit is also connected to the second communication proxy unit. The second privacy area includes a global model training unit, and the global model training unit is also connected to the second encryption / decryption unit;
[0067] The second communication proxy unit is used for data transmission between the service center and the outside world;
[0068] The second encryption / decryption unit is used for encrypting and decrypting data;
[0069] The second encryption / decryption key pool is used to provide keys for encryption and decryption operations;
[0070] The global model training unit is used to train the global model in the system.
[0071] Advantages of the present invention:
[0072] 1. The key distribution of the present invention is carried out through a quantum network, and the generated key file is used for subsequent communication processes, enabling the entire system to have no communication distance limit. Moreover, the communication rate of the transmission parameters matches the network transmission rate in actual applications, eliminating the need to wait for the key to be generated when communication is required, improving communication efficiency, and solving the problems of limited transmission distance and low coding rate in actual applications of the QKD method, thus having practical application value.
[0073] 2. The hardware device structures of the communication area, isolation area, and privacy area possessed by the service center and the client enable the transmission parameters involved in the working process of the communication system to be transmitted in an encrypted and secure manner. Any attack by a third-party device on the service center or the client can only reach its communication area and cannot successfully pass through the encryption and decryption units in the isolation area. Therefore, the attack cannot reach the privacy area and thus cannot affect the operation of the system.
[0074] 3. The present invention combines federated learning with a quantum network, protecting the data privacy of each client (i.e., the data does not leave the client) while achieving the effect of using the data of the client for machine learning to obtain an accurate learning model.
[0075] 4. In the process of federated learning, compared with the traditional federated learning process that directly regards the user data ratios of each client as the same, the present invention takes into account the different user data ratios of each client and introduces a parameter making the method proposed by the present invention closer to actual applications, obtaining a more accurate machine learning model, and having higher application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0077] Figure 1 It is a schematic structural diagram of the quantum-secure data processing system of the present invention;
[0078] Figure 2 It is a schematic structural diagram of each client in the client cluster of the present invention;
[0079] Figure 3 It is a schematic structural diagram of the service center of the present invention;
[0080] Figure 4 It is a schematic flow diagram of the quantum-secure data processing method of the present invention;
[0081] Figure 5Randomly select T sub-client clusters S from the client cluster for the present invention t Schematic diagram;
[0082] Figure 6 Schematic diagram of re-numbering the clients in the sub-client cluster S of the present invention 1 ;
[0083] Figure 7 Schematic diagram of the key file corresponding to each client in the sub-client cluster S of the present invention 1 ; Specific implementation manner
[0084] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0085] As described in the background art, there is a risk of privacy leakage in current federated learning, and its security threats include poisoning attacks, adversarial attacks and backdoor attacks. These attacks will not only affect the accuracy of the model, but may also be used for malicious purposes. Therefore, how to ensure the security of the global data in machine learning using federated learning in practical applications is an important issue faced by the current deep learning field.
[0086] In view of this, the present invention proposes a quantum-secure data processing system based on federated learning. The system includes a service center 1, a key center 2 and a client cluster 3, and the client cluster 3 is composed of multiple clients; among them, as Figure 1 shown, the service center 1 is connected to the key center 2, and the service center 1 and the key center 2 are also respectively connected to each client in the client cluster 3;
[0087] The service center 1 is used to train the global model in the quantum-secure data processing system; the key center 2 is used to generate and distribute the keys involved in the communication process for the service center 1 and each client in the client cluster 3; each client in the client cluster 3 has its own local data set, and the local data set is a data set in a quantum state, and each client trains a local model based on the local data set.
[0088] As Figure 2As shown, each client in the client cluster 3 includes a first communication area 31, a first isolation area 32, and a first privacy area 33. The first communication area 31 includes a first communication proxy unit 311. The first isolation area 32 includes a first encryption / decryption unit 321 and a first encryption / decryption key pool 322 that are interconnected. The first encryption / decryption unit 321 is also connected to the first communication proxy unit 311. The first privacy area 33 includes a local model training key pool 331, a local model training unit 332, and a data storage unit 333 that are connected in sequence. The local model training key pool 331 and the local model training unit 332 are also respectively connected to the first encryption / decryption unit 321;
[0089] The first communication proxy unit 311 is used for data transmission between the client and the outside world; the first encryption / decryption unit 321 is used for encrypting and decrypting data; the first encryption / decryption key pool 322 is used to provide keys for the encryption and decryption operations; the local model training key pool 331 is used to provide keys for the local model training unit 332; the local model training unit 332 is used to train the local model on the local data set in the client; the data storage unit 333 is used to store data, specifically such as the local data set of the client, and this local data set is a data set in a quantum state;
[0090] As Figure 3 shown, the service center 1 includes a second communication area 11, a second isolation area 12, and a second privacy area 13. The second communication area 11 includes a second communication proxy unit 111. The second isolation area 12 includes a second encryption / decryption unit 121 and a second encryption / decryption key pool 122 that are interconnected. The second encryption / decryption unit 121 is also connected to the second communication proxy unit 111. The second privacy area 13 includes a global model training unit 131, and the global model training unit 131 is also connected to the second encryption / decryption unit 121;
[0091] The second communication proxy unit 111 is used for data transmission between the service center 1 and the outside world; the second encryption / decryption unit 121 is used for encrypting and decrypting data; the second encryption / decryption key pool 122 is used to provide keys for the encryption and decryption operations; the global model training unit 131 is used to train the global model in this quantum-secure data processing system.
[0092] In the present invention, the hardware device structures of the three areas of the communication area, isolation area, and privacy area possessed by the service center 1 and the client enable the transmission parameters involved in the working process of the communication system to be transmitted in an encrypted and secure manner; any attack by a third-party device on the service center and the client can only reach its communication area and cannot successfully pass through the encryption / decryption unit in the isolation area. Therefore, the attack cannot reach the privacy area and thus cannot affect the operation of the system.
[0093] As Figure 4As shown in the figure, the present invention further includes a data processing method based on the above quantum-secure data processing system, comprising the following steps:
[0094] (1) The key center 2 distributes a set of encryption and decryption key files to the client cluster 3 and the service center 1 respectively; the encryption and decryption key files received by the client cluster 3 and the service center 1 can be the same encryption and decryption key files, which are used for the encryption and decryption processes during the parameter transmission process of federated learning. The encryption and decryption key files are stored in the first encryption and decryption key pool 322 in the first isolation area 32 of the client and the second encryption and decryption key pool 122 in the second isolation area 12 of the service center 1 respectively.
[0095] Then, the service center 1 sets the number of iterations T of federated learning. According to the number of iterations T, the service center 1 randomly selects T sub-client clusters from the client cluster, as Figure 5 shown; among them, the sub-client cluster is denoted as The number of clients in each sub-client cluster is N;
[0096] The determination criterion for the number of iterations T can be the number of rounds when federated learning converges. For example, if federated learning converges around the 40th round, at this time, the value of T can be any integer greater than 40, such as 50. In practical applications, the value here can be simply determined according to empirical values. Or, the determination criterion can be that the global model reaches the performance standard.
[0097] The sub-client cluster is denoted as Each sub-client cluster S among them t contains the same number of clients, and the number of clients in each sub-client cluster is N, as Figure 5 shown.
[0098] (2) The service center 1 selects the sub-client cluster S 1 to perform the parameter iteration of the first round of federated learning. After completion, it then selects the sub-client cluster S 2 to perform the parameter iteration of the second round of federated learning, and so on, until it selects the sub-client cluster S T to perform the parameter iteration of the Tth round of federated learning, and finally completes the parameter iteration of T rounds of federated learning; the specific process is as follows:
[0099] 1) The key center 2 assigns symmetric keys to each pair of clients in the sub-client cluster S 1 ; the specific process of assigning symmetric keys is as follows:
[0100] A: As Figure 6 shown, all the clients in the sub-client cluster S 1 are renumbered to obtain the first to the Nth clients;
[0101] Communication may occur between every two clients. Therefore, to ensure the smooth progress of communication between each pair of clients, each pair of clients needs a symmetric key. The key center 2 distributes keys to each client in the following manner:
[0102] B: The key center 2 selects the first client from the sub-client cluster S 1 and counts the total number N of clients in the sub-client cluster S 1 and distributes N - 1 key files to the first client; the first client numbers the N - 1 received key files. For example, the numbers can be 1-2, 1-3,..., 1-N, and feeds back the corresponding relationship between these numbers and the key files to the key center 2;
[0103] C: The key center 2 selects the second client from the sub-client cluster S 1 distributes the key file numbered 1-2 to the second client, and then distributes N - 2 key files to the second client; the second client changes the number of the received key file numbered 1-2 to 2-1. At the same time, numbers the N - 2 received key files as 2-3, 2-4,..., 2-N, and feeds back the corresponding relationship between these numbers and the key files to the key center 2;
[0104] D: The key center 2 selects the third client from the sub-client cluster S 1 distributes the key files numbered 1-3 and 2-3 to the third client, and then distributes (N - 3) key files to the third client; the third client changes the number of the received key file numbered 1-3 to 3-1, changes the number of the received key file numbered 2-3 to 3-2. At the same time, numbers the N - 3 received key files as 3-4, 3-5,..., 3-N; and feeds back the corresponding relationship between these numbers and the key files to the key center 2;
[0105] E: And so on, until the key center 2 selects the Nth client from the sub-client cluster S 1 distributes the key files numbered 1-N, 2-N,...,(N - 1)-N to the Nth client. The Nth client can modify the numbers according to the numbering method in step C or D, and changes the numbers of the received key files to N-1, N-2,..., N-(N - 1); thus, the key center 2 completes the key distribution for all clients in the entire sub-client cluster S 1 The key files are transmitted to the local model training key pool 331 in the first privacy area 33 for storage via the first communication proxy unit 311 and the first encryption / decryption unit 321.
[0106] The key in the key file allocated in this process is used for the parameter update process in federated learning. The key files in each client in the final sub-client cluster S 1 are as follows Figure 7 shown. In the present invention, the key distribution is carried out through a quantum network, and the generated key file is used for subsequent communication processes, so that there is no limitation on the communication distance in the whole system, and the communication rate of the transmission parameters matches the network transmission rate in actual applications, without waiting for the key to be generated when communication is required, improving the communication efficiency, solving the problems of limited transmission distance and low coding rate of the QKD method in actual applications, and having practical application value;
[0107] 2) Then, each client in the sub-client cluster S 1 encodes its local data set into a data set in a quantum state, specifically referring to:
[0108] For example, taking the k-th client as an example, select the k-th client in the sub-client cluster S 1 , k ∈ {1, 2,..., N}, and the k-th client encodes its local data set to obtain a data set in a quantum state The data set D k is stored in the data storage unit 333 of the first privacy area 33, representing the quantum state of each data in the data set, being the corresponding label for each data, n k representing the total number of data in the data set of the k-th client, and i representing the i-th data in the data set; the k-th client encodes the data set by way, being the data in the local data set of the k-th client, representing the d-dimensional real number space;
[0109]
[0110] Among them, is the unitary embedding operator, is the conjugate, and the data is encoded into an n-bit qubit quantum state data through the above encoding method.
[0111] 3) The service center 1 randomly initializes the global model parameters representing the M-dimensional real number space, and M represents the size of the global model parameters; then the global model parameters θ 0 are encrypted and sent to each client in the sub-client cluster S 1 ;
[0112] The initialization process is executed in the global model training unit 131 of the second privacy area 13. The global model training unit 131 sends the global model parameter θ 0 to the second encryption / decryption unit 121. The second encryption / decryption unit 121 obtains an encryption key from the encryption / decryption key file (the encryption / decryption key file sent in step (1)) in the local second encryption / decryption key pool 122, records the position information addr1 of the encryption key, and encrypts the global model parameter θ using the encryption key 0 to obtain the ciphertext key1. The ciphertext key1 and the position information addr1 are sent to each client in the sub-client cluster S via the second communication proxy unit 111 in the second communication area 11 1 in the sub-client cluster S
[0113] 4) Each client in the sub-client cluster S 1 performs decryption, and based on the received initialized global model parameter θ 0 trains the local local model to obtain local model optimization update parameters, and encrypts and feeds back the local model optimization update parameters to the service center 1;
[0114] Among them, each client obtains a decryption key from the encryption / decryption key file (the encryption / decryption key file sent in step (1)) of the local first encryption / decryption key pool 322 according to the received position information addr1, and uses the decryption key to decrypt the ciphertext key1 to obtain the initialized global model parameter θ 0 。
[0115] In this step, based on the received initialized global model parameter θ 0 the specific process of training the local local model is as follows:
[0116] S1: Taking the training of the local local model by the k-th client in the sub-client cluster S 1 as an example, first select the k-th client in the sub-client cluster S 1 where k ∈ {1, 2,..., N}, and then randomly select a client in the sub-client cluster S 1 as the communication peer. The client of the communication peer is the r-th client, r ∈ {1, 2,..., N}; the k-th client and the r-th client compare the initialized global model parameters θ they decrypted respectively 0 to see if they are the same. If they are the same, the training process continues; if they are not the same, the service center 1 re-sends the global model parameter θ 0 ;
[0117] S2: According to the rule of distributing the symmetric key during the key file distribution process, the local model training unit 332 on the k-th client obtains the key file numbered (k - r) from the local model training key pool 331, and obtains the key from this key file.
[0118] S3: The local model training unit 332 of the k-th client obtains the data set D from the data storage unit 333. k , and based on this data set D k By minimizing the loss function to obtain the basic parameters of the local model and calculate the basic update parameters of the local model in the first round where the loss function represents the difference between the predicted value of the local model and the true data value of the data set;
[0119] S4: The k-th client uses the obtained key to process the basic update parameters of the local model as follows to obtain the optimized update parameters of the local model
[0120]
[0121] where, represents taking (-1) when k > r; otherwise taking 1; Q q represents the quantization function, q represents the quantization length:
[0122]
[0123] This quantization function can quantize a scalar s into a signed integer in the range of [-(2 q-1 - 1), 2 q-1 - 1]; where, represents the set of real numbers, and the value range of s is [-β, β]; sgn() in the quantization function represents the sign function, abs() represents the absolute value, and Round() represents mapping the input to the nearest integer value;
[0124] S5: The k-th client encrypts and feeds back the optimized update parameters of the local model in the first round to the service center 1. Similarly, the communication process of the k-th client feeding back the optimized update parameters of the local model can also adopt the encrypted transmission method, and this method is the same as the process of encrypting and sending the global model parameter θ 0 , so it will not be elaborated here.
[0125] 5) The service center 1 receives and decrypts to obtain the sub-client cluster S 1The local model optimization update parameters of all clients in [[]], and weighted aggregation is performed on these local model optimization update parameters to obtain the global model optimization update parameters Δθ of the first round. 1 The service center 1 processes the global model optimization update parameters of the first round to obtain the global model parameters θ of the first round. 1 and sends the global model parameters θ 1 to each client in the sub-client cluster S 2 specifically referring to:
[0126] The service center 1 receives and decrypts to obtain the local model optimization update parameters of all clients in the sub-client cluster S 1 in, and performs weighted aggregation on the local model optimization update parameters of all clients in the sub-client cluster S 1 to obtain the global model optimization update parameters Δθ of the first round. 1 :
[0127]
[0128] Applying the basic attributes of to the above formula, the above formula can be simplified to obtain:
[0129]
[0130] The service center 1 processes the global model optimization update parameters of the first round to obtain the global model parameters θ of the first round. 1 θ 1 = θ 0 + D q (Δθ 1 ); The global model parameters θ 1 are sent to each client in the sub-client cluster S 2 ;
[0131] where D q is the dequantization function corresponding to the quantization function Q q :
[0132]
[0133] The service center 1 performs the following adjustment on the dequantization function: If v > 2 q-1 - 1, then update v to v - 2 q ; Otherwise, v remains unchanged.
[0134] 6) And so on, the service center 1 selects the next sub-client cluster S 2 , and the sub-client cluster S 2 to the sub-client cluster S T adopt the same method as the sub-client cluster S 1The same method is used to generate the corresponding local model optimization update parameters, and the service center processes the global model parameters for the corresponding round based on the local model optimization update parameters for each round; finally, through the parameter iteration of T rounds of federated learning, the final global model parameter θ is obtained. T 。
[0135] Among them, through the parameter iteration of T rounds of federated learning, the final global model parameter θ is obtained. T It means:
[0136] For the t-th round of federated learning, the service center 1 selects the sub-client cluster S t to perform the parameter iteration of the t-th round of federated learning: each client in the sub-client cluster S t trains the local local model based on the global model parameter θ of the previous round received, and obtains the global model parameter θ of the t-th round t-1 : t :
[0137] θ t = θ t-1 + D q (Δθ t )
[0138]
[0139] Finally, after T times of parameter iteration, the final global model parameter θ is obtained T :
[0140] θ T = θ T-1 + D q (Δθ T )
[0141]
[0142] (3) The service center 1 sends the parameter result θ of the T-round federated learning parameter iteration T to each client in the client cluster, and each client uses this parameter result as the parameter of the local federated learning model to perform prediction on the input data received by the client.
[0143] The present invention combines federated learning with a quantum network, while protecting the data privacy of each client (i.e., the data does not leave the client), and also realizes the effect of using the data of the client for machine learning to obtain an accurate learning model; moreover, in the process of federated learning, compared with the traditional federated learning process that directly regards the user data ratios of each client as the same, the present invention takes into account the different user data ratios of each client and introduces the parameter This makes the method proposed by the present invention closer to actual applications, the obtained machine learning model more accurate, and the application value higher.
Claims
1. A quantum secure data processing method based on federated learning, characterized in that: The following steps are involved: (1) The key center distributes a set of encryption and decryption key files to the client cluster and the service center respectively. Then, the service center sets the number of iterations T of federated learning. According to the number of iterations T, the service center randomly selects T sub-client clusters from the client cluster. The sub-client cluster is denoted as The number of clients in each sub-client cluster is N; (2) The service center selects a sub-client cluster S 1 Execute the first round of federated learning parameter iteration, and then select the sub-client cluster S after completion 2 Perform parameter iteration for the second round of federated learning, and so on, until a subclient cluster S is selected T Execute the parameter iteration of the Tth round of federated learning, and finally complete the parameter iteration of the Tth round of federated learning; (3) The service center sends the parameter results of T rounds of federated learning parameter iterations to each client in the client cluster. Each client uses the parameter results as the parameters of the local federated learning model to perform predictions on the input data received by the client.
2. A method for quantum secure data processing based on federated learning according to claim 1, characterized in that: The encryption and decryption key files in the client cluster and the service center are the same.
3. A method for quantum secure data processing based on federated learning according to claim 1, characterized in that: The specific process of step (2) is as follows: 1) The key center is the sub-client cluster S 1 Each pair of clients in the symmetric key distribution; 2) Then, the sub-client cluster S 1 Each client in the system encodes the local data set into a data set of quantum states; 3) The service center randomly initializes the global model parameters Represents an M-dimensional real number space, where M represents the size of the global model parameter; then the global model parameter θ 0 Encrypted and sent to sub-client cluster S 1 Every client in; 4) Sub-client cluster S 1 Each client in decrypts and receives the initialized global model parameter θ 0 Train the local model to obtain the local model optimization update parameters, and encrypt and feed back the local model optimization update parameters to the service center; 5) The service center receives and decrypts the sub-client cluster S 1 The local model optimization update parameters of all clients in the , weighted aggregation is performed on these local model optimization update parameters to obtain the first round of global model optimization update parameters Δθ 1 ; The service center optimizes and updates the parameters based on the first round of global model to obtain the first round of global model parameters θ 1 , and the global model parameter θ 1 Sent to sub-client cluster S 2 Every client in; 6) Similarly, the sub-client cluster S 2 To subclient cluster S T Use the subclient cluster S 1 The same method is used to generate the corresponding local model optimization update parameters. The service center obtains the global model parameters of the corresponding round according to each round of local model optimization update parameters. Finally, through T rounds of federated learning parameter iteration, the final global model parameters θ are obtained. T .
4. A method for quantum secure data processing based on federated learning according to claim 3, characterized in that: The key center is a sub-client cluster S 1 The specific process of assigning symmetric keys to each pair of clients in is as follows: A: Cluster the subclients S 1 All clients are renumbered to obtain the first to Nth clients; B: The key center is from the sub-client cluster S 1 Select the first client and count the sub-client cluster S 1 The total number of clients N in the first client is N-1 key files are issued to the first client; the first client numbers the received N-1 key files as 1-2, 1-3, ..., 1-N, and feeds back the corresponding relationship between these numbers and key files to the key center; C: Key center from subclient cluster S 1 Select the second client, send the key file numbered 1-2 to the second client, and then send N-2 key files to the second client; The second client modifies the number of the received key file numbered 1-2 to 2-1, and at the same time, numbers the received N-2 key files as 2-3, 2-4, ..., 2-N, and feeds back the corresponding relationship between these numbers and key files to the key center; D: The key center is from the sub-client cluster S 1 The third client is selected, the key files numbered 1-3 and 2-3 are sent to the third client, and (N-3) key files are sent to the third client; the third client modifies the number of the received key file numbered 1-3 to 3-1, modifies the number of the received key file numbered 2-3 to 3-2, and at the same time, the received N-3 key files are numbered 3-4, 3-5, ..., 3-N; and the corresponding relationship between these numbers and key files is fed back to the key center; E: And so on, until the key center is from the subclient cluster S 1 The Nth client is selected, and the key files numbered 1-N, 2-N, ..., (N-1)-N are sent to the Nth client. The Nth client modifies the number of the received key files to N-1, N-2, ..., N-(N-1). At this point, the key center completes the entire sub-client cluster S 1 Key distribution for all clients in the .
5. A method for quantum secure data processing based on federated learning according to claim 3, characterized in that: The sub-client cluster S 1 Each client in encodes the local data set into a quantum state data set: Select the subclient cluster S 1 The kth client in the set, k∈{1,2,…,N}, encodes the local data set to obtain the quantum state data set Represents the quantum state of each data in the data set, For each data, the corresponding label, n k represents the total number of data in the dataset of the k-th client, i represents the i-th data in the dataset; the k-th client passes The dataset is encoded in the following way: is the data in the local data set of the kth client, represents d-dimensional real number space; in, is the unitary embedding operator, is conjugate, and the data is encoded in the above way Encoded as n-bit quantum bit quantum state data.
6. A method for quantum secure data processing based on federated learning according to claim 5, characterized in that: The global model parameters θ based on the received initialization 0 The specific process of training the local local model is: S1: Select subclient cluster S 1 The kth client in the cluster S, k∈{1,2,…,N}, then 1 A client is randomly selected as the communication peer, and the client of the communication peer is the r-th client, r∈{1,2,…,N}; the k-th client and the r-th client compare the initial global model parameters θ obtained by their respective decryption 0 Are they the same? If they are the same, continue the training process; If they are not the same, the service center will re-send the global model parameters θ 0 ; S2: The kth client obtains the key file numbered (kr) from the local computer and obtains the key from the key file. S3: The kth client obtains the dataset D k , and based on the data set D k By minimizing the loss function To get the basic parameters of the local model And calculate the basic update parameters of the local model in the first round Among them, the loss function Represents the difference between the predicted value of the local model and the true data value of the dataset; S4: The kth client uses the obtained key Basic update parameters of the local model Perform the following processing to obtain the local model optimization update parameters in, (-1) k>r When k>r, take (-1); otherwise take 1; Q q Represents the quantization function, and q represents the quantization length: This quantization function quantizes a scalar s to [-(2 q-1 -1),2 q-1 -1]; where Represents a set of real numbers, and the value range of s is [-β, β]; sgn() in the quantization function represents the sign function, abs() represents the absolute value, and Round() represents mapping the input to the nearest integer value; S5: The kth client optimizes and updates the parameters of the first round of local model Encrypted feedback to the service center.
7. A method for quantum secure data processing based on federated learning according to claim 6, characterized in that: The step 5) refers to: The service center receives and decrypts the sub-client cluster S 1 The local model optimization update parameters of all clients in the sub-client cluster S 1 The local model optimization update parameters of all clients in the process are weighted and aggregated to obtain the global model optimization update parameter Δθ of the first round. 1 : Will The basic properties of are applied to the above formula, and the above formula is simplified to obtain: The service center optimizes and updates the parameters based on the first round of global model optimization to obtain the first round of global model parameters θ 1 ,θ 1 =θ 0 +D q (Δθ 1 );The global model parameter θ 1 Sent to sub-client cluster S 2 Every client in; Among them, D q is the quantization function Q q The corresponding inverse quantization function: The service center performs the following adjustments on the dequantization function: If v>2 q-1 -1, then update v to v-2 q ; otherwise, v remains unchanged.
8. A method for quantum secure data processing based on federated learning according to claim 7, characterized in that: In step 6), the final global model parameter θ is obtained by T rounds of federated learning parameter iteration. T means: For round t, the service center selects sub-client cluster S t Perform parameter iteration for round t of federated learning: subclient cluster S t Each client in the previous round receives the global model parameter θ t-1 Train the local local model to obtain the global model parameters θ of the tth round t : i t =θ t-1 +D q (Dth t ) Finally, after T parameter iterations, the final global model parameter θ is obtained. T : i T =θ T-1 +D q (Dth T ) 9. A system based on the data processing method according to any one of claims 1 to 8, characterized in that: The system includes a service center, a key center and a client cluster, wherein the client cluster is composed of a plurality of clients; wherein the service center is connected to the key center, and the service center and the key center are also respectively connected to each client in the client cluster; The service center is used to train the global model in the system; The key center is used to generate and issue keys involved in the communication process to the service center and each client in the client cluster; Each client in the client cluster has a local data set, and the local data set is a data set of quantum states. Each client trains a local model based on the local data set.
10. The system according to claim 9, characterized in that: Each client in the client cluster includes a first communication area, a first isolation area and a first privacy area, the first communication area includes a first communication agent unit, the first isolation area includes a first encryption and decryption unit and a first encryption and decryption key pool connected to each other, the first encryption and decryption unit is also connected to the first communication agent unit, the first privacy area includes a local model training key pool, a local model training unit and a data storage unit connected in sequence, the local model training key pool and the local model training unit are also respectively connected to the first encryption and decryption unit; The first communication agent unit is used for data transmission between the client and the outside world; The first encryption and decryption unit is used to perform encryption and decryption operations on data; The first encryption and decryption key pool is used to provide keys for encryption and decryption operations; The local model training key pool is used to provide keys for the local model training unit; The local model training unit is used to train the local model on the local data set; The data storage unit is used to store a local data set, and the local data set is a data set of quantum states; The service center includes a second communication area, a second isolation area, and a second privacy area, the second communication area includes a second communication agent unit, the second isolation area includes a second encryption and decryption unit and a second encryption and decryption key pool connected to each other, the second encryption and decryption unit is also connected to the second communication agent unit, and the second privacy area includes a global model training unit, which is also connected to the second encryption and decryption unit; The second communication agent unit is used for data transmission between the service center and the outside world; The second encryption and decryption unit is used to perform encryption and decryption operations on data; The second encryption and decryption key pool is used to provide keys for encryption and decryption operations; The global model training unit is used to train the global model in the system.