Federal distributed aggregation method and device, electronic equipment and storage medium

By splitting global secrets into multiple secret fragments in federated distributed learning and encrypting model parameters, the privacy and security issues of client data being attacked are solved, and an efficient and secure model training process is achieved.

CN120180461APending Publication Date: 2025-06-20JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510109184.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the distributed federated learning process of machine learning, the client's data is easily maliciously attacked, making it difficult to guarantee the privacy and security of the data.

Method used

A federated distributed aggregation method is proposed, which ensures that even if a part of the client is attacked, the attacker cannot obtain complete secret information. At the same time, the model parameters of the submodel trained by each client are encrypted, and the encrypted model parameters are divided into multiple shares for aggregation, improving the efficiency and security of data transmission.

Benefits of technology

Through this method, the privacy and security of model training are improved, the security of data during transmission is ensured, and the efficiency of model training is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180461A_ABST
    Figure CN120180461A_ABST
Patent Text Reader

Abstract

The invention provides a federated distributed aggregation method and device, electronic equipment and a storage medium, the method is applied to the technical field of machine learning, and the method comprises the following steps: segmenting a global secret s in a target global model training process into n secret segments, and sending the n secret segments to n model training devices; for any model training device, determining a secret key of the model training device and sending the secret key to the model training device, so that the model training device encrypts the original model parameters to obtain encrypted model parameters; segmenting the encryption model parameters to obtain m encryption shares; aggregating the m encrypted shares to obtain an aggregated share of the model training device; and updating the target global model based on the aggregated share of the n model training devices. According to the method, the data in the training process can be encrypted in the model training process of the federated distributed method, so that the privacy and security in the model training process are improved, and the success rate of model training is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and more specifically, to a federated distributed aggregation method, device, electronic device, and storage medium in the field of machine learning technology. Background Art

[0002] During the model training process, the traditional model training method is centralized learning (CL). Since all data in centralized learning is concentrated on the central server, once the server is breached, all private data is at risk of leakage.

[0003] Based on the problems brought by centralized learning, currently in the process of machine learning, researchers have also provided a new type of distributed federated learning (FL). Distributed federated learning means that each client retains its own data on its local device, without uploading the original data to the central server, but only sending the updated model parameters to the central server for aggregation.

[0004] However, during the process of implementing model aggregation based on distributed federated learning, it is inevitable that the data of the client will be maliciously attacked, and there is a lack of privacy and security protection for the data during the learning process. Summary of the Invention

[0005] This application provides a federated distributed aggregation method, device, electronic device, and storage medium. This method can encrypt the data during the model training process when using the federated distributed method to train the model, increasing the privacy and security during the model training process and ensuring the success rate of model training.

[0006] In a first aspect, a federated distributed aggregation method is provided. The method includes: splitting a global secret s during the training process of a target global model into n secret fragments, and sending the n secret fragments to n model training devices, where n≥1 and n is an integer; for any one of the n model training devices, determining the key of the model training device and sending the key to the model training device, where the key is used to encrypt the original model parameters obtained by training the model training device to obtain encrypted model parameters; splitting the encrypted model parameters to obtain m encrypted shares, where m≥1 and m is an integer; aggregating the m encrypted shares to obtain an aggregated share of the model training device; and updating the target global model based on the aggregated shares of the n model training devices.

[0007] In the above technical solution, when training a model using the federated distributed method, the present application proposes a federated distributed aggregation method. During the model training process, this method divides the global secret into multiple secret fragments and distributes them to multiple model training devices (clients), which can ensure that even if a part of the clients are attacked, the attacker cannot obtain the complete secret information. Further, for the model parameters of the sub-model trained by each client, specific keys are used for encryption to ensure that no sensitive information is leaked during the process of the client transmitting data to the central server. After receiving the encrypted model parameters sent by the client, the encrypted model parameters are divided into multiple shares, which improves the efficiency during the data transmission process, significantly reduces the transmission time, and at the same time can increase the difficulty of attack. Even if the attacker obtains some shares, the complete model parameters cannot be restored. Thus, during the model training process, the present application improves the training efficiency of the model by parallel processing of training tasks by multiple model training devices, compared with the centralized learning method, and ensures the privacy and security of the data during the model training process.

[0008] In combination with the first aspect, in some possible implementation manners, dividing the global secret s in the target global model training process into n secret fragments includes: obtaining a threshold value t, where the threshold value t is used to represent the number of secret fragments required to recover the global secret s; determining a polynomial of degree t - 1 according to the threshold value t; and determining the n secret fragments according to the polynomial of degree t - 1 and the global secret s.

[0009] In combination with the first aspect and the above implementation manners, in some possible implementation manners, the steps of the model training device encrypting the originally trained model parameters to obtain encrypted model parameters include: determining an initial state matrix according to the originally trained model parameters; performing a key expansion operation on the key to obtain round keys; performing a bitwise exclusive OR operation on the state matrix and the initial round key in the round keys to obtain a first state matrix; and repeating the operations of byte substitution, row shift, column mixing, and round key addition based on a preset number of rounds and the first state matrix to obtain the encrypted model parameters.

[0010] In combination with the first aspect and the above implementation manners, in some possible implementation manners, before aggregating the m encrypted shares to obtain the aggregated share of the model training device, the method further includes: generating a pair of public key d pk and private key d sk ; signing the m encrypted shares based on the private key d sk to obtain m signature results; verifying the m encrypted shares based on the public key d pk and the m signature results; wherein, the formula for signing the m encrypted shares based on the private key d sk to obtain m signature results is as follows:

[0011] SIG.sign(d sk , m) → σ;

[0012] m represents the encrypted share; σ represents the signature result of the encrypted share;

[0013] Among them, based on the public key d pk and the m signature results, the representation formula for verifying the m encrypted shares is as follows:

[0014] SIG.ver(d pk , m, σ) → 0, 1;

[0015] And, aggregating the m encrypted shares to obtain the aggregated share of the model training device, including: when the verification of the m encrypted shares is successful, aggregating the m encrypted shares to obtain the aggregated share of the model training device.

[0016] Combined with the first aspect and the above implementation, in some possible implementations, updating the target global model based on the aggregated shares of the n model training devices includes: decrypting the aggregated shares of the n model training devices based on the keys of the n model training devices to obtain the original aggregated shares of the n model training devices; performing feature extraction on the original aggregated shares of the n model training devices to obtain the feature data of the n sub-models corresponding to the n model training devices; updating the target global model based on the feature data of the n sub-models.

[0017] Combined with the first aspect and the above implementation, in some possible implementations, updating the target global model based on the feature data of the n sub-models includes: performing feature matching processing on the feature data of the n sub-models to obtain the processed feature data of the n sub-models, and the processed feature data of the n sub-models has the same dimension; fusing the processed feature data of the n sub-models based on the feature fusion method to obtain the fused feature data; constructing a target meta-learner according to the fused feature data; determining the final decision based on the comprehensive feature decision of the target meta-learner and the sub-model decisions of the n model training devices; updating the target global model according to the final decision.

[0018] Combined with the first aspect and the above implementation manners, in some possible implementation manners, updating the target global model according to the final decision includes: sending the final decision to the n model training devices, so that the n model training devices update the model parameters of the n model training devices according to the final decision; obtaining the updated model parameters of the n model training devices; aggregating the updated model parameters of the n model training devices to obtain aggregated model parameters; and applying the aggregated model parameters to the target global model to update the model parameters of the target global model.

[0019] In a second aspect, a federated distributed aggregation device is provided. The device includes: a first splitting module, configured to split a global secret s in the process of training a target global model into n secret fragments, and send the n secret fragments to n model training devices, where n≥1 and n is an integer; a key generation module, configured to, for any one of the n model training devices, determine a key of the model training device and send the key to the model training device, where the key is used to enable the model training device to encrypt the originally obtained model parameters to obtain encrypted model parameters; a second splitting module, configured to split the encrypted model parameters to obtain m encrypted shares, where m≥1 and m is an integer; an aggregation module, configured to aggregate the m encrypted shares to obtain an aggregated share of the model training device; and a model update module, configured to update the target global model based on the aggregated shares of the n model training devices.

[0020] Combined with the second aspect, in some possible implementation manners, the first splitting module is specifically configured to: obtain a threshold t, where the threshold t is used to represent the number of secret fragments required to recover the global secret s; determine a polynomial of degree t - 1 according to the threshold t; and determine the n secret fragments according to the polynomial of degree t - 1 and the global secret s.

[0021] Combined with the second aspect and the above implementation manners, in some possible implementation manners, the step of enabling the model training device to encrypt the originally obtained model parameters to obtain encrypted model parameters includes: determining an initial state matrix according to the originally obtained model parameters; performing a key expansion operation on the key to obtain round keys; performing a bitwise exclusive OR operation on the state matrix and an initial round key in the round keys to obtain a first state matrix; and repeating byte substitution, row shift, column mixing, and round key addition operations based on a preset number of rounds and the first state matrix to obtain the encrypted model parameters.

[0022] Combined with the second aspect and the above implementation manners, in some possible implementation manners, before aggregating the m encrypted shares to obtain an aggregated share of the model training device, the key generation module is further configured to: generate a pair of public key d pk and private key dsk ;

[0023] Based on the private key d sk , sign the m encrypted shares to obtain m signature results;

[0024] Based on the public key d pk and the m signature results, verify the m encrypted shares;

[0025] Among them, the expression formula for signing the m encrypted shares based on the private key d sk to obtain m signature results is as follows:

[0026] SIG.sign(d sk , m) → σ;

[0027] m represents the encrypted share; σ represents the signature result of the encrypted share;

[0028] Among them, the expression formula for verifying the m encrypted shares based on the public key d pk and the m signature results is as follows:

[0029] SIG.ver(d pk , m, σ) → 0, 1;

[0030] In addition, the aggregation module is specifically used for: when the verification of the m encrypted shares is successful, aggregating the m encrypted shares to obtain the aggregated share of the model training device.

[0031] Combined with the second aspect and the above implementation, in some possible implementation manners, the model update module is specifically used for: decrypting the aggregated shares of the n model training devices based on the keys of the n model training devices to obtain the original aggregated shares of the n model training devices; performing feature extraction on the original aggregated shares of the n model training devices to obtain the feature data of the n sub-models corresponding to the n model training devices; updating the target global model based on the feature data of the n sub-models.

[0032] Combined with the second aspect and the above implementation, in some possible implementation manners, the model update module is further used for: performing feature matching processing on the feature data of the n sub-models to obtain the processed feature data of the n sub-models, and the processed feature data of the n sub-models has the same dimension; fusing the processed feature data of the n sub-models based on a feature fusion method to obtain the fused feature data; constructing a target meta-learner according to the fused feature data; determining a final decision based on the comprehensive feature decision of the target meta-learner and the sub-model decisions of the n model training devices; updating the target global model according to the final decision.

[0033] Combined with the second aspect and the above implementation manners, in some possible implementation manners, the model update module is further configured to: send the final decision to the n model training devices, so that the n model training devices update the model parameters of the n model training devices according to the final decision; obtain the updated model parameters of the n model training devices; aggregate the updated model parameters of the n model training devices to obtain aggregated model parameters; and apply the aggregated model parameters to the target global model to update the model parameters of the target global model.

[0034] In a third aspect, an electronic device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, so that the electronic device executes the method in the first aspect or any possible implementation manner of the first aspect described above.

[0035] In a fourth aspect, a computer program product is provided, including: computer program code, when the computer program code runs on a computer, enabling the computer to execute the method in the first aspect or any possible implementation manner of the first aspect described above.

[0036] In a fifth aspect, a computer-readable storage medium is provided, storing computer program code, when the computer program code runs on a computer, enabling the computer to execute the method in the first aspect or any possible implementation manner of the first aspect described above. Description of the Drawings

[0037] Figure 1 is a schematic structural diagram of a federated distributed aggregation system provided by an embodiment of the present application;

[0038] Figure 2 is a schematic flowchart of a federated distributed aggregation method provided by an embodiment of the present application;

[0039] Figure 3 is an interaction scenario diagram of a federated distributed training model provided by an embodiment of the present application;

[0040] Figure 4 is a schematic structural diagram of a federated distributed aggregation device provided by an embodiment of the present application;

[0041] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0042] The technical solutions in the present application will be clearly and elaborately described below in conjunction with the accompanying drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" in the text is merely a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0043] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0044] Before introducing the embodiments of the present application, first, a glossary of professional terms that may be involved in the embodiments of the present application will be given.

[0045] Distributed Federated Learning FL: A machine learning technique that allows multiple participants (such as mobile devices, edge servers, or client nodes) to collaboratively train a shared global model without sharing local data.

[0046] Lightweight Model: Also known as a small model or a compact model, it refers to a machine learning model designed to operate efficiently in a resource-constrained environment.

[0047] Byzantine device: A device that exhibits abnormal behavior in a distributed system. Abnormal behavior includes, but is not limited to, fail-stop or malicious operation behavior.

[0048] The application scenarios of the embodiments of the present application will be introduced below.

[0049] Currently, in the training process of various models, traditional training methods are mostly centralized learning. However, during the training of traditional centralized learning, since all training data is concentrated on the central server, once the central server is attacked or leaked, all privacy data will be threatened; in addition, a large amount of network bandwidth is required to transmit a large amount of raw data to the central server, increasing the communication cost of the data. Especially in the case of a large amount of data, the time for data upload and download will be greatly delayed, affecting the response speed of the system.

[0050] Based on the various problems brought about by the above-mentioned centralized training method, current researchers have proposed a new type of training method, namely the distributed federated learning method. During the process of training a model using the distributed federated learning method, although the original data is not shared among clients, the model parameters sent by clients to the central server may also contain information about the underlying data and key information during the model training process, making it vulnerable to attacks by attackers or the influence of Byzantine devices, increasing the risk of leakage of key information during the distributed training process.

[0051] Based on the above problems, the embodiments of the present application propose a federated distributed aggregation method, which can encrypt the data during the process of training a model using the federated distributed method, increasing the privacy and security during the model training process and ensuring the success rate of model training.

[0052] After introducing the application scenarios of the embodiments of the present application, it should be understood that a federated distributed aggregation method provided by the embodiments of the present application mainly depends on a federated distributed aggregation system architecture provided by the embodiments of the present application during implementation. Therefore, in order to facilitate understanding of the implementation process of the embodiments of the present application, before introducing the method of the embodiments of the present application, the structure and working principle of the federated distributed aggregation system are first described.

[0053] Figure 1 It is a schematic structural diagram of a federated distributed aggregation system provided by the embodiments of the present application.

[0054] Exemplarily, as Figure 1 shown, a federated distributed aggregation system 100 provided by the embodiments of the present application mainly includes multiple devices participating in model training and a central server 105 according to different structures. Multiple devices participating in model training are, for example Figure 1 shown as client 101, client 102, client 103, and client 104.

[0055] Among them, during the federated distributed learning process, the role of the central server 105 is mainly to initialize the global model parameters and distribute the global model parameters to client 101, client 102, client 103, and client 104.

[0056] For any one of client 101, client 102, client 103, and client 104, when the client receives the global model parameters, it performs a certain number of training iterations on the local dataset of the client to update the model parameters. The client determines the number of training rounds or batch size according to its own computing power and data volume. When the client completes local training, it can send the updated model parameters or gradient differences to the central server 105.

[0057] Based on the updated model parameters or gradient differences sent by all clients, the central server 105 can use an aggregation algorithm to update the global model parameters.

[0058] The central server 105 redistributes the updated global model parameters to each client again, enabling each client to start a new round of model training, and so on until the training process reaches a preset convergence condition or number of training times. The training ends, and the central server 105 saves the final global model.

[0059] In a possible implementation, the central server 105 provided in the embodiments of the present application can be further divided into the following units according to functions: a sharding unit, a key management unit, an aggregation unit, and a collection unit. Among them, the role played by each unit in the model training process is also different. Specifically:

[0060] The sharding unit is used to divide the key information (global secret) that needs to be saved during the model training process into multiple secret fragments, and distribute the multiple secret fragments to multiple clients respectively, so that the multiple clients can perform local model training.

[0061] The key management unit is used to generate a corresponding key for each client and send it to each client and the collection unit, so that after the local model training process of each client ends, the trained model parameters are encrypted to obtain encrypted model parameters and sent to the central server 105.

[0062] In response to the encrypted model parameters sent by each client received, taking client 101 as an example, after client 101 sends the encrypted model parameters to the central server 105, it can be specifically sent to the sharding unit in the central server 105. The sharding unit performs a measurement process on the encrypted model parameters and divides them into multiple shares of the encrypted model parameters.

[0063] At the same time, the key management unit can also generate a pair of private key and public key, send the private key to the sharding unit, and send the public key to the aggregation unit.

[0064] When the sharding unit divides the encrypted model parameters in client 101 and obtains multiple shares, it can use the private key to sign each share to obtain the signature result of each share. Then, the sharding unit can send the multiple shares and the signature results of the multiple shares to the aggregation unit.

[0065] In response to the received multiple shares and the signature results of the multiple shares, the aggregation unit first verifies the signature results of the multiple shares based on the public key to determine whether the multiple shares are legal. When the multiple shares are legal, the aggregation unit can aggregate the multiple shares to obtain the aggregated share of the client 101 and send it to the collection unit.

[0066] The collection unit receives the aggregated shares of multiple clients, and the aggregated shares are encrypted aggregated shares. The collection unit can decrypt the encrypted aggregated shares through the keys of multiple clients sent by the key management unit to obtain the original aggregated shares. Based on the original aggregated shares of multiple clients, the collection unit can update the model parameters of the global model.

[0067] Thus, after obtaining the model parameters of the updated global model, the purpose of updating the global model is achieved. Further, during the training process, the central server can send the updated global model to all clients so that each client can use the updated model parameters in the next round of training process. This cycle continues until the model training reaches certain conditions and the training process ends, obtaining the final global model.

[0068] The above is the structure and working principle of a federated distributed aggregation system provided by an embodiment of the present application. Next, the federated distributed aggregation method of the embodiment of the present application will be introduced.

[0069] Figure 2 is a schematic flowchart of a federated distributed aggregation method provided by an embodiment of the present application. It should be understood that this method can be applied to Figure 1 the federated distributed aggregation system shown in Figure 1 the central server 105 in the system 100.

[0070] Exemplarily, as Figure 2 shown, the method 200 includes:

[0071] S201, splitting the global secret s during the training process of the target global model into n secret fragments, and sending the n secret fragments to n model training devices, where n≥1 and n is an integer.

[0072] It should be understood that during the process of training the model in federated distributed learning, multiple clients (i.e., model training devices) cooperate to train the same shared global model. In the initialization phase, the central server will customize the initial machine learning model structure and set hyperparameters (such as learning rate, batch size, etc.). After the central server completes the initialization, it can send the initialized global model parameters to multiple clients participating in the training.

[0073] It should also be understood that for each of the multiple clients, during the model training process, each client is equivalent to training a sub-model. Specifically, since the local datasets of each client are different, when each client uses its own local data to update a copy of the global model, this update process can be regarded as training a local model or a sub-model.

[0074] Among them, the target global model is the global model to be trained in the embodiments of the present application. The global secret s refers to the secret information used to protect the global model parameters or communication security, such as authentication credentials and other sensitive information. If these information are leaked, it may endanger the security or privacy of the federated distributed system.

[0075] In the process of federated distributed learning, in order to protect the security of the communication process, the central server can divide the global secret s into multiple secret fragments through the sharding unit and distribute them to different holders (clients). In the embodiments of the present application, only when a sufficient number of secret fragments are combined can the original global secret s be restored.

[0076] Specifically, in the embodiments of the present application, when dividing the global secret s, the method adopted is Shamir's Secret Sharing Scheme.

[0077] Shamir's Secret Sharing Scheme is a secret sharing method proposed by the Israeli cryptographer Adi Shamir in 1979. This algorithm allows a secret to be divided into multiple parts (shares), and these parts can be distributed to different participants. Only when a predetermined number of shares (i.e., the threshold k) are collected can the original secret be reconstructed.

[0078] In a possible implementation, the global secret s during the training process of the target global model is divided into n secret fragments, including:

[0079] Obtain the threshold t, where the threshold t is used to represent the number of secret fragments required to restore the global secret s;

[0080] According to the threshold t, determine a polynomial of degree t - 1;

[0081] According to the polynomial of degree t - 1 and the global secret s, determine n secret fragments.

[0082] Specifically, when Shamir's Secret Sharing Scheme divides a secret, it mainly includes the following steps: constructing a polynomial, dividing the secret, and distributing the secret.

[0083] Among them, when constructing the polynomial, the central server (fragmentation unit) needs to first set a threshold value t, which represents the minimum number of secret fragments required to recover the original global secret s. In the embodiments of the present application, the number of secret fragments to be divided is n.

[0084] After selecting the threshold value t, a polynomial f(x) of degree t - 1 is randomly selected. Among them, f(0) = s, that is, the constant term of the polynomial is the global secret s. The expression of the polynomial f(x) of degree t - 1 is shown in the following formula (1).

[0085] f(x) = s + a1x + a2x 2 +…+ a t-1 x t-1 Formula (1)

[0086] Among them, in formula (1):

[0087] a1, a2,... a t-1 are the coefficients of the polynomial respectively, and are randomly selected;

[0088] x: the independent variable in the polynomial, which is related to the number n of secret fragments, and is usually n consecutive integers (for example, 1, 2, 3... n);

[0089] f(x): the secret fragment. For each selected x, calculate f(x) to obtain the fragment (x, f(x)).

[0090] Exemplarily, as Figure 1 shown, assuming that the global secret s is 12345, the threshold value t is set to 3, and 5 secret fragments need to be generated. Assuming that the two randomly selected coefficients are a1 = 789 and a2 = 456 respectively, then the expression of the constructed polynomial is: f(x) = 12345 + 789x + 456x 2 .

[0091] The value of x is 1, 2, 3, 4. Then the corresponding f(1) = 12345 + 789×1 + 465×1, and so on. f(2), f(3), and f(4) can be calculated to obtain 4 secret fragments.

[0092] After obtaining the 4 secret fragments, the central controller can send the 4 secret fragments to the client 101, client 102, client 103, and client 104 respectively through the fragmentation unit.

[0093] When sending n secret fragments to n clients respectively, in the embodiments of the present application, a unique identity document number (Identity Document, ID) can be assigned to each client in advance, and the ID of each client is associated with the x value in the polynomial, so as to obtain the mapping relationship between the client and the secret fragment and store it.

[0094] After determining the n secret fragments, the central server can send the n secret fragments to the corresponding clients respectively based on the mapping relationship between the secret fragments and the clients. Among them, the client is the model training device in the embodiments of the present application.

[0095] Through the above process, each client participating in the model training can receive a part of the secret fragments. In the federated distributed learning process, the secret sharing mechanism is used to prevent Byzantine device attacks. Even if there are malicious clients trying to damage the model, as long as the number of legitimate clients exceeds the threshold, the global model can still be correctly restored and updated. In addition, each client can also receive the initial global model parameters of the target global model sent by the central server. After each client receives the initial global model parameters, the model is trained using the local dataset of the client to obtain the model parameters (such as weights or gradients) of the sub-model corresponding to the client. These model parameters are unencrypted model parameters and are called "original model parameters" in the embodiments of the present application.

[0096] S202. For any one of the n model training devices, determine the key of the model training device and send the key to the model training device. The key is used to enable the model training device to encrypt the obtained original model parameters to obtain encrypted model parameters.

[0097] Combined with the foregoing description, after each client receives the initial global model parameters, it can use the local dataset to perform model training to obtain the original model parameters of the corresponding sub-model.

[0098] When the client needs to send the original model parameters to the central server, in order to ensure the security of communication during the data transmission process between the client and the central server, the central server can generate a key for encrypting the original model parameters for each client through the key management unit before the client sends the original model parameters, so that the client encrypts the original model parameters based on the key to obtain encrypted model parameters.

[0099] Optionally, the type of the key in the embodiments of the present application can be a symmetric key.

[0100] Exemplarily, the key management unit can use a Cryptographically Secure Pseudo-Random Number Generator (CSPRNG) to generate a corresponding key for each client, and use the Hypertext Transfer Protocol Secure (HTTPS) encrypted by the Transport Layer Security (TLS) and Secure Sockets Layer (SSL) protocols to send the key of each client to the corresponding client.

[0101] In another example, when each client already has a public key, after the key management unit generates the key of each client through the CSPRNG, the generated key can be encrypted using the public key of the client, and then the encrypted key is sent to the client. The client decrypts it using its own private key to obtain the key.

[0102] It should be noted that after the key management unit generates a unique key for each client, when sending the key corresponding to each client to each client, the ID of each client can be associated with the key to obtain the mapping relationship between the client and the key, and based on the above mapping relationship, multiple keys are sent to the corresponding multiple clients.

[0103] For any one of the multiple clients, after receiving the key sent by the key management unit, the original model parameters of itself can be encrypted using the key to obtain encrypted model parameters.

[0104] In a possible implementation, the steps for the client to encrypt the original model parameters to obtain encrypted model parameters are as follows:

[0105] Determine the initial state matrix according to the original model parameters;

[0106] Perform a key expansion operation on the key to obtain round keys;

[0107] Perform a bitwise exclusive OR operation on the state matrix and the initial round key in the round keys to obtain the first state matrix;

[0108] Based on the preset number of rounds and the first state matrix, repeatedly perform byte substitution, row shift, column mixing, and round key addition operations to obtain the encrypted model parameters.

[0109] Optionally, when the client encrypts the original model parameters using a secret key (symmetric key), the encryption algorithms used include, but are not limited to, the Advanced Encryption Standard (AES), the Data Encryption Standard (DES), the Triple Data Encryption Algorithm (TDEA), the International Data Encryption Algorithm (IDEA), etc. In the embodiments of this application, the AES encryption algorithm is taken as an example to introduce the encryption process of the original model parameters in the embodiments of this application.

[0110] Specifically, the encryption process of the AES encryption algorithm can be summarized by the following formula (2).

[0111] X = AE.enc(c, x) Formula (2)

[0112] Wherein, in Formula (2):

[0113] x represents the plaintext, which is specifically the original model parameters in the embodiments of this application;

[0114] X represents the ciphertext, which is specifically the encrypted model parameters in the embodiments of this application;

[0115] c represents the secret key of the model training device.

[0116] When using the AES encryption algorithm to encrypt the original model parameters, the basic steps of the AES encryption algorithm can be summarized as: the initial round, multiple rounds of iteration, and the final round.

[0117] The following introduces and explains the content of these three parts.

[0118] (1) Initial round

[0119] In the initial round part, the client can perform plaintext preparation, key expansion operation, and initial round key addition operation.

[0120] It should be understood that the key lengths supported by the AES encryption algorithm are mainly 128 bits, 192 bits, and 256 bits. Based on different key lengths, the number of rounds of multiple rounds of iteration also varies. Generally, when the key length is 128 bits, the number of rounds of multiple rounds of iteration is 10 rounds; when the key length is 192 bits, the number of rounds of multiple rounds of iteration is 12 rounds; when the key length is 256 bits, the number of rounds of multiple rounds of iteration is 14 rounds. In the embodiments of this application below, the key length of 128 bits is taken as an example for illustration.

[0121] In the plaintext preparation phase, the client can organize the original model parameters into a 4×4 byte matrix, i.e., the initial state matrix. Each byte contains 8 bits of data, and there are 16 bytes in total.

[0122] In the key expansion phase, the client can use the key expansion algorithm to generate a series of round keys based on the original key length (128 bits).

[0123] In the initial round key addition phase, the client can perform a bitwise Exclusive OR (XOR) operation on the initial state matrix and the initial round key in the series of round keys (i.e., the first sub-key obtained by expanding the original key), resulting in the first state matrix. This first state matrix will be used as the input for the second round of operations.

[0124] (2) Multiple rounds of iteration

[0125] In the multiple rounds of iteration phase, when the number of rounds is determined based on the key length, for the specified number of rounds, the following four operations are repeatedly executed: Sub Bytes, Shift Rows, Mix Columns, and Add Round Key.

[0126] Among them, Sub Bytes means replacing each byte in the state matrix with an S-box; Shift Rows means performing a cyclic left shift on each row in the state matrix; Mix Columns means applying a linear transformation to each column in the state matrix; Add Round Key means performing an XOR operation on the current round's sub-key and the state matrix.

[0127] (3) Final round

[0128] The above process of multiple rounds of iteration continues until the second-to-last round. In the last round, the client only performs the operations of Sub Bytes, Shift Rows, and Add Round Key.

[0129] After the operations in the final round are completed, the resulting final state matrix is the encrypted model parameters, and the encrypted model parameters are usually converted into a string of consecutive byte form.

[0130] And so on, each client can perform the AES encryption operation on its own original model parameters using its own key in the above manner to obtain the encrypted model parameters.

[0131] After obtaining the encrypted model parameters, each client can further send its own encrypted model parameters to the central server. Thus, the central server can obtain the encrypted model parameters sent by each client.

[0132] S203. Split the encrypted model parameters to obtain m encrypted shares, where m≥1 and m is an integer.

[0133] In response to the n encrypted model parameters sent by n clients, for the encrypted model parameters of any one of the clients, to ensure the privacy and security of the data and the transmission efficiency of the data, the central server can split the encrypted model parameters through a sharding unit to obtain m encrypted shares of the current encrypted model parameters.

[0134] Specifically, when the sharding unit splits the encrypted model parameters, it can perform a measurement process on the encrypted model parameters to obtain m encrypted shares.

[0135] The measurement process refers to any data splitting and distribution technology for ensuring data security and privacy protection. Optionally, the splitting methods adopted in the measurement process in the embodiments of the present application include, but are not limited to, the Shamir secret sharing method, Blakley's geometric secret sharing, Asmuth-Bloom secret sharing, homomorphic encryption method, or additive homomorphic encryption method, etc. In the embodiments of the present application, taking the Shamir secret sharing method as an example, when splitting the encrypted model parameters, similar to the process of splitting the global secret s as described above, the sharding unit can construct a polynomial by determining a threshold value, and calculate each secret share. The specific splitting process is the same as that of splitting the global secret s and will not be elaborated here.

[0136] Through the above process, the central server can perform a measurement process on the encrypted model parameters of each client to obtain m encrypted shares corresponding to each encrypted model parameter. Therefore, there are a total of n×m encrypted shares corresponding to n model training devices.

[0137] S204. Aggregate the m encrypted shares to obtain the aggregated share of the model training device.

[0138] For the m encrypted shares corresponding to any one client, since the sharding unit splits the encrypted model parameters, it is necessary to re-aggregate the m encrypted splits.

[0139] After the sharding unit splits an encrypted model parameter into m shares, the m shares can be sent to the aggregation unit corresponding to the client.

[0140] Optionally, in the embodiments of the present application, one client corresponds to one aggregation unit, and those skilled in the art can preset the corresponding relationship between the client and the aggregation unit in advance and store it. After the sharding unit divides the encrypted model parameters of a certain client and obtains m shares, based on the corresponding relationship between the client and the aggregation unit, the current m shares can be correspondingly sent to the aggregation unit associated with the current client.

[0141] It should be understood that after any aggregation unit receives the m shares obtained by dividing the corresponding client by the sharding unit, in order to ensure that the m shares are not maliciously tampered with during the transmission process and the authenticity of the source of the m shares (indeed sent by the sharding unit), the embodiments of the present application also propose a data verification strategy when the sharding unit sends the m shares to the aggregation unit, specifically by verifying the data exchange process between the sharding unit and the aggregation unit through a digital signature algorithm.

[0142] In a possible implementation manner, before aggregating the m encrypted shares to obtain the aggregated shares of the model training device, the method further includes:

[0143] Generating a pair of public key d pk and private key d sk ;

[0144] Based on the private key d sk , signing the m encrypted shares to obtain m signature results;

[0145] Based on the public key d pk and the m signature results, verifying the m encrypted shares;

[0146] Among them, based on the private key d sk , the formula for signing the m encrypted shares to obtain m signature results is as follows:

[0147] SIG.sign(d sk , m) → σ;

[0148] m represents the encrypted share; σ represents the signature result of the encrypted share;

[0149] Among them, the formula for verifying the m encrypted shares based on the public key d pk and the m signature results is as follows:

[0150] SIG.ver(d pk , m, σ) → 0, 1;

[0151] Moreover, aggregating the m encrypted shares to obtain the aggregated shares of the model training device includes:

[0152] When all of the m encrypted shares are successfully verified, the m encrypted shares are aggregated to obtain the aggregated share of the model training device.

[0153] In the embodiment of the present application, the key management unit may generate a pair of public key d pk and private key d sk . Specifically, the key management unit may use the key generation function SIG.gen(k) of the digital signature algorithm to generate a pair of public key d pk and private key d sk based on a random number or a seed k, and this process can be represented by the following formula (3).

[0154] SIG.gen(k) → (d pk , d sk ) Formula (3)

[0155] Wherein, in formula (3):

[0156] d pk : Public key;

[0157] d sk : Private key.

[0158] After obtaining a pair of public key d pk and private key d sk , the key management unit may send the private key d sk to the sharding unit, and the sharding unit signs each of the m encrypted shares corresponding to each client based on the private key d sk . Among them, the process of the sharding unit signing each of the m encrypted shares for each client based on the private key d sk can be represented by the following formula (4).

[0159] SIG.sign(d sk , m) → σ Formula (4)

[0160] Wherein, in formula (4):

[0161] d sk : Private key;

[0162] m: Any encrypted share;

[0163] σ: The signature result of this encrypted share.

[0164] Optionally, during the signature process of any encrypted share, the digital signature algorithms include Rivest-Shamir-Adleman (RSA), Digital Signature Algorithm (DSA), Elliptic Curve Digital Signature Algorithm (ECDSA), Edwards-curve Digital Signature Algorithm (EdDSA), Schnorr algorithm, etc. In the following embodiments of the present application, the digital signature algorithm is taken as RSA as an example to introduce the signature process of the encrypted share.

[0165] When using the digital signature algorithm to sign the encrypted share, it specifically includes the following steps: calculating the hash value of the encrypted share, signing the hash value with the private key, and outputting the signature result.

[0166] Exemplarily, in the case where the digital signature algorithm is RSA, first, the sharding unit can determine the hash function (for example, Secure Hash Algorithm 256-bit, or Secure Hash Algorithm 3). In the embodiments of the present application, SHA-256 is taken as an example of the hash function.

[0167] For any one of the m encrypted shares, the sharding unit can perform a hash operation on the encrypted share using the selected hash function to obtain the hash value. Further, regarding the hash value as an integer, use the private key d sk Encrypt the hash value to generate the signature σ and output it.

[0168] After signing each of the m encrypted shares, m signatures can be obtained. The sharding unit can send the m signatures and the m encrypted shares to the aggregation unit.

[0169] After receiving the m encrypted shares and the m signatures, the aggregation unit can use the public key provided by the key management unit and the m signature results to verify the m encrypted shares to determine whether the m encrypted shares have been maliciously tampered with during the transmission process.

[0170] Specifically, the verification process of the aggregation unit includes the following steps: calculating the hash value of the message, decrypting the signature with the public key, and comparing the hash value and the decryption result.

[0171] For any one of the m received encrypted shares, the aggregation unit can calculate the hash value of the encrypted share and use the public key dpk Decrypt the signature corresponding to the encrypted share to obtain the decrypted value. Further, the aggregation unit can compare the decrypted value with the hash value calculated by the aggregation unit. When the decrypted value is the same as the hash value calculated by the aggregation unit, it indicates that the current encrypted share is successfully verified, and the aggregation unit outputs result 1; on the contrary, when the decrypted value is different from the hash value calculated by the aggregation unit, it indicates that the current encrypted share verification fails, and the aggregation unit outputs result 0.

[0172] When the aggregation unit successfully verifies all m encrypted shares corresponding to the current client, it indicates that the m encrypted shares have not been tampered with during the transmission process. The aggregation unit can further aggregate the m encrypted shares to obtain the aggregated share corresponding to the current m encrypted shares.

[0173] Among them, the aggregated share specifically refers to the original encrypted model parameters before segmentation.

[0174] Exemplarily, in the case of using the Shamir secret sharing algorithm to divide the encrypted model parameters of each client into m encrypted shares, when the aggregation unit aggregates the m encrypted shares of any client, the Lagrange interpolation method can be used to reconstruct the original encrypted model parameters from the m encrypted shares.

[0175] Specifically, the aggregation unit collects at least t encrypted shares, uses the Lagrange interpolation formula to reconstruct the polynomial f(x), and calculates the value of the polynomial x = 0, f(0), to obtain the encrypted model parameters.

[0176] Thus, each aggregation unit corresponding to a client can perform the above reconstruction operation on the m encrypted shares received, and then obtain the encrypted model parameters of each client, that is, the aggregated share.

[0177] S205, update the target global model based on the aggregated shares of n model training devices.

[0178] Exemplarily, as Figure 1 shown, after the n aggregation units corresponding to n clients respectively obtain the n aggregated shares of the n clients, the n aggregation units can send the n aggregated shares to the collection unit.

[0179] Combined with the foregoing description, when the key management unit generates keys for each client and sends the keys to each client respectively, it can also send them to the collection unit at the same time.

[0180] Based on the keys of each client, the collection unit can decrypt the n aggregated shares respectively to obtain the original model parameters of the n clients, and then update the target global model based on the original model parameters.

[0181] In a possible implementation, updating the target global model based on the aggregated shares of n model training devices includes:

[0182] Decrypting the aggregated shares of the n model training devices based on the keys of the n model training devices to obtain the original aggregated shares of the n model training devices;

[0183] Performing feature extraction on the original aggregated shares of the n model training devices to obtain the characterization data of the n sub-models corresponding to the n model training devices;

[0184] Updating the target global model based on the characterization data of the n sub-models.

[0185] Specifically, the process of the collection unit decrypting the n aggregated shares can be represented by the following formula (5).

[0186] AE.dec(c, AE.enc(c, x)) = x Formula (5)

[0187] Where, in Formula (5):

[0188] x: Original model parameters;

[0189] c: Key.

[0190] Specifically, when the collection unit decrypts the encrypted model parameters through the key, the process can be regarded as the inverse process of AES encryption. First, the collection unit expands the key according to the key length to generate a set of round keys. During the decryption process, the collection unit first performs the inverse operations of the inverse byte substitution (Inv SubBytes) operation, the inverse row shift (Inv Shift Rows) operation, the inverse column mixing (Inv Mix Columns) operation, and the round key addition operation on the encrypted model parameters using the last round of round keys. Correspondingly, the inverse byte substitution means restoring the replaced bytes during encryption using the S-box; the inverse row shift means restoring the shifted rows during encryption; the inverse column mixing means restoring the mixed columns during encryption.

[0191] Thus, through the above decryption process, the collection unit can decrypt the n aggregated shares one by one to obtain n original model parameters, that is, n original aggregated shares.

[0192] After obtaining the n original model parameters, the collection unit can perform feature extraction on any one of the n original model parameters to obtain the characterization data of the n sub-models corresponding to the n clients. The characterization data mainly refers to the main features of the input data of the sub-models.

[0193] After obtaining the representation data of the n sub-models, the collection unit may update the target global model based on the representation data of the n sub-models.

[0194] In a possible implementation, updating the target global model based on the representation data of the n sub-models includes:

[0195] Performing representation matching processing on the representation data of the n sub-models to obtain the processed representation data of the n sub-models, and the processed representation data of the n sub-models has the same dimension;

[0196] Based on the representation fusion method, fusing the processed representation data of the n sub-models to obtain the fused representation data;

[0197] Constructing a target meta-learner according to the fused representation data;

[0198] Determining the final decision based on the comprehensive representation decision of the target meta-learner and the sub-model decisions of the n model training devices;

[0199] Updating the target global model according to the final decision.

[0200] For the representation data of the n sub-models, the collection unit may perform representation matching processing on it to unify the representation data of the n sub-models into the same dimensional space. The steps of the representation matching processing include the following: feature alignment, normalization processing, and similarity calculation.

[0201] Among them, for feature alignment, the collection unit may map the representation data of different dimensions to a common space through linear or non-linear transformation.

[0202] For normalization processing, the collection unit may normalize the representation data (for example, normalization or Z-score normalization) to eliminate the differences between dimensions.

[0203] For similarity calculation, the methods for calculating similarity in the embodiments of the present application include but are not limited to Euclidean distance, cosine similarity, etc.

[0204] Thus, through the above process, the collection unit can obtain the processed representation data of the n sub-models.

[0205] In the representation fusion stage, for the representation data of the n sub-models, the collection unit may use the representation fusion method (for example, weighted average, deep neural network fusion layer, etc.) to fuse the representation data of the n sub-models.

[0206] Exemplarily, when the representation fusion method is weighted average, the collection unit may assign weights according to the performance of each sub-model and perform weighted average on all the representation data to obtain the fused representation data.

[0207] Exemplarily, when the characterization fusion method is a deep neural network fusion method, in the embodiments of the present application, a multi-input multi-output neural network model structure can be used. The processed characterization data of multiple sub-models is used as the input, and a fusion layer is trained as the output layer to synthesize this characterization data to obtain the fused characterization data.

[0208] After obtaining the fused characterization data, a person skilled in the art can divide the fused characterization data into a training set and a validation set, select a suitable model structure and perform hyperparameter tuning, use the training set to train a meta-learner, and evaluate the performance of the meta-learner on the validation set, and adjust the meta-learner until the best performance is achieved, so as to obtain the target meta-learner. The purpose of constructing the target meta-learner in the embodiments of the present application is to synthesize the knowledge of each sub-model, so as to provide more powerful decision-making capabilities.

[0209] Optionally, the target meta-learner can be any type of machine learning model, including but not limited to a random forest model, a support vector machine, or a neural network, etc.

[0210] After obtaining the target meta-learner, in the inference stage, the collection unit can predict the newly input data through multiple sub-models to obtain the decisions of the sub-models. At the same time, the collection unit can predict the newly input data through the meta-learner to generate a comprehensive characterization decision.

[0211] After obtaining the decisions of the sub-models and the comprehensive characterization decision, the meta-learner can adopt a majority voting method, or a weighted voting method, or train a stacking model to obtain the final decision, which refers to the comprehensive decision of the distributed large model, and it fuses the prediction results of the meta-learner and other each sub-model.

[0212] After obtaining the final decision, the collection unit can update the target global model according to the final decision.

[0213] In a possible implementation manner, updating the target global model according to the final decision includes:

[0214] Sending the final decision to n model training devices, so that the n model training devices update the model parameters of the n model training devices according to the final decision;

[0215] Obtaining the updated model parameters of the n model training devices;

[0216] Aggregating the updated model parameters of the n model training devices to obtain the aggregated model parameters;

[0217] Applying the aggregated model parameters to the target global model to update the model parameters of the target global model.

[0218] Specifically, the collection unit sends the obtained final decision to each client, and each client updates its local sub-model according to the final decision.

[0219] After the update is completed, the client sends the updated model parameters to the collection unit, and the collection unit aggregates the model parameters of all clients using an appropriate aggregation algorithm (e.g., Federated Averaging (FedAvg)). Specifically, the collection unit can allocate weights according to the data volume or model performance of each client, and perform weighted averaging on the model parameters to obtain the aggregated model parameters.

[0220] Furthermore, the collection unit can directly update the model parameters of the target global model to the aggregated model parameters, thereby completing the update of the target global model.

[0221] After updating the target global model, the collection unit can also distribute the updated global model to all clients so that each client can use the latest global model parameters in the next round of training.

[0222] In addition, when there is data transmission between multiple clients, the embodiments of the present application can also provide a communication step to improve the data privacy between clients.

[0223] Specifically, assuming there is data exchange between client A and client B, client A and client B can ensure the security of their communication through the Diffie-Hellman (DH) key exchange method.

[0224] Exemplarily, client A and client B first negotiate public parameters, including prime number q and generator g. Among them, 1 < g < q.

[0225] The process of client A and client B negotiating public parameters can be represented by the following formula (6).

[0226] KA.param(k)→(G′, g, q, H) Formula (6)

[0227] Among them, in formula (6):

[0228] g: Generator;

[0229] q: Prime number;

[0230] H: Hash function;

[0231] G′: Group related to g and q.

[0232] After obtaining g and q, for client A, a random private key x can be selectedA , calculate the public key g xA , and send the public key g xA to client B. Similarly, for client B, a random private key x B can be selected, calculate the public key g xB , and send the public key g xB to client A. Its specific formula is shown in formula (7) below.

[0233] KA.gen(G′, g, q, H) → (x.g x ) Formula (7)

[0234] Where, in formula (7):

[0235] x: the private key of client A, or the private key of client B;

[0236] g x : the public key generated by client A, or the public key generated by client B.

[0237] After client A receives the public key g xB sent by client B, use its own private key x A and the public key g xB of client B to calculate the shared key, denoted as "S A,B ". Similarly, after client B receives the public key g xA sent by client A, use its own private key x B and the public key g xA of client B to calculate the shared key, denoted as "S B,A ".

[0238] It should be understood that S A,B and S B,A should be the same, so that client A and client B can obtain the shared key required for their communication, facilitating the communication between client A and client B.

[0239] In summary, when training a model using the federated distributed method, the present application proposes a federated distributed aggregation method. During the model training process, this method divides the global secret into multiple secret fragments and distributes them to multiple model training devices (clients), which can ensure that even if a part of the clients are attacked, the attacker cannot obtain the complete secret information. Further, for the model parameters of each sub-model trained by the client, specific keys are used for encryption to ensure that no sensitive information is leaked during the process of the client transmitting data to the central server. After receiving the encrypted model parameters sent by the client, the encrypted model parameters are divided into multiple shares, which improves the efficiency during the data transmission process, significantly reduces the transmission time, and at the same time can increase the difficulty of attack. Even if the attacker obtains some shares, the complete model parameters cannot be restored. Thus, during the model training process, the present application improves the training efficiency of the model by parallel processing of training tasks by multiple model training devices, compared with the centralized learning method, and ensures the privacy and security of the data during the model training process.

[0240] To facilitate the understanding of the method of the embodiments of the present application, the following Figure 3 introduces and illustrates the distributed training process in the embodiments of the present application.

[0241] Figure 3 is an interaction scenario diagram of a federated distributed training model provided by an embodiment of the present application.

[0242] Exemplarily, as Figure 3 shown, in combination with Figure 1 , during the implementation of the federated distributed aggregation process, when each client receives the global model parameters, a certain number of training iterations are performed on the local dataset of the client to update the model parameters. After the client completes the local training, the updated model parameters or gradient differences can be sent to the sharding unit in the central server 105.

[0243] Before the client sends to the sharding unit, the key management unit in the central server 105 can generate a unique key for each client, such as Figure 3 k1, k2, and k3 in

[0244] The sharding unit divides each encrypted model parameter to obtain m encrypted shares. To ensure the security of the encrypted shares during transmission, the key management unit can also generate a pair of public and private keys and send the private key to the sharding unit, enabling the sharding unit to sign each of the m encrypted shares of the encrypted model parameter to obtain a signature result. In addition, the key management unit can send the public key to the aggregation unit for verification.

[0245] For any client, the sharding unit further sends the signatures of the m encrypted shares and the m encrypted shares to the aggregation unit corresponding to the client, enabling the aggregation unit to verify the m encrypted shares based on the signatures of the m encrypted shares. After the m encrypted shares pass the verification, the aggregation unit can verify the m encrypted shares to obtain aggregation share 1, aggregation share 2, and aggregation share 3.

[0246] Aggregation unit 1, aggregation unit 2, and aggregation unit 3 can respectively send aggregation share 1, aggregation share 2, and aggregation share 3 to the collection unit. The collection unit decrypts the aggregation shares based on the keys corresponding to each client to obtain the original aggregation shares, i.e., the original model parameters.

[0247] Finally, the collection unit performs feature extraction, feature matching, and feature fusion processing based on the original model parameters of multiple clients to update the global model parameters of the target global model.

[0248] Figure 4 It is a schematic structural diagram of a federated distributed aggregation device provided by an embodiment of the present application.

[0249] Exemplarily, as Figure 4 shown, the device 400 includes:

[0250] A first splitting module 401, configured to split the global secret s during the training process of the target global model into n secret fragments and send the n secret fragments to n model training devices, where n≥1 and n is an integer;

[0251] A key generation module 402, configured to determine the key of any one of the n model training devices and send the key to the model training device, where the key is used to encrypt the originally trained model parameters by the model training device to obtain encrypted model parameters;

[0252] A second splitting module 403, configured to split the encrypted model parameters to obtain m encrypted shares, where m≥1 and m is an integer;

[0253] An aggregation module 404, configured to aggregate the m encrypted shares to obtain the aggregation share of the model training device;

[0254] A model update module 405, configured to update the target global model based on the aggregation shares of the n model training devices.

[0255] In a possible implementation, the first splitting module 401 is specifically configured to: obtain a threshold value t, where the threshold value t is used to represent the number of secret fragments required to recover the global secret s; determine a polynomial of degree t - 1 according to the threshold value t; and determine the n secret fragments according to the polynomial of degree t - 1 and the global secret s.

[0256] In a possible implementation, the step of encrypting the originally trained model parameters by the model training device to obtain encrypted model parameters includes: determining an initial state matrix according to the originally trained model parameters; performing a key expansion operation on the key to obtain round keys; performing a bitwise exclusive OR operation on the state matrix and the initial round key in the round keys to obtain a first state matrix; and repeatedly performing byte substitution, row shift, column mixing, and round key addition operations based on a preset number of rounds and the first state matrix to obtain the encrypted model parameters.

[0257] In a possible implementation, before aggregating the m encrypted shares to obtain the aggregation share of the model training device, the key generation module 402 is further configured to: generate a pair of public key d pk and private key d sk ;

[0258] sign the m encrypted shares based on the private key d sk to obtain m signature results;

[0259] verify the m encrypted shares based on the public key d pk and the m signature results;

[0260] wherein, the formula for signing the m encrypted shares based on the private key d sk to obtain m signature results is as follows:

[0261] SIG.sign(d sk , m) → σ;

[0262] m represents the encrypted share; σ represents the signature result of the encrypted share;

[0263] wherein, the formula for verifying the m encrypted shares based on the public key d pk and the m signature results is as follows:

[0264] SIG.ver(d pk , m, σ) → 0, 1;

[0265] Moreover, the aggregation module 403 is specifically configured to: when all of the m encrypted shares are successfully verified, aggregate the m encrypted shares to obtain the aggregated share of the model training device.

[0266] In a possible implementation, the model update module 405 is specifically configured to: decrypt the aggregated shares of the n model training devices based on the keys of the n model training devices to obtain the original aggregated shares of the n model training devices; perform feature extraction on the original aggregated shares of the n model training devices to obtain the feature data of the n sub-models corresponding to the n model training devices; and update the target global model based on the feature data of the n sub-models.

[0267] In a possible implementation, the model update module 405 is further configured to: perform feature matching processing on the feature data of the n sub-models to obtain the processed feature data of the n sub-models, where the processed feature data of the n sub-models have the same dimension; fuse the processed feature data of the n sub-models based on a feature fusion method to obtain the fused feature data; construct a target meta-learner according to the fused feature data; determine a final decision based on the comprehensive feature decision of the target meta-learner and the sub-model decisions of the n model training devices; and update the target global model according to the final decision.

[0268] In a possible implementation, the model update module 405 is further configured to: send the final decision to the n model training devices, so that the n model training devices update the model parameters of the n model training devices according to the final decision; obtain the updated model parameters of the n model training devices; aggregate the updated model parameters of the n model training devices to obtain the aggregated model parameters; and apply the aggregated model parameters to the target global model to update the model parameters of the target global model.

[0269] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0270] Exemplarily, as Figure 5 shown, the electronic device 500 includes: a memory 501 and a processor 502, where an executable program code 5011 is stored in the memory 501, and the processor 502 is configured to call and execute the executable program code 3011 to execute a federated distributed aggregation method.

[0271] In this embodiment, the functional modules of the electronic device can be divided according to the above method examples. For example, each functional module can be corresponded, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0272] In the case of dividing each functional module according to each function, the electronic device may include: a first segmentation module, a key generation module, a second segmentation module, an aggregation module, a model update module, etc. It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here.

[0273] The electronic device provided in this embodiment is used to execute the above method of federated distributed aggregation, so the same effect as the above implementation method can be achieved.

[0274] In the case of adopting an integrated unit, the electronic device may include a processing module and a storage module. Among them, the processing module can be used to control and manage the actions of the electronic device. The storage module can be used to support the electronic device to execute relevant program codes and data, etc.

[0275] Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logical blocks, modules and circuits shown in combination with the disclosure of this application. The processor can also be a combination that realizes computing functions, such as including a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module can be a memory.

[0276] This embodiment also provides a computer-readable storage medium, in which computer program codes are stored. When the computer program codes run on a computer, the computer is enabled to execute the above relevant method steps to implement a federated distributed aggregation method in the above embodiment.

[0277] This embodiment also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above relevant steps to implement a federated distributed aggregation method in the above embodiment.

[0278] In addition, the electronic device provided in the embodiment of this application can specifically be a chip, a component or a module. The electronic device may include a processor and a memory connected thereto; among them, the memory is used to store instructions. When the electronic device runs, the processor can call and execute the instructions to enable the chip to execute a federated distributed aggregation method in the above embodiment.

[0279] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0280] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0281] In the embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0282] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A federated distributed aggregation method, characterized in that: The method comprises: Splitting the global secret s in the target global model training process into n secret fragments, and sending the n secret fragments to n model training devices, where n≥1 and n is an integer; For any model training device among the n model training devices, determining a key of the model training device and sending the key to the model training device, wherein the key is used to enable the model training device to encrypt original model parameters obtained through training to obtain encrypted model parameters; The encryption model parameters are divided to obtain m encryption shares, where m≥1 and m is an integer; Aggregating the m encrypted shares to obtain an aggregated share of the model training device; Based on the aggregated shares of the n model training devices, the target global model is updated.

2. The method according to claim 1, characterized in that The method of dividing the global secret s in the target global model training process into n secret fragments includes: Obtaining a threshold value t, where the threshold value t is used to represent the number of secret fragments required to recover the global secret s; Determine a t-1 degree polynomial according to the threshold value t; The n secret fragments are determined according to the t-1 degree polynomial and the global secret s.

3. The method according to claim 1, characterized in that The model training device encrypts the original model parameters obtained through training, and the step of obtaining the encrypted model parameters includes: Determine an initial state matrix according to the original model parameters; Performing a key expansion operation on the key to obtain a round key; Performing a bitwise XOR operation on the state matrix and an initial round key in the round keys to obtain a first state matrix; Based on a preset number of rounds and the first state matrix, byte substitution, row shift, column mixing and round key addition operations are repeatedly performed to obtain the encryption model parameters.

4. The method according to claim 1 or 2, characterized in that: Before aggregating the m encrypted shares to obtain the aggregated share of the model training device, the method further includes: Generate a pair of public keys d pk With private key d sk ; Based on the private key d sk , sign the m encrypted shares to obtain m signature results; Based on the public key d pk and the m signature results, verifying the m encrypted shares; Wherein, the private key d sk , the m encrypted shares are signed, and the expression formula for obtaining the m signature results is as follows: SIG.sign(d sk ,m)→σ; m represents the encrypted share; σ represents the signature result of the encrypted share; Among them, based on the public key d pk The formula for verifying the m encrypted shares using the m signature results is as follows: SIG.ver(d pk ,m,σ)→0.1; And, aggregating the m encrypted shares to obtain the aggregated share of the model training device includes: When the m encrypted shares are all verified successfully, the m encrypted shares are aggregated to obtain the aggregated share of the model training device.

5. The method according to claim 4, characterized in that The updating of the target global model based on the aggregated shares of the n model training devices includes: Decrypting the aggregated shares of the n model training devices based on the keys of the n model training devices to obtain original aggregated shares of the n model training devices; Performing characterization extraction on the original aggregated shares of the n model training devices to obtain characterization data of the n sub-models corresponding to the n model training devices; The target global model is updated based on the representation data of the n sub-models.

6. The method according to claim 1, characterized in that The updating of the target global model based on the representation data of the n sub-models includes: Performing a representation matching process on the representation data of the n sub-models to obtain the processed representation data of the n sub-models, wherein the processed representation data of the n sub-models have the same dimension; Based on the representation fusion method, the representation data of the n sub-models after the processing are fused to obtain fused representation data; Constructing a target meta-learner according to the fused representation data; Determining a final decision based on the comprehensive representation decision of the target meta-learner and the sub-model decisions of the n model training devices; According to the final decision, the target global model is updated.

7. The method according to claim 6, characterized in that The updating of the target global model according to the final decision includes: Sending the final decision to the n model training devices, so that the n model training devices update the model parameters of the n model training devices according to the final decision; Obtaining updated model parameters of the n model training devices; Aggregating the updated model parameters of the n model training devices to obtain aggregated model parameters; The aggregated model parameters are applied to the target global model to update the model parameters of the target global model.

8. A federated distributed aggregation device, characterized in that: The device comprises: A first segmentation module is used to segment the global secret s in the target global model training process into n secret fragments, and send the n secret fragments to n model training devices, where n≥1 and n is an integer; A key generation module, used for determining a key of any one of the n model training devices and sending the key to the model training device, wherein the key is used to enable the model training device to encrypt the original model parameters obtained through training to obtain encrypted model parameters; A second segmentation module is used to segment the encryption model parameters to obtain m encryption shares, where m≥1 and m is an integer; an aggregation module, used to aggregate the m encrypted shares to obtain an aggregated share of the model training device; A model updating module is used to update the target global model based on the aggregated shares of the n model training devices.

9. An electronic device, characterized in that: The electronic device comprises: A memory for storing executable program codes; A processor, configured to call and run the executable program code from the memory, so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 is implemented.