Distributed machine learning method and device based on fault-tolerant learning, equipment and medium

By adopting a fault-tolerant learning-based encryption method in distributed machine learning, data is encrypted and processed and model parameters are securely transmitted, the problem of data privacy leakage in distributed machine learning is solved, and efficient and secure distributed machine learning training is achieved.

CN120218277APending Publication Date: 2025-06-27GUANGZHOU FEISHU DATA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510140325.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Distributed machine learning faces the risk of data privacy leakage when processing large-scale data, and traditional methods are difficult to balance between data privacy protection and computing efficiency.

Method used

A distributed machine learning method based on fault-tolerant learning is adopted to encrypt the original data on the central node by encrypting the fault-tolerant learning problem, encrypted data is generated, and distributed to multiple computing nodes for training. The trained model parameters are transmitted to the central node through encryption, decrypted and aggregated to determine the final machine learning model.

Benefits of technology

This method not only improves the security of data privacy and prevents data leakage during transmission, but also realizes efficient distributed machine learning training, meeting the dual needs of computing efficiency and data privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218277A_ABST
    Figure CN120218277A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed machine learning method and device based on fault-tolerant learning, equipment and a medium, and relates to the technical field of machine learning, and the method comprises the steps: obtaining data used for training a machine learning model as original data; encrypting the original data based on the fault-tolerant learning problem to obtain encrypted data; transmitting the encrypted data to a plurality of computing nodes, so that each computing node respectively trains a machine learning model by using the encrypted data; acquiring an encryption result transmitted by each computing node; the encryption result is obtained by encrypting model parameters of the trained machine learning model based on a fault-tolerant learning problem; and determining a trained machine learning model according to an encryption result. Original data is encrypted through a fault-tolerant learning problem to obtain encrypted data, and each computing node uses the encrypted data to train a machine learning model instead of plaintext original data, so that the data security is improved; and the central node decrypts the encryption result transmitted by the computing node to obtain model parameters, so that efficient distributed machine learning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and in particular to a distributed machine learning method, device, equipment and medium based on fault-tolerant learning. Background Art

[0002] With the advent of the big data era, the rapid growth of data scale and complexity has posed unprecedented challenges to machine learning models. When dealing with large-scale data sets, traditional single-machine learning models are often limited by computing resources and storage capabilities, and it is difficult to meet the requirements of high efficiency and real-time. Therefore, distributed machine learning has emerged. It distributes computing tasks across multiple computers to achieve parallel processing of large-scale data, significantly improving computing efficiency and scalability. However, while bringing computing advantages, distributed machine learning also faces the risk of data privacy leakage. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a distributed machine learning method, device, equipment and medium based on fault-tolerant learning to improve the security of data privacy and perform distributed machine learning efficiently.

[0004] To achieve the above object, on the one hand, an embodiment of this application proposes a distributed machine learning method based on fault-tolerant learning. The method is applied to a central node and includes the following steps:

[0005] Obtain data for training a machine learning model as original data;

[0006] Encrypt the original data based on a fault-tolerant learning problem to obtain encrypted data;

[0007] Transmit the encrypted data to multiple computing nodes for each of the computing nodes to train the machine learning model using the encrypted data respectively;

[0008] Obtain the encrypted results transmitted by each of the computing nodes; wherein, the encrypted results are obtained by encrypting the model parameters of the trained machine learning model based on the fault-tolerant learning problem;

[0009] Determine the trained machine learning model according to the encrypted results.

[0010] In some embodiments, encrypting the original data based on a fault-tolerant learning problem to obtain encrypted data includes the following steps:

[0011] Generate a public key and a private key;

[0012] Calculate the ciphertext corresponding to the original data based on the definition of the fault-tolerant learning problem and the public key to encrypt the original data to obtain the encrypted data; wherein, the private key is used to decrypt the encrypted data;

[0013] The expression for calculating the ciphertext is as follows:

[0014] c = As + er + e′;

[0015] Where c represents the ciphertext, A represents the public key matrix corresponding to the public key, s represents the original data, e represents the error vector, r represents the random vector, and e′ represents the additional error vector.

[0016] In some embodiments, determining the trained machine learning model based on the encryption result includes the following steps:

[0017] Perform model update aggregation using the encryption result and the private key of the fault-tolerant learning problem to decrypt and aggregate each encryption result to obtain the model parameters;

[0018] The expression for the model parameters is:

[0019]

[0020] Where w t+1 represents the updated model parameters obtained by decryption, w t represents the model parameters before update, η represents the learning rate, Dec represents the decryption function of the fault-tolerant learning problem, Enc represents the encryption function of the fault-tolerant learning problem; L represents the total prediction loss value of the machine learning model, represents the gradient;

[0021] The expression for the total prediction loss value includes:

[0022] Enc(L) = ∑ i Enc(L i );

[0023]

[0024] Where L i represents the prediction loss value of the i-th computing node; represents the predicted value of the j-th feature in the i-th computing node; represents the square of the true value of the j-th feature in the i-th computing node; the encrypted data includes each feature;

[0025] Configure the model parameters into the machine learning model to obtain the trained machine learning model.

[0026] In some embodiments, the method further includes at least one of the following steps:

[0027] Receive the model prediction results of the encrypted states transmitted by each of the computing nodes; wherein, the model prediction results are obtained by encrypting the prediction results output by each of the computing nodes based on the fault-tolerant learning problem for the trained machine learning model; decrypt the model prediction results of the encrypted state using the private key of the fault-tolerant learning problem to obtain the model prediction results in the plaintext state;

[0028] Alternatively, verify whether the model parameters are within a preset numerical range;

[0029] Alternatively, verify whether the trained machine learning model satisfies preset constraint conditions or rules.

[0030] To achieve the above object, another aspect of the embodiments of the present application proposes a distributed machine learning method based on fault-tolerant learning. The method is applied to a computing node in the distributed machine learning method based on fault-tolerant learning as described above, and the method includes the following steps:

[0031] Obtain the encrypted data transmitted by the central node;

[0032] Train a machine learning model using the encrypted data to obtain an encrypted result; wherein, the encrypted result is obtained by encrypting the model parameters of the trained machine learning model based on a fault-tolerant learning problem;

[0033] Transmit the encrypted result to the central node.

[0034] In some embodiments, the training of the machine learning model using the encrypted data to obtain an encrypted result includes the following steps:

[0035] Randomly initialize the parameters of the machine learning model to obtain an initial model, or obtain a pre-trained model as the initial model;

[0036] Calculate the predicted value of the initial model for the encrypted data;

[0037] The expression for calculating the predicted value includes:

[0038]

[0039] wherein, represents the predicted value corresponding to the j-th feature of the encrypted data in the t-th training round, is the model parameter in the t-th training round, ∈ is the added noise; x tj represents the sample value corresponding to the j-th feature of the encrypted data in the t-th training round; Enc represents the encryption function of the fault-tolerant learning problem;

[0040] Calculate the predicted loss value of the encrypted state corresponding to the predicted value by using the properties of homomorphic encryption; wherein, the predicted loss value of the encrypted state is used as the encryption result;

[0041] The expression of the predicted loss value of the encrypted state is:

[0042]

[0043] wherein, L i represents the predicted loss value of the i-th computing node; represents the square of the true value of the j-th feature in the i-th computing node.

[0044] To achieve the above object, on the other hand, an embodiment of the present application proposes a distributed machine learning device based on fault-tolerant learning. The device is applied to a central node, and the device includes:

[0045] A first data acquisition unit, configured to acquire data for training a machine learning model as original data;

[0046] A data encryption unit, configured to encrypt the original data based on a fault-tolerant learning problem to obtain encrypted data;

[0047] A data transmission unit, configured to transmit the encrypted data to a plurality of computing nodes for each of the computing nodes to respectively train the machine learning model by using the encrypted data;

[0048] A result acquisition unit, configured to acquire encrypted results transmitted by each of the computing nodes; wherein, the encrypted results are obtained by encrypting the model parameters of the trained machine learning model based on the fault-tolerant learning problem;

[0049] A model determination unit, configured to determine the trained machine learning model according to the encrypted results.

[0050] To achieve the above object, on the other hand, an embodiment of the present application proposes a distributed machine learning device based on fault-tolerant learning. The device is applied to a computing node in the distributed machine learning method based on fault-tolerant learning as described above, and the device includes:

[0051] A second data acquisition unit, configured to acquire encrypted data transmitted by a central node;

[0052] A model training unit, configured to train a machine learning model by using the encrypted data to obtain an encrypted result; wherein, the encrypted result is obtained by encrypting the model parameters of the trained machine learning model based on a fault-tolerant learning problem;

[0053] A result transmission unit, configured to transmit the encrypted result to the central node.

[0054] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned distributed machine learning method based on fault-tolerant learning is implemented.

[0055] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned distributed machine learning method based on fault-tolerant learning is implemented.

[0056] The embodiments of the present application at least include the following beneficial effects:

[0057] The present application can obtain data for training a machine learning model as original data; encrypt the original data based on a fault-tolerant learning problem to obtain encrypted data; transmit the encrypted data to multiple computing nodes for each computing node to train the machine learning model using the encrypted data respectively; obtain the encrypted results transmitted by each computing node; where the encrypted results are obtained by encrypting the model parameters of the trained machine learning model based on the fault-tolerant learning problem; and determine the trained machine learning model according to the encrypted results. By encrypting the original data based on the fault-tolerant learning problem to obtain encrypted data, each computing node uses the encrypted data to train the machine learning model instead of the plaintext original data, improving data security; moreover, after each computing node trains the machine learning model, it encrypts the model parameters and transmits them to the central node as the decryption result, and the central node can obtain the model parameters by decrypting the encrypted result, realizing efficient distributed machine learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0059] Figure 1 It is a flowchart of the distributed machine learning method based on fault-tolerant learning provided by the embodiment of the present application;

[0060] Figure 2 It is a flowchart of another distributed machine learning method based on fault-tolerant learning provided by the embodiment of the present application;

[0061] Figure 3 It is a flowchart of the LWE algorithm provided by the embodiment of the present application;

[0062] Figure 4Schematic diagram of the structure of the distributed machine learning device based on fault-tolerant learning provided by the embodiment of the present application;

[0063] Figure 5 Schematic diagram of the structure of another distributed machine learning device based on fault-tolerant learning provided by the embodiment of the present application;

[0064] Figure 6 Schematic diagram of the hardware structure of an electronic device provided by the embodiment of the present application. Detailed implementation manners

[0065] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application described in detail in the appended claims.

[0066] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".

[0067] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0069] Before the embodiments of the present application are described in detail, some terms and related technologies involved in the embodiments of the present application are described as follows:

[0070] LWE: Learning With Errors, fault-tolerant learning.

[0071] As a public-key encryption scheme based on lattice cryptography, the Learning with Errors (LWE) algorithm provides a new approach to solving this problem. The security of the LWE algorithm depends on the mathematical problem of recovering linear relationships in the presence of random errors, which gives it a natural advantage in protecting data privacy. As a public-key encryption scheme, LWE has a high level of security. In the context of distributed machine learning, each participating party can use the LWE public key to encrypt gradients or model parameters and then transmit them to the central node or other participating parties for aggregation. Due to the security of the encryption process, the original data will not be leaked to unauthorized third parties during transmission. The security of LWE is based on number-theoretic problems, namely the difficulty of recovering linear relationships in the presence of random errors. This construction method gives LWE a solid theoretical foundation and broad application prospects in cryptography.

[0072] By applying the LWE algorithm to the field of distributed machine learning, it is possible to achieve encrypted protection of training data and prevent data leakage during transmission and storage. The right to be forgotten or the right to data deletion is an important aspect of data privacy, which allows individuals to request online platforms or search engines to delete or erase their personal information, giving individuals control over their personal data and its availability on the Internet. In machine learning, a large amount of data is used to train models, and model parameters contain feature representations of the data, and such feature representations may implicitly contain sensitive data information. This data usually contains personal information, and individuals may wish to delete or forget it. The right to be forgotten in the field of machine learning allows individuals to have the right to control the retention and use of their personal data during the machine learning process to ensure that their privacy is respected. Data forgetting also helps to improve the usability of machine learning models and reduce machine learning bias. Machine learning models may inherit bias and discrimination patterns existing in the data when trained on historical data. If individuals have the right to be forgotten, they can request the deletion of their data from the training set, thereby reducing the risk of persistent bias and discrimination in the model's predictions and decisions.

[0073] However, retraining the model will consume a large amount of computational overhead. Consider a scenario where multiple participating parties train a global model in a distributed machine learning manner by sharing the model parameters trained on their local data with a central server, and the central server is responsible for aggregating the model parameters of the participating parties. After repeated training, a global model is obtained. Then, one of the participating parties requests to forget the model parameters containing its local data from this global model, and the central server needs to comply with the implementation of data forgetting. However, retraining the global model requires re-coordinating the remaining participating parties to train a new global model from scratch, which is obviously a huge computational overhead.

[0074] Therefore, the present application provides a distributed machine learning solution with controllable results based on LWE. During the LWE encryption process, the fluctuation range of the decrypted results can be controlled by adjusting the error distribution. This controllability feature enables the adjustment of encryption parameters to influence the training process and final performance of the model in the distributed machine learning solution. For example, a suitable error distribution range can be set so that the decrypted gradient vectors fluctuate within the desired range, thereby avoiding the problems of overfitting or underfitting of the model. Combining the controllability feature of LWE, an adaptive model optimization algorithm can be designed. This algorithm can dynamically adjust the error distribution parameters during the LWE encryption process according to the training status and performance of the model to achieve more refined control of the model training results.

[0075] The distributed machine learning solution with controllable results based on LWE aims to combine the data encryption capabilities of the LWE algorithm and the computational advantages of distributed machine learning to construct a machine learning framework that is both secure and efficient. In the solution of this embodiment, the original data is first encrypted through the LWE algorithm to ensure the privacy protection of the data in the distributed environment. Subsequently, the encrypted data is distributed to each computing node for parallel training, and necessary data exchange and result aggregation are carried out between the nodes through an encryption protocol. After the training is completed, the central node decrypts the encrypted result using the private key and outputs the final learning model or prediction result.

[0076] In summary, the distributed machine learning solution with controllable results based on LWE not only solves the computational bottleneck problem in large-scale data processing but also effectively protects data privacy through an innovative encryption mechanism, providing strong support for the application of machine learning in more sensitive fields.

[0077] The embodiments of the present application provide a distributed machine learning method, device, equipment, and medium based on fault-tolerant learning. The technical solution of the present application includes: obtaining the data for training a machine learning model as the original data; encrypting the original data based on the fault-tolerant learning problem to obtain encrypted data; transmitting the encrypted data to multiple computing nodes for each computing node to train the machine learning model using the encrypted data respectively; obtaining the encrypted results transmitted by each computing node; wherein the encrypted results are obtained by encrypting the model parameters of the trained machine learning model based on the fault-tolerant learning problem; and determining the trained machine learning model according to the encrypted results. By encrypting the original data based on the fault-tolerant learning problem to obtain encrypted data, each computing node trains the machine learning model using the encrypted data instead of the plaintext original data, improving data security; moreover, after each computing node trains the machine learning model, the model parameters are encrypted and transmitted to the central node as the decryption result, and the central node can obtain the model parameters by decrypting the encrypted result, realizing efficient distributed machine learning.

[0078] The embodiments of the present application provide a distributed machine learning method based on fault-tolerant learning, which relates to the technical field of machine learning. The distributed machine learning method based on fault-tolerant learning provided by the embodiments of the present application can be applied to a terminal, or can be applied to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the distributed machine learning method based on fault-tolerant learning, etc., but is not limited to the above forms.

[0079] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0080] Referring to Figure 1 , the embodiments of the present application provide a distributed machine learning method based on fault-tolerant learning. The method can be applied to a central node, and the method can include but is not limited to S100 to S140, specifically as follows:

[0081] S100: Obtain data for training a machine learning model as original data.

[0082] Specifically, in this embodiment, data for training a machine learning model can be obtained from multiple data sources and data preprocessing can be performed, such as removing outliers, de-duplicating, denoising, normalizing, etc.

[0083] S110: Encrypt the original data based on a fault-tolerant learning problem to obtain encrypted data.

[0084] It is understandable that the encrypted original data can protect data privacy compared to before encryption, and Learning With Errors (LWE) is a public-key encryption scheme based on lattice cryptography, whose security depends on number theory problems, that is, recovering linear relationships in the presence of random errors. Through LWE encryption, the privacy and security of the original data during transmission and storage can be improved.

[0085] Furthermore, S110 may include S111 to S112:

[0086] S111: Generate a public key and a private key;

[0087] S112: Calculate the ciphertext corresponding to the original data based on the definition of the learning with errors problem and the public key to encrypt the original data to obtain the encrypted data; wherein, the private key is used to decrypt the encrypted data;

[0088] The expression for calculating the ciphertext is:

[0089] c = As + er + e';

[0090] wherein, c represents the ciphertext, A represents the public key matrix corresponding to the public key, s represents the original data, e represents the error vector, r represents the random vector, and e' represents the additional error vector.

[0091] S120: Transmit the encrypted data to multiple computing nodes for each of the computing nodes to use the encrypted data to train the machine learning model respectively.

[0092] Exemplarily, in this embodiment, the encrypted data can be transmitted to each computing node through a network, or the encrypted data can be stored in the form of a storage medium, and then the computing node reads the storage medium to achieve the transmission of the encrypted data.

[0093] S130: Obtain the encrypted results transmitted by each of the computing nodes; wherein, the encrypted results are obtained by encrypting the model parameters of the trained machine learning model based on the learning with errors problem.

[0094] Specifically, each computing node trains a machine learning model locally, and then transmits the data obtained from training the machine learning model, including model parameters, etc., to the central node; wherein, the data transmitted by the computing node is encrypted data, which is used as the encrypted result.

[0095] S140: Determine the trained machine learning model according to the encrypted results.

[0096] It can be understood that the encrypted result contains data generated by each computing node in training a machine learning model. In this embodiment, the parameters of the machine learning model can be configured according to the encrypted result, that is, the distributed training result is configured into the machine learning model to implement model training.

[0097] Further, S140 may include S141 to S142:

[0098] S141: Use the encrypted result and the private key of the fault-tolerant learning problem to perform model update aggregation to decrypt and aggregate each of the encrypted results to obtain the model parameters;

[0099] The expression of the model parameters is:

[0100]

[0101] where w t+1 represents the updated model parameters obtained by decryption, w t represents the model parameters before update, η represents the learning rate, Dec represents the decryption function of the fault-tolerant learning problem, Enc represents the encryption function of the fault-tolerant learning problem; L represents the total prediction loss value of the machine learning model, represents the gradient;

[0102] The expression of the total prediction loss value includes:

[0103] Enc(L) = ∑ i Enc(L i );

[0104]

[0105] where L i represents the prediction loss value of the i-th computing node; represents the predicted value of the j-th feature in the i-th computing node; represents the square of the true value of the j-th feature in the i-th computing node; the encrypted data includes each of the features;

[0106] S142: Configure the model parameters into the machine learning model to obtain the trained machine learning model.

[0107] Further, the embodiments of the present application may further include at least one of S151 to S153:

[0108] S151: Receive the model prediction results of the encrypted state transmitted by each of the computing nodes; wherein, the model prediction results are obtained by encrypting the prediction results output by the trained machine learning model by each of the computing nodes based on the fault-tolerant learning problem; decrypt the model prediction results of the encrypted state by using the private key of the fault-tolerant learning problem to obtain the model prediction results in the plaintext state.

[0109] It can be understood that in this embodiment, the trained machine learning model can be used for prediction in the computing nodes, and the central node can directly receive the prediction results of each computing node to implement model inference, which can release the computing resources of the central node locally and improve the inference efficiency.

[0110] S152: Verify whether the model parameters are within a preset numerical range.

[0111] In order to verify whether the machine learning model after distributed training can accurately predict, this embodiment can verify whether the model parameters are within a preset numerical range. If so, it can be considered that the performance of the trained machine learning model is good and it can accurately predict.

[0112] S153: Verify whether the trained machine learning model meets the preset constraint conditions or rules.

[0113] This embodiment can verify whether the trained machine learning model makes predictions under the preset constraint conditions or rules to improve the prediction accuracy.

[0114] Refer to Figure 2 , the embodiments of the present application provide a distributed machine learning method based on fault-tolerant learning. The method can be applied to the computing nodes in the aforementioned distributed machine learning method based on fault-tolerant learning. The method can include but is not limited to S200 to S220, specifically as follows:

[0115] S200: Obtain the encrypted data transmitted by the central node.

[0116] S210: Use the encrypted data to train a machine learning model to obtain an encrypted result; wherein, the encrypted result is obtained by encrypting the model parameters of the trained machine learning model based on the fault-tolerant learning problem.

[0117] Further, S210 can include S211 to S212:

[0118] S211: Randomly initialize the parameters of the machine learning model to obtain an initial model, or obtain a pre-trained model as the initial model;

[0119] Calculate the predicted value of the initial model for the encrypted data;

[0120] The expression for calculating the predicted value includes:

[0121]

[0122] where represents the predicted value corresponding to the j-th feature of the encrypted data in the t-th training round, is the model parameter in the t-th training round, ∈ is the added noise; x tj represents the sample value corresponding to the j-th feature of the encrypted data in the t-th training round; Enc represents the encryption function for the fault-tolerant learning problem;

[0123] S212: Calculate the predicted loss value of the encrypted state corresponding to the predicted value using the properties of homomorphic encryption; wherein, the predicted loss value of the encrypted state is used as the encrypted result;

[0124] The expression for the predicted loss value of the encrypted state is:

[0125]

[0126] where L i represents the predicted loss value of the i-th computing node; represents the square of the true value of the j-th feature in the i-th computing node.

[0127] S220: Transmit the encrypted result to the central node.

[0128] It can be understood that the implementation manner of the computing node corresponds to that of the central node, and the intended effects achieved can also be the same.

[0129] Next, the specific implementation manner of the method of the present application will be introduced.

[0130] This embodiment can provide a distributed machine learning method with controllable results based on LWE, which can not only solve the computational bottleneck problem in large-scale data processing, but also effectively protect data privacy through an innovative encryption mechanism, providing strong support for the application of machine learning in more sensitive fields.

[0131] This embodiment can be implemented through the following technical solutions:

[0132] S1. Data Preprocessing and Encryption. Collect raw data from various data sources and perform necessary data cleaning and preprocessing to improve data quality and consistency. LWE Encryption: Encrypt the preprocessed data using LWE encryption technology. LWE is a public-key encryption scheme based on lattice cryptography, and its security depends on number theory problems, that is, recovering linear relationships in the presence of random errors. Through LWE encryption, the privacy protection of data during transmission and storage can be ensured. Encryption Process: Usually includes generating a public-key and private-key pair, encrypting the data using the public key to generate ciphertext. The ciphertext will be used for subsequent distributed machine learning training.

[0133] S2. Distributed Machine Learning Training. Encrypted Data Distribution: Distribute the encrypted data to each computing node, and each computing node can be an independent server, cloud resource, or edge device. Design Encryption Protocols: To support machine learning training on encrypted data, specific encryption protocols need to be designed. These protocols need to be able to handle computational operations on encrypted data, such as addition, multiplication, etc., while maintaining the encrypted state of the data to avoid data leakage. Homomorphic Encryption Support: LWE encryption itself or combined with other homomorphic encryption technologies (such as Leveled Homomorphic Encryption LHE) can support computational operations on ciphertext to a certain extent, making it possible to perform machine learning training on encrypted data. Distributed Training: On each computing node, use the designed encryption protocols and machine learning algorithms to train the encrypted data. During the training process, the encrypted intermediate results can be exchanged between nodes to collaborate on completing the training task.

[0134] S3. Result Aggregation and Decryption. Secure Aggregation: After training, each computing node sends the encrypted intermediate results or model parameters to the central node or aggregation server. The central node uses a specific secure aggregation protocol to aggregate these encrypted results to ensure data security and privacy protection during the aggregation process. Decryption and Result Output: The aggregated encrypted results are decrypted using the private key to obtain the final machine learning model or prediction results. These results can be output under control to meet specific application requirements.

[0135] S4. Result Controllability. Access Control: Ensure that only authorized users or systems can access the decrypted results by implementing strict access control policies. Result Verification: Verify the decrypted results to ensure their accuracy and reliability. Methods such as cross-validation and model evaluation can be used to evaluate the quality of the results.

[0136] S5. Encrypted Data Preparation: Each computing node i holds the LWE-encrypted dataset {Enc(x ij ), Enc(y ij )}, where x ij is the feature vector and y ijis the corresponding tag, and Enc represents the LWE encryption function.

[0137] S6. Model Initialization: Select an initial model parameter w0 (which may be randomly initialized or based on a pre-trained model).

[0138] S7. Encrypted Computation: At each node i, compute the encrypted model prediction value where is the model parameter of the current round, ∈ is the added noise (used to enhance the security of LWE), and Enc represents LWE encryption of the corresponding value.

[0139] S8. Loss Computation and Gradient Update: Since the data is encrypted, the loss function cannot be directly computed. Therefore, a secure loss computation protocol needs to be designed. A possible approach is to use the properties of homomorphic encryption to compute the encrypted loss value

[0140] Then, each node sends the encrypted loss value to the central node (or through a secure aggregation mechanism), and the central node aggregates to obtain the total encrypted loss value Enc(L) = ∑ i Enc(L i ).

[0141] Next, use the encrypted gradient descent algorithm (such as secure backpropagation) to update the model parameters. Since gradient computation involves complex chain rules, a series of secure encrypted computation protocols need to be designed to achieve this. The updated model parameter is where η is the learning rate, and Dec represents the LWE decryption function (in actual operations, since decryption may involve private keys, this step is usually performed by a trusted third party or the central node).

[0142] S9. Iterative Training: Repeat steps S7 to S8 until a predetermined number of training rounds is reached or the convergence condition is satisfied.

[0143] S10. Result Output: Finally, obtain the trained model parameter w * . When the model prediction result needs to be output, the new input data can be encrypted using LWE, then encrypted prediction can be performed using the trained model, and the prediction result can be obtained through a secure decryption mechanism.

[0144] S11. Result Aggregation: After each computing node completes local model training, it will obtain encrypted model updates (such as gradient or weight updates), denoted as Enc(Δw i ), where i represents the index of the node, and Δw i represents the model update of this node, and Enc represents using LWE encryption.

[0145] S12. Decryption: Finally, the trained model parameters w are obtained. * . When the model prediction result needs to be output, the new input data can be encrypted using LWE, and then the encrypted prediction can be performed using the trained model, and the prediction result can be obtained through a secure decryption mechanism. These encrypted model updates will be securely aggregated, usually achieved through the properties of homomorphic encryption, i.e., textEnc(Deltaw textagg ) = prod.textEnc(Deltaw i ) quadtext or quadtextEnc(Deltaw textagg ) = sum.textEnc(Deltaw i ). Where w agg represents the aggregated model update, text represents the input ciphertext, and depending on the specific encryption scheme and aggregation method, multiplicative or additive homomorphicity may be used. In the expression textEnc(Deltaw textagg ) = prod.textEnc(Deltaw i ), prod means multiplying the ciphertext results of multiple textEnc(Deltaw i ) to obtain the encrypted value of Deltaw textagg . And in textEnc(Deltaw textagg ) = sum.textEnc(Deltaw i ), sum means adding the ciphertext results of multiple textEnc(Deltaw i ) to obtain the encrypted value of Deltaw textagg .

[0146] The decryption process requires access to the private key in the LWE encryption scheme. In LWE encryption, the private key is usually a short random vector where Z is the vector space, q is a large prime number, and n is the dimension of the vector. At the same time, the noise distribution parameters used during encryption also need to be known, although these parameters are usually not directly used during the decryption process. The decryption entity receives the encrypted data or model update, denoted as Enc(m) = (A, b), where is a random matrix, is a vector, satisfying b ≈ As + e mod q, and e is the noise vector. The decryption algorithm uses the private key S to calculate the decryption result. Specifically, the decryption algorithm may calculate a vector y such that y ≈ A Ts modq, and then use this vector y and the encrypted vector b to recover the plaintext message m. However, in the actual application of LWE encryption, the decryption process usually does not directly calculate m, but verifies the correctness of the encrypted data in some way and extracts useful information (such as model update) from it.

[0147] Furthermore, in the LWE encryption scheme based on Regev, the decryption process is as follows: Assume that the public key in the encryption scheme is (A, b) and the private key is s, where b=As+e mod q, is the noise vector.

[0148] When decrypting, the decryption entity calculates:

[0149] y=A T smodq;

[0150] Then, a reconstruction technique is used to recover the plaintext message m from y and b. However, in most LWE encryption applications, especially in distributed machine learning scenarios, this embodiment does not actually directly recover the plaintext message m, but uses the encrypted data (A, b) and the private key s to perform some calculations (such as aggregation of model updates) without decrypting it to plaintext. The LWE algorithm workflow diagram of this embodiment can be referenced Figure 3 .

[0151] S13. Controllability of results: After decryption, the resulting plaintext model updates can be further verified and controlled to ensure that the results are correct and meet expectations. This may include checking whether the updates are within a reasonable range, or applying additional constraints and rules.

[0152] S14. Finally, the verified and controlled model update is applied to the global model to complete a training iteration.

[0153] More specifically, this embodiment can also be implemented by the following steps:

[0154] A1: Data preprocessing and encryption.

[0155] 1. Data collection and cleaning: Collect raw data from multiple data sources (such as devices, sensors, databases, etc.). Clean the data to remove duplicate data, outliers, noise, etc. to ensure data accuracy and consistency. Select features that have an important impact on model training and remove irrelevant or redundant features.

[0156] 2. Feature selection and scaling: Scale the feature data according to certain rules, such as normalization or standardization, so that they are at the same order of magnitude to improve the efficiency of model training.

[0157] III. LWE Encryption. The public key is used for encryption, and the private key is used for decryption. Each data point or data batch is encrypted using the public key through LWE to generate ciphertext. The encryption process is based on the definition of the LWE problem, where random vectors, error vectors, and the public key matrix are selected to calculate the ciphertext. The ciphertext is securely transmitted to each computing node in the distributed system or stored in a secure storage medium.

[0158] Key Generation: Generate the public key and the private key where q is a large prime number, and m, n are integers.

[0159] Encryption Process: For each data point encrypt using the public key A and the perturbation vector to generate the ciphertext y = Ax + e mod q. Note that the ciphertext generation formula in this embodiment does not directly include the private key s because public key encryption does not require the participation of the private key. However, the private key s is used for subsequent decryption.

[0160] A2: Distributed Machine Learning Training.

[0161] I. Encrypted Data Distribution: Distribute the encrypted data y to each computing node, which can be independent servers, cloud resources, or edge devices.

[0162] II. Model Initialization: Select an initial model parameter θ0, which can be randomly initialized or based on a pre-trained model.

[0163] III. Encrypted Computation: On each computing node, calculate the predicted value for the encrypted data. Usually, this requires designing a specific encryption protocol. Since directly encrypting the model function is not practical, it is assumed here that there is a certain protocol P that can indirectly calculate the encrypted predicted value:

[0164] EncPred i = P(theta t , y i );

[0165] Note: EncPred i in this embodiment is the representation of the encrypted predicted value, and P: This is the function representation of the encryption algorithm. It receives some inputs (such as the model parameter theta t and the data point y i ) and outputs the encrypted result EncPred i . theta t This is the parameter of the machine learning model, the value at time t, and the actual calculation depends on the specific protocol.

[0166] IV. Loss Calculation and Gradient Update: Design a secure loss calculation protocol and calculate the encrypted loss using homomorphic encryption:

[0167] EncLoss i = textEnc(L(theta i , textDec textapprox (textEncPred i ), y i ));

[0168] Note: EncLoss i represents the encrypted loss value on the i-th computing node. In this context, the inputs to the loss function L include:

[0169] theta i : The parameters of the machine learning model on the i-th computing node.

[0170] textDec textapprox (textEncPred i ), y i ): The approximate value after decrypting the encrypted prediction result on the i-th computing node. Here, textEncPred i represents the encrypted prediction result, and textDec textapprox is a function that decrypts and approximately restores the original prediction value.

[0171] y i : The corresponding true label or result on the i-th computing node.

[0172] Note: Dec approx here represents an approximate decryption operation used to estimate the loss in a homomorphic encryption environment. In practice, it may not be necessary to fully decrypt.

[0173] Using an encrypted gradient descent algorithm to update the model parameters, assuming there is an encrypted gradient calculation protocol G:

[0174] EncGrad = G(textEncLoss ii , theta t );

[0175] θ t+1 = theta t - eta * textDec(textAgg(textEncGrad));

[0176] Here, EncGrad represents the encrypted gradient, where G is a function that takes two parameters:

[0177] textEncLoss i : The encrypted representation of the loss on the i-th computing node.

[0178] theta t : Model parameters at time step t.

[0179] eta: Learning rate, a hyperparameter used to control the step size of parameter updates.

[0180] textEnc: Decryption function used to restore encrypted data to plaintext.

[0181] textAgg: Aggregation function used to aggregate encrypted gradients from different computing nodes.

[0182] (textEncGrad): Represents the set of encrypted gradients, which is the summary of encrypted gradients calculated by multiple computing nodes.

[0183] V. Iterative Training: Repeat the above steps of encrypted calculation, loss calculation, and gradient update until a predetermined number of training rounds is reached or convergence conditions are met.

[0184] A3: Result aggregation and decryption.

[0185] After each computing node completes local model training, it sends the encrypted model updates (such as gradient or weight updates) to the central node. The central node uses the properties of homomorphic encryption to securely aggregate these encrypted model updates to obtain the aggregated encrypted model updates.

[0186] EncUpdate textagg = sum i textEncUpdate i ;

[0187] The above equation uses additive homomorphicity for aggregation.

[0188] Meanwhile, use the private key to decrypt the aggregated encrypted model updates to obtain the final plaintext model parameters. When predicting results are needed, encrypt the new input data using LWE, then perform encrypted prediction using the trained model, and obtain the prediction results through a secure decryption mechanism.

[0189] Update textagg = textDec(textEncUpdate textagg , s);

[0190] The decryption process in the above equation may depend on specific LWE decryption algorithms, such as the Regev decryption algorithm. textEncUpdate textagg: This represents "encrypted text aggregation update". In distributed machine learning, especially when dealing with text data, it may be necessary to aggregate updates (such as gradient updates or model parameter updates) from different nodes. textEnc indicates that these updates have been encrypted before transmission or storage. Update textagg This represents the result of the decryption operation, i.e., the output of textDec(textEncUpdate textagg , s). It represents the decrypted text aggregation update. This decrypted data can be used for subsequent machine learning calculations, such as updating model parameters or performing gradient descent, etc.

[0191] Encrypted update aggregation: EncUpdate textagg = sum i textEncUpdate i ;

[0192] Decryption: In the Regev decryption algorithm, decryption may involve complex operations such as solving linear equations. The formula is not directly given here, but the core is to use the private key s to recover the plaintext information.

[0193] A4: Result controllability.

[0194] Implement strict access control policies so that only authorized users or systems can access the decrypted results.

[0195] Furthermore, verify the decrypted results, including checking the rationality of model updates, model performance evaluation, etc., to ensure the accuracy and reliability of the results.

[0196] Accuracy = (TP + TN) / (TP + TN + FP + FN);

[0197] Where TP (True Positives): True positives, referring to the number of positive samples correctly predicted by the model.

[0198] TN (True Negatives): True negatives, referring to the number of negative samples correctly predicted by the model.

[0199] FP (False Positives): False positives, referring to the number of negative samples wrongly predicted as positive samples by the model.

[0200] FN (False Negatives): False negatives, referring to the number of positive samples wrongly predicted as negative samples by the model.

[0201] Further, when a data forgetting request is received, the affected group-level global model is retrained according to the request. The process is similar to steps A2 and A3, but is limited to specific group data, so that the data forgetting request is handled in compliance.

[0202] Referring to Figure 4 , an embodiment of the present application further provides a distributed machine learning device based on fault-tolerant learning. The device is applied to a central node and includes:

[0203] A first data acquisition unit, configured to acquire data for training a machine learning model as original data;

[0204] A data encryption unit, configured to encrypt the original data based on a fault-tolerant learning problem to obtain encrypted data;

[0205] A data transmission unit, configured to transmit the encrypted data to multiple computing nodes for each of the computing nodes to train the machine learning model using the encrypted data respectively;

[0206] A result acquisition unit, configured to acquire encrypted results transmitted by each of the computing nodes; wherein the encrypted results are obtained by encrypting model parameters of the trained machine learning model based on the fault-tolerant learning problem;

[0207] A model determination unit, configured to determine the trained machine learning model according to the encrypted results.

[0208] Referring to Figure 5 , an embodiment of the present application further provides a distributed machine learning device based on fault-tolerant learning. The device is applied to a computing node and includes:

[0209] A second data acquisition unit, configured to acquire encrypted data transmitted by the central node;

[0210] A model training unit, configured to train a machine learning model using the encrypted data to obtain encrypted results; wherein the encrypted results are obtained by encrypting model parameters of the trained machine learning model based on the fault-tolerant learning problem;

[0211] A result transmission unit, configured to transmit the encrypted results to the central node.

[0212] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0213] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned distributed machine learning method based on fault-tolerant learning is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0214] It can be understood that the content in the above method embodiments is applicable to this device embodiment. The functions specifically implemented by this device embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0215] Please refer to Figure 6 , Figure 6 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0216] A processor 601, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0217] A memory 602, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 602 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 602, and the processor 601 is used to call and execute the distributed machine learning method based on fault-tolerant learning in the embodiments of the present application;

[0218] An input / output interface 603, which is used to implement information input and output;

[0219] A communication interface 604, which is used to implement communication and interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0220] A bus 605, which transmits information between various components of the device (such as the processor 601, the memory 602, the input / output interface 603, and the communication interface 604);

[0221] Among them, the processor 601, the memory 602, the input / output interface 603, and the communication interface 604 are communicatively connected to each other inside the device through the bus 605.

[0222] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-described distributed machine learning method based on fault-tolerant learning is implemented.

[0223] It can be understood that the content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0224] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0225] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0226] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0227] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present embodiment.

[0228] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0229] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0230] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0231] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0232] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0233] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0234] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0235] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A distributed machine learning method based on fault-tolerant learning, characterized in that: The method is applied to a central node and comprises the following steps: Get the data used to train the machine learning model as raw data; Encrypting the original data based on the fault-tolerant learning problem to obtain encrypted data; Transmitting the encrypted data to a plurality of computing nodes, so that each computing node uses the encrypted data to train the machine learning model respectively; Obtaining encryption results transmitted by each of the computing nodes; wherein the encryption results are obtained by encrypting model parameters of the trained machine learning model based on the fault-tolerant learning problem; The trained machine learning model is determined according to the encryption result.

2. The distributed machine learning method based on fault-tolerant learning according to claim 1, characterized in that: The method of encrypting the original data based on the fault-tolerant learning problem to obtain encrypted data comprises the following steps: Generate public and private keys; Calculating the ciphertext corresponding to the original data based on the definition of the fault-tolerant learning problem and the public key to encrypt the original data to obtain the encrypted data; wherein the private key is used to decrypt the encrypted data; The expression for calculating the ciphertext is: c=As+er+e′; Wherein, c represents the ciphertext, A represents the public key matrix corresponding to the public key, s represents the original data, e represents the error vector, r represents the random vector, and e′ represents the additional error vector.

3. The distributed machine learning method based on fault-tolerant learning according to claim 1, characterized in that: Determining the trained machine learning model according to the encryption result comprises the following steps: Performing model update aggregation using the encryption result and the private key of the fault-tolerant learning problem to decrypt and aggregate each of the encryption results to obtain the model parameters; The expressions of the model parameters are: Among them, w t+1 represents the updated model parameters obtained after decryption, w t represents the model parameters before updating, η represents the learning rate, Dec represents the decryption function of the fault-tolerant learning problem, Enc represents the encryption function of the fault-tolerant learning problem; L represents the total prediction loss value of the machine learning model, represents the gradient; The expression of the total prediction loss value includes: Enc(L)=∑ i Enc(L i ); Among them, L i represents the predicted loss value of the i-th computing node; represents the predicted value of the j-th feature in the i-th computing node; represents the square of the true value of the jth feature in the i-th computing node; the encrypted data includes each of the features; The model parameters are configured to the machine learning model to obtain the trained machine learning model.

4. The distributed machine learning method based on fault-tolerant learning according to any one of claims 1 to 3, characterized in that: The method further comprises at least one of the following steps: Receive the model prediction results in an encrypted state transmitted by each of the computing nodes; wherein the model prediction results are obtained by each of the computing nodes encrypting the prediction results output by the trained machine learning model based on the fault-tolerant learning problem; decrypt the model prediction results in an encrypted state using the private key of the fault-tolerant learning problem to obtain the model prediction results in a plaintext state; Alternatively, verifying whether the model parameter is within a preset value range; Alternatively, verify whether the trained machine learning model satisfies preset constraints or rules.

5. A distributed machine learning method based on fault-tolerant learning, characterized in that: The method is applied to a computing node in a distributed machine learning method based on fault-tolerant learning as claimed in claim 1, and the method comprises the following steps: Obtain the encrypted data transmitted by the central node; Using the encrypted data to train a machine learning model to obtain an encrypted result; wherein the encrypted result is obtained by encrypting model parameters of the trained machine learning model based on a fault-tolerant learning problem; The encryption result is transmitted to the central node.

6. The distributed machine learning method based on fault-tolerant learning according to claim 5, characterized in that: The method of using the encrypted data to train the machine learning model to obtain an encrypted result includes the following steps: Randomly initialize the parameters of the machine learning model to obtain an initial model, or obtain a pre-trained model as the initial model; Calculating a predicted value of the initial model for the encrypted data; The expression for calculating the predicted value includes: in, represents the predicted value corresponding to the j-th feature of the encrypted data in the t-th training round, is the model parameter in the tth training round, ∈ is the added noise; x tj represents the sample value corresponding to the jth feature of the encrypted data in the tth training round; Enc represents the encryption function of the fault-tolerant learning problem; The predicted loss value of the encrypted state corresponding to the predicted value is calculated using the property of homomorphic encryption; wherein the predicted loss value of the encrypted state is used as the encryption result; The expression of the predicted loss value in the encrypted state is: Among them, L i represents the predicted loss value of the i-th computing node; Represents the square of the true value of the j-th feature in the i-th computing node.

7. A distributed machine learning device based on fault-tolerant learning, characterized in that: The device is applied to a central node, and the device: A first data acquisition unit, used to acquire data for training a machine learning model as raw data; A data encryption unit, used for encrypting the original data based on the fault-tolerant learning problem to obtain encrypted data; A data transmission unit, used to transmit the encrypted data to a plurality of computing nodes, so that each computing node can use the encrypted data to train the machine learning model respectively; A result acquisition unit, used to acquire the encryption results transmitted by each of the computing nodes; wherein the encryption results are obtained by encrypting the model parameters of the trained machine learning model based on the fault-tolerant learning problem; A model determination unit is used to determine the trained machine learning model according to the encryption result.

8. A distributed machine learning device based on fault-tolerant learning, characterized in that: The device is applied to a computing node in the distributed machine learning method based on fault-tolerant learning as claimed in claim 1, and the device comprises: A second data acquisition unit, used to acquire the encrypted data transmitted by the central node; A model training unit, used to train a machine learning model using the encrypted data to obtain an encrypted result; wherein the encrypted result is obtained by encrypting model parameters of the trained machine learning model based on a fault-tolerant learning problem; A result transmission unit is used to transmit the encryption result to the central node.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 6 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.