Training Method, Device, Electronic Device, and Storage Medium for Distributed Machine Learning Model

By constructing and processing node relationship diagrams, the privacy and convergence problems of machine learning model training in scattered data sources are solved, and efficient training and model convergence are achieved without meeting independent and same distribution conditions.

CN115146788BActive Publication Date: 2025-07-22BEIJING TOPSEC NETWORK SECURITY TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210770991.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-07-22
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

Traditional machine learning model training methods are difficult to perform under the premise of protecting user privacy in multiple scattered data sources, and it is difficult to converge when the scattered training data does not meet the independent and same distribution conditions.

Method used

By constructing a node relationship graph, deleting weight elements and edges that do not meet the preset conditions, obtaining a second node relationship graph, and training the second machine learning model based on this graph to ensure data privacy and model convergence.

Benefits of technology

On the premise of protecting user data privacy, the convergence and accuracy of the model without meeting independent and same distribution conditions are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115146788B_ABST
    Figure CN115146788B_ABST
Patent Text Reader

Abstract

The present application provides a training method, device, electronic device, and storage medium for a distributed machine learning model. Among them, the training method for the distributed machine learning model includes deleting weight elements in the weight set that do not meet the preset conditions, and deleting the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph; obtaining one or more third node relationship graphs based on the second node relationship graph; training a second machine learning model based on the third node relationship graph, and other steps. The present application can complete the training of the machine learning model while protecting the privacy of user data. At the same time, the present application can overcome the problem that it is difficult for the distributed machine learning model to converge when the scattered training data does not meet the independent and identically distributed conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and specifically, to a training method, apparatus, and electrical equipment storage medium for a distributed machine learning model. Background Art

[0002] At present, with the rapid development of artificial intelligence technology, the application of artificial intelligence technology is becoming more and more extensive, and has been applied to many fields such as image analysis, speech recognition, text processing, intelligent recommendation, security detection, etc. Artificial intelligence technology represented by machine learning is strongly dependent on the data for training models. When the model training data comes from multiple scattered data sources, and the data provided by the data source involves user privacy, business secrets and other non-public data, the traditional machine learning model training method will be difficult to deploy and implement. At the same time, since the email sample data involves user privacy and the email sample data has the characteristics of non-publicity, it is difficult to find out the characteristic distribution of the email sample data before model training, and it is difficult to confirm whether the email sample data meets the assumption of independent and identical distribution, and it is difficult to select and design a suitable machine learning model in a targeted manner. Summary of the invention

[0003] The purpose of the embodiments of the present application is to provide a training method, device, and electrical equipment storage medium for a distributed machine learning model, so as to complete the training of the machine learning model while protecting the privacy of user data. At the same time, the present application can overcome the problem that the distributed machine learning model is difficult to converge when the distributed training data does not meet the independent and identically distributed conditions.

[0004] To this end, the first aspect of the present application discloses a training method for a distributed machine learning model, the method being applied to a server node in a distributed system, the method comprising:

[0005] Acquire a plurality of model accuracies for a first machine learning model sent by a plurality of user end nodes, wherein the model accuracies are obtained by the user end nodes testing models of other user end nodes;

[0006] Constructing a first node relationship graph based on a plurality of the model accuracies, wherein the first node relationship graph includes a weight set, a vertex set, and an edge set, the vertices of the vertex set represent the user-end nodes, the edges of the edge set represent the model test relationships between the user-end nodes, and the weight elements in the weight set represent the average value of the model accuracies between two of the user-end nodes;

[0007] Deleting the weight elements that do not meet the preset conditions in the weight set, and deleting the edges corresponding to the weight elements that do not meet the preset conditions from the edge set, so as to obtain a second node relationship graph;

[0008] Obtain one or more third node relationship graphs based on the second node relationship graph;

[0009] Train a second machine learning model based on the third node relationship graph, where each third node relationship graph corresponds to a second machine learning model.

[0010] In the first aspect of the present application, during the training process of the machine learning model, each client node trains the first machine learning model based on local data and sends the model accuracy of the first machine learning model to the server node. In this way, the server node does not need to obtain the local data of each client node, and thus can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, the present application can delete the weight elements in the weight set that do not meet the preset conditions based on the first node relationship graph, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain the second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, the present application can obtain the third node relationship graph based on the second node relationship graph, and thus can train the second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be separately trained based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0011] In the first aspect of the present application, as an optional implementation manner, the training the second machine learning model based on the third node relationship graph includes:

[0012] Send the first model parameters of the second machine learning model to each client node in the third node relationship graph, so that each client node in the third node relationship graph trains the second machine learning model based on the first model parameters of the second machine learning model and local sample data until the second model parameters when the second machine learning model meets the first preset training condition are obtained;

[0013] Obtain the second model parameters sent by each client node;

[0014] Calculate the third model parameters of the second machine learning model based on the second model parameters of each client node;

[0015] Determine whether the third model parameter satisfies the second preset training condition. If it does not satisfy the second preset training condition, iteratively train the second machine learning model until the third model parameter satisfies the second preset training condition.

[0016] In this alternative embodiment, by sending the first model parameter of the second machine learning model to each client node in the third node relationship graph, each client node in the third node relationship graph can train the second machine learning model based on the first model parameter of the second machine learning model and local sample data until the second model parameter when the second machine learning model satisfies the first preset training condition is obtained. At the same time, by obtaining the second model parameters sent by each client node, the third model parameter of the second machine learning model can be calculated based on the second model parameters of each client node. Furthermore, by determining whether the third model parameter satisfies the second preset training condition, the second machine learning model is iteratively trained until the third model parameter satisfies the second preset training condition. That is, through the above steps, the training of the second machine learning model can be completed in a distributed training manner.

[0017] In the first aspect of this application, as an alternative embodiment, calculating the third model parameter of the second machine learning model based on the second model parameters of each of the client nodes includes:

[0018] Determine the total number of the second model parameters;

[0019] Based on the total number of the second model parameters and each of the second model parameters, calculate the average value of all the second model parameters;

[0020] Take the average value of all the second model parameters as the third model parameter.

[0021] In this alternative embodiment, by determining the total number of the second model parameters, the average value of all the second model parameters can be calculated based on the total number of the second model parameters and each of the second model parameters, so that the average value of all the second model parameters can be taken as the third model parameter.

[0022] In the first aspect of this application, as an alternative embodiment, after constructing the first node relationship graph based on several model accuracies and before deleting the weight elements in the weight set that do not satisfy the preset conditions, the method further includes:

[0023] Sort the weight elements in the weight set and determine the weight elements whose values are less than the preset threshold;

[0024] Determine the weight elements whose values are less than the preset threshold as the weight elements that do not satisfy the preset conditions.

[0025] In this alternative embodiment, by sorting the weight elements in the weight set, it is possible to determine the weight elements whose values are less than a preset threshold, and thus the weight elements whose values are less than the preset threshold can be determined as the weight elements that do not meet the preset conditions.

[0026] In a first aspect of the present application, as an alternative embodiment, obtaining one or more third node relationship graphs based on the second node relationship graph includes:

[0027] Determine whether the second node relationship graph is a connected graph. If so, determine one third node relationship graph;

[0028] When the second node relationship graph is not a connected graph, the second node relationship graph is divided into a plurality of third node relationship graphs, and the third node relationship graph is a connected subgraph.

[0029] In this alternative embodiment, by determining whether the second node relationship graph is a connected graph, when the second node relationship graph is not a connected graph, the second node relationship graph can be divided into a plurality of third node relationship graphs, and the third node relationship graph is a connected subgraph, so that at least two second machine learning models can be separately trained.

[0030] In a first aspect of the present application, as an alternative implementation method, before obtaining the model accuracies of a first machine learning model, the method further includes:

[0031] Send the initialization model parameters of the first machine learning model to each of the client nodes, so that each client node trains the first machine learning model based on the initialization model parameters of the first machine learning model until the first machine learning model meets the third preset training condition, where the parameters corresponding to the first machine learning model under the third preset training condition are the model parameters of the client node.

[0032] In this alternative embodiment, by sending the initialization model parameters of the first machine learning model to each client node, each client node can train the first machine learning model based on the initialization model parameters of the first machine learning model until the first machine learning model meets the third preset training condition.

[0033] A second aspect of the present application discloses a method for training a distributed machine learning model, which is applied to a plurality of client nodes in a distributed system. The method includes:

[0034] Send the model accuracy of the first machine learning model to the server node, so that the server node constructs a first node relationship graph based on a number of the model accuracies, deletes weight elements in the weight set that do not meet the preset conditions, deletes edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph, obtains one or more third node relationship graphs based on the second node relationship graph, and trains a second machine learning model based on the third node relationship graph, where the first node relationship graph includes a weight set, a vertex set, and an edge set, vertices in the vertex set represent the client nodes, edges in the edge set represent model test relationships between the client nodes, weight elements in the weight set represent the average model accuracy between two client nodes, each third node relationship graph corresponds to a second machine learning model, and the model accuracy is obtained by the client node testing the models of other client nodes.

[0035] In the second aspect of the present application, during the training process of the machine learning model, each client node trains the first machine learning model based on local data and sends the model accuracy of the first machine learning model to the server node. In this way, the server node does not need to obtain the local data of each client node, and thus can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the present application can delete weight elements in the weight set that do not meet the preset conditions and delete edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph. In this way, the influence of model accuracies that do not meet the preset conditions on model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, based on the second node relationship graph, the present application can obtain a third node relationship graph and thus train a second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be trained separately based on the third node relationship graph, enabling the second machine learning model to have better model convergence.

[0036] In the second aspect of the present application, as an optional implementation manner, the method further includes:

[0037] Receive the first model parameters of the second machine learning model sent by the server node, where the client node is the client node in the third node relationship graph;

[0038] Train the second machine learning model based on the first model parameters of the second machine learning model and local sample data until the second model parameters when the second machine learning model meets the first preset training condition are obtained;

[0039] Send the second model parameters to the server node, so that the server node calculates the third model parameters of the second machine learning model based on the second model parameters of each client node, determines whether the third model parameters meet the second preset training condition, and iteratively trains the second machine learning model when the second preset training condition is not met until the third model parameters meet the second preset training condition.

[0040] By sending the first model parameters of the second machine learning model to each client node in the third node relationship graph, each client node in the third node relationship graph can train the second machine learning model based on the first model parameters of the second machine learning model and local sample data until the second model parameters when the second machine learning model meets the first preset training condition are obtained. At the same time, by obtaining the second model parameters sent by each client node, the third model parameters of the second machine learning model can be calculated based on the second model parameters of each client node. Furthermore, by determining whether the third model parameters meet the second preset training condition, the second machine learning model is iteratively trained until the third model parameters meet the second preset training condition. That is, through the above steps, the training of the second machine learning model can be completed in a distributed training manner.

[0041] In a second aspect of the present application, as an optional implementation manner, the third model parameter is the average value of all the second model parameters, where the average value of all the second model parameters is calculated by the server node based on the total number of the second model parameters and each of the second model parameters after determining the total number of the second model parameters.

[0042] In this optional implementation manner, by determining the total number of the second model parameters, the average value of all the second model parameters can be calculated based on the total number of the second model parameters and each of the second model parameters, so that the average value of all the second model parameters can be used as the third model parameter.

[0043] In a second aspect of the present application, as an optional implementation manner, the weight elements that do not meet the preset condition are the weight elements with values less than the preset threshold, where the weight elements with values less than the preset threshold are determined by the server node after sorting the weight elements in the weight set.

[0044] In this alternative embodiment, by sorting the weight elements in the weight set, it is possible to determine the weight elements whose values are less than a preset threshold, and thus the weight elements whose values are less than the preset threshold can be determined as the weight elements that do not meet the preset conditions.

[0045] In a second aspect of the present application, as an alternative embodiment, when the second node relationship graph is a connected graph, a third node relationship graph is determined. When the second node relationship graph is a connected graph, the second node relationship graph is divided into a plurality of the third node relationship graphs, and the third node relationship graph is a connected subgraph.

[0046] In this alternative embodiment, by determining whether the second node relationship graph is a connected graph, when the second node relationship graph is not a connected graph, the second node relationship graph can be divided into a plurality of the third node relationship graphs, where the third node relationship graph is a connected subgraph, so that at least two second machine learning models can be separately trained.

[0047] In a second aspect of the present application, as an alternative embodiment, before receiving the first model parameters of the second machine learning model sent by the server node, the method further includes:

[0048] Receiving the initialization model parameters of the first machine learning model sent by the server node;

[0049] Training the first machine learning model based on the initialization model parameters of the first machine learning model until the first machine learning model meets the third preset training condition, where the parameters corresponding to the first machine learning model under the third preset training condition are the model parameters of the client node.

[0050] In this alternative embodiment, by sending the initialization model parameters of the first machine learning model to each client node, each client node can train the first machine learning model based on the initialization model parameters of the first machine learning model until the first machine learning model meets the third preset training condition.

[0051] A third aspect of the present application discloses a training device for a distributed machine learning model. The device is located in a server node, and the device includes:

[0052] An acquisition module, configured to acquire a plurality of model accuracies of a first machine learning model sent by a plurality of client nodes, where the model accuracy is obtained by the client node testing the models of other client nodes;

[0053] A construction module for constructing a first node relationship graph based on a plurality of the model accuracies, wherein the first node relationship graph includes a weight set, a vertex set, and an edge set, vertices in the vertex set represent the client nodes, edges in the edge set represent the model test relationships between the client nodes, and weight elements in the weight set represent the average model accuracy between two client nodes;

[0054] A deletion module for deleting weight elements in the weight set that do not meet a preset condition, and deleting edges corresponding to the weight elements that do not meet the preset condition from the edge set, so as to obtain a second node relationship graph;

[0055] A splitting module for obtaining one or more third node relationship graphs based on the second node relationship graph;

[0056] A training module for training a second machine learning model based on the third node relationship graph, wherein each third node relationship graph corresponds to a second machine learning model.

[0057] In this application, during the training process of the machine learning model, each client node trains the first machine learning model based on local data and sends the model accuracy for the first machine learning model to the server node. In this way, the server node does not need to obtain the local data of each client node, and thus can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, this application can delete weight elements in the weight set that do not meet the preset condition, and delete edges corresponding to the weight elements that do not meet the preset condition from the edge set, so as to obtain a second node relationship graph. In this way, the influence of model accuracies that do not meet the preset condition on model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed condition. At the same time, based on the second node relationship graph, this application can obtain a third node relationship graph, and thus can train the second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be separately trained based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0058] A fourth aspect of this application discloses a training device for a distributed machine learning model. The device is applied to a plurality of client nodes in a distributed system, and the device includes:

[0059] A sending module, configured to send the model accuracy of a first machine learning model to a server node, so that the server node constructs a first node relationship graph based on a plurality of the model accuracies, deletes weight elements in the weight set that do not meet a preset condition, deletes edges corresponding to the weight elements that do not meet the preset condition from the edge set to obtain a second node relationship graph, obtains one or more third node relationship graphs based on the second node relationship graph, and trains a second machine learning model based on the third node relationship graph. Wherein, the first node relationship graph includes a weight set, a vertex set, and an edge set. The vertices of the vertex set represent the client nodes, the edges of the edge set represent the model test relationships between the client nodes, the weight elements in the weight set represent the average value of the model accuracies between two client nodes, and each third node relationship graph corresponds to a second machine learning model. The model accuracy is obtained by the client node testing the models of other client nodes.

[0060] In this application, during the training process of the machine learning model, each client node trains the first machine learning model based on local data and sends the model accuracy of the first machine learning model to the server node. In this way, the server node does not need to obtain the local data of each client node, and thus can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, this application can delete the weight elements in the weight set that do not meet the preset condition, and delete the edges corresponding to the weight elements that do not meet the preset condition from the edge set to obtain a second node relationship graph. In this way, it can avoid the influence of the model accuracy that does not meet the preset condition on model training, and thus can improve the convergence of the model accuracy on the premise that the data used for training the model does not meet the independent and identically distributed condition. At the same time, based on the second node relationship graph, this application can obtain a third node relationship graph, and then can train a second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be trained separately based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0061] A fifth aspect of this application discloses an electronic device, which is applied to a server node. The device includes: at least one processor and a memory;

[0062] The memory stores computer execution instructions;

[0063] The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the method of the first aspect of this application.

[0064] By executing the training method of the distributed machine learning model, the electronic device of the present application can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the present application can delete the weight elements in the weight set that do not meet the preset conditions, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, based on the second node relationship graph, the present application can obtain a third node relationship graph, and then can train a second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be separately trained based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0065] A sixth aspect of the present application discloses an electronic device, which is applied to a client node. The device includes: at least one processor and a memory;

[0066] The memory stores computer execution instructions;

[0067] The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the method of the second aspect of the present application.

[0068] By executing the training method of the distributed machine learning model, the electronic device of the present application can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the present application can delete the weight elements in the weight set that do not meet the preset conditions, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, based on the second node relationship graph, the present application can obtain a third node relationship graph, and then can train a second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be separately trained based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0069] A seventh aspect of the present application discloses a storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the methods of the first and second aspects of the present application.

[0070] By executing the training method of the distributed machine learning model, the storage medium of the present application can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the present application can delete the weight elements in the weight set that do not meet the preset conditions and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on the model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, based on the second node relationship graph, the present application can obtain a third node relationship graph, and then can train a second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, the second machine learning model can be separately trained based on the third node relationship graph, making the second machine learning model have better model convergence. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0072] Figure 1 is a schematic flowchart of a training method for a distributed machine learning model disclosed in Embodiment 1 of the present application;

[0073] Figure 2 is a schematic flowchart of a training method for a distributed machine learning model disclosed in Embodiment 2 of the present application;

[0074] Figure 3 is a schematic structural diagram of a training device for a distributed machine learning model disclosed in Embodiment 3 of the present application;

[0075] Figure 4 is a schematic structural diagram of a training device for a distributed machine learning model disclosed in Embodiment 4 of the present application;

[0076] Figure 5 is a schematic structural diagram of an electronic device disclosed in Embodiment 4 of the present application;

[0077] Figure 6 It is a schematic structural diagram of an electronic device disclosed in Embodiment 4 of the present application. Detailed implementation manners

[0078] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0079] Embodiment 1

[0080] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a method for training a distributed machine learning model disclosed in an embodiment of the present application. Among them, this method is applied to a server node in a distributed system. As Figure 1 shown, the method of the embodiment of the present application includes the following steps:

[0081] 101. Obtain a plurality of model accuracies of a first machine learning model sent by a plurality of client nodes, where the model accuracy is obtained by the client nodes testing the models of other client nodes;

[0082] 102. Construct a first node relationship graph based on the plurality of model accuracies. The first node relationship graph includes a weight set, a vertex set, and an edge set. The vertices in the vertex set represent client nodes, the edges in the edge set represent the model test relationships between client nodes, and the weight elements in the weight set represent the average model accuracy between two client nodes;

[0083] 103. Delete the weight elements in the weight set that do not meet the preset conditions, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph;

[0084] 104. Obtain one or more third node relationship graphs based on the second node relationship graph;

[0085] 105. Train a second machine learning model based on the third node relationship graph, where each third node relationship graph corresponds to a second machine learning model.

[0086] In the embodiments of the present application, during the training process of the machine learning model, each client node trains the first machine learning model based on local data and sends the model accuracy of the first machine learning model to the server node. In this way, the server node does not need to obtain the local data of each client node, and thus can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, in the embodiments of the present application, based on the first node relationship graph, the weight elements in the weight set that do not meet the preset conditions can be deleted, and the edges corresponding to the weight elements that do not meet the preset conditions can be deleted from the edge set to obtain the second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on the model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, in the embodiments of the present application, the third node relationship graph can be obtained based on the second node relationship graph, and thus the second machine learning model can be trained based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be trained separately based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0087] In the embodiments of the present application, for step 101, the first machine learning model may be a neural network model.

[0088] In the embodiments of the present application, for step 101, the model accuracy is obtained by the client node testing the models of other client nodes, which means that each client node tests the models of other client nodes to obtain the model accuracy. For example, assuming there are M client nodes, then M client nodes can obtain M model parameters by training the first machine learning model. Then each client node tests the model parameters of the other M - 1 client nodes to obtain M*(M - 1) model accuracies.

[0089] In the embodiments of the present application, for step 102, the first node relationship graph is an undirected weighted graph, where the undirected weighted graph is represented by G(V, E, W), where the vertex set V = {C i |1 ≤ i ≤ M, i ∈ N}, the edge set E = {(C i , C j )|C i , C j ∈ V, i ∈ N, 1 ≤ i ≤ M, j ∈ N, 1 ≤ j ≤ M}, the weight set W = {w ij |w ij = (P ij + P ji ) / 2}, and M represents the total number of client nodes. The weights w of all the edges in the graph G(V, E, W)ij Perform size sorting, and the sorting results form a set W o .

[0090] In the embodiment of the present application, for step 102, the second machine learning model may be a neural network model. Further, each third node relationship graph corresponding to a second machine learning model means that each third node relationship graph corresponds to a second machine learning model. For example, if two third node relationship graphs can be segmented based on the second node relationship graph, then there are two second machine learning models.

[0091] In the embodiment of the present application, for step 103, the weight elements that do not meet the preset conditions refer to the weight elements with values less than the preset value. For example, the weight elements with values less than 0.4 in the weight set are the weight elements that do not meet the preset conditions.

[0092] In the embodiment of the present application, for step 104, by deleting the weight elements that do not meet the preset conditions in the weight set and deleting the edges corresponding to the weight elements that do not meet the preset conditions from the edge set, a second node relationship graph can be obtained. Among them, if the first node relationship graph is a connected graph, then after deleting the edges corresponding to the weight elements that do not meet the preset conditions from the first node relationship graph, the second node relationship graph obtained has only one connected subgraph. At this time, only one third node relationship graph can be obtained based on the second node relationship graph, that is, only one connected subgraph can be obtained based on the second node relationship graph. On the other hand, if the first node relationship graph is not a connected graph, then after deleting the edges corresponding to the weight elements that do not meet the preset conditions from the first node relationship graph, the second node relationship graph obtained can be segmented into multiple third node relationship graphs, that is, the second node relationship graph can be segmented into multiple connected subgraphs. It should be noted that the third node relationship graph refers to a connected subgraph, and multiple third node relationship graphs refer to two or more third node relationship graphs.

[0093] In the embodiment of the present application, as an alternative implementation, training the second machine learning model based on the third node relationship graph includes the following sub-steps:

[0094] Send the first model parameters of the second machine learning model to each client node in the third node relationship graph, so that each client node in the third node relationship graph trains the second machine learning model based on the first model parameters of the second machine learning model and local sample data until the second model parameters when the second machine learning model meets the first preset training condition are obtained;

[0095] Obtain the second model parameters sent by each client node;

[0096] Calculate the third model parameters of the second machine learning model based on the second model parameters of each client node;

[0097] Determine whether the third model parameter meets the second preset training condition. If it does not meet the second preset training condition, then iteratively train the second machine learning model until the third model parameter meets the second preset training condition.

[0098] In this alternative embodiment, by sending the first model parameter of the second machine learning model to each client node in the third node relationship graph, each client node in the third node relationship graph can train the second machine learning model based on the first model parameter of the second machine learning model and local sample data until the second model parameter when the second machine learning model meets the first preset training condition is obtained. At the same time, by obtaining the second model parameters sent by each client node, the third model parameter of the second machine learning model can be calculated based on the second model parameters of each client node. Furthermore, by determining whether the third model parameter meets the second preset training condition, the second machine learning model is iteratively trained until the third model parameter meets the second preset training condition. That is, through the above steps, the training of the second machine learning model can be completed in a distributed training manner.

[0099] In this alternative embodiment, the first preset training condition may be that the second machine learning model converges, or it may mean that the number of training times of the second machine learning model reaches a preset number of times, such as reaching 5 times.

[0100] In this alternative embodiment, the second preset training condition may be that the second machine learning model converges, or it may mean that the number of iterations in the cyclic training process of the second machine learning model reaches a preset number of times, such as reaching 20 times.

[0101] In an embodiment of the present application, as an alternative embodiment, calculating the third model parameter of the second machine learning model based on the second model parameters of each client node includes the following sub-steps:

[0102] Determine the total number of second model parameters;

[0103] Based on the total number of second model parameters and each second model parameter, calculate the average value of all second model parameters;

[0104] Take the average value of all second model parameters as the third model parameter.

[0105] In this alternative embodiment, by determining the total number of second model parameters, the average value of all second model parameters can be calculated based on the total number of second model parameters and each second model parameter, so that the average value of all second model parameters can be taken as the third model parameter.

[0106] In an embodiment of the present application, as an alternative implementation, after constructing the first node relationship graph based on several model accuracies and before deleting the weight elements in the weight set that do not meet the preset conditions, the method of the embodiment of the present application further includes the following steps:

[0107] Sort the weight elements in the weight set and determine the weight elements whose values are less than the preset threshold;

[0108] Determine the weight elements whose values are less than the preset threshold as the weight elements that do not meet the preset conditions.

[0109] In this alternative implementation, by sorting the weight elements in the weight set, it is possible to determine the weight elements whose values are less than the preset threshold, and thus it is possible to determine the weight elements whose values are less than the preset threshold as the weight elements that do not meet the preset conditions.

[0110] In an embodiment of the present application, as an alternative implementation, obtaining one or more third node relationship graphs based on the second node relationship graph includes the following sub-steps:

[0111] Judge whether the second node relationship graph is a connected graph. If so, determine one third node relationship graph;

[0112] When the second node relationship graph is not a connected graph, divide the second node relationship graph into multiple third node relationship graphs. The third node relationship graph is a connected sub-graph.

[0113] In this alternative implementation, by judging whether the second node relationship graph is a connected graph, when the second node relationship graph is not a connected graph, the second node relationship graph can be divided into multiple third node relationship graphs. The third node relationship graph is a connected sub-graph, so that at least two second machine learning models can be trained separately.

[0114] In an embodiment of the present application, as an alternative implementation method, before obtaining several model accuracies for the first machine learning model, the method further includes:

[0115] Send the initial model parameters of the first machine learning model to each client node, so that each client node trains the first machine learning model based on the initial model parameters of the first machine learning model until the first machine learning model meets the third preset training condition. Among them, the parameters corresponding to the first machine learning model under the third preset training condition are the model parameters of the client node.

[0116] In this alternative implementation, the server node initializes the parameters of the first machine learning model using the Gaussian random initialization method to obtain the initial model parameters of the first machine learning model.

[0117] In this alternative embodiment, the third preset training condition may refer to the convergence of the first machine learning model, or it may refer to the number of training times of the first machine learning model reaching a preset number, such as reaching 80 times.

[0118] It can be seen that in this alternative embodiment, by sending the initial model parameters of the first machine learning model to each client node, each client node can train the first machine learning model based on the initial model parameters of the first machine learning model until the first machine learning model meets the third preset training condition.

[0119] It should be noted that the training method of the distributed machine learning model in the embodiments of the present application can be applied to the scenario of email analysis. The training method of the distributed machine learning model in the embodiments of the present application will be described below in combination with this scenario.

[0120] As an example, the server node S initializes the parameters of the machine learning model (such as a neural network model, etc.) it uses by the Gaussian random initialization method, and sends the initialized values of the model parameters to all distributed email sample data analysis client nodes (i.e., client nodes).

[0121] Furthermore, each distributed email sample data analysis client node C i uses the email sample data stored locally as the training data set, and iteratively trains the first machine learning model using the stochastic gradient descent algorithm. When the first machine learning model converges or the number of training iterations reaches I = 80 times, the training of the first machine learning model stops.

[0122] Furthermore, each distributed email sample data analysis client node C i sends its own model parameters θ ik (for the model parameters of the first machine learning model) to the server node S, and the server node S forwards θ ik to the other 9 email sample data analysis client nodes (M = 10).

[0123] Furthermore, each distributed email sample data analysis client node C i tests the accuracy P of the 9th first machine learning model received from the server node S on the local training data ij , where 1 ≤ j ≤ 10 and j ≠ i.

[0124] Furthermore, each distributed email sample data analysis client node C i uploads the 9 accuracies P tested locally ij to the server node S.

[0125] Further, the server node S aggregates the data analysis of 10 email samples to obtain 90 precisions uploaded by the client nodes. Taking the email sample data analysis client nodes as vertices, the precisions as edges, and the average of the two precisions between two vertices as the weight of the edge between the two vertices, an undirected weighted graph (i.e., the first node relationship graph) G(V, E, W) is constructed. Among them, the vertex set V = {C i |1 ≤ i ≤ 10, i ∈ N}, the edge set E = {(C i , C j )|C i , C j ∈ V, i ∈ N, 1 ≤ i ≤ 10, j ∈ N, 1 ≤ j ≤ 10}, and the weight set W = {w ij |w ij = (P ij + P ji ) / 2}.

[0126] Further, the weights w ij of all edges in the graph G(V, E, W) are sorted by size, and the sorting results form a set W o . The minimum weight value w o is selected from W min in sequence, the edge corresponding to this weight value is deleted from G(V, E, W), and the set W o is updated as W o - {w min}. The above process is repeated cyclically until w min ≥ TH = 0.6. At this time, the second node relationship graph is obtained.

[0127] Further. If G(V, E, W) is a connected graph, then G(V, E, W) after deleting some edges has only one subgraph. If G(V, E, W) is not a connected graph, then G(V, E, W) after deleting some edges will be split into multiple subgraphs, that is, split into multiple third node relationship graphs. The third node relationship graph is a connected subgraph.

[0128] Further, for each subgraph G′, the server node S will separately construct and maintain a distributed machine learning model M′ (i.e., the second machine learning model). In the kth training cycle, the server node S sends the current parameter values of the model M′ to all nodes within the subgraph G′, and all nodes update the parameters of their respective local machine learning models with the received model parameter values. Each node C i uses the email sample data stored locally as the training data set, and adopts the stochastic gradient descent algorithm to iteratively train the current model. When the model converges or the number of training iterations reaches I′ = 5 times, the training stops. Each node C i sends the model parameter θ ikSend it back to the server node S. For each sub-graph G′, the server node S updates the parameters of the model M′ to θ according to the following formula k .

[0129]

[0130] where, if the parameters θ of the model M′ k converge or k≥E TH = 20, then terminate the training process, otherwise update the value of k to k = k + 1, and repeat training the second machine learning model.

[0131] By applying the method of the embodiments of the present application to the scenario of email analysis, it can solve the problem that it is difficult for a distributed machine learning model to converge when the scattered training data does not meet the independent and identically distributed condition. At the same time, under the condition of protecting the privacy of node email data, it is possible to cluster distributed nodes with highly similar email data features (that is, divide the second node graph into several third node relationship graphs, so as to classify similar nodes in the same third node relationship graph). On the other hand, under the condition of protecting the privacy of node email data, the similarity and relevance between the email data of multiple nodes can be evaluated through the second machine learning model obtained by training.

[0132] Embodiment 2

[0133] Please refer to Figure 2 . Figure 2 is a schematic flowchart of a method for training a distributed machine learning model disclosed in the embodiments of the present application, where this method is applied to the user end node in a distributed system. As Figure 2 shown, the method of the embodiments of the present application includes the following steps:

[0134] 201. Send the model accuracy of the first machine learning model to the server node, so that the server node constructs a first node relationship graph based on a plurality of model accuracies, deletes the weight elements in the weight set that do not meet the preset conditions, deletes the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph, obtains one or more third node relationship graphs based on the second node relationship graph, and trains a second machine learning model based on the third node relationship graph. Among them, the first node relationship graph includes a weight set, a vertex set, and an edge set. The vertices of the vertex set represent user end nodes, the edges of the edge set represent the model test relationships between user end nodes, the weight elements in the weight set represent the average model accuracy between two user end nodes, each third node relationship graph corresponds to a second machine learning model, and the model accuracy is obtained by the user end node testing the models of other user end nodes.

[0135] In the embodiments of the present application, during the training process of the machine learning model, each client node trains the first machine learning model based on local data and sends the model accuracy of the first machine learning model to the server node. In this way, the server node does not need to obtain the local data of each client node, and thus can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the present application can delete the weight elements in the weight set that do not meet the preset conditions, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain the second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on the model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, based on the second node relationship graph, the present application can obtain the third node relationship graph, and then train the second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model is trained separately based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0136] In the embodiments of the present application, as an optional implementation manner, the method of the embodiments of the present application further includes the following steps:

[0137] Receiving the first model parameters of the second machine learning model sent by the server node, where the client node is the client node in the third node relationship graph;

[0138] Training the second machine learning model based on the first model parameters of the second machine learning model and local sample data until the second model parameters when the second machine learning model meets the first preset training condition are obtained;

[0139] Sending the second model parameters to the server node, so that the server node calculates the third model parameters of the second machine learning model based on the second model parameters of each client node, determines whether the third model parameters meet the second preset training condition, and when the third model parameters do not meet the second preset training condition, iteratively trains the second machine learning model until the third model parameters meet the second preset training condition.

[0140] By sending the first model parameters of the second machine learning model to each client node in the third node relationship graph, each client node in the third node relationship graph can train the second machine learning model based on the first model parameters of the second machine learning model and local sample data until the second model parameters when the second machine learning model meets the first preset training condition are obtained. At the same time, by obtaining the second model parameters sent by each client node, the third model parameters of the second machine learning model can be calculated based on the second model parameters of each client node. Furthermore, by determining whether the third model parameters meet the second preset training condition, the second machine learning model is iteratively trained until the third model parameters meet the second preset training condition. That is, through the above steps, the training of the second machine learning model can be completed in a distributed training manner.

[0141] In an embodiment of the present application, as an alternative implementation, the third model parameters are the average values of all the second model parameters. Among them, the average values of all the second model parameters are calculated by the server node based on the total number of the second model parameters and each second model parameter after determining the total number of the second model parameters.

[0142] In this alternative implementation, by determining the total number of the second model parameters, the average values of all the second model parameters can be calculated based on the total number of the second model parameters and each second model parameter, so that the average values of all the second model parameters can be used as the third model parameters.

[0143] In an embodiment of the present application, as an alternative implementation, the weight elements that do not meet the preset conditions are the weight elements with values less than the preset threshold. Among them, the weight elements with values less than the preset threshold are determined by the server node after sorting the weight elements in the weight set.

[0144] In this alternative implementation, by sorting the weight elements in the weight set, the weight elements with values less than the preset threshold can be determined, so that the weight elements with values less than the preset threshold can be determined as the weight elements that do not meet the preset conditions.

[0145] In an embodiment of the present application, as an alternative implementation, when the second node relationship graph is a connected graph, a third node relationship graph is determined. When the second node relationship graph is a connected graph, the second node relationship graph is divided into multiple third node relationship graphs. The third node relationship graph is a connected subgraph.

[0146] In this alternative embodiment, by determining whether the second node relationship graph is a connected graph, when the second node relationship graph is not a connected graph, the second node relationship graph can be divided into multiple third node relationship graphs, and the third node relationship graph is a connected subgraph, so that at least two second machine learning models can be separately trained.

[0147] In an embodiment of the present application, as an alternative embodiment, before receiving the first model parameters of the second machine learning model sent by the server node, the method of the embodiment of the present application further includes the following steps:

[0148] Receiving the initialization model parameters of the first machine learning model sent by the server node;

[0149] Training the first machine learning model based on the initialization model parameters of the first machine learning model until the first machine learning model meets the third preset training condition, where the parameters corresponding to the first machine learning model under the third preset training condition are the model parameters of the user - end node.

[0150] In this alternative embodiment, by sending the initialization model parameters of the first machine learning model to each user - end node, each user - end node can train the first machine learning model based on the initialization model parameters of the first machine learning model until the first machine learning model meets the third preset training condition.

[0151] Embodiment III

[0152] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a training device for a distributed machine learning model disclosed in an embodiment of the present application, where the device is located in the server node. As Figure 3 shown, the device includes:

[0153] An obtaining module 301, configured to obtain several model accuracies of the first machine learning model sent by several user - end nodes, where the model accuracy is obtained by the user - end node testing the models of other user - end nodes;

[0154] A constructing module 302, configured to construct a first node relationship graph based on the several model accuracies, where the first node relationship graph includes a weight set, a vertex set, and an edge set, the vertices of the vertex set represent user - end nodes, the edges of the edge set represent the model test relationships between user - end nodes, and the weight elements in the weight set represent the average model accuracy between two user - end nodes;

[0155] A deletion module 303, configured to delete weight elements in the weight set that do not meet a preset condition, and delete edges corresponding to the weight elements that do not meet the preset condition from the edge set, so as to obtain a second node relationship graph;

[0156] A splitting module 304, configured to obtain one or more third node relationship graphs based on the second node relationship graph;

[0157] A training module 305, configured to train a second machine learning model based on the third node relationship graph, where each third node relationship graph corresponds to a second machine learning model.

[0158] In the embodiments of the present application, during the training process of the machine learning model, each client node trains the first machine learning model based on local data and sends the model accuracy of the first machine learning model to the server node. In this way, the server node does not need to obtain the local data of each client node, and thus can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, in the embodiments of the present application, based on the first node relationship graph, weight elements in the weight set that do not meet the preset condition can be deleted, and edges corresponding to the weight elements that do not meet the preset condition can be deleted from the edge set, so as to obtain a second node relationship graph. In this way, the influence of model accuracy that does not meet the preset condition on model training can be avoided, and thus the convergence of model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed condition. At the same time, in the embodiments of the present application, based on the second node relationship graph, a third node relationship graph can be obtained, and thus the second machine learning model can be trained based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be trained separately based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0159] It should be noted that for other detailed descriptions of the device in the embodiments of the present application, please refer to the relevant descriptions in Embodiment 1 of the present application, and the embodiments of the present application will not elaborate on this.

[0160] Embodiment 4

[0161] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a training device for a distributed machine learning model disclosed in the embodiments of the present application. Among them, the device is applied to several client nodes in a distributed system. As Figure 4 shown, the device in the embodiments of the present application includes the following functional modules:

[0162] A sending module 401, configured to send the model accuracy of the first machine learning model to a server node, so that the server node constructs a first node relationship graph based on a plurality of model accuracies, deletes weight elements in a weight set that do not meet a preset condition, deletes edges corresponding to the weight elements that do not meet the preset condition from an edge set to obtain a second node relationship graph, obtains one or more third node relationship graphs based on the second node relationship graph, and trains a second machine learning model based on the third node relationship graph. The first node relationship graph includes a weight set, a vertex set, and an edge set. The vertices in the vertex set represent client nodes, the edges in the edge set represent model test relationships between client nodes, and the weight elements in the weight set represent the average model accuracy between two client nodes. Each third node relationship graph corresponds to a second machine learning model, and the model accuracy is obtained by a client node testing a model of other client nodes.

[0163] In the embodiment of the present application, during the training process of the machine learning model, each client node trains the first machine learning model based on local data and sends the model accuracy of the first machine learning model to the server node. In this way, the server node does not need to obtain the local data of each client node, and thus can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the embodiment of the present application can delete weight elements in the weight set that do not meet the preset condition, and delete edges corresponding to the weight elements that do not meet the preset condition from the edge set to obtain a second node relationship graph. In this way, the influence of model accuracies that do not meet the preset condition on model training can be avoided, and thus the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed condition. At the same time, based on the second node relationship graph, the embodiment of the present application can obtain a third node relationship graph, and thus can train a second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model is trained separately based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0164] It should be noted that for other detailed descriptions of the device in the embodiment of the present application, please refer to the relevant descriptions in Embodiment 1 of the present application, and the embodiment of the present application will not elaborate on this.

[0165] Embodiment Five

[0166] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device disclosed in the embodiment of the present application. The electronic device is applied to a server node. As Figure 5As shown in the figure, the device according to the embodiment of the present application includes: at least one processor 501 and a memory 502;

[0167] The memory 502 stores computer-executable instructions;

[0168] At least one processor 501 executes the computer-executable instructions stored in the memory 502, so that at least one processor executes the method of Embodiment 1 of the present application.

[0169] By executing the training method of the distributed machine learning model, the electronic device according to the embodiment of the present application can, while completing the training of the machine learning model, ensure the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the embodiment of the present application can delete the weight elements in the weight set that do not meet the preset conditions, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on model training can be avoided, and further, on the premise that the data used for training the model does not meet the independent and identically distributed conditions, the convergence of the model accuracy can be improved. At the same time, based on the second node relationship graph, the embodiment of the present application can obtain a third node relationship graph, and further can train a second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model is separately trained based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0170] Embodiment 6

[0171] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device disclosed in the embodiment of the present application. Among them, the electronic device is applied to a client node. As Figure 6 shown, the device according to the embodiment of the present application includes: at least one processor 601 and a memory 602;

[0172] The memory 602 stores computer-executable instructions;

[0173] At least one processor 601 executes the computer-executable instructions stored in the memory 602, so that at least one processor 601 executes the method of Embodiment 2 of the present application.

[0174] By executing the training method of the distributed machine learning model, the electronic device according to the embodiment of the present application can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the present application can delete the weight elements in the weight set that do not meet the preset conditions, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain the second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on the model training can be avoided, and then the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, based on the second node relationship graph, the present application can obtain the third node relationship graph, and then can train the second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be separately trained based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0175] Embodiment VII

[0176] The embodiment of the present application discloses a storage medium storing computer-executable instructions, which are used to implement the methods of Embodiment 1 and Embodiment 2 of the present application when executed by a processor.

[0177] By executing the training method of the distributed machine learning model, the storage medium according to the embodiment of the present application can complete the training of the machine learning model while ensuring the privacy of the local data of each client node. On the other hand, based on the first node relationship graph, the present application can delete the weight elements in the weight set that do not meet the preset conditions, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain the second node relationship graph. In this way, the influence of the model accuracy that does not meet the preset conditions on the model training can be avoided, and then the convergence of the model accuracy can be improved on the premise that the data used for training the model does not meet the independent and identically distributed conditions. At the same time, based on the second node relationship graph, the present application can obtain the third node relationship graph, and then can train the second machine learning model based on the third node relationship graph. In this way, since each third node relationship graph corresponds to a second machine learning model, each second machine learning model can be separately trained based on the third node relationship graph, so that the second machine learning model has better model convergence.

[0178] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0179] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0180] Furthermore, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0181] It should be noted that if the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other various media that can store program codes.

[0182] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0183] The above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for a distributed machine learning model, characterized in that The method is applied to a server node in a distributed system, and the method includes: Obtaining a plurality of model accuracies of a first machine learning model sent by a plurality of client nodes, where the model accuracy is obtained by the client nodes testing the models of other client nodes; Constructing a first node relationship graph based on the plurality of model accuracies, where the first node relationship graph includes a weight set, a vertex set, and an edge set, the vertices in the vertex set represent the client nodes, the edges in the edge set represent the model test relationships between the client nodes, and the weight elements in the weight set represent the average model accuracy between two client nodes; Deleting the weight elements in the weight set that do not meet the preset conditions, and deleting the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph; Obtaining one or more third node relationship graphs based on the second node relationship graph; Training a second machine learning model based on the third node relationship graph, where each third node relationship graph corresponds to a second machine learning model; Processing email data based on the second machine learning model.

2. The method according to claim 1, wherein The training of the second machine learning model based on the third node relationship graph includes: Sending first model parameters of the second machine learning model to each client node in the third node relationship graph, so that each client node in the third node relationship graph trains the second machine learning model based on the first model parameters of the second machine learning model and local sample data until second model parameters when the second machine learning model meets the first preset training condition are obtained; Obtaining the second model parameters sent by each client node; Calculating third model parameters of the second machine learning model based on the second model parameters of each client node; Determining whether the third model parameters meet the second preset training condition. If the third model parameters do not meet the second preset training condition, iteratively train the second machine learning model until the third model parameters meet the second preset training condition.

3. The method according to claim 2, wherein The calculating the third model parameters of the second machine learning model based on the second model parameters of each client node includes: Determining the total number of the second model parameters; Calculating the average value of all the second model parameters based on the total number of the second model parameters and each second model parameter; Taking the average value of all the second model parameters as the third model parameters.

4. The method according to claim 1, wherein After constructing the first node relationship graph based on the plurality of model accuracies and before deleting the weight elements in the weight set that do not meet the preset conditions, the method further includes: Sorting the weight elements in the weight set and determining the weight elements with values less than a preset threshold; Determining the weight elements with values less than the preset threshold as the weight elements that do not meet the preset conditions.

5. The method according to claim 1, characterized in that, The obtaining one or more third node relationship graphs based on the second node relationship graph includes: Determine whether the second node relationship graph is a connected graph. If so, determine one of the third node relationship graphs; When the second node relationship graph is not a connected graph, divide the second node relationship graph into multiple third node relationship graphs, where the third node relationship graph is a connected subgraph.

6. The method according to claim 1, wherein Before obtaining the model accuracies of a number of user client nodes for the first machine learning model, the method further includes: Send the initial model parameters of the first machine learning model to each user client node, so that each user client node trains the first machine learning model based on the initial model parameters of the first machine learning model until the first machine learning model meets the third preset training condition, where the parameters corresponding to the first machine learning model under the third preset training condition are the model parameters of the user client node.

7. A training method for a distributed machine learning model, characterized in that, The method is applied to a number of user client nodes in a distributed system. The method includes: Send the model accuracy for the first machine learning model to the server node, so that the server node constructs a first node relationship graph based on a number of the model accuracies, and deletes the weight elements in the weight set that do not meet the preset conditions, and deletes the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph, and obtains one or more third node relationship graphs based on the second node relationship graph, and trains a second machine learning model based on the third node relationship graph, and processes the email data based on the second machine learning model, where the first node relationship graph includes the weight set, the vertex set, and the edge set, the vertices of the vertex set represent the user client nodes, the edges of the edge set represent the model test relationships between the user client nodes, the weight elements in the weight set represent the average model accuracy between two user client nodes, each third node relationship graph corresponds to a second machine learning model, and the model accuracy is obtained by the user client node testing the models of other user client nodes.

8. The method according to claim 7, wherein The method further includes: Receive the first model parameters of the second machine learning model sent by the server node, where the user client node is the user client node in the third node relationship graph; Train the second machine learning model based on the first model parameters of the second machine learning model and local sample data until the second model parameters when the second machine learning model meets the first preset training condition are obtained; Send the second model parameters to the server node, so that the server node calculates the third model parameters of the second machine learning model based on the second model parameters of each user client node, and determines whether the third model parameters meet the second preset training condition, and when the second preset training condition is not met, iteratively train the second machine learning model until the third model parameters meet the second preset training condition.

9. The method according to claim 8, wherein The third model parameter is the average value of all the second model parameters, where the average value of all the second model parameters is calculated by the server node based on the total number of the second model parameters and each of the second model parameters after determining the total number of the second model parameters.

10. The method according to claim 7, characterized in that, The weight elements that do not meet the preset conditions are the weight elements with values less than the preset threshold, where the weight elements with values less than the preset threshold are determined by the server node after sorting the weight elements in the weight set.

11. The method according to claim 7, wherein When the second node relationship graph is a connected graph, the number of the third node relationship graphs is one. When the second node relationship graph is not a connected graph, a plurality of the third node relationship graphs are obtained based on the second node relationship graph, and the third node relationship graph is a connected subgraph.

12. The method according to claim 8, wherein Before receiving the first model parameter of the second machine learning model sent by the server node, the method further includes: Receiving the initialization model parameter of the first machine learning model sent by the server node; Training the first machine learning model based on the initialization model parameter of the first machine learning model until the first machine learning model meets the third preset training condition, where the parameters corresponding to the first machine learning model under the third preset training condition are the model parameters of the client node.

13. A training device for a distributed machine learning model, characterized in that, The device is located in the server node, and the device includes: An acquisition module, configured to acquire a plurality of model accuracies of a first machine learning model sent by a plurality of client nodes, where the model accuracy is obtained by the client node testing the models of other client nodes; A construction module, configured to construct a first node relationship graph based on the plurality of model accuracies, where the first node relationship graph includes a weight set, a vertex set, and an edge set, the vertices in the vertex set represent the client nodes, the edges in the edge set represent the model test relationships between the client nodes, and the weight elements in the weight set represent the average value of the model accuracies between two client nodes; A deletion module, configured to delete the weight elements in the weight set that do not meet the preset conditions, and delete the edges corresponding to the weight elements that do not meet the preset conditions from the edge set to obtain a second node relationship graph; A splitting module, configured to obtain one or more third node relationship graphs based on the second node relationship graph; A training module, configured to train a second machine learning model based on the third node relationship graph, where each of the third node relationship graphs corresponds to a second machine learning model; The device is further configured to: process the email data based on the second machine learning model.

14. A training device for a distributed machine learning model, characterized in that, The device is applied to a plurality of client nodes in a distributed system, and the device includes: A sending module, configured to send the model accuracy of a first machine learning model to a server node, so that the server node constructs a first node relationship graph based on a plurality of the model accuracies, and deletes weight elements in a weight set that do not meet a preset condition, and deletes edges corresponding to the weight elements that do not meet the preset condition from an edge set to obtain a second node relationship graph, and obtains one or more third node relationship graphs based on the second node relationship graph, and trains a second machine learning model based on the third node relationship graph, and processes email data based on the second machine learning model, wherein the first node relationship graph includes the weight set, a vertex set, and the edge set, vertices in the vertex set represent the client nodes, edges in the edge set represent model test relationships between the client nodes, weight elements in the weight set represent the average model accuracy between two client nodes, each third node relationship graph corresponds to a second machine learning model, and the model accuracy is obtained by a client node testing a model of other client nodes.

15. An electronic device, characterized in that, The electronic device is applied to a server node, and the device includes: at least one processor, a memory; The memory stores computer execution instructions; The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the method according to any one of claims 1-6.

16. An electronic device, characterized in that, The electronic device is applied to a client node, and the device includes: at least one processor, a memory, and; The memory stores computer execution instructions; The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the method according to any one of claims 7-12.

17. A storage medium, characterized in that, Computer execution instructions are stored in the storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1-6 or any one of claims 7-12.

Citation Information

Patent Citations

  • Transmission optimization method for distributed machine learning data rearrangement in data center

    CN111917648A

  • Online federated learning task allocation method and device and federated learning method and system

    CN113887748A