A Global Gradient Double-Tracking Distributed Method Based on Heterogeneous Mixed Data
The global gradient double tracking method addresses the challenge of non-uniform HBL problems by using local and global tracking variables to converge models efficiently, reducing communication overhead and enhancing accuracy in medical diagnosis.
Patent Information
- Application Number
- CN202211610706.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-14
AI Technical Summary
Existing distributed learning methods cannot effectively solve the most generalized non-uniform HBL problem including all scenarios, especially in non-convex federated learning, how to overcome the incompleteness of local data and reduce communication overhead.
A global gradient dual-tracking distributed method based on heterogeneous hybrid data is adopted. By constructing a local data matrix and a global data matrix, a dual random interaction matrix is introduced, the gradient of the loss function is estimated using two tracking variables, and a multi-step stochastic gradient descent is performed in a multi-node network to realize information interaction between distributed nodes, and finally obtain consistent model parameters.
It effectively solves the most complex non-uniform HBL problem, reduces computational complexity and communication overhead, and improves the accuracy of medical diagnosis.
Smart Images

Figure CN115906362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a global gradient double-tracking distributed method based on heterogeneous mixed data. Background Art
[0002] With the continuous acceleration of the digitalization process of human society, the volume of data has grown explosively. This phenomenon makes it difficult to store all data in one device or processor, and further renders centralized algorithms gradually inapplicable. In view of this, distributed optimization has attracted great attention in various signal processing fields. Moreover, people are paying more and more attention to data privacy and security, which requires that distributed optimization algorithms cannot directly exchange raw data. To address the above challenges, various distributed learning methods have emerged.
[0003] According to the distribution of data at network nodes, distributed learning tasks can be divided into three categories, namely horizontal learning (HL) problems, vertical learning (VL) problems, and hybrid learning (HBL) problems; in the HL framework, researchers usually construct it into a problem of finite average sum, where each network node has part of the samples but the entire feature set. In the VL system, each node can observe all the samples but can only obtain part of the feature subset; compared with the HL and VL systems, the HBL scenario is the most complex, where each network node only has part of the sample subset and data feature subset, as Figure 1 shown, taking four organizations, namely a general hospital, a dental clinic, an orthopedic hospital, and a research institution as examples. All four are studying gene multi-omics data to analyze the lifespan of patients, but each organization has part of the omics samples or part of the omics features. In addition, HBL can be further divided into uniform and non-uniform modes. For the uniform mode, the data samples owned by customers have the same features, while in the non-uniform mode, they have different feature subsets.
[0004] For non-convex federated learning (FL) problems, the prior art proposes a stochastic gradient tracking algorithm to estimate the global average gradient using auxiliary variables. However, the existing distributed methods of this kind are only applicable to FL problems and are respectively for the HL, VL, and HBL scenarios in the uniform mode, and cannot solve the most general non-uniform HBL problem including all the above scenarios. Therefore, how to overcome the incompleteness of local data and reduce communication overhead is an urgent challenge for distributed HBL problems.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The object of the present invention is to overcome the shortcomings of the prior art, and provides a global gradient double-tracking distributed method based on heterogeneous mixed data, which solves the problem that traditional distributed methods cannot solve the most general non-uniform HBL problem including all the above scenarios.
[0007] The object of the present invention is achieved by the following technical solutions: a global gradient double-tracking distributed method based on heterogeneous mixed data, the distributed method includes:
[0008] Construct a local data matrix, and then determine the global data matrix B i , and determine the system optimization objective;
[0009] System network environment configuration: Considering a multi-node distributed network scenario, set N network nodes, and the nodes follow a certain connected topological structure, and each node exchanges information with neighboring nodes, and select the connection relationship between distributed nodes;
[0010] Construct a network model: Model the multi-node network as an undirected graph, and construct a doubly stochastic interaction matrix;
[0011] Construct a non-uniform HBL problem, and determine the local optimization variable x coupled with the data n and the ordinary optimization variable θ n , and there is also a loss function f used to optimize them. Then introduce two tracking variables to estimate the parameters x n and θ n of the loss function f with respect to each node n, that is, each node n introduces a global linear combination tracking variable z n,i , i ∈ [S] (where i ∈ [S] is a short representation of i = 1,..., S), estimate the global linear combination B i x n ; Based on the tracking variable z n,i , i ∈ [S], each node n continues to introduce a global gradient tracking variable u n , estimate the global gradient where represents the transpose of B i , represents the gradient of the function with respect to z n,i , i ∈ [S];
[0012] Update each node locally Q times respectively to complete the selection of model parameters and realize the information interaction between distributed nodes;
[0013] Local model update: At the r = 1, …, T rounds of iteration, each node n = 1, …, N performs Q-step local updates respectively. Assume that the initial variables of each node are the same and are randomly initialized as: the parameters coupled with the data Other network parameters Global linear combination tracking variable Global gradient tracking variable where for any node m ≠ n;
[0014] Repeat the information interaction and local model update steps between distributed nodes for T rounds, so that each distributed network node obtains a consistent model based on the parameters θ and x obtained in the T-th round;
[0015] Each node inputs the sample data containing the entire feature set into the convergent and consistent distributed model, so as to obtain a medical diagnosis result with high accuracy.
[0016] The steps of constructing the local data matrix include the following:
[0017] Obtain useful data: Obtain the medical record information of patients recorded in general hospitals, specialized hospitals, and scientific research institutions, and process the obtained data information to screen out the useful medical records of patients;
[0018] Construct a data matrix: Take each institution including the hospital as a distributed node, each patient as a sample, and each medical record of the patient as a feature to construct a data matrix with samples as rows and features as columns, and the feature sets of each node do not overlap;
[0019] Represent the global data matrix as where S and J are the total number of samples and features respectively, Set node n to have the local data matrix Set the data that cannot be observed to zero, and then the global data matrix is represented as
[0020] The steps of constructing the network model specifically include the following:
[0021] Model the multi-node network as an undirected graph where ε represents the set of edges connecting nodes, represents the set of nodes. If (m, n) ∈ ε, then nodes m and n can interact information;
[0022] Construct an interaction matrix: Set the interaction matrix as where, for any (m, n) ∈ ε, W n,m > 0, otherwise W n,m= 0, and the interaction matrix W satisfies the double - stochastic condition W1 = 1, 1 T W = 1, where 1 represents the all - ones vector, λ w represents the second - largest eigenvalue of W, and
[0023] The construction of the non - uniform HBL problem specifically includes the following:
[0024] Set the convergence - consistency problem of the multi - node network as:
[0025]
[0026]
[0027] where x n represents the local optimization variable coupled with the data, θ n represents the ordinary optimization variable, represents the set of all neighbor nodes of node n.
[0028] The introduction of two tracking variables to estimate the parameters x and θ of the loss function f specifically includes the following:
[0029] Pre - determine that the gradient forms of the selected loss function f with respect to x and θ are respectively and
[0030] Each node n introduces the global linear - combination tracking variable z n,i , i ∈ [S], where i ∈ [S] means i = 1,..., S, to estimate the global linear combination B i x n ; Based on the tracking variable z n,i , i ∈ [S], each node n continues to introduce the global - gradient tracking variable u n , to estimate the global gradient
[0031] The completion of the selection of model parameters by updating each node locally Q times specifically includes the following:
[0032] Set the initial as the 0 - th round of iteration, and the random initial variables of each node are and The initial tracking variable of each node n is Calculate the derivative Then the initial tracking variable
[0033] Set each node to update locally Q times respectively, and then interact with neighbor nodes for the following model parameters: the parameter x coupled with the data n, other network parameters θ n , global linear combination tracking variable z n , global gradient tracking variable u n , perform one iteration, and go through a total of T loop iterations until the optimization variables and tracking variables of each network node converge to be consistent.
[0034] The information interaction between the distributed nodes specifically includes the following content:
[0035] In the r = 1, …, T rounds of iteration, node n sends to its neighbor nodes
[0036] After node n receives the variables sent by all its weighted neighbor nodes m, it weights them again as
[0037] The local model update specifically includes the following content:
[0038] A1. Randomly select a batch of sample sets denoted as where the number of samples is q = 1, …, Q;
[0039] A2. When q = 1:
[0040] A21. Calculate the derivative of the loss function f with respect to the ordinary optimization variable θ where represents any sample in the sample set randomly selected by node n in the r-th round in;
[0041] A22. Calculate the derivative of the loss function f with respect to the local optimization variable x coupled with the data
[0042] A23. Based on the information received from the neighbor nodes, perform stochastic gradient descent on the ordinary optimization variable θ and the local optimization variable x coupled with the data with step sizes α and β respectively to obtain and
[0043] A24. Based on the information received from the neighbor nodes, update the global linear combination tracking variable by adding the difference between the current and the previous round of data variables multiplied by the total number of nodes
[0044] A25. Calculate the current derivative of the loss function f with respect to the local optimization variable x coupled with the data where represents any sample in the sample set randomly selected by node n in the r-th round in;
[0045] A26. Update the global gradient tracking variable of the local optimization variable parameter x coupled with the data for the loss function f That is, based on the information obtained from neighbors, multiply the difference between the current and the previous round of derivatives by N to estimate the true gradient;
[0046] A3. When q = 2,..., Q:
[0047] A31. Calculate the derivative of the loss function f with respect to the ordinary optimization variable θ where represents any sample in the sample set randomly selected by node n in the r-th round in;
[0048] A32. Calculate the derivative of the loss function f with respect to the local optimization variable x coupled with the data
[0049] A33. Based on the local updates of the previous (q - 1) steps, perform stochastic gradient descent on the ordinary optimization variable θ and the local optimization variable x coupled with the data with step sizes α and β respectively to obtain and
[0050] A34. Based on the local updates of the (q - 1) steps, update the linear combination tracking variable by adding the difference between the current and the previous round of data variables multiplied by the total number of nodes
[0051] A35. Calculate the current derivative of the loss function f with respect to the local optimization variable x coupled with the data where represents any sample in the sample set randomly selected by node n in the r-th round in;
[0052] A36. Update the tracking variable of the gradient of the loss function f with respect to the local optimization variable x coupled with the data That is, based on the local updates of the (q - 1) steps, multiply the difference between the current and the previous round of derivatives by N times to estimate the true gradient;
[0053] A4. After Q steps of local updates, update all variables: the parameters coupled with the data other network parameters global linear combination tracking variable global gradient tracking variable and the selected batch sample set
[0054] The present invention has the following advantages: A global gradient double-tracking distributed method based on heterogeneous mixed data, and the global gradient double-tracking distributed algorithm proposed based on heterogeneous mixed data can effectively solve the most complex non-uniform HBFL problem. Description of the Drawings
[0055] Figure 1 It is a schematic diagram of the HBL distributed learning scenario based on heterogeneous mixed omics data;
[0056] Figure 2 It is a schematic diagram of the process structure of the present invention;
[0057] Figure 3 It is a schematic diagram of the simulation result of the training loss function value of the present invention;
[0058] Figure 4 It is a schematic diagram of the simulation result of the test classification accuracy of the present invention;
[0059] Figure 5 It is a schematic diagram of the convergence error result of the optimized variable x of the present invention;
[0060] Figure 6 It is a schematic diagram of the convergence error result of the tracking variable z of the present invention. Detailed Embodiment
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the protection scope of the present application claimed, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present application. The present invention will be further described below in conjunction with the accompanying drawings.
[0062] The present invention specifically relates to a global gradient double-tracking distributed method based on heterogeneous mixed data. Specifically, in a network system, nodes are distributed at various locations without a central server to allocate information. Each node only has partial samples and partial feature data. It is assumed that the sum of the data of each node is the global data, and each node can only interact with its neighbor nodes to exchange relevant information that is quite different from the original data. By introducing two auxiliary variables, the linear coupling sum of the data and the variable is estimated globally, as well as the gradient based on the global data variable. Each node performs multi-step stochastic gradient descent locally, which not only reduces the computational complexity but also saves communication overhead. Each node exchanges information with its neighbors and finally converges to obtain a global model.
[0063] As Figure 1 shown, it specifically includes the following steps:
[0064] Step 1: Construct a local data matrix;
[0065] 1. Obtain useful data:
[0066] (1) Data collection: For example, general hospitals, specialized hospitals, research institutions, etc. record the medical record information of patients, such as blood pressure, medication situation, injection situation, bone density, etc. They each have partial medical records of partial patients.
[0067] (2) Data cleaning: Such as removing damaged or unclear records, screening important medical records of patients, etc.
[0068] 2. Construct a data matrix:
[0069] (1) Take each institution including hospitals as a distributed node, each patient as a sample, and each medical record of a patient as a feature to construct a data matrix with samples as rows and features as columns, and the feature sets of each node do not overlap;
[0070] (2) Represent the global data matrix as where S and J are the total number of samples and the total number of features respectively, Set node n to have a local data matrix Set the data that cannot be observed to zero, and then the global data matrix is represented as
[0071] Step 2: Determine the system optimization objective;
[0072] Assume that each medical institution wants to provide a relatively reliable diagnosis and treatment advice plan for doctors by exchanging some information that does not expose patient privacy based on local medical record data.
[0073] Step 3: Configure the system network environment;
[0074] Consider a multi - node distributed network scenario. Set N network nodes, where the nodes follow a certain connected topological structure, and each node interacts with its neighboring nodes. Also, select the connection relationships between the distributed nodes, such as line graphs, ring graphs, and fully connected graphs, etc.
[0075] Step 4: Construct a network model;
[0076] 1. Model the multi - node network as an undirected graph where ε represents the set of edges connecting the nodes, and represents the set of nodes. If (m, n) ∈ ε, then nodes m and n can interact with each other. Additionally, assume that the undirected graph is a connected graph, that is, after a long - time interaction, each node can obtain the information of any other node.
[0077] 2. Construct an interaction matrix: Set the interaction matrix as where, for any (m, n) ∈ ε, W n,m > 0, otherwise W n,m = 0, and the interaction matrix W satisfies the doubly - stochastic condition W1 = 1, 1 T W = 1, where 1 represents the all - ones vector, λ w represents the second - largest eigenvalue of W, and
[0078] Step 5: Construct a non - uniform HBL problem;
[0079] Set the convergence - consistency problem of the multi - node network as:
[0080]
[0081]
[0082] where x n is the local optimization variable coupled with the data, θ n is the general optimization variable, represents the set of all neighbor nodes of node n.
[0083] From the above problem, it can be seen that all nodes do not know the complete and only know the local data matrix B n,i . How to enable the distributed nodes to quickly estimate the global matrix with a small communication overhead is the main problem to be solved by the present invention.
[0084] Step 6: Introduce a tracking variable to estimate the gradient;
[0085] 1. Predetermine that the gradient forms of the selected loss function f with respect to x and θ are respectively and
[0086] As can be seen from the above formula, since the complete data matrix B of each node position i , it is impossible to solve the gradient locally. There are two types of unknown terms in the two gradients, including the global linear combination B i x n , and
[0087] 2. Introduce two tracking variables to estimate the above two position terms to track and the two global gradients in;
[0088] (1) Each node n introduces a global linear combination tracking variable z n,i , i ∈ [S] (where i ∈ [S] is a shorthand for i = 1,..., S), to estimate the global linear combination B i x n ;;
[0089] (2) Based on the tracking variable z n,i , i ∈ [S], each node n continues to introduce a global gradient tracking variable u n , to estimate the global gradient where represents the transpose of B i , represents the gradient of the function with respect to z n,i , i ∈ [S].
[0090] Step 7: Selection of parameters;
[0091] 1. Initialize the model, assuming it is the 0th round of iteration at the beginning;
[0092] (1) Each node randomly initializes the variables as and
[0093] (2) Each node n initializes the tracking variable as
[0094] (3) Calculate the derivative
[0095] (4) Initial tracking variable
[0096] 2. Set each node to update Q times locally, and then interact with neighbor nodes for the following model parameters: the parameter x coupled with data n , other network parameters θ n , the global linear combination tracking variable z n , the global gradient tracking variable un To achieve one iteration, a total of T loop iterations are performed until the optimization variables and tracking variables of each network node converge to be consistent.
[0097] Step 8: Information interaction between distributed nodes;
[0098] 1. At the r = 1, …, T rounds of iteration, node n sends to its neighbor nodes
[0099] 2. After node n receives the variables sent by all its weighted neighbor nodes m, it re - weights them to be
[0100] Step 9: Local model update;
[0101] At the r = 1, …, T rounds of iteration, each node n = 1, …, N performs Q - step local updates respectively. For node n, the 0 - th step of the initial local multi - step iteration is
[0102] Assume q = 1, …, Q;
[0103] A1. Randomly select a batch of sample sets denoted as where the number of samples is q = 1, …, Q;
[0104] A2. When q = 1:
[0105] A21. Calculate the derivative of the loss function f with respect to the ordinary optimization variable θ where represents any sample in the sample set randomly selected by node n in the r - th round in;
[0106] A22. Calculate the derivative of the loss function f with respect to the local optimization variable x coupled with the data where represents any sample in the sample set randomly selected by node n in the r - th round in;
[0107] A23. Based on the information received from the neighbor nodes, perform stochastic gradient descent on the ordinary optimization variable θ and the local optimization variable x coupled with the data with step sizes α and β respectively to obtain and
[0108] A24. Based on the information received from the neighbor nodes, update the global linear combination tracking variable by adding the difference between the current and the previous round of data variables multiplied by the total number of nodes
[0109] A25. Calculate the derivative of the loss function f with respect to the local optimization variable x coupled with the data where represents any sample in the sample set randomly selected by node n in the r-th round in;
[0110] A26. Update the global gradient tracking variable of the loss function f with respect to the parameter x of the local optimization variable coupled with the data That is, based on the information obtained from neighbors, multiply the difference between the current and the previous round of derivatives by N to estimate the true gradient;
[0111] A3. When q = 2,..., Q:
[0112] A31. Calculate the derivative of the loss function f with respect to the ordinary optimization variable θ where represents any sample in the sample set randomly selected by node n in the r-th round in;
[0113] A32. Calculate the derivative of the loss function f with respect to the optimization variable x where represents any sample in the sample set randomly selected by node n in the r-th round in;
[0114] A33. Based on the local updates in the previous (q - 1) steps, perform stochastic gradient descent on the ordinary optimization variable θ and the local optimization variable x coupled with the data with step sizes α and β respectively to obtain and
[0115] A34. Based on the local updates in (q - 1) steps, update the linear combination tracking variable by adding the difference between the current and the previous round of data variables multiplied by the total number of nodes
[0116] A35. Calculate the current derivative of the loss function f with respect to the local optimization variable x coupled with the data where represents any sample in the sample set randomly selected by node n in the r-th round in;
[0117] A36. Update the global gradient tracking variable of the loss function f with respect to the local optimization variable x coupled with the data That is, based on the local updates in (q - 1) steps, multiply the difference between the current and the previous round of derivatives by N times to estimate the true gradient;
[0118] A4. After Q steps of local updates, update all variables: the parameters coupled with the data Other network parameters Global linear combination tracking variable Global gradient tracking variable and the selected batch sample set
[0119] Step 10: Repeat the content of Step 8 and Step 9 for T rounds;
[0120] Finally, each distributed network node obtains a consistent model based on the parameters θ and x obtained in the T-th round.
[0121] Step 11: Test the model;
[0122] In the test phase, each node inputs the sample data containing the entire feature set (such as all the medical records of a certain patient) into the converged and consistent distributed model, so as to obtain a medical diagnosis with high accuracy.
[0123] The simulation experiment of the present invention is as follows:
[0124] Dataset selection: Considering the real medical dataset, the MIMIC-III (Medical Information Mart for Intensive Care, MIMIC) dataset. It is a large publicly available intensive care medicine information database provided by the Massachusetts Institute of Technology in the United States. This database records the relevant data of patients in the intensive care unit of Beth Israel Deaconess Medical Center from 2001 to 2012, with the medical health data and records of more than 40,000 patients.
[0125] Among them, the training dataset: 17,903 samples, 714 features; the test dataset: 3,236 samples, 714 features.
[0126] Generate a distributed dataset: Considering 17 medical institutions, randomly allocate non-uniform medical data, where each medical institution has only part of the samples and part of the features.
[0127] Optimization objective: In order to predict the survival tendency (death or survival) of patients during hospitalization, construct the following simple non-convex regularized logistic regression problem:
[0128]
[0129] where, l i represents the label (0 or 1, death or survival) of sample b n,i . At this time, θ is not included in the model.
[0130] Parameter setting: Let λ = 0.1, c = 0.5, consider the step size β = 0.1, and update locally 10 times.
[0131] Network topology:
[0132] 1. Consider the network distribution diagrams of three types of medical structures: line graph, ring graph, and fully connected graph. For these three types of topologies, the connectivity between medical institutions is in the order of: line graph < ring graph < fully connected graph.
[0133] 2. Random connections are made between medical institutions to meet the requirements of a connected graph.
[0134] 3. The interaction matrix is generated as:
[0135]
[0136] where d i represents the number of neighbors of a node.
[0137] Performance evaluation metrics: As Figures 3 - 6 shown, the loss function value, test classification accuracy, optimization variables, and the convergence error (Consensus error, CE) of the linear combination tracking variable of data variables are as follows:
[0138]
[0139]
[0140] The global gradient double tracking algorithm proposed by the present invention can effectively solve the non-uniform HBL problem, and the performance is better when the connectivity of network nodes is stronger.
[0141] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in related fields. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A global gradient double-tracking distributed method based on heterogeneous mixed data, characterized in that: The distributed method includes the following: Construct a local data matrix, and then determine the global data matrix B i , and determine the system optimization objective; System network environment configuration: Considering the multi-node distributed network scenario, set N network nodes. The nodes follow a certain connected topological structure, and each node interacts with adjacent nodes, and select the connection relationship between distributed nodes; Construct a network model: Model the multi-node network as an undirected graph and construct a doubly stochastic interaction matrix; Construct a non-uniform HBL problem, determine the local optimization variable x coupled with the data n and the ordinary optimization variable θ n , and there is also a loss function f used to optimize them. Then introduce two tracking variables to estimate the parameters x n and θ n of the loss function f with respect to each node n, that is, each node n introduces a global linear combination tracking variable z n,i , i ∈ [S] (where i ∈ [S] is a shorthand for i = 1,..., S), estimate the global linear combination B i x n ; based on the tracking variable z n,i , i ∈ [S], each node n continues to introduce a global gradient tracking variable u n to estimate the global gradient where represents the transpose of B i , represents the gradient of the function with respect to z n,i , i ∈ [S]; Update each node locally Q times respectively to complete the selection of model parameters and realize the information interaction between distributed nodes; Local model update: At the r = 1, …, T-th iteration, each node n = 1, …, N performs Q-step local updates respectively. Assume that the initial variables of each node are the same and are randomly initialized as: parameters coupled with data Other network parameters Global linear combination tracking variable Global gradient tracking variable where for any node m ≠ n; Repeat the information interaction between distributed nodes and the local model update step for T rounds, so that each distributed network node obtains a consistent model based on the parameters θ and x obtained in the T-th round; Each node inputs the sample data containing the entire feature set into the converged and consistent distributed model to obtain a medical diagnosis result with high accuracy.
2. The global gradient double-tracking distributed method based on heterogeneous mixed data according to claim 1, characterized in that: The steps of constructing the local data matrix include the following: Obtain useful data: Obtain the medical record information of patients recorded in general hospitals, specialized hospitals, and research institutions, and process the obtained data information to screen out the useful medical records of patients; Construct a data matrix: Take each institution including a hospital as a distributed node, each patient as a sample, and each medical record of a patient as a feature, and construct a data matrix with samples as rows and features as columns, and the feature sets of each node do not overlap; Represent the global data matrix as where S and J are the total number of samples and features respectively, Set node n to have a local data matrix Set the data that cannot be observed to zero, and then the global data matrix is represented as 3. A global gradient double-tracking distributed method based on heterogeneous mixed data according to claim 2, characterized in that: The steps of constructing the network model specifically include the following: Model a multi-node network as an undirected graph where ε represents the set of edges connecting nodes, represents the set of nodes. If (m, n) ∈ ε, then nodes m and n can exchange information; Construct an interaction matrix: Set the interaction matrix as where, for any (m, n) ∈ ε, W n,m > 0, otherwise W n,m = 0, and the interaction matrix W satisfies the double-stochastic condition where 1 represents a vector of all 1s, λ w represents the second largest eigenvalue of W, and 4. A global gradient double-tracking distributed method based on heterogeneous mixed data according to claim 2, characterized in that: The construction of the non-uniform HBL problem specifically includes the following: Set the convergence and consistency problem of the multi-node network as: Among them, x n represents a locally optimized variable coupled with data, and θ n represents a general optimization variable, represents the set of all neighbor nodes of node n.
5. A global gradient double-tracking distributed method based on heterogeneous mixed data according to claim 4, characterized in that: The introduction of two tracking variables to estimate the parameters x and θ of the loss function f specifically includes the following: The pre-determined gradient forms of the selected loss function f with respect to x and θ are respectively and Each node n introduces a global linear combination tracking variable z n,i , i ∈ [S], where i ∈ [S] means i = 1,..., S, and estimates the global linear combination B i x n ; Based on the tracking variable z n,i , i ∈ [S], each node n continues to introduce a global gradient tracking variable u n , and estimates the global gradient 6. A global gradient double-tracking distributed method based on heterogeneous mixed data according to claim 5, characterized in that: The step of updating each node locally Q times respectively to complete the selection of model parameters specifically includes the following: Set the initial iteration to the 0th round, and randomly initialize the variables of each node as and The initial tracking variable of each node n is Calculate the derivative Then the initial tracking variable Set each node to update Q times locally and then interact with its neighbor nodes to exchange the following model parameters: the parameter x coupled with data n , and other network parameters θ n , the global linear combination tracking variable z n , the global gradient tracking variable u n , to complete one iteration. A total of T loop iterations are performed until the optimization variables and tracking variables of each network node converge to be consistent.
7. A global gradient double-tracking distributed method based on heterogeneous mixed data according to claim 6, characterized in that: The information interaction between distributed nodes specifically includes the following: At the r-th iteration, where r = 1, …, T, node n sends to its neighbor nodes After node n receives the variables sent by all its neighboring nodes m with weights, it is weighted again to be 8. A global gradient double-tracking distributed method based on heterogeneous mixed data according to claim 1, characterized in that: The local model update specifically includes the following: A1. Randomly select a batch of sample sets denoted as where the number of samples is q = 1, …, Q; A2. When q = 1: A21. Calculate the derivative of the loss function f with respect to the ordinary optimization variable θ where represents any sample in the sample set randomly selected by node n in the r-th round in; A22. Compute the derivative of the loss function f with respect to the local optimization variable x that is coupled with the data A23. Based on the information received from neighbor nodes, perform stochastic gradient descent on the general optimization variable θ and the local optimization variable x coupled with data with step sizes α and β respectively to obtain and A24. Update the global linear combination tracking variable by adding the difference between the current and the previous round of data variables multiplied by the total number of nodes, based on the information received from neighboring nodes A25. Calculate the current derivative of the loss function f with respect to the local optimization variable x coupled with the data where denotes any sample in the sample set randomly selected by node n in the r-th round in; A26. Update the global gradient tracking variable of the local optimization variable parameter x coupled with the data for the loss function f That is, based on the information obtained from neighbors, multiply the difference between the current and previous round derivatives by N to estimate the true gradient; A3. When q = 2, …, Q: A31. Calculate the derivative of the loss function f with respect to the ordinary optimization variable θ where represents a sample randomly selected from the sample set of node n in the r-th round any sample in; A32. Compute the derivative of the loss function f with respect to the local optimization variable x coupled with the data Based on the local updates in the previous (q - 1) steps, perform stochastic gradient descent on the ordinary optimization variable θ and the locally optimized variable x coupled with the data with step sizes α and β respectively to obtain and A34. Based on the local update of (q - 1) steps, update the linear combination tracking variable by adding the difference between the current and the previous round of data variables multiplied by the total number of nodes A35. Calculate the current derivative of the loss function f with respect to the local optimization variable x coupled with the data where denotes any sample in the sample set randomly selected by node n in the r-th round in; A36. Tracking variable for updating the gradient of the loss function f with respect to the local optimization variable x coupled with the data That is, based on the local update of (q - 1) steps, the difference between the current and the previous round of derivatives is multiplied by N times to estimate the true gradient; After locally updating for Q steps, update all variables: the parameters coupled with data Other network parameters Global linear combination tracking variables Global gradient tracking variables And the selected batch sample set
Citation Information
Patent Citations
A Multi-Step Strategy with Stochastic Averaging Gradient for Distributed Optimization
AU2020100182A4
Big data binary classification distributed optimization method based on stochastic gradient tracking technology
CN111950611A