Decentralized federated learning processing method and device, terminal equipment and product

By selecting the optimal mask data for category prediction in a distributed system, the problems of model bias and single point failure of the central server in traditional federated learning are solved, and category prediction with high accuracy and robustness is achieved.

CN120671775APending Publication Date: 2025-09-19HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510660082.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When traditional federated learning faces non-independent and identically distributed client data, model bias seriously affects data prediction results, and the central server is prone to single point failures that cause system paralysis, resulting in low robustness.

Method used

Based on the decentralized computing node, the original mask data of each client is obtained, the optimal mask data is selected for category prediction, and the dot product operation is performed between the optimal mask data and the initialized model to obtain the target category probability distribution, and the category corresponding to the maximum probability is determined as the target category.

Benefits of technology

The accuracy of category prediction is improved in the case of data heterogeneity, the robustness of the system is enhanced, and the risk of single point failure of the central node is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671775A_ABST
    Figure CN120671775A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of machine learning, and provides a decentralized federated learning processing method and device, terminal equipment and a product, and the method comprises the steps: obtaining the original mask data of each client, the original mask data of each client being obtained by training an initialization model based on the local data of each client, the local data of each client presents a non-independent identical distribution characteristic; selecting optimal mask data from the original mask data of each client; and performing category prediction processing on the to-be-predicted data based on the optimal mask data to obtain a target category of the to-be-predicted data. According to the method, for any client of a distributed system, the heterogeneous characteristics of local data of each client can be fully mined by obtaining original mask data obtained by each client based on local data training on the basis of not needing a central computing node, so that the prediction target data category has relatively high accuracy, and meanwhile, the prediction efficiency is improved. And the robustness of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of machine learning technology, and in particular relates to a processing method, apparatus, terminal equipment and product for decentralized federated learning. Background Art

[0002] Against the backdrop of the rapid development of artificial intelligence technology, Federated Learning (FL), as a distributed machine learning paradigm, achieves collaborative training without the need for centralized datasets by aggregating local model updates from multiple clients.

[0003] However, when client data exhibits non-independent and identically distributed (non-IID) characteristics (also known as heterogeneity), traditional federated learning approaches face the problem of model bias caused by client drift, which seriously affects data prediction results. Furthermore, in traditional federated learning approaches, the central server, as the sole aggregation node, faces the risk of a single point of failure. Any hardware or network failure can cause the system to crash, making the system less robust. Summary of the Invention

[0004] The embodiments of the present application provide a processing method, apparatus, terminal device, and product for decentralized federated learning. For any client in a distributed system, the heterogeneous characteristics of local data of each client can be fully exploited without the need for a central computing node, so that the target category obtained by category prediction processing has a high accuracy rate while improving the robustness of the system.

[0005] In a first aspect, an embodiment of the present application provides a decentralized federated learning processing method, which is applied to any client in a distributed system, and the method includes:

[0006] In response to the processing instruction for the prediction data, obtaining original mask data of each client, wherein the original mask data of each client is obtained by training the initialization model based on the local data of each client, the local data of each client exhibits non-independent and identically distributed characteristics, and the original mask data and the initialization model are the same type of tensors;

[0007] Selecting optimal mask data from the original mask data of each client, where the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the data to be predicted;

[0008] The category prediction process of the data to be predicted is performed based on the optimal mask data to obtain the target category of the data to be predicted.

[0009] In a possible implementation of the first aspect, selecting optimal mask data from original mask data of each client includes:

[0010] Perform weighted summation on the original mask data of each client to obtain the global mask data;

[0011] According to the global mask data and the initialization model, the optimal mask data is selected from the original mask data of each client.

[0012] In a possible implementation of the first aspect, selecting optimal mask data from original mask data of each client according to global mask data and an initialization model includes:

[0013] Perform a dot product operation on the global mask data and the initialized model to obtain the first dot product data;

[0014] According to the output entropy minimization strategy and the first dot product data, the optimal mask data is selected from the original mask data of each client.

[0015] In a possible implementation of the first aspect, selecting optimal mask data from original mask data of each client according to the output entropy minimization strategy and the first dot product data includes:

[0016] Based on the first dot product data, the category prediction processing of the data to be predicted is performed to obtain the initial category probability distribution of the data to be predicted Among them, f is the category prediction function, X is the data to be predicted, W is the initialization model, Mi represents the original mask data of client Ci, the subscript i represents the identity number of each client, and a i represents the weight value corresponding to the client Ci, N is the total number of clients, and ⊙ is the dot product operation;

[0017] According to the output entropy minimization strategy and the initial category probability distribution, the optimal mask data is selected from the original mask data of each client, where the output entropy minimization strategy is Where H is the entropy of the initial category probability distribution, and j is the identity number of the client corresponding to the optimal mask data.

[0018] In a possible implementation of the first aspect, performing category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category of the data to be predicted includes:

[0019] Based on the optimal mask data, the category prediction processing of the data to be predicted is performed to obtain the target category probability distribution of the data to be predicted;

[0020] The category corresponding to the maximum probability in the target category probability distribution is determined as the target category of the data to be predicted.

[0021] In a possible implementation of the first aspect, performing category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category probability distribution of the data to be predicted includes:

[0022] Perform a dot product operation on the optimal mask data and the initialized model to obtain the second dot product data;

[0023] The data to be predicted is subjected to category prediction processing based on the second dot product data to obtain a target category probability distribution of the data to be predicted.

[0024] In a possible implementation manner of the first aspect, the original mask data is obtained by each client using an edge-popup (EP) algorithm.

[0025] In a second aspect, an embodiment of the present application provides a decentralized federated learning processing device, which is configured on any client in a distributed system, and includes:

[0026] an acquisition module, configured to obtain original mask data of each client in response to a processing instruction for the prediction data, wherein the original mask data of each client is obtained by training an initialization model based on the local data of each client, the local data of each client exhibits a non-independent and identically distributed characteristic, and the original mask data and the initialization model are tensors of the same type;

[0027] A selection module is used to select the optimal mask data from the original mask data of each client, wherein the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the data to be predicted;

[0028] The category prediction module is used to perform category prediction processing on the data to be predicted based on the optimal mask data to obtain the target category of the data to be predicted.

[0029] In a third aspect, an embodiment of the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the methods of the first aspect when executing the computer program.

[0030] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method of any one of the first aspects.

[0031] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute any one of the methods in the first aspect above.

[0032] The embodiments of the present application provide a processing method, apparatus, terminal device and product for decentralized federated learning, which is applied to any client in a distributed system, including: in response to a processing instruction for the data to be predicted, obtaining the original mask data of each client, wherein the original mask data of each client is obtained by each client training an initialization model based on their own local data, the local data of each client presents non-independent and identically distributed characteristics, and the original mask data and the initialization model are the same type of tensors; selecting the optimal mask data from the original mask data of each client, wherein the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the data to be predicted; performing category prediction processing on the data to be predicted based on the optimal mask data to obtain the target category of the data to be predicted. By utilizing the above technical solution, for any client in the distributed system, it is possible to directly obtain the original mask data obtained by each client based on their own local data training without the need for a central computing node, by fully exploiting the heterogeneous characteristics of the local data of each client, so that the target category obtained by the category prediction processing has a higher accuracy while improving the robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0034] Figure 1 This is a flowchart of a decentralized federated learning processing method provided in one embodiment of the present application;

[0035] Figure 2 This is a flowchart of a decentralized federated learning processing method provided by another embodiment of the present application;

[0036] Figure 3 This is a schematic diagram of a process for obtaining original mask data provided by an embodiment of the present application;

[0037] Figure 4 This is a flow chart of a category prediction process provided by an embodiment of the present application;

[0038] Figure 5 This is a structural block diagram of a decentralized federated learning processing device provided in one embodiment of the present application;

[0039] Figure 6 This is a structural diagram of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0040] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0041] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0042] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0043] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0044] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0045] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0046] Client drift can be considered to refer to the phenomenon in which client models deviate from the global optimal solution due to adaptation to local data distribution in heterogeneous data scenarios. For example, when client data has label skew (such as the uneven distribution of digits in the MNIST dataset) or feature skew (such as the color distribution differences in CIFAR-10 images), the federated averaging algorithm based on simple averaging can fall into a local optimum due to inconsistent model update directions.

[0047] To alleviate this problem, researchers have proposed methods such as proximal term regularization and variance reduction to suppress client drift by limiting the amplitude of local updates. However, these methods essentially treat data heterogeneity as noise and fail to fully tap its potential value.

[0048] In addition, in federated learning, the centralized architecture has significant technical flaws. For example, the central server, as the only aggregation node, faces the risk of single point failure. Any hardware or network failure may cause the system to crash. The central server bears all computing tasks and can easily become a performance bottleneck when deployed on a large scale.

[0049] In summary, how to design a decentralized federated learning method that can effectively handle heterogeneous data is an urgent problem that needs to be solved.

[0050] Based on this, the embodiments of this application provide a decentralized federated learning processing method that can fully exploit the heterogeneous characteristics of local data on each client in the case of data heterogeneity, and has the significant advantage of high accuracy in prediction, providing an innovative and effective solution for federated learning in heterogeneous data environments. Furthermore, the embodiments of this application do not require a central node, which provides good robustness.

[0051] Figure 1 This is a flowchart of a decentralized federated learning processing method provided in one embodiment of the present application. As an example and not a limitation, this method can be applied to any client in a distributed system, such as Figure 1 As shown, the method includes:

[0052] S101 : Responding to a processing instruction for data to be predicted, obtaining original mask data of each client.

[0053] The original mask data for each client is obtained by training the initialization model based on its own local data. The local data of each client is not independent and identically distributed, and the original mask data and the initialization model are the same type of tensors. The data to be predicted can be considered as the data for which category prediction is required.

[0054] In this embodiment, the initialization model can be understood as each client using the same initialization method to obtain the same model, and then each client can be trained on its own local data based on the initialization model to obtain its own super mask (i.e., original mask data). The super mask can be composed of the numbers 0 and 1, and the shape is consistent with the initialization model, that is, the original mask data and the initialization model can be considered to be isotype tensors. Specifically, it can be understood that the dimensions between the original mask data and the initialization model are the same, and the size of each dimension (i.e., the number of elements) is the same. For example, when the initialization model is a two-dimensional 2*3 matrix, the original mask data is also a two-dimensional 2*3 matrix. The difference is that the values ​​of the elements in the matrix are different. Among them, the specific training means are not limited, and each client can obtain its own original mask data based on the same or different algorithms.

[0055] Optionally, the original mask data is obtained by each client using an edge-popup (EP) algorithm.

[0056] The EP algorithm can be an algorithm for random sparse neural networks. It learns subnetwork masks through backpropagation, significantly improving the usability of random sparse networks. More specifically, the EP algorithm can be an optimization method for finding supermasks in large, randomly initialized neural networks (called supernetworks), with performance approaching that of fully trained supernetworks. The EP algorithm does not train the network's weights (θw), but instead only determines the set of edges to retain and removes (pops) the remaining edges. Specifically, the EP algorithm assigns a positive score to each edge (θs) in the supernetwork. In the forward pass, the top k percent of edges with the highest scores are selected, where k is the percentage of the total number of edges in the supernetwork that will be retained in the final subnetwork. In the backward pass, the scores are updated using a straight-through gradient estimator. This improves the network's representational capabilities, enabling the network to effectively represent data. It also gradually reduces the number of unique parameter values ​​in the network, improving network compression and reducing the storage and transmission overhead of the neural network.

[0057] Furthermore, each client can exchange the original mask data obtained through their own training through the communication link. Taking any client in the distributed system as an example, the client can train the initialization model based on its own local data to obtain its own original mask data, and can also obtain the original mask data of other clients through communication interaction, so that the original mask data of all clients can be obtained.

[0058] S102: Select optimal mask data from the original mask data of each client.

[0059] Among them, the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the data to be predicted.

[0060] After obtaining the original mask data from each client, the original mask data with the smallest entropy of the probability distribution for class prediction of the data to be predicted can be selected as the optimal mask data. This embodiment does not limit the specific process of selecting the optimal mask data. For example, based on a preset model, the original mask data of each client can be input into the preset model to directly output the corresponding optimal mask data. The preset model can be a pre-trained neural network model that can be used to select the optimal mask data. Alternatively, the corresponding optimal mask data can be selected by performing a series of calculations on each original mask data. The specific calculation process is not further elaborated here.

[0061] S103 : performing category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category of the data to be predicted.

[0062] After selecting the optimal mask data through the above steps, the optimal mask data can be used to perform category prediction on the data to be predicted. For example, a decision tree model can be constructed to achieve category prediction for the data to be predicted. Alternatively, a federated support vector machine can be integrated to achieve accurate category prediction for the data to be predicted while protecting the data privacy of all parties. Alternatively, the data to be predicted can be first subjected to category prediction based on the optimal mask data to obtain a target category probability distribution for the data to be predicted. The category corresponding to the maximum probability in the target category probability distribution is then determined as the target category for the data to be predicted. The target category probability distribution can be used to characterize the probability of the data to be predicted belonging to each category.

[0063] In some embodiments, performing category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category probability distribution of the data to be predicted includes:

[0064] Perform a dot product operation on the optimal mask data and the initialized model to obtain the second dot product data;

[0065] The data to be predicted is subjected to category prediction processing based on the second dot product data to obtain a target category probability distribution of the data to be predicted.

[0066] In a specific embodiment, a dot product operation can be performed between the optimal mask data M1 and the initialization model W to obtain second dot product data W⊙M1. Subsequently, the target category probability distribution of the data to be predicted can be obtained using a model prediction function. For example, the target category probability distribution of the data to be predicted X can be predicted using f(X, W⊙M1). f can be a model prediction function that outputs the probability distribution of the category prediction.

[0067] Furthermore, the final prediction result Y = arg max(f(X, W⊙M1)) is the target category of the data to be predicted.

[0068] This embodiment provides a decentralized federated learning processing method, which responds to a processing instruction for the data to be predicted, obtains the original mask data of each client, wherein the original mask data of each client is obtained by training an initialization model based on the local data of each client, and the local data of each client exhibits non-independent and identically distributed characteristics, and the original mask data and the initialization model are the same type of tensors; selects the optimal mask data from the original mask data of each client, wherein the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the data to be predicted; performs category prediction processing on the data to be predicted based on the optimal mask data to obtain the target category of the data to be predicted. Utilizing this method, for any client in a distributed system, it is possible to directly obtain the original mask data obtained by each client based on their own local data training without the need for a central computing node, by fully exploiting the heterogeneous characteristics of the local data of each client, so that the target category obtained by the category prediction processing has a high accuracy rate while improving the robustness of the system.

[0069] Figure 2 This is a flowchart of a decentralized federated learning processing method provided by another embodiment of the present application. This embodiment further optimizes the selection of optimal mask data from the original mask data of each client by performing weighted summation on the original mask data of each client to obtain global mask data; and selects the optimal mask data from the original mask data of each client based on the global mask data and the initialization model. Figure 2 As shown, the method includes:

[0070] S201 : Responding to a processing instruction for data to be predicted, obtaining original mask data of each client.

[0071] S202: Perform weighted summation on the original mask data of each client to obtain global mask data.

[0072] S203 : Selecting optimal mask data from the original mask data of each client according to the global mask data and the initialization model.

[0073] In the specific calculation, Mi can represent the original mask data obtained by each client Ci through training and learning, and the subscript i represents the client's identity number. First, the global mask can be generated by weighted summation of the mask At this time, the weight value ai=1 / N, where N is the total number of clients.

[0074] Then, based on the generated global mask data and the initialization model, the optimal mask data can be selected from the original mask data of each client. For example, a neural network model can be used to achieve the selection of the optimal mask data, or further calculations can be performed on the global mask data and the initialization model to select the optimal mask data. For example, a dot product operation can be performed on the global mask data and the initialization model to obtain the first dot product data; then, based on the output entropy minimization strategy and the first dot product data, the optimal mask data can be selected from the original mask data of each client. The output entropy minimization strategy can be used to select the optimal mask data that minimizes the entropy of the category probability distribution for category prediction of the data to be predicted.

[0075] As an example, according to the output entropy minimization strategy and the first dot product data, the optimal mask data is selected from the original mask data of each client, including:

[0076] Based on the first dot product data, the category prediction processing of the data to be predicted is performed to obtain the initial category probability distribution of the data to be predicted Among them, f is the category prediction function, X is the data to be predicted, W is the initialization model, Mi represents the original mask data of client Ci, the subscript i represents the identity number of each client, and a i represents the weight value corresponding to the client Ci, N is the total number of clients, and ⊙ is the dot product operation;

[0077] According to the output entropy minimization strategy and the initial category probability distribution, the optimal mask data is selected from the original mask data of each client, where the output entropy minimization strategy is Where H is the entropy of the initial category probability distribution, and j is the identity number of the client corresponding to the optimal mask data.

[0078] Assume that there are three clients in a distributed system. Each client i has local data Di, and the local data of the clients exhibit non-IID characteristics, that is, D1, D2, and D3 are heterogeneous data.

[0079] Taking the current client 1 as the perspective, the processing method of decentralized federated learning provided in the embodiment of the present application is exemplarily described.

[0080] Figure 3 This is a flow chart of obtaining original mask data provided by an embodiment of the present application. Figure 3As shown, first, each client (such as client 1, 2, and 3) can use the same initialization method to obtain the same initialization model W, and then each client can use the EP algorithm on its own local data based on the initialization model W to obtain a super mask, and exchange its own super mask with other clients. For example, client 1 can use the EP algorithm on local data D1 to obtain super mask M1, client 2 can use the EP algorithm on local data D2 to obtain super mask M2, and client 3 can use the EP algorithm on local data D3 to obtain super mask M3. Client 1 can obtain super mask M2 of client 2 and super mask M3 of client 3 through communication interaction.

[0081] Figure 4 This is a flow chart of a category prediction process provided by an embodiment of the present application. Figure 4 As shown, client 1 can generate a global mask by performing mask weighted summation on super masks M1, M2 and M3, and perform a dot product operation with the initialized model W to obtain Client 1 can then predict the function based on the model Get "probability distribution 1 of prediction results"; based on the output minimization entropy strategy It can be seen that by increasing a1, the output entropy can be reduced, that is, by increasing a1 and reducing a2 and a3, the "probability distribution 2 of the prediction result" can be obtained, thereby selecting the most appropriate super mask M1 as the optimal mask data.

[0082] Finally, when client 1 uses super mask M1 for category prediction processing, a1=1, a2=a3=0 can be set. By performing a dot product operation on the super mask M1 and the initialized model W, the model W⊙M1 is obtained. The "probability distribution 3 of the prediction result" (that is, the target category probability distribution) is predicted by f(X, W⊙M1) to obtain the probability distribution of the prediction result, and the final prediction category result Y=arg max(f(X, W⊙M1)) is determined as the target category of the data X to be predicted.

[0083] S204 : performing category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category of the data to be predicted.

[0084] This embodiment provides a decentralized federated learning processing method, which obtains global mask data by weighted summation of the original mask data of each client, and then selects the optimal mask data based on the global mask data and the initialization model, providing an accurate data basis for subsequent category prediction processing, further improving the accuracy of the target category.

[0085] Corresponding to the decentralized federated learning processing method in the above embodiment, Figure 5 This is a structural block diagram of a decentralized federated learning processing device provided in one embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0086] Reference Figure 5 , the device comprises:

[0087] An acquisition module 301 is configured to obtain original mask data of each client in response to a processing instruction for the prediction data, wherein the original mask data of each client is obtained by training an initialization model based on the local data of each client, the local data of each client exhibits a non-independent and identically distributed characteristic, and the original mask data and the initialization model are tensors of the same type;

[0088] A selection module 302 is configured to select optimal mask data from the original mask data of each client, wherein the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the data to be predicted;

[0089] The category prediction module 303 is configured to perform category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category of the data to be predicted.

[0090] This embodiment provides a decentralized federated learning processing device, which obtains the original mask data of each client in response to the processing instruction of the predicted data through an acquisition module, wherein the original mask data of each client is obtained by training the initialization model based on the local data of each client, and the local data of each client exhibits non-independent and identically distributed characteristics, and the original mask data and the initialization model are the same type of tensors; the selection module selects the optimal mask data from the original mask data of each client, wherein the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the predicted data; the category prediction module performs category prediction processing on the predicted data based on the optimal mask data to obtain the target category of the predicted data. Using this device, for any client in the distributed system, it is possible to directly obtain the original mask data obtained by each client based on the local data training without the need for a central computing node, by fully exploiting the heterogeneous characteristics of the local data of each client, so that the target category obtained by the category prediction processing has a high accuracy rate while improving the robustness of the system.

[0091] Optionally, the selection module includes:

[0092] A weighting unit, configured to perform weighted summation on the original mask data of each client to obtain global mask data;

[0093] The selection unit is used to select the optimal mask data from the original mask data of each client according to the global mask data and the initialization model.

[0094] Optionally, the selection unit includes:

[0095] A dot product subunit, configured to perform a dot product operation on the global mask data and the initialized model to obtain first dot product data;

[0096] The selection subunit is used to select optimal mask data from the original mask data of each client according to the output entropy minimization strategy and the first dot product data.

[0097] Optionally, the subunit is selected to be specifically used for:

[0098] Based on the first dot product data, the category prediction processing of the data to be predicted is performed to obtain the initial category probability distribution of the data to be predicted Among them, f is the category prediction function, X is the data to be predicted, W is the initialization model, Mi represents the original mask data of client Ci, the subscript i represents the identity number of each client, and a i represents the weight value corresponding to the client Ci, N is the total number of clients, and ⊙ is the dot product operation;

[0099] According to the output entropy minimization strategy and the initial category probability distribution, the optimal mask data is selected from the original mask data of each client, where the output entropy minimization strategy is Where H is the entropy of the initial category probability distribution, and j is the identity number of the client corresponding to the optimal mask data.

[0100] Optionally, the category prediction module includes:

[0101] A category prediction unit is used to perform category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category probability distribution of the data to be predicted;

[0102] The determination unit is used to determine the category corresponding to the maximum probability in the target category probability distribution as the target category of the data to be predicted.

[0103] Optionally, the category prediction unit is specifically configured to:

[0104] Perform a dot product operation on the optimal mask data and the initialized model to obtain the second dot product data;

[0105] The data to be predicted is subjected to category prediction processing based on the second dot product data to obtain a target category probability distribution of the data to be predicted.

[0106] Optionally, the original mask data is obtained by each client using an edge-popup (EP) algorithm.

[0107] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0108] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0109] The embodiment of the present application also provides a terminal device, Figure 6 This is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Figure 6 As shown, the terminal device includes: at least one processor 401, a memory 402, an input device 403, an output device 404, and a computer program stored in the memory 402 and executable on at least one processor 401. When the processor 401 executes the computer program, the steps in any of the above-mentioned method embodiments are implemented.

[0110] The input device 403 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the terminal device. The output device 404 may include a display device such as a display screen.

[0111] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor 401, the steps in the above-mentioned method embodiments can be implemented.

[0112] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 401, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include at least: any entity or device that can carry the computer program code to the device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, a computer-readable storage medium cannot be an electric carrier signal or a telecommunication signal.

[0114] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0115] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0116] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal devices and methods can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0117] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0118] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A decentralized federated learning processing method, characterized in that: Applied to any client in a distributed system, the method includes: In response to a processing instruction for the prediction data, obtaining original mask data of each client, wherein the original mask data of each client is obtained by training an initialization model based on the local data of each client, the local data of each client exhibits a non-independent and identically distributed characteristic, and the original mask data and the initialization model are tensors of the same type; Selecting optimal mask data from the original mask data of each of the clients, wherein the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the data to be predicted; A category prediction process is performed on the data to be predicted based on the optimal mask data to obtain a target category of the data to be predicted.

2. The decentralized federated learning processing method according to claim 1, characterized in that: The selecting the optimal mask data from the original mask data of each client includes: Performing weighted summation on the original mask data of each client to obtain global mask data; According to the global mask data and the initialization model, optimal mask data is selected from the original mask data of each client.

3. The decentralized federated learning processing method according to claim 2, characterized in that: The selecting optimal mask data from the original mask data of each client according to the global mask data and the initialization model includes: Performing a dot product operation on the global mask data and the initialization model to obtain first dot product data; According to the output entropy minimization strategy and the first dot product data, optimal mask data is selected from the original mask data of each of the clients.

4. The decentralized federated learning processing method according to claim 3, characterized in that: The selecting optimal mask data from the original mask data of each of the clients according to the output entropy minimization strategy and the first dot product data includes: Based on the first dot product data, the category prediction processing of the data to be predicted is performed to obtain the initial category probability distribution of the data to be predicted Wherein, f is the category prediction function, X is the data to be predicted, W is the initialization model, Mi represents the original mask data of the client Ci, the subscript i represents the identity number of each client, and a i represents the weight value corresponding to the client Ci, N is the total number of clients, and ⊙ is the dot product operation; According to the output entropy minimization strategy and the initial category probability distribution, the optimal mask data is selected from the original mask data of each client, wherein the output entropy minimization strategy is Wherein, H is the entropy of the initial category probability distribution, and j is the identity number of the client corresponding to the optimal mask data.

5. The decentralized federated learning processing method according to claim 1, characterized in that: The performing category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category of the data to be predicted includes: performing category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category probability distribution of the data to be predicted; The category corresponding to the maximum probability in the target category probability distribution is determined as the target category of the data to be predicted.

6. The decentralized federated learning processing method according to claim 5, characterized in that: The performing category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category probability distribution of the data to be predicted includes: Performing a dot product operation on the optimal mask data and the initialization model to obtain second dot product data; A category prediction process is performed on the data to be predicted based on the second dot product data to obtain a target category probability distribution of the data to be predicted.

7. The decentralized federated learning processing method according to any one of claims 1 to 6, characterized in that: The original mask data is obtained by each of the clients using an edge-popup (EP) algorithm.

8. A decentralized federated learning processing device, characterized in that: The device is configured on any client in a distributed system, and includes: an acquisition module, configured to acquire original mask data of each client in response to a processing instruction for the prediction data, wherein the original mask data of each client is obtained by training an initialization model based on the local data of each client, the local data of each client exhibits a non-independent and identically distributed characteristic, and the original mask data and the initialization model are tensors of the same type; A selection module, configured to select optimal mask data from the original mask data of each client, wherein the optimal mask data has the smallest entropy of the category probability distribution for category prediction of the data to be predicted; The category prediction module is used to perform category prediction processing on the data to be predicted based on the optimal mask data to obtain a target category of the data to be predicted.

9. A terminal device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the terminal device implements the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that When the computer program product is run on a terminal device, the terminal device is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Privacy protection decentralized federated learning method based on directed cooperation and adaptive aggregation

    CN121745224A