Method, device and system for determining contribution degree of participation end in federated learning

By evaluating the contribution in federated learning using the distilled sample set and uncertainty at the participant end, the problems of high computational complexity and low efficiency in the prior art are solved, and efficient and accurate contribution calculations are achieved.

CN119990357APending Publication Date: 2025-05-13HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311512181.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In existing federated learning techniques, methods to evaluate participant contribution require an additional test sample set, which increases the cost of data collection and processing, and the calculation complexity of Shapley values ​​leads to low computational efficiency of contribution.

Method used

By obtaining the distilled sample set at the participant end, the contribution of each participant end is evaluated based on the uncertainty of the distilled sample. The specific steps include: pre-training the distilled sample set at each participating end, determining its uncertainty, and calculating the contribution degree based on the uncertainty.

Benefits of technology

On the basis of accurately evaluating the contribution of participants, the computational complexity is reduced, the computational efficiency of contribution is improved, and it can efficiently handle large-scale federated learning tasks, with good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990357A_ABST
    Figure CN119990357A_ABST
Patent Text Reader

Abstract

The invention discloses a contribution degree determination method, device and system of a participation end in federated learning, and relates to the technical field of federated learning. According to the method, on the basis of a distillation sample set provided by a plurality of participation ends in federated learning, the information amount brought by distillation samples for the training process of a machine learning model is reflected through the uncertainty of the distillation samples in the distillation sample set; therefore, the contribution degree of each participation end is determined according to the uncertainty of the distillation sample provided by each participation end, the calculation complexity is effectively reduced on the basis of accurately evaluating the contribution degree of each participation end, and the contribution degree calculation efficiency is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of federated learning, and in particular to a method, device and system for determining the contribution of a participant in federated learning. Background Art

[0002] With the widespread popularity of IoT devices, the application scenarios of federated learning are constantly expanding. In federated learning, participants often need to upload local model parameters to the central server for parameter aggregation. The central server evaluates the contribution of each participant to federated learning in order to better coordinate cooperation and resource allocation between participants.

[0003] In related technologies, the contribution of each participant to federated learning is usually evaluated based on the Shapley value. For example, the participant uses a local sample set to train a local model, and uploads the model parameters of the local model to the central server to aggregate the model parameters. The central server uses the Shapley value calculation method to evaluate the contribution of each participant to federated learning based on the test sample set.

[0004] However, the above method requires the introduction of additional test sample sets to evaluate the contribution of the participating terminals, which increases the cost of data collection and processing. In addition, the calculation complexity of the Shapley value is high, resulting in low calculation efficiency of the contribution. Summary of the invention

[0005] The embodiment of the present application provides a method, device and system for determining the contribution of a participant in federated learning, which can effectively reduce the computational complexity and improve the efficiency of contribution calculation based on accurate evaluation of the contribution of the participant. The technical solution is as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for determining the contribution of a participant in federated learning, which is applied to a central server in a federated learning system, wherein the system further includes a plurality of participants, and the method includes:

[0007] Acquire distilled sample sets of the multiple participating terminals, where the distilled sample set of each participating terminal includes at least one distilled sample obtained after the corresponding participating terminal performs data set distillation on the local sample set;

[0008] Pre-training a machine learning model based on at least one first distilled sample in the distilled sample set of each participating end, and determining the uncertainty of at least one second distilled sample in the distilled sample set of each participating end based on the pre-trained machine learning model, wherein the uncertainty indicates the amount of information brought by the corresponding second distilled sample to the training process of the machine learning model;

[0009] Based on the uncertainty of at least one second distillation sample in the distillation sample set of each participating end, the contribution of each participating end is determined, and the contribution of each participating end is sent to the corresponding participating end, and the contribution is used to instruct the corresponding participating end to adjust the distillation sample set of the corresponding participating end.

[0010] In the above method, based on the distilled sample set provided by multiple participants in federated learning, the uncertainty of the distilled samples in the distilled sample set is used to reflect the amount of information brought by the distilled samples to the training process of the machine learning model, so as to determine the contribution of each participant according to the uncertainty of the distilled samples provided by each participant. On the basis of accurately evaluating the contribution of the participating terminals, the computational complexity is effectively reduced, thereby improving the efficiency of contribution calculation. Moreover, the method provided in this application can efficiently handle large-scale federated learning tasks and has good scalability.

[0011] In some embodiments, the pre-training of the machine learning model based on at least one first distilled sample in the distilled sample set of each participating terminal, and determining the uncertainty of at least one second distilled sample in the distilled sample set of each participating terminal based on the pre-trained machine learning model, includes:

[0012] Determine at least one first distillation sample and at least one second distillation sample corresponding to each participating end from the distillation sample set of each participating end;

[0013] Aggregating at least one first distilled sample corresponding to each participating end to obtain a first aggregated sample set;

[0014] Pre-training the machine learning model based on the first aggregated sample set;

[0015] Based on the pre-trained machine learning model, the uncertainty of at least one second distillation sample corresponding to each participating end is determined.

[0016] In some embodiments, determining the uncertainty of at least one second distilled sample corresponding to each participating terminal based on the pre-trained machine learning model includes:

[0017] Inputting at least one second distilled sample corresponding to a first participating end into the pre-trained machine learning model to obtain a prediction result corresponding to each second distilled sample, wherein the first participating end refers to any participating end;

[0018] The uncertainty of each second distillation sample is determined based on the prediction result corresponding to each second distillation sample.

[0019] In the above method, by dividing the distilled sample set of each participating end into two parts, namely the first distilled sample and the second distilled sample, it is equivalent to taking one part of the distilled samples as training samples and the other part of the distilled samples as test samples. Therefore, after training the basic model according to the first distilled sample, the trained basic model is used to evaluate the labeling benefit of the second distilled sample. The uncertainty of the second distilled sample is used to reflect its potential contribution to federated learning, and then the contribution of the participating end is determined. The calculation complexity is low and the calculation efficiency of the contribution is high.

[0020] In some embodiments, determining the contribution of each participating end based on the uncertainty of at least one second distilled sample in the distilled sample set of each participating end includes:

[0021] Based on at least one of an average value and a variance of uncertainties of a plurality of second distillation samples corresponding to the first participating end, a contribution of the first participating end is determined, where the first participating end refers to any one participating end.

[0022] In some embodiments, the method further comprises:

[0023] The distilled sample sets of the multiple participating terminals are aggregated to obtain a target aggregated sample set, where the target aggregated sample set is used by the multiple participating terminals to locally train the machine learning model.

[0024] Through the above method, any participating terminal can obtain the global distilled samples, so that the model training can be realized locally without uploading the model parameters of the trained model to the central server. Moreover, by performing data set distillation on the local sample set, that is, compressing and distilling the low-rank high-dimensional data into a small amount of data, the visible part of the local sample set is difficult to be understood by humans, and is stored in a small amount of distilled samples in a high-density manner, thereby achieving privacy protection from the training input.

[0025] In some embodiments, the method further comprises:

[0026] A distilled sample set of at least one second participating end is sent to a first participating end, and the distilled sample set of at least one second participating end is used for the first participating end to locally train the machine learning model, the first participating end refers to any participating end, and the second participating end refers to a participating end among the multiple participating ends except the first participating end.

[0027] Through the above method, the sample aggregation task originally performed by the central server is offloaded to each participating terminal, which can save the computing resources of the central server. Moreover, for any participating terminal, the participating terminal can obtain the distilled samples of other participating terminals. In this way, model training can be implemented locally without uploading the model parameters of the trained model to the central server.

[0028] In a second aspect, an embodiment of the present application provides a method for determining the contribution of a participant in federated learning, which is applied to a participant in a federated learning system, wherein the system further includes a central server, and the method includes:

[0029] Sending the distilled sample set of the participating terminal to the central server, the distilled sample set including at least one distilled sample obtained by the participating terminal after performing data set distillation on the local sample set;

[0030] Acquire the contribution of the participating end, wherein the contribution is used to instruct the participating end to adjust the distilled sample set; wherein the method for determining the contribution includes: pre-training the machine learning model based on at least one first distilled sample in the distilled sample set of each participating end in the federated learning system, and determining the uncertainty of at least one second distilled sample in the distilled sample set of each participating end based on the pre-trained machine learning model, wherein the uncertainty indicates the amount of information brought by the corresponding second distilled sample to the training process of the machine learning model; determine the contribution of each participating end based on the uncertainty of at least one second distilled sample in the distilled sample set of each participating end.

[0031] In some embodiments, the method further includes: performing data set distillation on the local sample set of the participating end based on a random model obtained by randomly sampling the model parameters of the machine learning model to obtain a distilled sample set of the participating end.

[0032] In some embodiments, the method further comprises:

[0033] Obtaining a target aggregated sample set, where the target aggregated sample set is obtained by aggregating distilled sample sets of multiple participating terminals in the federated learning system;

[0034] The machine learning model is trained locally based on the target aggregate sample set.

[0035] In a third aspect, an embodiment of the present application provides a device for determining the contribution of a participating end in federated learning, the device comprising at least one functional unit for implementing a method for determining the contribution of a participating end in federated learning as involved in the aforementioned first aspect or any possible implementation method of the first aspect.

[0036] In a fourth aspect, an embodiment of the present application provides a device for determining the contribution of a participating end in federated learning, the device comprising at least one functional unit for implementing a method for determining the contribution of a participating end in federated learning as involved in the aforementioned second aspect or any possible implementation method of the second aspect.

[0037] In a fifth aspect, an embodiment of the present application provides a federated learning system, which includes multiple participating terminals and a central server of federated learning; the central server is used to execute the method for determining the contribution of the participating terminals in federated learning as involved in the first aspect or any possible implementation of the first aspect, and the participating terminals are used to execute the method for determining the contribution of the participating terminals in federated learning as involved in the second aspect or any possible implementation of the second aspect.

[0038] In a sixth aspect, an embodiment of the present application provides a computing device, comprising a processor and a memory, wherein the memory is used to store at least one piece of program code, and the at least one piece of program code is loaded by the processor and executes a method for determining the contribution of a participating terminal in federated learning as involved in the first aspect or any possible implementation of the first aspect, or executes a method for determining the contribution of a participating terminal in federated learning as involved in the second aspect or any possible implementation of the second aspect.

[0039] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium is used to store at least one program code, the at least one program code is used to execute the method for determining the contribution of a participating terminal in federated learning as described in the first aspect or any possible implementation of the first aspect, or execute the method for determining the contribution of a participating terminal in federated learning as described in the second aspect or any possible implementation of the second aspect. The storage medium includes but is not limited to a volatile memory, such as a random access memory, a non-volatile memory, such as a flash memory, a hard disk drive (HDD), and a solid state drive (SSD).

[0040] In an eighth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computing device, enables the computing device to execute the method for determining the contribution of a participating terminal in federated learning as described in the first aspect or any possible implementation of the first aspect, or to execute the method for determining the contribution of a participating terminal in federated learning as described in the second aspect or any possible implementation of the second aspect. The computer program product may be a software installation package, and when the method for determining the contribution of a participating terminal in federated learning needs to be implemented, the computer program product may be downloaded and executed on a computing device. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0042] Figure 2 is a schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application;

[0043] Figure 3 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0044] Figure 4 This is a schematic diagram of a connection method of a computing device cluster provided in an embodiment of the present application;

[0045] Figure 5 It is a schematic diagram of a method for determining the contribution of a participating terminal in federated learning provided in an embodiment of the present application;

[0046] Figure 6 is a schematic diagram of a federated learning system provided in an embodiment of the present application;

[0047] Figure 7 is a schematic diagram of a federated learning method provided in an embodiment of the present application;

[0048] Figure 8 It is a flowchart of a method for determining the contribution of a participating terminal in federated learning provided in an embodiment of the present application;

[0049] Fig. 9 It is a schematic diagram of a method for determining the contribution of a participating terminal in federated learning provided in an embodiment of the present application;

[0050] Fig.10 This is a comparison chart of experimental results provided in the embodiment of the present application;

[0051] Fig.11 It is a structural schematic diagram of a device for determining the contribution of a participating terminal in federated learning provided in an embodiment of the present application;

[0052] Fig.12 It is a structural diagram of another device for determining the contribution of a participating end in federated learning provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below in conjunction with the accompanying drawings. It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the machine learning models, local sample sets, etc. involved in this application are all obtained with full authorization.

[0054] For ease of understanding, the key terms and key concepts involved in this application are explained below.

[0055] Federated learning (FL) is a machine learning method of multi-party collaborative computing, which is mainly used to solve the data island problem faced by artificial intelligence (AI) algorithms when they are implemented in industry. Each participant in federated learning can train their own models with the help of data from other parties (such as model parameters), that is, each participant jointly trains the machine learning model by exchanging encrypted model parameters, so that data can participate in multi-party collaborative modeling without leaving the local area, thereby realizing data sharing. In an embodiment of the present application, each participant in federated learning can train their own models with the help of the distilled sample set of the other party, participate in multi-party collaborative modeling on the basis of realizing privacy protection, and realize data sharing.

[0056] Dataset distillation is a method of compressing a dataset. It forms a synthetic dataset including a small number of samples by distilling a dataset. It is usually used to reduce the size of the dataset to reduce the computational complexity of the machine learning model. Taking the dataset distillation for a sample image set as an example, the knowledge of the entire sample image set is encapsulated into a small number of synthetic images, such as synthesizing the sample images of each category into one image, and training models with similar performance on these synthetic images (its role is to make the synthetic images and sample images have a similar effect on the model to be trained in the end). These synthetic images are also called distilled images. In the embodiment of the present application, the samples after dataset distillation are called distilled samples.

[0057] Active learning is a machine learning method that actively selects the most valuable samples for annotation. Its purpose is to use as few high-quality sample annotations as possible to achieve the best possible performance of the model.

[0058] The following is an introduction to the application scenarios and implementation environment involved in this application.

[0059] The technical solution provided in the embodiments of the present application can be applied to scenarios where data is shared between multiple entities based on federated learning, such as Internet of Vehicles, smart homes, medical care, etc., and the present application does not limit this.

[0060] Figure 1 Schematic diagram of an implementation environment provided by the embodiment of the present application. Figure 1 As shown, the implementation environment includes a federated learning system 100, which includes multiple participating terminals 101 of federated learning and a central server 102. Multiple participating terminals 101 and the central server 102 are directly or indirectly connected through a wired network or a wireless network. Schematically, the participating terminal 101 can also be called a participating node, a participating device, and the central server 102 can also be called a central node, a central device, an aggregation node, etc., but the present application is not limited thereto.

[0061] For any participating terminal 101, the participating terminal 101 can be a terminal with model training capabilities (such as an Internet of Things device, a laptop computer, a vehicle-mounted terminal, etc.), or an independent physical server, or a server cluster or distributed system composed of multiple physical servers, etc., and the present application is not limited thereto. Each participating terminal 101 has a local sample set, such as training samples and test samples, and can perform data set distillation on the local sample set to obtain a corresponding distilled sample set, and send the distilled sample set to the central server 102 to participate in federated learning. In addition, the number of participating terminals 101 shown in the figure is for illustrative purposes only, and the number of participating terminals 101 can be more or less, and the present application is not limited thereto.

[0062] The central server 102 is used to determine the contribution of each participating terminal 101 based on the distilled sample sets provided by multiple participating terminals 101, so as to evaluate the potential contribution of each participating terminal to federated learning. In addition, the central server 102 is also used to aggregate the distilled sample sets provided by multiple participating terminals 101 to obtain a target aggregated sample set, and distribute the target aggregated sample set to each participating terminal 101, so that each participating terminal 101 can perform model training locally based on the target aggregated sample set, and then participate in multi-party collaborative modeling to achieve data sharing.

[0063] Schematically, the central server 102 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers. In some embodiments, the central server 102 can also be deployed on a cloud platform to improve the scalability of data storage and computing in the federated learning system 100. Schematically, the cloud platform is the abbreviation of the cloud computing platform, which refers to a service based on hardware resources and software resources, providing computing, network and storage capabilities. Through the network "cloud", the huge data computing process is processed and analyzed at the remote end and returned to the user, with the characteristics of large-scale, distributed, virtualized, high availability, scalability, on-demand service and security. The cloud platform can realize the rapid issuance and release of configurable computing resources with a small management cost or a low interaction complexity between users and service providers. Schematically, the cloud platform is a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In some embodiments, the cloud platform can also be a virtual machine instance, a container instance, etc., and the present application is not limited thereto.

[0064] In some embodiments, the wireless network or wired network uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any network, including but not limited to any combination of local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, wired or wireless network, private network or virtual private network. In some implementations, the data exchanged through the network is represented by technology and / or format including hypertext markup language (HTML), extensible markup language (XML), etc. In addition, conventional encryption technologies such as secure socket layer (SSL), transport layer security (TLS), virtual private network (VPN), Internet protocol security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technology can also be used to replace or supplement the above data communication technology.

[0065] The hardware structure involved in the above implementation environment is introduced below.

[0066] An embodiment of the present application provides a computing device that can be configured as a participant terminal 101 or a central server 102 in the above-mentioned federated learning system 100.

[0067] refer to Figure 2 , Figure 2 Schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application. Figure 2 As shown, the computing device 200 includes a memory 201, a processor 202, a communication interface 203 and a bus 204. The memory 201, the processor 202 and the communication interface 203 are connected to each other through the bus 204.

[0068] The memory 201 may include a volatile memory, such as a random access memory (RAM). The memory 201 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD), but the present application is not limited thereto. Schematically, an executable program code is stored in the memory 201, and the processor 202 executes the executable program code to implement the method for determining the contribution of the participating terminal in the federated learning provided in the present application. That is, the memory 201 stores instructions for executing the method for determining the contribution of the participating terminal in the federated learning.

[0069] The processor 202 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP). The processor 202 may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The number of processors 202 may be one or more, and the present application is not limited thereto.

[0070] The memory 201 and the processor 202 may be provided separately or integrated together.

[0071] The communication interface 203 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 200 and other devices or a communication network. For example, data can be obtained through the communication interface 203 .

[0072] The bus 204 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 The bus 204 is represented by only one line, but does not mean that there is only one bus or one type of bus. The bus 204 may include a path for transmitting information between various components of the computing device 200 (eg, the memory 201, the processor 202, and the communication interface 203).

[0073] The present application also provides a computing device cluster, which includes multiple computing devices. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. The computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone. Figure 3 , Figure 3 Schematic diagram of a computing device cluster provided in an embodiment of the present application. Figure 3 As shown, the computing device cluster includes multiple computing devices 200. The memory 201 in the multiple computing devices 200 in the computing device cluster may store the same instructions for executing the method for determining the contribution of the participating terminal in federated learning. In some embodiments, the memory 201 of the multiple computing devices 200 in the computing device cluster may also respectively store partial instructions for executing the method for determining the contribution of the participating terminal in federated learning. In other words, the combination of multiple computing devices 200 can jointly execute the instructions of the method for determining the contribution of the participating terminal in federated learning.

[0074] In some embodiments, multiple computing devices 200 in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 4 Schematic diagram of a connection method of a computing device cluster provided in an embodiment of the present application. Figure 4 As shown, two computing devices 200 are connected via a network. Specifically, the network is connected via the communication interface 203 in each computing device 200. It should be understood that Figure 4 The functionality of the computing device 200 shown in FIG. 2 may also be performed by multiple computing devices 200 .

[0075] Based on the above introduction to the implementation environment of this application, the technical solution provided by this application is introduced below.

[0076] Based on the above introduction to the implementation environment, it can be seen that the present application provides a federated learning system, including multiple participating terminals and a central server. For this federated learning system, the present application provides a participating terminal contribution determination solution that can effectively reduce the computational complexity.

[0077] First, refer to Figure 5 , a brief description of the scheme is given. Figure 5 Schematic diagram of a method for determining the contribution of a participant in federated learning provided in an embodiment of the present application. Figure 5 As shown, taking any round of federated learning process as an example, the participating end performs data set distillation on the local sample set to obtain a distilled sample set, and uploads the distilled sample set to the central server. The central server aggregates the distilled sample sets received from multiple participating ends to obtain the target aggregated sample set, and based on the active learning method, the distilled samples provided by each participating end are labeled and evaluated for benefit, that is, the uncertainty of these distilled samples is determined. For example, the uncertainty is reflected by the information entropy of the distilled samples. It should be understood that information entropy is a concept to measure the amount of information. The larger the information entropy, the greater the uncertainty, and the richer the amount of information contained. Therefore, the greater the uncertainty of a distilled sample, the richer the amount of information it contains, and the richer the amount of information it can bring to the training process of the machine learning model, and the greater its potential contribution to federated learning, or the higher the training value of the sample. In this way, the central server can accurately determine the contribution of each participating terminal based on the uncertainty of the distilled samples provided by each participating terminal, distribute the contribution of each participating terminal to the corresponding participating terminal, and send the target aggregated sample set to each participating terminal, so that the participating terminal can update its own local sample set based on the target aggregated sample set and re-distill the data set, and adjust its own sample selection according to the contribution, and perform local model training according to the target aggregated sample set to enter the next round of federated learning process, or, after multiple rounds of interaction with the central server, train the model locally according to the target aggregated sample set obtained in the last round, and so on.

[0078] Next, continue to refer to Figure 6 and Figure 7 , Figure 6 is a schematic diagram of a federated learning system provided in an embodiment of the present application. Figure 7 Schematic diagram of a federated learning method provided in an embodiment of the present application. Figure 6 and Figure 7As shown, for the participating end, its software layer is deployed with a data set distillation module, which is used to implement the aforementioned data set distillation process for the local sample set, and its hardware layer is deployed with GPU, CPU and other processors to provide hardware support for the data set distillation process for the local sample set. For the central server, its software layer is deployed with a contribution determination module and a federated learning algorithm module, wherein the contribution determination module is used to implement the aforementioned contribution determination process for the participating end, and the federated learning algorithm module is used to implement the aforementioned aggregation process for the distilled samples, the aggregated sample distribution process, etc., and its hardware layer is deployed with GPU, CPU and other processors to provide hardware support for the contribution determination, distilled sample aggregation and other processes.

[0079] In the above process, on the one hand, by distilling the local sample set, that is, compressing and distilling the low-rank high-dimensional data into a small amount of data, the visible part of the local sample set is difficult to be understood by humans, and is stored in a small amount of distilled samples in a high-density manner, thereby achieving privacy protection; on the other hand, this distillation method will reduce the amount of data, thereby reducing the complexity of the calculation, and thus improving the efficiency of the overall contribution evaluation calculation; on the other hand, the active learning method is used to determine the potential contribution of the distilled samples provided by each participant to federated learning, and then accurately determine the contribution of each participant, effectively improving the efficiency of the contribution calculation.

[0080] Reference below Figure 8 , combined with the method embodiments, the technical solution provided by this application is introduced in detail.

[0081] Figure 8 1 is a flow chart of a method for determining the contribution of a participant in federated learning provided by an embodiment of the present application. Figure 8 As shown, the method is applied to a federated learning system, and is introduced by taking the interaction between multiple participating terminals and a central server in the federated learning system as an example. The method includes the following steps 801 to 807.

[0082] 801. Multiple participating terminals perform data set distillation on their respective local sample sets to obtain a distilled sample set of each participating terminal, where the distilled sample set of each participating terminal includes at least one distilled sample.

[0083] In an embodiment of the present application, each participant in the federated learning has its own local sample set, and the local sample set includes training samples and test samples, wherein the samples are annotated with label information. Schematically, taking any participating end as an example, the participating end performs data set distillation on the local sample set of the participating end based on the random model obtained by randomly sampling the model parameters of the machine learning model, and obtains the distilled sample set of the participating end, and the distilled sample set includes at least one distilled sample. Among them, the distilled sample is annotated with distillation label information, and the machine learning model is also the target model of the federated learning, such as the image detection model, the image classification model, the text detection model, etc. The present application does not limit the functions implemented by the target model of the federated learning and its model structure. Accordingly, the local sample can be a sample image, a sample text, etc. Random sampling can be, for example, Gaussian random sampling, which is not limited in the present application. For example, taking the local sample as a sample image as an example, the local sample is also a sample image annotated with label information, and accordingly, the distilled sample is also a distilled image annotated with distillation label information.

[0084] In some embodiments, the central server initializes the machine learning model based on the number of participating terminals in the federated learning, obtains multiple random models, and distributes the multiple random models to the multiple participating terminals. For example, taking the number of multiple participating terminals as K (K is a positive integer), the central server initializes the machine learning model to obtain K random models, and distributes these K random models to K participating terminals. In this way, each participating terminal will obtain a random model, and based on this random model, perform data set distillation on their respective local sample sets to obtain a distilled sample set for each participating terminal. Of course, the model initialization process can also be performed by the participating terminals, and this application does not limit this.

[0085] Taking any participating end as an example, the following introduces the process of performing data set distillation on the local sample set by the participating end to obtain a distilled sample set.

[0086] Schematically, the participating end initializes the distilled sample set and the learning rate based on the local sample set, and trains the random model based on the local sample set, the initialized distilled sample set and the initialized learning rate. In this process, the participating end traverses the local sample set and takes a batch of local samples each time. Then, the random model is trained using the initialized distilled sample set, wherein each time a distilled sample is taken, the model parameters of the random model are updated using the gradient descent method, and the error is calculated using the distilled label information of the distilled sample, multiplied by a learning rate. Then, the error is calculated based on the label information of a batch of local samples. At this time, the model parameters are the latest model parameters updated by the gradient descent. Back propagation is performed based on the updated model parameters to update the distilled sample set and the learning rate. The overall training goal is to minimize the error under the local sample set. It should be understood that after traversing the local sample set once, a batch of distilled samples and learning rates that have undergone multiple gradient descents are obtained, that is, the distilled sample set of the participating end is obtained. Of course, in the actual local training process, the learning rate can be adjusted manually, and this application does not limit this. There are many common learning rate adjustment methods, such as StepLR, which multiplies a fixed step (epoch) by a fixed scaling factor, or ConsineAnnealingLR, which is dynamically adjusted according to the cosine function. Adjusting the learning rate can make the change of the entire training error more stable, and ultimately achieve higher task accuracy. This application does not limit this.

[0087] In some embodiments, the distillation label information of the distilled sample is a soft label. The soft label is a label type that is suitable for large-scale unsupervised distillation models by converting the label into a discrete value and performing secondary training. For example, the soft label is a one-hot vector, which is a one-dimensional vector containing C elements. C is a positive integer and is the total number of categories. The element corresponding to the category position is 1, and the other elements are 0. This application does not limit this.

[0088] In addition, the above process of obtaining the distilled sample set is only an example. In some embodiments, other data set distillation algorithms can also be used to perform data set distillation on the local sample set of the participating end. The present application is not limited to this.

[0089] 802. Multiple participating terminals send their respective distilled sample sets to the central server.

[0090] Among them, multiple participating terminals can send their respective distilled sample sets to the central server synchronously or asynchronously, and the present application is not limited to this. In some embodiments, multiple participating terminals send their respective distilled sample sets and learning rates to the central server, so that the central server can subsequently determine the contribution of the participating terminals in combination with the learning rate and improve the accuracy of the contribution.

[0091] 803. The central server obtains distillation sample sets from multiple participating terminals.

[0092] 804. The central server pre-trains the machine learning model based on at least one first distilled sample in the distilled sample set of each participating terminal, and determines the uncertainty of at least one second distilled sample in the distilled sample set of each participating terminal based on the pre-trained machine learning model.

[0093] In an embodiment of the present application, the uncertainty indicates the amount of information that the corresponding second distilled sample brings to the training process of the machine learning model. For example, the uncertainty is reflected by the information entropy of the second distilled sample. It should be understood that information entropy is a concept that measures the amount of information. The larger the information entropy, the greater the uncertainty, and the richer the amount of information contained. Therefore, the greater the uncertainty of a second distilled sample, the richer the amount of information it contains, and the richer the amount of information it can bring to the training process of the machine learning model, and the greater its potential contribution to federated learning, or the higher the training value of the sample. In this way, the contribution of each participating end can be accurately determined based on the uncertainty of the second distilled sample provided by each participating end.

[0094] The implementation of this step is described in detail below. Schematically, this implementation includes the following steps A to D:

[0095] Step A: Determine at least one first distillation sample and at least one second distillation sample corresponding to each participating end from the distillation sample set of each participating end.

[0096] Among them, for any participating end, at least one first distilled sample and at least one second distilled sample are respectively determined from the distilled sample set of the participating end. For example, the first distilled sample and the second distilled sample are determined from the distilled sample set by an average division method. Taking the distilled sample set of the participating end including 100 distilled samples as an example, the first 50 distilled samples are determined as the first distilled samples, and the last 50 distilled samples are determined as the second distilled samples. For another example, the first distilled sample and the second distilled sample are determined from the distilled sample set by a random sampling method. In addition, the number of the first distilled sample and the second distilled sample may be the same or different, and the present application does not limit this. Moreover, the method of determining these two types of distilled samples from the distilled sample set may also adopt other methods, and is not limited to the aforementioned method.

[0097] Step B: performing sample aggregation on at least one first distillation sample corresponding to each participating end to obtain a first aggregated sample set.

[0098] Among them, the first aggregated sample set includes at least one first aggregated sample. In some embodiments, the number of the at least one first aggregated sample is the same as the number of the at least one first distilled sample. That is, taking the multiple participating terminals including participating terminal-1 and participating terminal-2 as an example, 50 first distilled samples are determined from the distilled sample set of participating terminal-1, and 50 first distilled samples are determined from the distilled sample set of participating terminal-2, and the first distilled samples corresponding to the two participating terminals are aggregated to obtain 50 first aggregated samples, that is, the first aggregated sample set is obtained. Of course, in other embodiments, the number of first aggregated samples and first distilled samples may also be different. For example, after performing sample aggregation on at least one first distilled sample corresponding to each participating terminal, the aggregated samples are screened to obtain the first aggregated sample set. In addition, the present application does not limit the method of sample aggregation. For example, arithmetic average aggregation may be used, weighted aggregation may be used, and so on.

[0099] In some embodiments, based on the aforementioned step 802, multiple participating terminals can also send their respective learning rates to the central server. In this way, the central server can aggregate the learning rates corresponding to each participating terminal to obtain an aggregated learning rate set, which is convenient for the subsequent central server to combine the aggregated learning rate set to determine the contribution of the participating terminal and improve the accuracy of the contribution.

[0100] Step C: pre-train the machine learning model based on the first aggregated sample set.

[0101] Among them, the machine learning model is also the target model of federated learning. In this step, the central server initializes the model parameters of the machine learning model. The initialized machine learning model is referred to as the basic model below. Then, the central server pre-trains the basic model based on the first aggregated sample set using the gradient descent method to obtain the pre-trained machine learning model. In some embodiments, if multiple participating terminals send their respective learning rates to the central server, the central server pre-trains the basic model based on the first aggregated sample set and the aggregated learning rate set using the gradient descent method to obtain the pre-trained machine learning model. It should be noted that if multiple participating terminals do not send their respective learning rates to the central server, the central server can initialize the learning rate by itself, or set it manually, and this application does not limit this.

[0102] It should be understood that it is called pre-training here because this part of the training process is not the training process of federated learning, but a preparatory step performed by the central server to determine the contribution of each participating terminal.

[0103] Step D: Based on the pre-trained machine learning model, determine the uncertainty of at least one second distillation sample corresponding to each participating end.

[0104] Among them, taking any participating end as an example (hereinafter referred to as the first participating end), the central server inputs at least one second distilled sample corresponding to the first participating end into the pre-trained machine learning model, obtains the prediction result corresponding to each second distilled sample, and determines the uncertainty of each second distilled sample based on the prediction result corresponding to each second distilled sample.

[0105] In an embodiment of the present application, the uncertainty of the second distillation sample can be calculated using a variety of sample labeling methods involved in active learning. For example, based on entropy, learning rate, minimum confidence, variance, and minimum interval, etc., the present application does not limit this, that is, the present application uses the method of labeling the sample in active learning to determine the uncertainty of each second distillation sample. It should be understood that any method that uses active learning to label the sample can be applied to the present application.

[0106] The following is an example to illustrate this step.

[0107] For example, referring to the following formula (1), the uncertainty can be calculated based on the entropy value and the learning rate:

[0108]

[0109] In the above formula (1), is the calculation function of uncertainty, In the above equation, f represents the model structure of the machine learning model, θ represents the model parameters of the pre-trained machine learning model, For the second distillation sample, is the distillation label information of the second distillation sample, is the aggregate learning rate, w c is the weight parameter. The output of the machine learning model f, that is, the prediction result corresponding to the second distilled sample, is a vector containing C elements, namely o i represents the confidence level of participant i, In this way, the model output, soft label and learning rate are taken into account at the same time, which effectively characterizes the uncertainty generated by the distillation sample, thereby achieving accurate and efficient contribution calculation.

[0110] For another example, referring to the following formula (2), the uncertainty is calculated based on the entropy value:

[0111]

[0112] The meanings of the various terms in the above formula (2) refer to the above formula (1) and will not be repeated here.

[0113] 805. The central server determines the contribution of each participating terminal based on the uncertainty of at least one second distilled sample in the distilled sample set of each participating terminal, and the contribution is used to instruct the corresponding participating terminal to adjust the distilled sample set of the corresponding participating terminal.

[0114] In an embodiment of the present application, taking any participating end as an example (hereinafter referred to as the first participating end), the central server determines the contribution of the first participating end based on at least one of the mean and variance of the uncertainties of multiple second distillation samples corresponding to the first participating end.

[0115] The following uses the average value to calculate the contribution as an example to illustrate this step.

[0116] Schematically, taking the number of multiple participating terminals as K as an example (K is a positive integer), the contribution of the multiple participating terminals is represented as s1, s2, ..., s K , for the i-th participant, its contribution is calculated by the following formula (3):

[0117]

[0118] In the above formula (3), q refers to the ratio coefficient for pre-training the machine learning model (the ratio coefficient indicates that in the sequence data of length N, the first q×N distilled samples are selected to pre-train a basic model, that is, how many first distilled samples are selected in the distilled sample set of each participating terminal as the training samples for pre-training); N is the number of distilled samples in the distilled sample set uploaded by the participating terminal i; is the calculation function of uncertainty, refer to the above formula (1), and no further description is given, and j represents any second distillation sample. In this way, the contribution of each participating end can be determined quickly and accurately, effectively reducing the calculation complexity.

[0119] In the above steps 804 and 805, the distilled sample set of each participating end is divided into two parts, namely the first distilled sample and the second distilled sample, which is equivalent to using one part of the distilled samples as training samples and the other part of the distilled samples as test samples. Therefore, after training the basic model according to the first distilled sample, the trained basic model is used to evaluate the labeling benefit of the second distilled sample, and the uncertainty of the second distilled sample is used to reflect its potential contribution to federated learning, thereby determining the contribution of the participating end, with low computational complexity and high efficiency in calculating the contribution.

[0120] 806. The central server sends the contribution of each participating terminal to the corresponding participating terminal.

[0121] 807. Multiple participating terminals adjust their respective distillation sample sets based on their respective contributions.

[0122] In an embodiment of the present application, the participating end can adopt a variety of methods to adjust the distilled sample set based on the contribution degree. For example, if the contribution degree is greater than a preset threshold, some samples in the local sample set are selected for re-distillation. If the contribution degree is less than the preset threshold, the local sample set is expanded and then re-distilled, and so on. The present application does not limit this.

[0123] After the above steps 801 to 807, after the multiple participants of the federated learning send their respective distilled sample sets to the central server, the central server determines the contribution of each participant based on these distilled sample sets, and sends the contribution of each participant to the corresponding participant. In this way, each participant can adjust its own distilled sample set based on its own contribution.

[0124] In some embodiments, the central server can also perform sample aggregation on the distilled sample sets of multiple participating terminals to obtain a target aggregated sample set, and send the target aggregated sample set to multiple participating terminals, so that multiple participating terminals can perform local model training based on the target aggregated sample set. For example, taking the multiple participating terminals including participant-1 and participant-2 as an example, the distilled sample set of participant-1 includes 100 distilled samples, and the distilled sample set of participant-2 includes 100 distilled samples. The distilled sample sets corresponding to the two participating terminals are sample aggregated (such as using arithmetic average aggregation) to obtain 100 target aggregated samples, that is, the target aggregated sample set is obtained. Among them, the present application does not limit the method of sample aggregation. For example, arithmetic average aggregation can be used, or weighted aggregation can be used, etc. Through the above method, for any participating terminal, the participating terminal can obtain the global distilled sample, so that the model training can be realized locally, and the model parameters of the trained model do not need to be uploaded to the central server. Moreover, by performing data set distillation on the local sample set, that is, compressing and distilling the low-rank high-dimensional data into a small amount of data, the visible part of the local sample set is difficult to be understood by humans, and is stored in a small amount of distilled samples in a high-density manner, thereby achieving privacy protection from the training input. It should be noted that this application does not limit the timing of the central server performing sample aggregation, and the central server can perform sample aggregation at any time in the aforementioned steps 803 to 807.

[0125] In other embodiments, taking the first participant as an example, the central server sends the distilled sample set of at least one second participant to the first participant, so that the first participant performs local model training based on the distilled sample set of the first participant and the distilled sample set of at least one second participant, and the second participant refers to a participant other than the first participant among the multiple participants. In the above manner, the sample aggregation task originally performed by the central server is offloaded to each participant, which can save the computing resources of the central server, and for any participant, the participant can obtain the distilled samples of other participants, so that model training can be implemented locally without uploading the model parameters of the trained model to the central server.

[0126] Reference below Fig. 9 , taking any participating terminal as an example, the above steps 801 to 807 are illustrated. Fig. 9 Schematic diagram of a method for determining the contribution of a participant in federated learning provided in an embodiment of the present application. Fig. 9 As shown, the federated learning system includes K participating terminals. For participant-1, participant-1 performs data set distillation on the local sample set based on the random model-1 obtained by randomly sampling the model parameters of the machine learning model to obtain the distilled sample set-1, and uploads the distilled sample set-1 to the central server. The central server obtains the distilled sample sets of the K participating terminals, that is, distilled sample sets-1 to distilled sample sets-K. The central server evaluates the contribution of each participating terminal based on the basic model obtained by randomly sampling the model parameters of the machine learning model to obtain K contributions, that is, contribution-1 to contribution-K, and performs sample aggregation on the distilled sample sets of the K participating terminals to obtain the target aggregated sample set. Then the central server sends the K contributions and the target aggregated sample set to the K participating terminals. In this way, for participant-1, participant-1 can update its own local sample set based on the target aggregated sample set and re-distill the data set, and adjust its own sample selection according to the contribution, and perform local model training according to the target aggregated sample set, and so on.

[0127] In addition, it should be understood that in the process of federated learning, multiple participating terminals and the central server can perform multiple rounds of interactions based on the aforementioned steps 801 to 807. For example, the number of interaction rounds is set to T (T is a positive integer), and there are a total of K participating terminals. In each round of interaction, the K participating terminals perform data set distillation on the local sample set to obtain the corresponding distilled sample set, and send the distilled sample set to the central server. The central server obtains the K distilled sample sets, performs sample aggregation on the K distilled sample sets, and obtains the target aggregated sample set. In this process, the contribution of the K participating terminals is determined using an active learning method. Thereafter, the target aggregated sample set and the contribution are sent to each participating terminal, so that the participating terminal can perform model training locally based on the target aggregated sample set, and further adjust its own distilled sample set based on the contribution, and re-upload the adjusted distilled sample set to the central server to enter the next round of federated learning process, or, after T rounds of interaction with the central server, the model is trained locally based on the target aggregated sample set obtained in the Tth round, and so on.

[0128] In summary, in the method provided in the embodiment of the present application, based on the distilled sample set provided by multiple participating terminals in federated learning, the uncertainty of the distilled samples in the distilled sample set is used to reflect the amount of information brought by the distilled samples to the training process of the machine learning model, so as to determine the contribution of each participating terminal according to the uncertainty of the distilled samples provided by each participating terminal, and on the basis of accurately evaluating the contribution of the participating terminals, the computational complexity is effectively reduced, thereby improving the efficiency of the contribution calculation. Moreover, the method provided in the present application can efficiently handle large-scale federated learning tasks and has good scalability.

[0129] On the other hand, related technologies, such as the Shapley value calculation method, usually require the use of test sample sets to evaluate the contribution of participants. Therefore, on the one hand, it is necessary to introduce additional test sample sets, which increases the cost of data collection and processing. On the other hand, the calculation of the Shapley value requires considering all possible subsets, so the calculation complexity is very high, resulting in slow algorithm operation. For example, refer to Fig.10 , Fig.10 This is a comparison chart of experimental results provided in the embodiment of the present application. Fig.10As shown in Figure (a), Figure (a) shows the change in order accuracy of AlexNet under the Shapley value calculation method and the method of the present application under the CIFAR-10, CIFAR-100 and MNIST datasets, where the range of the order accuracy indicator order-acc value is -1 to 1, -1 means that the order of the two sequences is completely opposite, and 1 means that the two sequences are exactly the same. It can be seen that the Shapley value calculation method performs poorly on CIFAR-100 and MNIST, while the method of the present application performs well under various datasets. Furthermore, as Fig.10 As shown in Figure (b), Figure (b) shows the sequential accuracy of AlexNet under the Shapley value calculation method under the CIFAR-100 and MNIST data sets. It can be seen that in the later stage of training, as the model began to converge, the sequential accuracy dropped rapidly. In the method of the present application, since there is no need to rely on the test sample set, but pre-training and calculation of the distilled samples, the scalability of the algorithm is greatly improved. At the same time, as the training progresses, since it does not rely on the test sample set, the contribution of the distilled samples is measured by the uncertainty of the prediction results corresponding to the distilled samples, so the stability of the algorithm is also relatively high.

[0130] Fig.11 1 is a schematic diagram of a structure of a device for determining the contribution of a participant in federated learning provided by an embodiment of the present application. The device for determining the contribution of a participant in federated learning can be implemented as part or all of the functions of the aforementioned central server through software, hardware, or a combination of both. Fig.11 As shown, the device is applied to a central server in a federated learning system, the system also includes multiple participating terminals, and the device includes an acquisition unit 1101, a first determination unit 1102 and a second determination unit 1103.

[0131] An acquisition unit 1101 is used to acquire distilled sample sets of multiple participating terminals, where the distilled sample set of each participating terminal includes at least one distilled sample obtained after the corresponding participating terminal performs data set distillation on a local sample set;

[0132] A first determining unit 1102 is configured to pre-train a machine learning model based on at least one first distilled sample in the distilled sample set of each participating terminal, and determine the uncertainty of at least one second distilled sample in the distilled sample set of each participating terminal based on the pre-trained machine learning model, wherein the uncertainty indicates the amount of information brought by the corresponding second distilled sample to the training process of the machine learning model;

[0133] The second determination unit 1103 is used to determine the contribution of each participating end based on the uncertainty of at least one second distillation sample in the distillation sample set of each participating end, and send the contribution of each participating end to the corresponding participating end, and the contribution is used to instruct the corresponding participating end to adjust the distillation sample set of the corresponding participating end.

[0134] In some embodiments, the device is also used to implement other functions of the central server described in the above embodiments, which will not be repeated here.

[0135] In addition, in the above device, the acquisition unit 1101, the first determination unit 1102, and the second determination unit 1103 can be implemented by software or by hardware. Exemplarily, the implementation of the acquisition unit 1101 is described below by taking the acquisition unit 1101 as an example. Similarly, the implementation of other modules can refer to the implementation of the acquisition unit 1101.

[0136] As an example of a software functional unit, the acquisition unit 1101 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the acquisition unit 1101 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.

[0137] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0138] As an example of a hardware functional unit, the acquisition unit 1101 may include at least one computing device. Alternatively, the acquisition unit 1101 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0139] The multiple computing devices included in the acquisition unit 1101 can be distributed in the same region or in different regions. The multiple computing devices included in the acquisition unit 1101 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the acquisition unit 1101 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0140] It should be noted that: the contribution determination device of the participating end in federated learning provided in the above embodiment only uses the division of the above functional modules as an example when determining the contribution. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the contribution determination device of the participating end in federated learning provided in the above embodiment and the contribution determination method embodiment of the participating end in federated learning belong to the same concept. The specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0141] Fig.12 1 is a schematic diagram of another structure of a device for determining the contribution of a participant in federated learning provided by an embodiment of the present application. The device for determining the contribution of a participant in federated learning can be implemented as part or all of the functions of the aforementioned central server through software, hardware, or a combination of both. Fig.12 As shown, the device is applied to a participating end in a federated learning system, the system also includes a central server, and the device includes an acquisition unit 1201 and a sending unit 1202.

[0142] A sending unit 1201 is used to send a distilled sample set of a participating terminal to a central server, where the distilled sample set includes at least one distilled sample obtained after the participating terminal performs data set distillation on a local sample set;

[0143] The acquisition unit 1202 is used to acquire the contribution of the participating terminal, and the contribution is used to instruct the participating terminal to adjust the distilled sample set; wherein, the method for determining the contribution includes: pre-training the machine learning model based on at least one first distilled sample in the distilled sample set of each participating terminal in the federated learning system, and determining the uncertainty of at least one second distilled sample in the distilled sample set of each participating terminal based on the pre-trained machine learning model, wherein the uncertainty indicates the amount of information brought by the corresponding second distilled sample to the training process of the machine learning model; based on the uncertainty of at least one second distilled sample in the distilled sample set of each participating terminal, determining the contribution of each participating terminal.

[0144] In some embodiments, the device is also used to implement other functions of the participating end described in the above embodiments, which will not be repeated here.

[0145] It should be noted that: the contribution determination device of the participating end in federated learning provided in the above embodiment only uses the division of the above functional modules as an example when determining the contribution. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the contribution determination device of the participating end in federated learning provided in the above embodiment and the contribution determination method embodiment of the participating end in federated learning belong to the same concept. The specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0146] In this application, the terms "first", "second", etc. are used to distinguish between identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is the quantity and execution order limited. It should also be understood that although the following description uses the terms first, second, etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various described examples, the first participant terminal can be referred to as the second participant terminal, and similarly, the second participant terminal can be referred to as the first participant terminal. The first participant terminal and the second participant terminal can both be participant terminals, and in some cases, can be separate and different participant terminals.

[0147] In the present application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple participating terminals refer to two or more participating terminals.

[0148] The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0149] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of program structure information. The program structure information includes one or more program instructions. When the program instructions are loaded and executed on a computing device, all or part of the processes or functions in the embodiments of the present application are generated.

[0150] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0151] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for determining the contribution of a participant in federated learning, characterized in that: A central server applied to a federated learning system, wherein the system further comprises a plurality of participating terminals, and the method comprises: Acquire distilled sample sets of the multiple participating terminals, where the distilled sample set of each participating terminal includes at least one distilled sample obtained after the corresponding participating terminal performs data set distillation on the local sample set; Pre-training a machine learning model based on at least one first distilled sample in the distilled sample set of each participating end, and determining the uncertainty of at least one second distilled sample in the distilled sample set of each participating end based on the pre-trained machine learning model, wherein the uncertainty indicates the amount of information brought by the corresponding second distilled sample to the training process of the machine learning model; Based on the uncertainty of at least one second distillation sample in the distillation sample set of each participating end, the contribution of each participating end is determined, and the contribution of each participating end is sent to the corresponding participating end, and the contribution is used to instruct the corresponding participating end to adjust the distillation sample set of the corresponding participating end.

2. The method according to claim 1, characterized in that The method of pre-training the machine learning model based on at least one first distilled sample in the distilled sample set of each participating terminal, and determining the uncertainty of at least one second distilled sample in the distilled sample set of each participating terminal based on the pre-trained machine learning model, comprises: Determine at least one first distillation sample and at least one second distillation sample corresponding to each participating end from the distillation sample set of each participating end; Aggregating at least one first distilled sample corresponding to each participating end to obtain a first aggregated sample set; Pre-training the machine learning model based on the first aggregated sample set; Based on the pre-trained machine learning model, the uncertainty of at least one second distillation sample corresponding to each participating end is determined.

3. The method according to claim 2, characterized in that The step of determining the uncertainty of at least one second distilled sample corresponding to each participating terminal based on the pre-trained machine learning model includes: Inputting at least one second distilled sample corresponding to a first participating end into the pre-trained machine learning model to obtain a prediction result corresponding to each second distilled sample, wherein the first participating end refers to any participating end; The uncertainty of each second distillation sample is determined based on the prediction result corresponding to each second distillation sample.

4. The method according to any one of claims 1 to 3, characterized in that The step of determining the contribution of each participating end based on the uncertainty of at least one second distillation sample in the distillation sample set of each participating end includes: Based on at least one of an average value and a variance of uncertainties of a plurality of second distillation samples corresponding to the first participating end, a contribution of the first participating end is determined, where the first participating end refers to any one participating end.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The distilled sample sets of the multiple participating terminals are aggregated to obtain a target aggregated sample set, where the target aggregated sample set is used by the multiple participating terminals to locally train the machine learning model.

6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: A distilled sample set of at least one second participating end is sent to a first participating end, and the distilled sample set of at least one second participating end is used for the first participating end to locally train the machine learning model, the first participating end refers to any participating end, and the second participating end refers to a participating end among the multiple participating ends except the first participating end.

7. A method for determining the contribution of a participant in federated learning, characterized in that: Applied to a participant in a federated learning system, the system further includes a central server, and the method includes: Sending the distilled sample set of the participating terminal to the central server, the distilled sample set including at least one distilled sample obtained by the participating terminal after performing data set distillation on the local sample set; Acquire the contribution of the participating end, wherein the contribution is used to instruct the participating end to adjust the distilled sample set; wherein the method for determining the contribution includes: pre-training the machine learning model based on at least one first distilled sample in the distilled sample set of each participating end in the federated learning system, and determining the uncertainty of at least one second distilled sample in the distilled sample set of each participating end based on the pre-trained machine learning model, wherein the uncertainty indicates the amount of information brought by the corresponding second distilled sample to the training process of the machine learning model; determine the contribution of each participating end based on the uncertainty of at least one second distilled sample in the distilled sample set of each participating end.

8. The method according to claim 7, characterized in that The method further comprises: Based on the random model obtained by randomly sampling the model parameters of the machine learning model, the local sample set of the participating end is distilled to obtain the distilled sample set of the participating end.

9. The method according to claim 7 or 8, characterized in that: The method further comprises: Obtaining a target aggregated sample set, where the target aggregated sample set is obtained by aggregating distilled sample sets of multiple participating terminals in the federated learning system; The machine learning model is trained locally based on the target aggregate sample set.

10. A device for determining the contribution of a participant in federated learning, characterized in that: A central server applied to a federated learning system, wherein the system further comprises a plurality of participating terminals, and the device comprises: An acquisition unit, configured to acquire distilled sample sets of the plurality of participating terminals, wherein the distilled sample set of each participating terminal includes at least one distilled sample obtained after the corresponding participating terminal performs data set distillation on a local sample set; A first determining unit is configured to pre-train the machine learning model based on at least one first distilled sample in the distilled sample set of each participating terminal, and determine the uncertainty of at least one second distilled sample in the distilled sample set of each participating terminal based on the pre-trained machine learning model, wherein the uncertainty indicates an amount of information brought by the corresponding second distilled sample to the training process of the machine learning model; The second determination unit is used to determine the contribution of each participating end based on the uncertainty of at least one second distillation sample in the distillation sample set of each participating end, and send the contribution of each participating end to the corresponding participating end, wherein the contribution is used to instruct the corresponding participating end to adjust the distillation sample set of the corresponding participating end.

11. A device for determining the contribution of a participant in federated learning, characterized in that: Applied to a participant in a federated learning system, the system further includes a central server, and the device includes: A sending unit, configured to send the distilled sample set of the participating terminal to the central server, wherein the distilled sample set includes at least one distilled sample obtained by the participating terminal after performing data set distillation on a local sample set; An acquisition unit is used to acquire the contribution of the participating end, and the contribution is used to instruct the participating end to adjust the distilled sample set; wherein the method for determining the contribution includes: pre-training the machine learning model based on at least one first distilled sample in the distilled sample set of each participating end in the federated learning system, and determining the uncertainty of at least one second distilled sample in the distilled sample set of each participating end based on the pre-trained machine learning model, wherein the uncertainty indicates the amount of information brought by the corresponding second distilled sample to the training process of the machine learning model; based on the uncertainty of at least one second distilled sample in the distilled sample set of each participating end, determining the contribution of each participating end.

12. A federated learning system, characterized in that: The system includes multiple participating terminals and a central server of federated learning; the central server is used to execute the method for determining the contribution of the participating terminals in federated learning as described in any one of claims 1 to 6, and the participating terminals are used to execute the method for determining the contribution of the participating terminals in federated learning as described in any one of claims 7 to 9.

13. A computing device, characterized in that: The computing device includes a processor and a memory, the memory is used to store at least one piece of program code, and the at least one piece of program code is loaded by the processor and executes the method for determining the contribution of a participating end in federated learning as described in any one of claims 1 to 6, or executes the method for determining the contribution of a participating end in federated learning as described in any one of claims 7 to 9.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store at least one program code, and the at least one program code is used to execute the method for determining the contribution of a participant in federated learning as described in any one of claims 1 to 6, or to execute the method for determining the contribution of a participant in federated learning as described in any one of claims 7 to 9.

15. A computer program product, characterized in that When the computer program product runs on a computing device, the computing device executes the method for determining the contribution of a participant in federated learning as described in any one of claims 1 to 6, or executes the method for determining the contribution of a participant in federated learning as described in any one of claims 7 to 9.