Data summary generation method, model training method, related device and system

By adaptively generating targeted data digests and training the global model based on this data digest and non-local data digests, the problem of data heterogeneity and data digest selection in the prior art is solved, the effectiveness and reliability of distributed learning is improved, and the performance and accuracy of the global model are enhanced.

CN114358120BActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111452498.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-05-27
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

In the prior art, due to the problems such as data heterogeneity and untargeted selection of data digests in local data distributed in each client device, the effectiveness and accuracy of the global model trained based on data digests output by each client device are poor, which in turn significantly reduces the application effect and performance of the global model.

Method used

By adaptively generating a data summary corresponding to the local training data based on the model parameters of the received global model and the model training contribution level to which the local training data belongs, the global model is trained together based on the data summary and the data summary of multiple non-local training data. The method includes obtaining the feature parameters of the local training data, determining the weight of the feature parameters based on the entropy weight method, calculating the model training contribution score, and determining the contribution level of the training data through hierarchical clustering, and finally generating a data summary using the adaptive data condensation algorithm.

Benefits of technology

It improves the pertinence and adaptability of data digests, enhances the effectiveness and reliability of data digests in the distributed learning process, reduces the communication overhead of the distributed learning process, and ensures the privacy and security of the training data, thereby improving the performance and accuracy of the global model trained using data digests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358120B_ABST
    Figure CN114358120B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data digest generation method, a model training method, related devices, and a system. The data digest generation method includes: adaptively generating a data digest corresponding to local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs, where the local training data is distributed with multiple non-local training data, and the model training contribution level to which each piece of training data belongs is obtained by performing hierarchical clustering on each piece of training data separately in advance; outputting the data digest corresponding to the local training data to jointly train the global model based on the data digests corresponding to each piece of training data. The present application can effectively improve the pertinence and adaptability of generating data digests, and can improve the effectiveness and reliability of data digests in the distributed learning process, and further can effectively improve the application effectiveness and accuracy of the global model trained using data digests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a data digest generation method, a model training method, related devices and systems. Background Art

[0002] The rapid growth in the number of modern mobile devices leads to the generation of a large amount of discretely distributed data, which is very valuable for model training. Due to restrictions in data privacy and security and expensive communication overheads, it is difficult to share the data collected and stored in different devices and locations with a server or other devices. Therefore, it is necessary to study how to apply machine learning to distributed data scenarios with these limitations.

[0003] Currently, as one of the main methods of distributed learning, the method of transmitting data digests calculates the data digests of local data by each client respectively, and then sends these data digests to a central server for global model training, which can reduce communication overhead and protect data privacy to a certain extent. However, due to problems such as data heterogeneity in the local data scattered in each client device and lack of pertinence in data digest selection, the effectiveness and accuracy of the global model trained based on the data digests output by each client device are poor, and thus the application effect and performance of the global model are significantly reduced. Summary of the Invention

[0004] Aiming at the problems in the prior art, this application provides a data digest generation method and device, which can effectively improve the pertinence and self - adaptability of generated data digests, and can improve the effectiveness and reliability of data digests in the process of distributed learning, and further can effectively improve the application effectiveness and accuracy of the global model trained using data digests.

[0005] To solve the above - mentioned technical problems, this application provides the following technical solutions:

[0006] In a first aspect, this application provides a data digest generation method, including:

[0007] Adaptive generation of a data digest corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs, where the local training data is distributed with multiple non - local training data, and the model training contribution level to which each piece of training data belongs is obtained by pre - performing hierarchical clustering on each piece of training data respectively;

[0008] Outputting the data digest corresponding to the local training data to jointly train the global model based on the data digest corresponding to the local training data and the data digests corresponding to multiple non - local pieces of training data respectively.

[0009] Further, before adaptively generating a data summary corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs, the following steps are also included:

[0010] Obtain all the feature parameters corresponding to the local training data;

[0011] Based on the entropy weight method, determine the weight corresponding to each of the feature parameters, and determine the model training contribution score of the local training data according to the weights corresponding to each of the feature parameters;

[0012] Output the model training contribution score of the local training data to summarize the model training contribution scores corresponding to each portion of the training data, and perform hierarchical clustering on each portion of the training data based on the model training contribution scores corresponding to each portion of the training data to determine the model training contribution level to which each portion of the training data belongs.

[0013] Further, the feature parameters include: the number of samples, the degree of non-independent and identically distributed, and the balance offset degree;

[0014] Correspondingly, the determining the weight corresponding to each of the feature parameters based on the entropy weight method includes:

[0015] Input the number of samples, the degree of non-independent and identically distributed, and the balance offset degree corresponding to the local training data into a preset entropy weight comprehensive evaluation model, so that the entropy weight comprehensive evaluation model outputs the weights corresponding to the number of samples, the degree of non-independent and identically distributed, and the balance offset degree respectively.

[0016] Further, the adaptively generating a data summary corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs includes:

[0017] Determine the number of samples of the synthetic data set corresponding to the local training data according to the model training contribution level to which the local training data belongs to initialize the synthetic data set;

[0018] For the model parameters of the received global model, generate a data summary corresponding to the local training data based on a preset adaptive data condensation algorithm.

[0019] Further, the generating a data summary corresponding to the local training data based on a preset adaptive data condensation algorithm for the model parameters of the received global model includes:

[0020] For the model parameters of the received global model, a first neural network for a target parameter is trained according to the local training data, and a second neural network for the target parameter is trained based on an initialized synthetic data set;

[0021] Converge the second neural network according to the solution of the first neural network to obtain a converged synthetic data set, and determine the converged synthetic data set as the data summary corresponding to the local training data.

[0022] In a second aspect, the present application provides a model training method, including:

[0023] Output the model parameters of the global model to adaptively generate data summaries corresponding to each piece of training data based on the model parameters and the model training contribution levels to which each piece of distributed training data belongs, where the model training contribution levels to which each piece of training data belongs are obtained by pre-stratified clustering of each piece of training data;

[0024] Receive the data summaries corresponding to each piece of the distributed training data, and jointly train the global model based on the data summaries corresponding to each piece of the training data.

[0025] In a third aspect, the present application provides a data summary generation device, including:

[0026] A summary generation module, configured to adaptively generate a data summary corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs, where the local training data is distributed with multiple pieces of non-local training data, and the model training contribution levels to which each piece of training data belongs are obtained by pre-stratified clustering of each piece of training data;

[0027] A summary output module, configured to output the data summary corresponding to the local training data to jointly train the global model based on the data summary corresponding to the local training data and the data summaries corresponding to multiple pieces of non-local training data.

[0028] In a fourth aspect, the present application provides a model training device, including:

[0029] A parameter sending module, configured to output the model parameters of the global model to adaptively generate data summaries corresponding to each piece of training data based on the model parameters and the model training contribution levels to which each piece of distributed training data belongs, where the model training contribution levels to which each piece of training data belongs are obtained by pre-stratified clustering of each piece of training data;

[0030] An aggregation training module, configured to receive data summaries corresponding to respective portions of the training data set in a distributed manner, and jointly train the global model based on the data summaries corresponding to respective portions of the training data.

[0031] In a fifth aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the data summary generation method described above is implemented, or when the processor executes the computer program, the model training method described above is implemented.

[0032] In a sixth aspect, the present application provides a computer-readable storage medium, having a computer program stored thereon. When the computer program is executed by a processor, the data summary generation method described above is implemented, or when the computer program is executed by the processor, the model training method described above is implemented.

[0033] In a seventh aspect, the present application provides a distributed learning system, including: a central server and a plurality of clients respectively communicatively connected to the central server;

[0034] The central server is configured to execute the model training method described above;

[0035] The client is configured to execute the data summary generation method described above.

[0036] As can be seen from the above technical solutions, a data summary generation method provided by the present application adaptively generates a data summary corresponding to the local training data according to the received model parameters of the global model and the model training contribution level to which the local training data belongs. Among them, the local training data is distributed with multiple portions of non-local training data, and the model training contribution level to which each portion of the training data belongs is obtained by pre-stratified clustering of each portion of the training data; outputs the data summary corresponding to the local training data to jointly train the global model based on the data summary corresponding to the local training data and the data summaries corresponding to multiple portions of non-local training data; by pre-stratified clustering of each portion of the training data to generate the model training contribution levels to which the training data corresponding to each distributed client belongs, and then each client respectively adaptively generates a data summary corresponding to its own local training data according to the model training contribution level to which the local training data belongs, which can effectively improve the pertinence and adaptability of generating data summaries, and can improve the effectiveness and reliability of data summaries in the distributed learning process, and can reduce the communication overhead in the distributed learning process and ensure the privacy and security of training data; furthermore, it can effectively improve the performance of the global model trained using data summaries in the scenario where the client data is not independently and identically distributed, and improve the application effectiveness and accuracy of the global model. Brief Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0038] Figure 1 It is the first process schematic diagram of the data digest generation method in the embodiments of the present application.

[0039] Figure 2 It is the second process schematic diagram of the data digest generation method in the embodiments of the present application.

[0040] Figure 3 It is the third process schematic diagram of the data digest generation method in the embodiments of the present application.

[0041] Figure 4 It is the process schematic diagram of the model training method in the embodiments of the present application.

[0042] Figure 5 It is the structural schematic diagram of the data digest generation device in the embodiments of the present application.

[0043] Figure 6 It is the relationship schematic diagram between the model training device and the data digest generation device in the embodiments of the present application.

[0044] Figure 7 It is the structural schematic diagram of the distributed learning system in the embodiments of the present application.

[0045] Figure 8 It is the logical example schematic diagram of the distributed learning algorithm based on hierarchical clustering and adaptive dataset condensation provided by the application example of the present application.

[0046] Figure 9 It is the structural schematic diagram of the electronic device in the embodiments of the present application. Detailed Embodiments

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0048] Distributed learning is one of the most valuable and promising research directions in large-scale machine learning. Distributed learning methods are generally divided into two categories. One is the method of transmitting model parameters represented by federated learning, and the other is the method of transmitting data summaries represented by sharing core sets.

[0049] In the training process of the method of transmitting model parameters represented by federated learning, the central server first initializes the model parameters and distributes the parameters and the model to all clients. The edge clients send the model parameters updated using local data to the central server, and the central server aggregates these model parameters and then sends the global model to each client. The most commonly used parameter aggregation method in federated learning is to take the average of each parameter. Federated learning has the advantage of keeping the original data locally at the training nodes and only uploading gradients or model parameters during the training process, which can achieve the purpose of privacy protection to a certain extent. Based on the basic federated learning framework, many federated learning algorithms have been derived for the purpose of further reducing communication overhead or more secure data privacy. Some research has shown that most of the data in the gradients transmitted in federated learning is redundant, and an algorithm for significantly reducing communication bandwidth is proposed, which uses four methods: momentum correction, local gradient clipping, momentum factor masking, and pre-training to avoid the accuracy loss caused by reducing the communication volume. However, for the parameter transmission scheme represented by federated learning, some research has shown that the shared model parameters or gradients also contain information about the data, and attackers can obtain information in the original data through gradient inversion or membership inference attacks, which makes the security of traditional federated learning not well guaranteed. Therefore, the frequent transmission of a large number of parameters of modern deep neural networks should be avoided as much as possible.

[0050] In the training process of the method of transmitting data summaries represented by sharing core sets, each client calculates the data summaries of its local data respectively, and then sends these data summaries to the central server. The central server conducts model training to obtain the global model and shares it with all clients. In the distributed learning algorithm of transmitting core sets, the core set serves as the purpose of acting as a data summary. The core set is a weighted data set with a very small number of samples, which can be used as a proxy for the original data set and has provable approximation guarantees. The core set has the same sample feature space as the original data set, so it has certain advantages compared with other data summary methods. The above two methods have both achieved the purpose of reducing communication overhead and protecting data privacy to a certain extent.

[0051] Based on this, the present application considers adopting a data digest scheme, specifically obtaining a data digest by means of data set condensation. Among them, data set condensation is a new direction studied in recent years. The purpose is to condense a large data set into a small data set for a specific network. The condensed data set may only be 1% - 5% of the sample size of the original data set, but training the network with the condensed data set can produce an effect similar to that of using the original data set.

[0052] However, distributed learning often faces the problem of data heterogeneity. Since the data stored dispersedly is usually generated in different environments, this will lead to inconsistencies in data distribution, that is, there will be significant differences in the data distributions of the local data of different clients. Existing research has shown that non-independent and identically distributed data will seriously affect the effect of the model obtained by distributed learning. Most of the current algorithms only focus on the case of independent and identically distributed data or have only conducted very limited research on non-independent and identically distributed data.

[0053] That is to say, the data digest transfer scheme based on the core set has the following two disadvantages:

[0054] (1) The generation of the core set depends on the selection of the clustering center, and representative samples need to exist in the original data set. Otherwise, the selected core set will not be able to represent the original data set well, which makes it impossible to guarantee the effectiveness of the data digest generated by this method.

[0055] (2) When the data distribution is non-independent and identically distributed, this type of method will randomly determine the local center, resulting in a significant reduction in the model effect.

[0056] If a data subset that can be shared among various clients is created to reduce the overall non-independent and identically distributed degree of the system and it is proved through experiments that this method can improve the accuracy of the model, however, the increase in communication cost by this method is unacceptable in some cases.

[0057] Therefore, without increasing communication costs, this application aims to solve the problems in the existing distributed learning method based on data summaries. In this method, due to data heterogeneity in the local data scattered in various client devices and the lack of pertinence in data summary selection, etc., the effectiveness and accuracy of the global model trained based on the data summaries respectively output by each client device are poor, and further the application effect and performance of the global model are significantly reduced. A data summary generation method is provided. By calculating a score for each client and performing hierarchical clustering on each client according to the value of the score, and then using the dataset condensation technology based on the hierarchical clustering result to obtain the synthetic dataset of each client to form a data summary. Further, considering that the number of samples in the synthetic dataset obtained by the existing dataset condensation technology is mainly determined by experience, if it is directly applied to calculate synthetic data on each client in a distributed learning scenario, since the number of samples, data balance degree, and data non-independent and identically distributed degree of the data on each client are different, setting the size of the synthetic sample dataset of each client to a fixed value will not yield ideal results. Therefore, this application considers improving the dataset condensation technology to adaptive dataset condensation and then applying it to the distributed scenario.

[0058] Specifically, it will be described in detail through the following various embodiments and application examples respectively.

[0059] To solve the problems in the existing distributed learning method based on data summaries. In this method, due to data heterogeneity in the local data scattered in various client devices and the lack of pertinence in data summary selection, etc., the effectiveness and accuracy of the global model trained based on the data summaries respectively output by each client device are poor, and further the application effect and performance of the global model are significantly reduced. An embodiment of a data summary generation method is provided in this application. Refer to Figure 1 , the data summary generation method executed by the data summary generation device specifically includes the following content:

[0060] Step 100: Adaptively generate a data summary corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs, where the local training data is distributed with multiple non-local training data, and the model training contribution level to which each training data belongs is obtained by performing hierarchical clustering on each training data in advance.

[0061] It can be understood that the local training data refers to the data local to the data summary generation device that currently executes the data summary generation method, and the non-local training data refers to the data local to other data summary generation devices respectively. That is to say, local or non-local is only based on the client device that currently executes this method. For example, the local training data S1 of client device A, together with the non-local training data S2, S3 to S100 of client device A, jointly form a distributed set of 100 training data S1 to S100. Among them, the training data S2 can be the local training data of client device B. At this time, the training data S1 is non-local training data for client device B.

[0062] Step 200: Output the data summary corresponding to the local training data to jointly train the global model based on the data summary corresponding to the local training data and the data summaries corresponding to multiple non-local training data respectively.

[0063] It can be understood that the global model can be a machine learning model for data prediction, recognition, classification, etc. For example, it can be a deep neural network model for face recognition, so as to improve the accuracy of face recognition of the deep neural network model jointly trained based on the data summary corresponding to the local training data and the data summaries corresponding to multiple non-local training data through the methods of step 100 and step 200.

[0064] In the embodiments of the present application, the execution subject of step 100 and step 200 is any client in the distributed learning system. That is to say, each client in the distributed learning system can execute step 100 and step 200 synchronously or asynchronously. Specifically, a client k in the distributed learning system receives the model parameters of the global model sent by the central server, and then this client k retrieves the model training contribution tier 2 stored in advance locally, in the central server, in a third-party server or in a database, etc. 2 ; then, client k adaptively generates the data summary N corresponding to the local training data according to the model parameters of the global model and the model training contribution tier k to which the local training data belongs; client k then sends the data summary N corresponding to the local training data k to the central server, so that the central server can train the global model according to these data summaries after receiving the data summaries sent by all the clients in the distributed learning system respectively.

[0065] As can be seen from the above description, the data digest generation method provided by the embodiments of the present application pre-performs hierarchical clustering on each piece of training data to generate the model training contribution levels to which the training data corresponding to each distributed client belongs. Then, each client adaptively generates a data digest corresponding to its local training data according to the model training contribution level to which its local training data belongs, which can effectively improve the pertinence and adaptability of generating the data digest, and can improve the effectiveness and reliability of the data digest in the distributed learning process, and can reduce the communication overhead in the distributed learning process and ensure the privacy and security of the training data; furthermore, it can effectively improve the performance of the global model trained using the data digest in the scenario where the client data is not independently and identically distributed, and improve the application effectiveness and accuracy of the global model.

[0066] To effectively improve the application reliability of hierarchical clustering, in an embodiment of the data digest generation method provided by the present application, refer to Figure 2 , before step 100 of the data digest generation method, the following specific content is further included:

[0067] Step 010: Obtain all the feature parameters corresponding to the local training data.

[0068] Step 020: Determine the weight corresponding to each of the feature parameters based on the entropy weight method, and determine the model training contribution score of the local training data according to the weight corresponding to each of the feature parameters.

[0069] Step 030: Output the model training contribution score of the local training data to summarize the model training contribution scores corresponding to each piece of the training data, and perform hierarchical clustering on each piece of the training data based on the model training contribution scores corresponding to each piece of the training data to determine the model training contribution level to which each piece of the training data belongs.

[0070] In a specific application scenario of step 030, if client k retrieves the model training contribution level tier pre-stored locally 2, the client k can pre - perform hierarchical clustering on all the client training data in the distributed learning system locally. The specific method can be as follows. Assuming that the client k is the execution entity, after calculating the model training contribution scores of the local training data based on steps 010 and 020, the client k can send the model training contribution scores of the local training data to a local aggregation module. Then, it receives the model training contribution scores of other training data sent by other client devices and stores these model training contribution scores in the aggregation module as well. Then, the client k performs hierarchical clustering on each piece of the training data in the aggregation module based on the model training contribution scores corresponding to each piece of the training data to determine the model training contribution levels to which each piece of the training data belongs. Then, it sends the model training contribution levels to which the other training data belong to other clients respectively, and stores the model training contribution level to which the local training data belongs locally. Using this method, there is no need to add a new third - party server, etc. in the current distributed learning system. Only one or more existing client devices are selected for hierarchical clustering, which can effectively save hardware costs.

[0071] In another specific application scenario of step 030, if the client k retrieves or receives the model training contribution level tier pre - stored in a third - party server or a central server 2 , the third - party server or the central server can pre - perform hierarchical clustering on all the client training data in the distributed learning system locally. The specific method can be as follows. Assuming that the third - party server is the execution entity, after the client k calculates the model training contribution scores of the local training data based on steps 010 and 020, it sends the model training contribution scores of the local training data to the third - party server. Then, after the third - party server confirms that it has received the model training contribution scores of the training data sent by all the clients in the distributed learning system, it aggregates these model training contribution scores. Then, the third - party server performs hierarchical clustering on each piece of the training data based on the model training contribution scores corresponding to each piece of the training data to determine the model training contribution levels to which each piece of the training data belongs. Then, it stores locally the model training contribution levels to which each piece of the training data belongs and / or sends the model training contribution levels to which each piece of the training data belongs to each client respectively. Using this method, it is convenient for maintenance, data monitoring, and management based on an independent third - party device, which can effectively avoid problems such as poor accuracy of the hierarchical clustering result caused by incorrect calculation of individual clients, providing an accurate and reliable data basis for subsequent adaptive dataset condensation according to the hierarchical clustering result.

[0072] In order to effectively improve the comprehensiveness and effectiveness of the application of feature parameters, in an embodiment of the data summary generation method provided in this application, the feature parameters in the data summary generation method specifically include: the number of samples, the degree of non-independent and identically distributed, and the balance offset degree; the expression of the degree of non-independent and identically distributed EMD is as follows:

[0073]

[0074] where C is the number of categories in the classification task.

[0075] The expression of the balance offset degree BD is as follows:

[0076]

[0077] Correspondingly, step 020 of the data summary generation method specifically includes the following content:

[0078] Step 021: Input the number of samples, the degree of non-independent and identically distributed, and the balance offset degree corresponding to the local training data into a preset entropy weight comprehensive evaluation model, so that the entropy weight comprehensive evaluation model outputs the weights corresponding to the number of samples, the degree of non-independent and identically distributed, and the balance offset degree respectively.

[0079] In step 021, all clients calculate the feature parameters such as the degree of non-independent and identically distributed and the balance offset degree of the local data respectively, and then input these feature parameters into the entropy weight comprehensive evaluation model to obtain the weights corresponding to the number of samples, the degree of non-independent and identically distributed, and the balance offset degree of the client.

[0080] In a specific example, the entropy weight comprehensive evaluation model can be an entropy Topsis comprehensive analysis model, where entropy Topsis refers to combining the characteristics of the entropy value method and the Topsis method for evaluation. The Topsis method is a method for evaluating the quality levels of each sample by approaching the ideal solution. In the normalized original data matrix, the optimal solution and the worst solution among the limited solutions are found, and then the distances between the evaluation object and the optimal solution and the worst solution are calculated respectively, and these are used as the basis to evaluate the quality levels of the samples.

[0081] In order to construct a cross-client unified metric standard based on these three factors, we use the entropy Topsis method to calculate the weights of each factor and obtain a local data score S for each client k , representing the suitability of the local data of client k for the final model training. According to this parameter, we can divide the participating clients into multiple different layers, for example, at least three layers: {tier 1 , tier 2 , …, tier N}, where tier 1 the client in is most suitable for global model training; while tier N the client in is least suitable for global model training, which will affect the accuracy of the global model.

[0082] Step 022: Determine the model training contribution score of the local training data according to the weights corresponding to the sample quantity, the degree of non-independent and identically distributed, and the balance offset degree respectively.

[0083] In step 022, weighted processing can be performed on the sample quantity, the degree of non-independent and identically distributed, and the balance offset degree to obtain the model training contribution score of the local training data and the client score S k , and this score S k can represent the contribution degree of the local data of client k to the global model training.

[0084] In order to further improve the pertinence and self-adaptability of generating the data summary corresponding to the local training data, in an embodiment of the data summary generation method provided in this application, see Figure 3 , step 100 of the data summary generation method specifically includes the following content:

[0085] Step 110: Determine the sample quantity of the synthetic dataset corresponding to the local training data according to the model training contribution level to which the local training data belongs to initialize the synthetic dataset.

[0086] Step 120: Generate the data summary corresponding to the local training data based on the preset adaptive data condensation algorithm for the received model parameters of the global model.

[0087] In step 110 and step 120, the central server initializes the global model parameters and distributes them to all participating clients. Each client initializes the synthetic dataset N according to its position in the hierarchical structure k , and then uses the data condensation algorithm to implement the mapping of the local dataset D k →N k .

[0088] Specifically, the adaptive dataset condensation algorithm is improved from the dataset condensation algorithm, and can adaptively adjust the condensation process according to the position of the client in the hierarchical structure to be applicable to the non-independent and identically distributed scenario. The purpose of dataset condensation is to obtain a dataset N with a very small sample quantity k , so that it can achieve a training effect comparable to the original dataset D k for a specific model.

[0089] In order to further improve the application reliability and effectiveness of the adaptive data condensation algorithm, in an embodiment of the data summary generation method provided in this application, step 120 of the data summary generation method specifically includes the following content:

[0090] Step 121: For the model parameters of the received global model, train a first neural network of the target parameter according to the local training data, and train a second neural network of the target parameter based on the initialized synthetic data set;

[0091] Step 122: Converge the second neural network according to the solution of the first neural network to obtain a converged synthetic data set, and determine the converged synthetic data set as the data summary corresponding to the local training data.

[0092] Specifically, assume that the local data of client k consists of |D| pairs of training data and labels. This application uses this data to train a first neural network with parameters θ D Then θ D can be calculated by the following formula of the first neural network:

[0093]

[0094] where l is the loss function. The optimization goal is to obtain a synthetic data set N, which contains much fewer sample numbers than the original data set D. The parameter θ N obtained by training on the synthetic data set can be calculated by the following formula of the second neural network:

[0095]

[0096] This application requires that θ N and θ D are comparable. To achieve this goal, in the parameter space, the model obtained by training on the data set N should converge to a solution approximately the same as the model obtained by training on the data set D. Therefore, the training goal of data condensation can be expressed by the following formula:

[0097]

[0098]

[0099] where d(·, ·) is the distance function.

[0100] In this way, the synthetic sample set N of each client can be calculated k . It should be noted that the data set N k is the synthetic summary used to train a specific deep neural network, rather than the real sample D ka subset within. Therefore, during the transmission of these synthetic data samples, the customer's original data will not be exposed, fully protecting the customer's privacy.

[0101] All participating clients calculate a synthetic sample set N of their local data k and send it to the central server. The central server collects these samples and conducts training, and finally distributes the trained global model to each client, and the distributed algorithm process ends here.

[0102] To solve the problems in the existing distributed learning method based on data summaries, such as data heterogeneity in the local data scattered in each client device and the lack of pertinence in data summary selection, etc., which easily lead to poor effectiveness and accuracy of the global model trained based on the data summaries output by each client device respectively, and further reduce the application effect and performance of the global model significantly. This application provides an embodiment of a model training method. Refer to Figure 4 The model training method executed by the model training device specifically includes the following content:

[0103] Step 300: Output the model parameters of the global model, and adaptively generate the data summaries corresponding to each piece of the training data respectively based on the model parameters and the model training contribution levels to which each piece of the distributed training data belongs, where the model training contribution levels to which each piece of the training data belongs are obtained by pre-stratified clustering of each piece of the training data respectively.

[0104] Step 400: Receive the data summaries corresponding to each piece of the distributed training data respectively, and jointly train the global model based on the data summaries corresponding to each piece of the training data respectively.

[0105] Among them, step 300 is executed before step 100, and step 400 is executed after step 200.

[0106] As can be seen from the above description, the model training method provided by the embodiment of this application, by outputting the model parameters of the global model, adaptively generates the data summaries corresponding to each piece of the training data respectively based on the model parameters and the model training contribution levels to which each piece of the distributed training data belongs, where the model training contribution levels to which each piece of the training data belongs are obtained by pre-stratified clustering of each piece of the training data respectively, can effectively improve the pertinence and self-adaptability of generating data summaries, and can improve the effectiveness and reliability of data summaries in the distributed learning process, and can reduce the communication overhead in the distributed learning process and ensure the privacy and security of training data; and further can effectively improve the performance of the global model trained using data summaries in the scenario where client data is not independently and identically distributed, and improve the application effectiveness and accuracy of the global model.

[0107] Based on the above, the present application also provides a data digest generation device for implementing the data digest generation method provided in one or more embodiments of the present application. The specific implementation of this data digest generation device can be a client. In a specific example, refer to Figure 5 , the data digest generation device specifically includes the following:

[0108] A digest generation module 10, configured to adaptively generate a data digest corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs, where the local training data is distributed with multiple non-local training data, and the model training contribution level to which each piece of the training data belongs is obtained by performing hierarchical clustering on each piece of the training data separately in advance;

[0109] A digest output module 20, configured to output the data digest corresponding to the local training data to jointly train the global model based on the data digest corresponding to the local training data and the data digests corresponding to multiple non-local pieces of the training data respectively.

[0110] The embodiments of the data digest generation device provided by the present application can specifically be used to execute the processing flow of the embodiments of the data digest generation method in the above embodiments, and its functions will not be elaborated here. Reference can be made to the detailed description of the embodiments of the above data digest generation method.

[0111] It can be understood that in one or more embodiments of the present application, the client may specifically refer to a client device. The client device may include any mobile device capable of installing a DAPP, such as a smart phone, a tablet electronic device, a network set-top box, a portable computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Any mobile device capable of loading an application. Among them, the smart wearable device may include smart glasses, smart watches, smart bracelets, etc.

[0112] The above-mentioned client device may have a communication module (i.e., a communication unit), and may be communicatively connected to a central server, a third-party server, or a database.

[0113] As can be seen from the above description, the data digest generation device provided in the embodiments of the present application pre-performs hierarchical clustering on each piece of training data to generate the model training contribution levels to which the training data respectively corresponding to each distributed client belong. Then, each client adaptively generates a data digest corresponding to its local training data according to the model training contribution level to which its local training data belongs, which can effectively improve the pertinence and adaptability of generating data digests, and can improve the effectiveness and reliability of data digests in the distributed learning process, and can reduce the communication overhead in the distributed learning process and ensure the privacy and security of training data; furthermore, it can effectively improve the performance of the global model trained using data digests in the scenario where the client data is not independently and identically distributed, and improve the application effectiveness and accuracy of the global model.

[0114] Based on the above content, the present application further provides a model training device for implementing the model training method provided in one or more embodiments of the present application. The specific implementation of this model training device can be a central server. In a specific example, refer to Figure 6 , the model training device specifically includes the following:

[0115] A parameter sending module 30, configured to output the model parameters of the global model, so as to adaptively generate data digests respectively corresponding to each piece of the training data based on the model parameters and the model training contribution levels to which each piece of the distributed training data belongs, wherein the model training contribution levels to which each piece of the training data belongs are obtained by pre-performing hierarchical clustering on each piece of the training data respectively;

[0116] A summary training module 40, configured to receive the data digests respectively corresponding to each piece of the distributed training data, and jointly train the global model based on the data digests respectively corresponding to each piece of the training data.

[0117] The embodiments of the model training device provided in the present application can specifically be used to execute the processing flow of the embodiments of the model training method in the above embodiments, and its functions will not be elaborated here, and reference can be made to the detailed description of the embodiments of the above model training method.

[0118] As can be seen from the above description, the model training device provided in the embodiments of the present application outputs the model parameters of the global model, and adaptively generates data summaries corresponding to each piece of the training data respectively based on the model parameters and the model training contribution levels to which each piece of the distributed training data belongs. Among them, the model training contribution levels to which each piece of the training data belongs are obtained by pre-performing hierarchical clustering on each piece of the training data respectively, which can effectively improve the pertinence and adaptability of generating data summaries, and can improve the effectiveness and reliability of data summaries in the distributed learning process, and can reduce the communication overhead in the distributed learning process and ensure the privacy and security of training data; furthermore, it can effectively improve the performance of the global model trained using data summaries in the scenario where client data is not independently and identically distributed, and improve the application effectiveness and accuracy of the global model.

[0119] Based on the above embodiments of the data summary generation device and the model training device, or the embodiments of the data summary generation method and the model training method, the present application also provides an embodiment of a distributed learning system. Refer to Figure 7 , the distributed learning system specifically includes the following:

[0120] A central server and each client communicatively connected to the central server respectively;

[0121] The central server is used to execute the functions of the model training method or the model training device mentioned in one or more embodiments of the present application;

[0122] The client is used to execute the functions of the data summary generation method or the data summary generation device mentioned in one or more embodiments of the present application.

[0123] The embodiment of the distributed learning system provided by the present application can specifically be used to execute the processing flows of the data summary generation device and the model training device in the above embodiments, or the data summary generation method and the model training method. Its functions will not be elaborated here, and reference can be made to the detailed descriptions of the above device or method embodiments.

[0124] As can be seen from the above description, the distributed learning system provided in the embodiments of the present application can effectively improve the pertinence and adaptability of generating data summaries, and can improve the effectiveness and reliability of data summaries in the distributed learning process, and can reduce the communication overhead in the distributed learning process and ensure the privacy and security of training data; furthermore, it can effectively improve the performance of the global model trained using data summaries in the scenario where client data is not independently and identically distributed, and improve the application effectiveness and accuracy of the global model.

[0125] To further illustrate the present solution, the present application also provides a distributed learning algorithm based on hierarchical clustering and adaptive dataset condensation. Considering that the current distributed learning system based on data summaries mainly adopts the method of calculating the core sets of each client to implement, and cannot well replace the original dataset. Since the calculated core sets are universal, when training for a specific network structure, they cannot well replace the original dataset, and the global model cannot achieve the best effect. At the same time, due to the lack of special treatment for the non-independent and identically distributed degree of each client, the data heterogeneity problem commonly existing in the distributed learning scenario will have a great negative impact on the training results.

[0126] Therefore, the purpose of the application example of the present application is to propose a distributed learning algorithm based on hierarchical clustering and adaptive dataset condensation on the basis of the traditional distributed learning system that transmits data summaries. A comprehensive entropy Topsis analysis model is proposed. Based on the hierarchical model, an improved dataset condensation technology is used to adaptively obtain the synthetic dataset of each client. The central server collects these synthetic datasets and trains the global model. This training method comprehensively considers factors such as the sample number, non-independent and identically distributed degree, and balance deviation degree of the clients, and can effectively improve the performance of the model in the scenario where the data of each client is non-independent and identically distributed. And what the client transmits to the server is the synthetic small sample dataset instead of the subset of the original dataset, ensuring the communication efficiency and privacy security of the training process.

[0127] See Figure 8 , the application example of the present application proposes a distributed learning algorithm based on hierarchical clustering and adaptive dataset condensation. The application example of the present application relies on the traditional distributed learning architecture and deep neural network frameworks (such as PyTorch, TensorFlow) for model training. A client hierarchical clustering model is added during the training process, and then the data in each client is subjected to adaptive dataset condensation according to the hierarchical structure, improving the performance of the algorithm in the scenario where the data of each node is non-independent and identically distributed on the basis of ensuring low communication overhead and data privacy security.

[0128] The application example of the present application mainly consists of two important components: the comprehensive entropy Topsis analysis model responsible for obtaining the client hierarchical structure and the adaptive dataset condensation module responsible for calculating the data summaries of each client. Generally speaking, the algorithm mainly consists of the following three steps:

[0129] S1: Hierarchical clustering. All clients respectively calculate the characteristic parameters such as the non-independent and identically distributed degree and balance offset degree of the local data, and then input these characteristic parameters into the comprehensive entropy Topsis analysis model to obtain the client score S k , and this score can represent the contribution degree of the local data of client k to the training of the global model.

[0130] Specifically, in the hierarchical model, the application instances of the present application mainly consider the following factors to hierarchically classify each client:

[0131] (1) The number of samples in the client.

[0132] (2) The degree of non-independent and identically distributed EMD of the local data of the client, and the expression is as follows:

[0133]

[0134] where C is the number of categories in the classification task.

[0135] (3) The balance offset BD of the local data of the client, and the expression is as follows:

[0136]

[0137] In order to construct a unified metric across clients based on these three factors, we use the entropy Topsis method to calculate the weights of each factor and obtain a local data score S for each client k , which represents the suitability of the local data of client k for the final model training. According to this parameter, we can divide the participating clients into multiple different layers, such as at least three layers: {tier 1 , tier 2 , …, tier N}, where the clients in tier 1 are the most suitable for the global model training; while the clients in tier N are the least suitable for the global model training, which will affect the accuracy of the global model.

[0138] S2: Global model download and adaptive dataset condensation. The central server initializes the global model parameters and distributes them to all participating clients. Each client initializes the synthetic dataset N k according to its position in the hierarchical structure, and then uses the data condensation algorithm to implement the mapping of the local dataset D k → N k .

[0139] Specifically, assuming that the local data of client k consists of |D| pairs of training data and labels, the present application uses this data to train a first neural network with parameters θ D , then θ D can be calculated by the following formula of the first neural network:

[0140]

[0141] Among them, l is the loss function. The optimization goal is to obtain a synthetic dataset N, which contains a much smaller number of samples than the original dataset D. The parameter θ obtained by training on the synthetic dataset N , can be calculated by the following formula of the second neural network:

[0142]

[0143] This application requires that θ N and θ D be comparable. To achieve this goal, in the parameter space, the model obtained by training on the dataset N should converge to an approximate solution of the model obtained by training on the dataset D. Therefore, the training objective of dataset condensation can be expressed by the following formula:

[0144]

[0145]

[0146] where d(·, ·) is the distance function.

[0147] In this way, the synthetic sample set N of each client can be calculated k . It should be noted that the dataset N k is the synthetic summary for training a specific deep neural network, rather than a subset of the real samples D k . Therefore, during the transmission of these synthetic data samples, the original data of the clients will not be exposed, fully protecting the privacy of the clients.

[0148] S3: Dataset aggregation and model training. The central server collects the synthetic datasets of all clients and trains the global model, and finally distributes the trained model to all clients participating in the training.

[0149] Specifically, all participating clients calculate the synthetic sample set N of their local data k and send it to the central server. The central server collects these samples and conducts training, and finally distributes the trained global model to each client, and the distributed algorithm process ends here.

[0150] Based on this, in order to solve the data heterogeneity problem of distributed learning and improve the final accuracy of the global model on the premise of meeting low communication overhead and data privacy security, the main improvements of the application example of this application are as follows:

[0151] 1. A novel distributed learning algorithm based on hierarchical clustering and adaptive dataset condensation is proposed. A hierarchical model is constructed according to the characteristic parameters of each client, and the data on each client is adaptively condensed according to the hierarchical model. The larger the sample size, the lower the degree of non-independent and identically distributed, and the lower the balance deviation. The larger the sample size of the synthetic dataset obtained by the nodes with lower balance deviation through the adaptive condensation algorithm, and vice versa, so as to control the contribution degree of each node to the global model.

[0152] 2. An entropy Topsis comprehensive analysis model is proposed to establish the hierarchical structure of each client. The entropy Topsis method is used to determine the weights of the three factors that affect the sample size of the synthetic dataset, and each client is weighted to determine its contribution degree to the global model, so that the nodes with a high degree of non-independent and identically distributed have a smaller contribution to the global model training.

[0153] 3. An efficient adaptive dataset condensation algorithm is proposed, aiming to calculate a dataset with a small sample size for a specific network and enable it to obtain a training effect comparable to the original dataset. First, the size of the synthetic dataset is determined according to the client hierarchical model, the synthetic dataset is initialized, and the optimization goal is set to minimize the distance between the two sets of gradients obtained from the original dataset and the synthetic dataset respectively. The sample size of the synthetic dataset calculated by this algorithm is between 1% and 5% of the original dataset size, greatly reducing the communication overhead and realizing the privacy and security of data.

[0154] Based on this, the application example of this application proposes a novel distributed learning algorithm based on hierarchical clustering and adaptive dataset condensation. Compared with the traditional distributed learning algorithm, this algorithm focuses on the performance in the non-independent and identically distributed scenario, adds a hierarchical module for comprehensively analyzing node characteristic parameters, fully considers the differences between the local data of each client, changes their contribution degrees to the global training, and is applicable to the situation with a high degree of data heterogeneity and complexity. At the same time, the application example of this application overcomes the limitations of the original algorithm by improving the dataset condensation technology, enables it to adaptively adjust the process of dataset condensation, is applicable to the distributed scenario, and effectively improves the accuracy of the global model.

[0155] At the hardware level, this application provides an embodiment of an electronic device for implementing all or part of the content in the data summary generation method or the model training method. The electronic device specifically includes the following content:

[0156] Figure 9 It is a schematic block diagram of the system composition of the electronic device 9600 according to the embodiment of this application. As Figure 9As shown, the electronic device 9600 may include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It is noted that the Figure 9 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0157] In one embodiment, the data digest generation function may be integrated into the central processor. Among them, the central processor may be configured to perform the following controls:

[0158] Step 100: According to the model parameters of the received global model and the model training contribution level to which the local training data belongs, adaptively generate a data digest corresponding to the local training data, where the local training data is distributed with multiple non-local training data, and the model training contribution level to which each piece of the training data belongs is obtained by performing hierarchical clustering on each piece of the training data separately in advance.

[0159] Step 200: Output the data digest corresponding to the local training data to jointly train the global model based on the data digest corresponding to the local training data and the data digests corresponding to multiple non-local pieces of the training data.

[0160] From the above description, it can be seen that the electronic device provided in the embodiment of the present application may be a client.

[0161] In one embodiment, the model training function may be integrated into the central processor. Among them, the central processor may be configured to perform the following controls:

[0162] Step 300: Output the model parameters of the global model to adaptively generate data digests corresponding to each piece of the training data respectively based on the model parameters and the model training contribution levels to which the distributed pieces of the training data belong, where the model training contribution level to which each piece of the training data belongs is obtained by performing hierarchical clustering on each piece of the training data separately in advance.

[0163] Step 400: Receive the data digests corresponding to the distributed pieces of the training data respectively, and jointly train the global model based on the data digests corresponding to the distributed pieces of the training data respectively.

[0164] From the above description, it can be seen that the electronic device provided in the embodiment of the present application may be a central server.

[0165] In another embodiment, the data digest generation or model training device can be separately configured from the central processing unit 9100. For example, the data digest generation or model training device can be configured as a chip connected to the central processing unit 9100, and the data digest generation or model training function can be realized through the control of the central processing unit.

[0166] As Figure 9 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include Figure 9 all the components shown in Figure 9 ; in addition, the electronic device 9600 may further include

[0167] As Figure 9 shown, the central processing unit 9100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.

[0168] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above information related to failures can be stored, and in addition, programs for executing relevant information can also be stored. And the central processing unit 9100 can execute the programs stored in the memory 9140 to implement information storage or processing, etc.

[0169] The input unit 9120 provides inputs to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display can be, for example, an LCD display, but is not limited thereto.

[0170] The memory 9140 can be a solid-state memory. For example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It can also be such a memory that stores information even when powered off, can be selectively erased and has more data. Examples of this memory are sometimes referred to as EPROMs, etc. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 can include an application / function storage section 9142, and the application / function storage section 9142 is used to store application programs and functional programs or the processes for operating the electronic device 9600 through the central processing unit 9100.

[0171] The memory 9140 may further include a data storage unit 9143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (such as a messaging application, an address book application, etc.).

[0172] The communication module 9110 is a transmitter / receiver 9110 that transmits and receives signals via the antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as in the case of a conventional mobile communication terminal.

[0173] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module (transmitter / receiver) 9110 is also coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby implementing normal telecommunication functions. The audio processor 9130 may include any suitable buffers, decoders, amplifiers, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.

[0174] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps in the data summary generation method or the model training method in the above embodiments. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all steps in the data summary generation method or the model training method in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0175] Step 100: Adaptively generate a data summary corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs, where the local training data is distributed with multiple non-local training data, and the model training contribution level to which each piece of the training data belongs is obtained by performing hierarchical clustering on each piece of the training data separately in advance.

[0176] Step 200: Output the data summary corresponding to the local training data, and jointly train the global model based on the data summary corresponding to the local training data and the data summaries corresponding to multiple non-local training data respectively.

[0177] For another example, when the processor executes the computer program, the following steps are implemented:

[0178] Step 300: Output the model parameters of the global model, and adaptively generate the data summaries corresponding to the respective training data based on the model parameters and the model training contribution levels to which the respective training data in the distributed setting belong, where the model training contribution levels to which the respective training data belong are obtained by performing hierarchical clustering on the respective training data in advance.

[0179] Step 400: Receive the data summaries corresponding to the respective training data in the distributed setting, and jointly train the global model based on the data summaries corresponding to the respective training data.

[0180] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0181] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatuses), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes.

[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in Figure 1 one process or more processes and / or boxes Figure 1 the functions specified in one box or more boxes.

[0184] In the present invention, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only for helping to understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for generating a data digest, characterized in that, it includes: According to the model parameters of the received global model and the model training contribution level to which the local training data belongs, adaptively generate a data digest corresponding to the local training data, wherein the local training data is distributed with multiple non-local training data, and the model training contribution level to which each piece of training data belongs is obtained by respectively performing hierarchical clustering on each piece of training data in advance; Output the data digest corresponding to the local training data to jointly train the global model based on the data digest corresponding to the local training data and the data digests corresponding to each of the multiple non-local training data; The step of adaptively generating a data digest corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs includes: Determine the number of samples of the synthetic data set corresponding to the local training data according to the model training contribution level to which the local training data belongs to initialize the synthetic data set; For the model parameters of the received global model, train a first neural network for a target parameter according to the local training data, and train a second neural network for the target parameter based on the initialized synthetic data set; Converge the second neural network according to the solution of the first neural network to obtain a converged synthetic data set, and determine the converged synthetic data set as the data digest corresponding to the local training data.

2. The data digest generation method according to claim 1, characterized in that, before the step of adaptively generating a data digest corresponding to the local training data according to the model parameters of the received global model and the model training contribution level to which the local training data belongs, it further includes: Obtain all the feature parameters corresponding to the local training data; Determine the weight corresponding to each of the feature parameters based on the entropy weight method, and determine the model training contribution score of the local training data according to the weight corresponding to each of the feature parameters; Output the model training contribution score of the local training data to summarize the model training contribution scores corresponding to each piece of training data, and perform hierarchical clustering on each piece of training data based on the model training contribution scores corresponding to each piece of training data to determine the model training contribution level to which each piece of training data belongs.

3. The data digest generation method according to claim 2, characterized in that, the feature parameters include: the number of samples, the degree of non-independent and identically distributed, and the balance offset degree; Correspondingly, the step of determining the weight corresponding to each of the feature parameters based on the entropy weight method includes: Input the number of samples, the degree of non-independent and identically distributed, and the balance offset degree corresponding to the local training data into a preset entropy weight comprehensive evaluation model, so that the entropy weight comprehensive evaluation model respectively outputs the weights corresponding to the number of samples, the degree of non-independent and identically distributed, and the balance offset degree.

4. A model training method, characterized in that, it includes: Output the model parameters of the global model, and adaptively generate data summaries corresponding to each piece of the training data based on the model parameters and the model training contribution levels to which each piece of the distributed training data belongs, where the model training contribution levels to which each piece of the training data belongs are obtained by performing hierarchical clustering on each piece of the training data in advance; Receive the data summaries corresponding to each piece of the training data set up distributively, and jointly train the global model based on the data summaries corresponding to each piece of the training data; Among them, the step of adaptively generating data summaries corresponding to each piece of the training data based on the model parameters and the model training contribution levels to which each piece of the distributed training data belongs includes: Determine the number of samples of the synthetic data set corresponding to the local training data according to the model training contribution level to which the local training data of the data summary generation device belongs, so as to initialize the synthetic data set; For the received model parameters of the global model, train a first neural network with a target parameter according to the local training data, and train a second neural network with the target parameter based on the initialized synthetic data set; Converge the second neural network according to the solution of the first neural network to obtain a converged synthetic data set, and determine the converged synthetic data set as the data summary corresponding to the local training data.

5. A data summary generation device Characterized in that It includes: A summary generation module, configured to adaptively generate a data summary corresponding to the local training data according to the received model parameters of the global model and the model training contribution level to which the local training data belongs, where the local training data is distributively set with multiple pieces of non-local training data, and the model training contribution levels to which each piece of the training data belongs are obtained by performing hierarchical clustering on each piece of the training data in advance; A summary output module, configured to output the data summary corresponding to the local training data to jointly train the global model based on the data summary corresponding to the local training data and the data summaries corresponding to multiple pieces of non-local training data respectively; The step of adaptively generating a data summary corresponding to the local training data according to the received model parameters of the global model and the model training contribution level to which the local training data belongs includes: Determine the number of samples of the synthetic data set corresponding to the local training data according to the model training contribution level to which the local training data belongs, so as to initialize the synthetic data set; For the received model parameters of the global model, train a first neural network with a target parameter according to the local training data, and train a second neural network with the target parameter based on the initialized synthetic data set; Converge the second neural network according to the solution of the first neural network to obtain a converged synthetic data set, and determine the converged synthetic data set as the data summary corresponding to the local training data.

6. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, Characterized in that When the processor executes the computer program, it implements the data digest generation method described in any one of claims 1 to 3, or when the processor executes the computer program, it implements the model training method described in claim 4.

7. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, it implements the data digest generation method described in any one of claims 1 to 3, or when the computer program is executed by a processor, it implements the model training method described in claim 4.

8. A distributed learning system, characterized in that, comprising: a central server and each client that is communicatively connected to the central server; the central server is configured to execute the model training method described in claim 4; the client is configured to execute the data digest generation method described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Federal learning-based model training method and device, and federal learning system

    CN112862011A

  • Distributed text summarization method based on cloud data center

    CN112883722A