Federal learning method and device using high-order sub-model to assist training

By dividing multi-level high-order submodels in federated learning, setting up early exit classifiers, and using comparative learning to build a quantitative matrix, the problem of insufficient inference accuracy and generalization ability of low-order submodels is solved, and more accurate knowledge transfer and stronger generalization ability are achieved.

CN119940475APending Publication Date: 2025-05-06BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411851758.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the heterogeneous federated learning of systems, the inference accuracy and generalization ability of low-order submodels are not as good as those of high-order submodels, and the decision-making information assisted training of low-order submodels has knowledge transfer bias, and inconsistent feature dimensions lead to the problem of generalization ability improvement.

Method used

The global model is divided into multiple levels through the server. The client receives high-order submodels of different levels, sets up early exit classifiers, and uses the feature map of the highest-order submodel to build a quantitative matrix through comparison learning, assists in low-order submodel training, reduces the target loss of target category confidence, and optimizes the low-order submodel.

Benefits of technology

The reasoning accuracy and generalization ability of low-order submodels are improved, knowledge transfer deviations are avoided, feature dimensions of submodels at all levels are unified, and the computational cost of auxiliary training is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940475A_ABST
    Figure CN119940475A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning method and device using high-order sub-models to assist training, and the method comprises the steps: enhancing the reasoning accuracy of a low-order sub-model through integrating the decision information of multiple high-order sub-models; the low-order sub-model can assist in training by means of decision information of the high-order sub-model and a quantization matrix of the highest-order sub-model based on the feature distance, and the training effect is improved; a low-order sub-model is optimized by reducing target loss (information entropy) of target category confidence, and the problem of information offset caused by re-fitting of target category confidence of a high-order sub-model in traditional self-distillation is avoided; unifying feature dimensions of all levels of sub-models by utilizing comparative learning, constructing a quantization matrix of the highest-order sub-model based on feature distance as supervision information, and guiding the low-order sub-model to approach the quantization matrix of the highest-order sub-model by minimizing the distribution difference between the low-order sub-model and the distance matrix of the highest-order sub-model, so as to realize the supervision of the highest-order sub-model. And a low-order sub-model is guided to mine semantic relationships in different images, so that the generalization ability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and specifically relates to a federated learning method and device using high-order sub-models to assist training. Background Art

[0002] In federated learning with heterogeneous systems, the server meets the resource constraints of different clients by sending down multiple levels of sub-models, but the reasoning accuracy and generalization ability of low-order sub-models are not as good as those of high-order sub-models.

[0003] At present, the method of using a single highest-order sub-model decision information to assist in training low-order sub-models has knowledge transfer bias, so it is a challenge to fully utilize the decision information of multiple high-order sub-models to achieve more accurate knowledge transfer; the feature dimensions extracted by sub-models at each level are inconsistent, which makes it a challenge to use the feature information of the highest-order sub-model to enhance the generalization ability of low-order sub-models. At the same time, the computing resources of the client vary greatly, and reducing the computational cost of auxiliary training is a challenge. How to improve the generalization of low-order sub-models is a problem that needs to be solved. Summary of the invention

[0004] In view of this, the purpose of the present invention is to propose a federated learning method and device that utilizes high-order sub-models to assist in training, so as to solve or partially solve the above-mentioned technical problems.

[0005] Based on the above objectives, the present invention provides a federated learning method using a high-order sub-model to assist training, comprising:

[0006] The server divides the global model into J levels and defines the high-order sub-model M of level j received by the client. j Contains the model parameters and structures of j sub-models, is the i-th sub-model M i Model parameters of

[0007] An early exit classifier is set at the end of the sub-model so that the client can obtain the inference results of each level of sub-model to assist training; the client's local nth training data x n Different outputs are produced after being processed by different levels of sub-models The low-order sub-model uses the true labels of the data Training, where low-order sub-models use the decision information of higher-order sub-models to optimize performance;

[0008] By using the feature map extracted by the highest-order sub-model in the training data, the distances between different feature maps are obtained through comparative learning and a quantization matrix is ​​constructed. The constructed quantization matrix assists the low-order sub-model in enhancing the generalization ability.

[0009] As a preferred solution for the federated learning method using high-order sub-models to assist training, for the nth training data x in the client n Target category Low-order submodel M i Target category confidence for:

[0010]

[0011] By reducing the target loss of target category confidence Optimize low-order sub-models.

[0012] As a preferred solution for the federated learning method using high-order sub-models to assist training, for the nth training data x in the client n The set of non-target categories Non-target confidence obtained by each sub-model for:

[0013]

[0014] The confidences of the two non-target categories are normalized Nor(·) respectively, and the difference between the two distributions is obtained as the non-target loss of self-distillation under the guidance of the high-order sub-model.

[0015] As a preferred solution for federated learning methods that use high-order sub-models to assist training, the target loss and non-target loss As the overall loss of the low-order sub-model

[0016]

[0017] In the formula, Represents the low-order sub-model M i Target category confidence Perform normalization; Represents the low-order sub-model M j Target category confidence Perform normalization;

[0018] Low-order submodel M i By outputting sub-models with higher levels than itself As weak supervision.

[0019] As a preferred solution for the federated learning method using high-order sub-models to assist training, in the process of extracting feature graphs from training data using the highest-order sub-model:

[0020] A batch of data Input network, through the i-th level sub-model Mi Feature extraction layer After that, it is mapped into feature maps of different dimensions

[0021] As a preferred solution for the federated learning method using high-order sub-models to assist training, in the process of obtaining the distances of different feature maps and constructing the quantization matrix through comparative learning:

[0022] Calculate each feature map separately in the feature space of the i-level sub-model With feature map The distance vector between

[0023] The distance vectors of each feature map are concatenated and normalized to form the distance quantization matrix Q of the i-level sub-model i .

[0024] As a preferred solution for the federated learning method using high-order sub-models to assist in training, the constructed quantization matrix assists the low-order sub-models in enhancing their generalization capabilities:

[0025] By minimizing the low-order submodel M i With the highest order submodel M j Distribution differences of distance matrices The low-order sub-models are guided to approach the quantization matrix of the highest-order sub-model, and the low-order sub-models are guided to mine the semantic relationships in different images.

[0026] The present invention also provides a federated learning device using a high-order sub-model to assist in training, comprising:

[0027] The parameter configuration module is used to divide the global model into J levels through the server and define the high-order sub-model M of level j received by the client. j Contains the model parameters and structures of j sub-models, is the i-th sub-model M i Model parameters of

[0028] The sub-model performance optimization module is used to set an early exit classifier at the end of the sub-model so that the client can obtain the inference results of each level of sub-model to assist training; the client's local n-th training data x n Different outputs are produced after being processed by different levels of sub-models The low-order sub-model uses the true labels of the data Training, where low-order sub-models use the decision information of higher-order sub-models to optimize performance;

[0029] The sub-model generalization enhancement module is used to utilize the feature maps extracted by the highest-order sub-model in the training data, obtain the distances between different feature maps through comparative learning and construct a quantization matrix, and assist the low-order sub-model in enhancing its generalization capability through the constructed quantization matrix.

[0030] As a preferred solution for a federated learning device using a high-order sub-model to assist training, in the sub-model performance optimization module, for the nth local training data x n Target category Low-order submodel M i Target category confidence for:

[0031]

[0032] By reducing the target loss of target category confidence Optimize low-order sub-models;

[0033] In the sub-model performance optimization module, for the nth local training data x on the client, n The set of non-target categories Non-target confidence obtained by each sub-model for:

[0034]

[0035] The confidences of the two non-target categories are normalized Nor(·) respectively, and the difference between the two distributions is obtained as the non-target loss of self-distillation under the guidance of the high-order sub-model.

[0036] In the sub-model performance optimization module, the target loss and non-target loss As the overall loss of the low-order sub-model

[0037]

[0038] In the formula, Represents the low-order sub-model M i Target category confidence Perform normalization; Represents the low-order sub-model M j Target category confidence Perform normalization;

[0039] Low-order submodel M i By outputting sub-models with higher levels than itself As weak supervision.

[0040] As a preferred solution for a federated learning device using a high-order sub-model to assist in training, in the sub-model generalization enhancement module:

[0041] A batch of data Input network, through the i-th level sub-model M i Feature extraction layer After that, it is mapped into feature maps of different dimensions

[0042] In the sub-model generalization enhancement module:

[0043] Calculate each feature map separately in the feature space of the i-level sub-model With feature map The distance vector between

[0044] The distance vectors of each feature map are concatenated and normalized to form the distance quantization matrix Q of the i-level sub-model i ;

[0045] In the sub-model generalization enhancement module:

[0046] By minimizing the low-order submodel M i With the highest order submodel M j Distribution differences of distance matrices The low-order sub-models are guided to approach the quantization matrix of the highest-order sub-model, and the low-order sub-models are guided to mine the semantic relationships in different images.

[0047] From the above, it can be seen that the technical solution provided by the present invention enhances the reasoning accuracy of low-order sub-models by integrating the decision information of multiple high-order sub-models; the low-order sub-model can assist in training with the help of the decision information of the high-order sub-model and the quantization matrix of the highest-order sub-model based on the feature distance to improve the training effect; the low-order sub-model is optimized by reducing the target loss (information entropy) of the target category confidence, thereby avoiding the information deviation problem caused by re-fitting the target category confidence of the high-order sub-model in the traditional self-distillation; the feature dimensions of sub-models at all levels are unified by contrastive learning, and the quantization matrix of the highest-order sub-model based on the feature distance is constructed as the supervision information, and the distribution difference between the distance matrix of the low-order sub-model and the highest-order sub-model is minimized to guide the low-order sub-model to approach the quantization matrix of the highest-order sub-model, thereby guiding the low-order sub-model to mine the semantic relationships in different images and improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0049] Figure 1 A schematic diagram of a flow chart of a federated learning method using a high-order sub-model to assist training provided in an embodiment of the present invention;

[0050] Figure 2 A schematic diagram of a federated learning method using high-order sub-models to assist in training, wherein decision information of multiple high-order sub-models guides the joint assistance training of low-order sub-models in an embodiment of the present invention;

[0051] Figure 3 A schematic diagram of auxiliary training of low-order sub-models guided by feature information of the highest-order sub-model in a federated learning method using high-order sub-models for auxiliary training provided by an embodiment of the present invention;

[0052] Figure 4 An architecture diagram of a federated learning device using a high-order sub-model to assist training provided by an embodiment of the present invention;

[0053] Figure 5 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0055] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The words "include" or "comprise" and the like used in the embodiments of the present invention mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, but do not exclude other elements or objects.

[0056] In federated learning with heterogeneous systems, the server sends down multiple levels of sub-models to meet the resource constraints of different clients, but the reasoning accuracy and generalization ability of low-order sub-models are not as good as those of high-order sub-models. When sub-models have multiple exit points, the solution to this problem is to use the performance advantages of high-order sub-models to assist in optimizing low-order sub-models, so as to ensure that low-order sub-models deployed on small clients can have better performance. At the same time, high-order sub-models can exit early during reasoning to improve reasoning speed.

[0057] The high-order sub-models received by the big clients contain the structure and parameter information of the low-order sub-models, so these clients can assist in the training of the low-order sub-models when training the high-order sub-models. However, directly sharing some parameters between different models cannot effectively transfer the effective information in the high-order sub-models to the low-order sub-models. Knowledge distillation technology can reduce the difference in reasoning ability between high-order sub-models (teacher models) and low-order sub-models (student models). Its core idea is to use the reasoning logic value of the teacher model as a soft label to guide the student model. The traditional knowledge distillation method used in federated learning needs to rely on the global model as a guide to assist in training the client model, which brings additional forward propagation costs. Under the premise that the server sends a high-order sub-model with multiple exit points, self-distillation is used to distill the knowledge from the deep layer of the model to the shallow layer, thereby improving the accuracy of the sub-model at a low cost. The existing auxiliary training methods only transfer the knowledge of a single highest-order sub-model, and there is a knowledge transfer bias. Therefore, correcting the knowledge bias of auxiliary training is a problem that needs to be solved to achieve accurate guidance of low-order sub-models.

[0058] High-order sub-models can capture more effective data features through more complex network structures. Therefore, with the assistance of high-order sub-models, the generalization ability of low-order sub-models can be improved by maximizing the similarity between samples of the same category. In related technologies, methods based on contrastive learning can enhance the generalization ability of models, but they are limited to measuring distance and loss through feature maps of the same dimension. When the feature dimensions between models do not match, such methods fail. Therefore, when the output feature dimensions of sub-models at each level are different, improving the generalization of low-order sub-models is a problem that needs to be solved. The method of using a single highest-order sub-model decision information to assist in training low-order sub-models has knowledge transfer bias, so it is a challenge to make full use of multiple high-order sub-model decision information to achieve more accurate knowledge transfer; the feature dimensions extracted by sub-models at each level are inconsistent, which makes it a challenge to use the feature information of the highest-order sub-model to enhance the generalization ability of low-order sub-models. At the same time, the computing resources of clients vary greatly, and it is a challenge to reduce the computational cost of auxiliary training.

[0059] In view of this, in order to improve the inference accuracy and speed of low-order sub-models, correct the knowledge bias of auxiliary training to achieve precise guidance of low-order sub-models, and improve the generalization of low-order sub-models, an embodiment of the present invention provides a federated learning method and device using high-order sub-models to assist in training. The following is the specific content of the embodiment of the present invention.

[0060] See also Figure 1 , an embodiment of the present invention provides a federated learning method using a high-order sub-model to assist training, comprising the following steps:

[0061] S1. Divide the global model into J levels through the server and define the high-order sub-model M of level j received by the client jContains the model parameters and structures of j sub-models, is the i-th sub-model M i Model parameters of

[0062] S2. Set an early exit classifier at the end of the sub-model so that the client can obtain the inference results of each level of sub-model to assist training; the client's local n-th training data x n Different outputs are produced after being processed by different levels of sub-models The low-order sub-model uses the true labels of the data Training, where low-order sub-models use the decision information of higher-order sub-models to optimize performance;

[0063] S3. Using the feature map extracted by the highest-order sub-model in the training data, the distances between different feature maps are obtained through comparative learning and a quantization matrix is ​​constructed. The constructed quantization matrix is ​​used to assist the low-order sub-model in enhancing the generalization ability.

[0064] See also Figure 2 ,In this embodiment, a federated learning scenario with heterogeneous systems is simulated. Clients with sufficient system resources will train high-order ,sub-models, and clients with limited system resources will train low-order ,sub-models. The low-order sub-models can assist in training with the ,decision information of high-order sub-models to improve the training effect.

[0065] Specifically, the server divides the global model into J levels, and the client receives the j-level high-order sub-model M j contains the model parameters and structures of j sub-models, is the i-th sub-model M i If an early exit classifier is set at the end of the sub-model, the client will also get the inference results of each sub-model for auxiliary training. n Different outputs will be generated after being processed by different levels of sub-models In this case, the low-order sub-model can not only utilize the true labels of the data Training can also be performed while further utilizing the decision information of higher-order sub-models to optimize performance.

[0066] Among them, for the nth local training data x on the client n Target category (the true category of the data), the confidence of the high-order sub-model for the sample output is stable during training. However, the low-order sub-model cannot give a high-confidence inference result of the target category, so the difference in inference confidence between different data is small. By expanding the confidence difference of the low-order sub-model for each sample, the knowledge transfer to the low-order sub-model is accelerated.

[0067] Specifically, for the nth training data x on the client siden Target category Low-order submodel M i Target category confidence for:

[0068]

[0069] By reducing the target loss of target category confidence Optimize low-order sub-models by reducing the information entropy (target loss) of the target category confidence when calculating the inference bias The low-order sub-model is optimized to avoid the information deviation problem caused by refitting the target category confidence of the high-order sub-model in traditional self-distillation.

[0070] Among them, for the nth local training data x on the client n The set of non-target categories The non-target confidence obtained by each sub-model is but The inequality prevents the output distribution of the low-order sub-model from being close to that of the high-order sub-model. In this embodiment, the confidence of the two non-target categories is normalized Nor(·) respectively, and the difference between the two distributions is obtained as the non-target loss of self-distillation under the guidance of the high-order sub-model. The above target loss and non-target loss are used as the overall loss of the low-order sub-model This optimizes the model.

[0071] Among them, the target loss and non-target loss As the overall loss of the low-order sub-model

[0072]

[0073] In the formula, Represents the low-order sub-model M i Target category confidence Perform normalization; Represents the low-order sub-model M j Target category confidence Normalize.

[0074] For high-order sub-models containing multiple early exit classifiers, although the inference results of sub-models at each level do not have the high accuracy advantage of the highest-order sub-model, they still have certain information content and reference value. i By outputting sub-models with higher levels than itself As weak supervision, multiple models are used for ensemble reasoning, and the loss value is calculated as follows:

[0075]

[0076] Therefore, additional supervision information can be introduced during the training process, alleviating the bias caused by a single model and improving the training effect.

[0077] See also Figure 3 In this embodiment, a federated learning scenario with heterogeneous systems is simulated. Clients with sufficient system resources will train high-order sub-models, and clients with limited system resources will train low-order sub-models. The feature maps extracted by the highest-order sub-model in the training data are used to obtain the distances between different feature maps through comparative learning and construct a quantization matrix to assist the low-order sub-model in enhancing its generalization ability.

[0078] In this embodiment, in the process of extracting feature graphs from training data using the highest-order sub-model:

[0079] A batch of data Input network, through the i-th level sub-model M i Feature extraction layer After that, it is mapped into feature maps of different dimensions

[0080] The different dimensions of the feature graphs make it impossible for the lower-order sub-models to directly learn the feature extraction information of the highest-order sub-model. This embodiment measures the distance between feature graphs by comparative learning, that is, calculating the distance between each feature graph in the feature space of the i-level sub-model. With other feature maps The distance vector between The distance vectors of each feature map are concatenated and normalized to form the distance quantization matrix Q of the i-level sub-model i At this time, the feature maps of different dimensions of each sub-model are converted into matrices of the same dimension.

[0081] Among them, the highest-order sub-model has a stronger feature extraction capability due to its complex structure, so the distance of the feature map obtained can more accurately reflect the relationship between the data. i With the highest order submodel M j Distribution differences of distance matrices Guide the low-order sub-model to approach the quantization matrix of the highest-order sub-model, guide the low-order sub-model to mine the semantic relations in different images, and improve the generalization ability of the model.

[0082] In summary, the embodiment of the present invention divides the global model into J levels through the server, and defines the j-level high-order sub-model M received by the client. j Contains the model parameters and structures of j sub-models, is the i-th sub-model M iThe model parameters; set the early exit classifier at the end of the sub-model so that the client can obtain the inference results of each level of sub-model to assist training; the client local nth training data x n Different outputs are produced after being processed by different levels of sub-models The low-order sub-model uses the true labels of the data Training, the low-order sub-model optimizes performance by using the decision information of the higher-order sub-model; using the feature map extracted by the highest-order sub-model in the training data, the distance of different feature maps is obtained by contrastive learning and a quantization matrix is ​​constructed, and the constructed quantization matrix is ​​used to assist the low-order sub-model in enhancing the generalization ability. The present invention enhances the reasoning accuracy of the low-order sub-model by integrating the decision information of multiple high-order sub-models; the low-order sub-model can assist in training with the decision information of the high-order sub-model and the quantization matrix of the highest-order sub-model based on the feature distance to improve the training effect; the low-order sub-model is optimized by reducing the target loss (information entropy) of the target category confidence, avoiding the information offset problem caused by the target category confidence of the high-order sub-model being fitted again in the traditional self-distillation; the feature dimensions of sub-models at all levels are unified by contrastive learning, and the quantization matrix of the highest-order sub-model based on the feature distance is constructed as the supervision information, and the distribution difference between the low-order sub-model and the highest-order sub-model distance matrix is ​​minimized, and the low-order sub-model is guided to approach the quantization matrix of the highest-order sub-model, and the low-order sub-model is guided to mine the semantic relationship in different images, thereby improving the generalization ability of the model.

[0083] It should be noted that the method of the embodiment of the present invention can be performed by a single device, such as a computer or a server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present invention, and the multiple devices will interact with each other to complete the described method.

[0084] It should be noted that some embodiments of the present invention are described above. In some cases, the actions or steps recorded can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0085] See also Figure 4 Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, an embodiment of the present invention further provides a federated learning device using a high-order sub-model to assist in training, including:

[0086] The parameter configuration module 100 is used to divide the global model into J levels through the server and define the high-order sub-model M of the j-level received by the client. j Contains the model parameters and structures of j sub-models, is the i-th sub-model M i Model parameters of

[0087] The sub-model performance optimization module 200 is used to set an early exit classifier at the end of the sub-model so that the client can obtain the inference results of each level of sub-model to assist training; the client local nth training data x n Different outputs are produced after being processed by different levels of sub-models The low-order sub-model uses the true labels of the data Training, where low-order sub-models use the decision information of higher-order sub-models to optimize performance;

[0088] The sub-model generalization enhancement module 300 is used to utilize the feature graph extracted by the highest-order sub-model in the training data, obtain the distance between different feature graphs through comparative learning and construct a quantization matrix, and assist the low-order sub-model in enhancing the generalization ability through the constructed quantization matrix.

[0089] In this embodiment, in the sub-model performance optimization module 200, for the nth local training data x of the client n Target category Low-order submodel M i Target category confidence for:

[0090]

[0091] By reducing the target loss of target category confidence Optimize low-order sub-models;

[0092] In the sub-model performance optimization module 200, for the nth local training data x of the client, n The set of non-target categories Non-target confidence obtained by each sub-model for:

[0093]

[0094] The confidences of the two non-target categories are normalized Nor(·) respectively, and the difference between the two distributions is obtained as the non-target loss of self-distillation under the guidance of the high-order sub-model.

[0095] In the sub-model performance optimization module 200, the target loss and non-target loss As the overall loss of the low-order sub-model

[0096]

[0097] In the formula, Represents the low-order sub-model M i Target category confidence Perform normalization; Represents the low-order sub-model M j Target category confidence Perform normalization;

[0098] Low-order submodel M i By outputting sub-models with higher levels than itself As weak supervision.

[0099] In this embodiment, in the sub-model generalization enhancement module 300:

[0100] A batch of data Input network, through the i-th level sub-model M i Feature extraction layer After that, it is mapped into feature maps of different dimensions

[0101] In the sub-model generalization enhancement module 300:

[0102] Calculate each feature map separately in the feature space of the i-level sub-model With feature map The distance vector between

[0103] The distance vectors of each feature map are concatenated and normalized to form the distance quantization matrix Q of the i-level sub-model i ;

[0104] In the sub-model generalization enhancement module 300:

[0105] By minimizing the low-order submodel M i With the highest order submodel M j Distribution differences of distance matrices The low-order sub-models are guided to approach the quantization matrix of the highest-order sub-model, and the low-order sub-models are guided to mine the semantic relationships in different images.

[0106] The device of the above embodiment is used to implement a corresponding federated learning method using a high-order sub-model to assist training in any of the aforementioned embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0107] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a federated learning method using a high-order sub-model to assist training as described in any of the above embodiments is implemented.

[0108] Figure 5 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 410, a memory 420, an input / output interface 430, a communication interface 440, and a bus 450. The processor 410, the memory 420, the input / output interface 430, and the communication interface 440 are connected to each other through the bus 450 in the device.

[0109] The processor 410 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0110] The memory 420 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 420 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 420 and are called and executed by the processor 410.

[0111] The input / output interface 430 is used to connect the input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0112] The communication interface 440 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0113] The bus 450 comprises a pathway for transmitting information between the various components of the device (eg, the processor 410, the memory 420, the input / output interface 430, and the communication interface 440).

[0114] It should be noted that, although the above device only shows the processor 410, the memory 420, the input / output interface 430, the communication interface 440 and the bus 450, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0115] The electronic device of the above-mentioned embodiment is used to implement a corresponding federated learning method using a high-order sub-model to assist training in any of the aforementioned embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0116] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present invention also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute a federated learning method using high-order sub-models to assist training as described in any of the above embodiments.

[0117] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0118] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute a federated learning method using high-order sub-models to assist training as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0119] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention is limited to these examples. Under the concept of the present invention, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of simplicity.

[0120] In addition, to simplify the description and discussion, and in order not to obscure the embodiments of the present invention, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the system may be shown in the form of a block diagram so as to avoid obscuring the embodiments of the present invention, and this also takes into account the fact that the details of the implementation of these block diagram systems are highly dependent on the platform on which the embodiments of the present invention will be implemented (i.e., these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present invention, it will be apparent to those skilled in the art that embodiments of the present invention may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0121] Although the present invention has been described in conjunction with specific embodiments of the present invention, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0122] The embodiments of the present invention are intended to cover all such substitutions, modifications and variations that fall within the scope of the protection claimed. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the protection scope of the present invention.

Claims

1. A federated learning method using high-order sub-models to assist training, where: include: The server divides the global model into J levels and defines the high-order sub-model M of level j received by the client. j Contains the model parameters and structures of j sub-models, is the i-th sub-model M i Model parameters of An early exit classifier is set at the end of the sub-model so that the client can obtain the inference results of each level of sub-model to assist training; the client's local nth training data x n Different outputs are produced after being processed by different levels of sub-models The low-order sub-model uses the true labels of the data Training, where low-order sub-models use the decision information of higher-order sub-models to optimize performance; By using the feature map extracted by the highest-order sub-model in the training data, the distances between different feature maps are obtained through comparative learning and a quantization matrix is ​​constructed. The constructed quantization matrix assists the low-order sub-model in enhancing the generalization ability.

2. A federated learning method using high-order sub-models to assist training according to claim 1, wherein: For the nth piece of training data x on the client side n Target category Low-order submodel M i Target category confidence for: By reducing the target loss of target category confidence Optimize low-order sub-models.

3. A federated learning method using high-order sub-models to assist training according to claim 2, wherein: For the nth piece of training data x on the client side n The set of non-target categories Non-target confidence obtained by each sub-model for: The confidences of the two non-target categories are normalized Nor(·) respectively, and the difference between the two distributions is obtained as the non-target loss of self-distillation under the guidance of the high-order sub-model.

4. A federated learning method using high-order sub-models to assist training according to claim 3, wherein: The target loss and non-target loss As the overall loss of the low-order sub-model In the formula, Represents the low-order sub-model M i Target category confidence Perform normalization; Represents the low-order sub-model M j Target category confidence Perform normalization; Low-order submodel M i By outputting sub-models with higher levels than itself As weak supervision.

5. A federated learning method using high-order sub-models to assist training according to claim 4, wherein: Using the feature map extracted from the training data by the highest order sub-model: A batch of data Input network, through the i-th level sub-model M i Feature extraction layer After that, it is mapped into feature maps of different dimensions 6. A federated learning method using high-order sub-models to assist training according to claim 5, wherein: In the process of obtaining the distance of different feature maps and constructing the quantization matrix through comparative learning: Calculate each feature map V in the feature space of the i-level sub-model i x′ With feature map The distance vector between The distance vectors of each feature map are concatenated and normalized to form the distance quantization matrix Q of the i-level sub-model i .

7. A federated learning method using high-order sub-models to assist training according to claim 6, wherein: The constructed quantization matrix assists the low-order sub-model in enhancing the generalization ability: By minimizing the low-order submodel M i With the highest order submodel M j Distribution differences of distance matrices The low-order sub-models are guided to approach the quantization matrix of the highest-order sub-model, and the low-order sub-models are guided to mine the semantic relationships in different images.

8. A federated learning device using a high-order sub-model to assist training, wherein: include: The parameter configuration module is used to divide the global model into J levels through the server and define the high-order sub-model M of level j received by the client. j Contains the model parameters and structures of j sub-models, is the i-th sub-model M i Model parameters of The sub-model performance optimization module is used to set an early exit classifier at the end of the sub-model so that the client can obtain the inference results of each level of sub-model to assist training; the client's local n-th training data x n Different outputs are produced after being processed by different levels of sub-models The low-order sub-model uses the true labels of the data Training, where low-order sub-models use the decision information of higher-order sub-models to optimize performance; The sub-model generalization enhancement module is used to utilize the feature maps extracted by the highest-order sub-model in the training data, obtain the distances between different feature maps through comparative learning and construct a quantization matrix, and assist the low-order sub-model in enhancing its generalization capability through the constructed quantization matrix.

9. A federated learning device using high-order sub-models to assist training according to claim 8, wherein: In the sub-model performance optimization module, for the nth local training data x on the client, n Target category Low-order submodel M i Target category confidence for: By reducing the target loss of target category confidence Optimize low-order sub-models; In the sub-model performance optimization module, for the nth local training data x on the client, n The set of non-target categories Non-target confidence obtained by each sub-model for: The confidences of the two non-target categories are normalized Nor(·) respectively, and the difference between the two distributions is obtained as the non-target loss of self-distillation under the guidance of the high-order sub-model. In the sub-model performance optimization module, the target loss and non-target loss As the overall loss of the low-order sub-model In the formula, Represents the low-order sub-model M i Target category confidence Perform normalization; Represents the low-order sub-model M j Target category confidence Perform normalization; Low-order submodel M i By outputting sub-models with higher levels than itself As weak supervision.

10. A federated learning device using high-order sub-models to assist training according to claim 9, wherein: In the sub-model generalization enhancement module: A batch of data Input network, through the i-th level sub-model M i Feature extraction layer After that, it is mapped into feature maps of different dimensions In the sub-model generalization enhancement module: Calculate each feature map V in the feature space of the i-level sub-model i x′ With feature map The distance vector between The distance vectors of each feature map are concatenated and normalized to form the distance quantization matrix Q of the i-level sub-model i ; In the sub-model generalization enhancement module: By minimizing the low-order submodel M i With the highest order submodel M j Distribution differences of distance matrices The low-order sub-models are guided to approach the quantization matrix of the highest-order sub-model, and the low-order sub-models are guided to mine the semantic relationships in different images.

Citation Information

Cited By

  • Federal learning-based edge end model updating method and device, equipment and medium

    CN120124781A