Heterogeneous Federated Learning Method and Device Based on Structure-Aware Information Optimization and Adaptive Aggregation

By adopting structure-aware information optimization and adaptive aggregation methods in heterogeneous federated learning, the problem of low accuracy brought about by data and model heterogeneity is solved, and higher federated learning accuracy and personalized performance are achieved.

CN119670852BActive Publication Date: 2025-06-17ZHEJIANG GONGSHANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510180617.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-17
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

When the existing heterogeneous federated learning method faces data heterogeneity and model heterogeneity between clients, it is difficult to effectively improve the accuracy of federated learning, and fails to fully consider the personalized needs of clients.

Method used

The heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation is adopted. By building a complex multi-scale model, structure-aware information is extracted and gradient adjustment factors are generated to enhance the feature extraction capability of the client model. At the same time, personalized feature representations are extracted using an autoencoder and uploaded to the server through weighted average to optimize the global prediction head.

Benefits of technology

It significantly improves the accuracy of federated learning in heterogeneous data and model environments, enhances the generalization ability and personalized performance of client models, and solves the challenges of data and model heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670852B_ABST
    Figure CN119670852B_ABST
Patent Text Reader

Abstract

Heterogeneous Federated Learning Method and Device Based on Structure-Aware Information Optimization and Adaptive Aggregation. The method includes: The client constructs a complex multi-scale model and trains it using local data; obtains the structure-aware information of the multi-scale model and forms a gradient adjustment factor; the client incorporates the gradient adjustment factor into the local model for training, and extracts the features of all data through a feature extractor; the client extracts more personalized feature representations and records the corresponding labels, and after weighted averaging the features of each label, uploads them to the server together with the corresponding labels; the server receives the average features and corresponding labels of each client in a step, and trains the server's global prediction head; the server broadcasts the updated global prediction head to the clients, and each client's prediction head adaptively aggregates the global prediction head based on local data; repeat multiple rounds of training, and calculate the test metrics in each round; each client selects the optimal model parameters, imports them, and inputs the image to complete the classification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of federated learning and privacy protection, and relates to a heterogeneous federated learning method and device based on structure-aware information optimization and adaptive aggregation. Background Art

[0002] With the increasing demand for data privacy, federated learning, as a machine learning method that enables collaborative model training without centralized data, has gradually become a research hotspot. The core idea of federated learning is to train models locally on each client, ensuring that the data always remains local, and aggregating the trained models to a central server, thereby improving model performance while protecting data privacy. This feature makes federated learning show great application potential in highly privacy-sensitive fields such as healthcare and finance. However, in current practical applications, federated learning methods usually face the actual challenges of data heterogeneity and model heterogeneity among clients. Data heterogeneity is mainly manifested in that the data of each client is non-independent and identically distributed, while model heterogeneity refers to the fact that each client may adopt different model architectures due to reasons such as hardware and task requirements. The challenges brought by these heterogeneities make traditional federated learning methods such as FedAvg and FedProx difficult to be directly applicable to these complex scenarios, because they usually assume that all clients share the same model architecture and the data distribution is consistent.

[0003] Therefore, a series of heterogeneous federated learning methods have emerged in recent years, aiming to address the challenges of data and model heterogeneity. However, existing heterogeneous federated learning methods still have some defects in dealing with model and data heterogeneity. Current heterogeneous model federated learning algorithms usually use existing simple and general model architectures. However, general model architectures cannot fully exploit the unique data features of each client, and blindly using complex models will reduce the training efficiency of federated learning. Secondly, many existing heterogeneous federated learning methods do not fully consider the personalized needs of clients. Since the data distribution and task requirements of each client are different, it is difficult to meet the performance requirements of all clients by replacing the local model with the model trained by the server. Summary of the Invention

[0004] Aiming at the above technical problems existing in the prior art, the present invention provides a heterogeneous federated learning method and device based on structure-aware information optimization and adaptive aggregation, aiming to improve the accuracy of federated learning in heterogeneous data and heterogeneous model environments.

[0005] The present invention adopts the following technical solutions:

[0006] The first aspect of the present invention relates to a heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation, including the following steps:

[0007] Structure-Aware Federated Learning Feature Enhancement Phase:

[0008] Step 1: Construct a complex multi-scale model and train it using local data;

[0009] Step 2: Obtain the structure-aware information of the multi-scale model in Step 1 and generate a gradient adjustment factor using the structure-aware information;

[0010] Step 3: Each client incorporates the gradient adjustment factor into the local model for training. After training is completed, each client extracts the features and labels of all data through a feature extractor;

[0011] Personalized Adaptive Aggregation Phase Based on Autoencoder Optimization:

[0012] Step 4: Each client uses the features and labels extracted in Step 3 as a dataset and uses an autoencoder to extract more personalized feature representations and corresponding labels; after each client performs weighted averaging on the features of each label, it uploads them to the server together with the corresponding labels;

[0013] Step 5: The server receives the average features and corresponding labels of each client in Step 4 and uses this data to train the server's global prediction head;

[0014] Step 6: After the server's global prediction head is trained, the server broadcasts the updated global prediction head to the clients; each client's prediction head adaptively aggregates the global prediction head based on local data;

[0015] Step 7: Repeat Steps 3 to 6 until a preset number of training rounds is reached, and calculate the test metrics for each round;

[0016] Step 8: Each client selects the optimal model parameters, imports them, and inputs the image to complete the classification task.

[0017] Furthermore, Step 1 includes:

[0018] Step 1-1: Construct a complex multi-scale model. For the layers that need to improve feature extraction ability locally, add an additional multi-scale feature enhancer MFE beside the original convolution, and add a scale regulator SC after the original convolution Conv and the newly added multi-scale feature enhancer MFE;

[0019] Step 1-2: After the multi-scale model is constructed, train it using local data.

[0020] Furthermore, Step 2 includes:

[0021] Extract the scale regulator parameter sd of the trained multi-scale model, which is the structure-aware information, and obtain the gradient adjustment factor GAF through calculation.

[0022] Furthermore, Step 3 includes:

[0023] Step 3-1: Incorporate the gradient adjustment factor into the local original model training, and adjust the gradient update process of the model to achieve a training process equivalent to that of a multi-scale model;

[0024] Step 3-2: After the training is completed, use the feature extractor to extract features from each training sample in the local dataset to generate features .

[0025] Furthermore, Step 4 includes:

[0026] Step 4-1: Each client uses an encoder to compress the local features generated from the local extractor to remove redundant information and noise;

[0027] Step 4-2: Restore the compressed features to the original feature dimension size through a decoder to ensure the dimensional consistency of the data uploaded to the client;

[0028] Step 4-3: Each client calculates the average features of all samples belonging to the same class, and uploads the average feature representation and the corresponding label to the server.

[0029] Furthermore, Step 5 includes:

[0030] The server receives the average features and the corresponding labels s from all participating clients, uses these as a dataset to train the server global prediction head; calculates the cross-entropy loss between the prediction result of the server global prediction head and the true class label s; then uses the gradient descent method to update the parameters of the global prediction head.

[0031] Furthermore, this Step 6 includes:

[0032] The server broadcasts the server global prediction head after the training is completed. According to the local data of each client, the local client prediction head adaptively aggregates the downloaded server global prediction head.

[0033] Furthermore, Step 7 includes:

[0034] Repeat Step 3 to Step 6 until a preset number of training rounds is reached, and calculate the global average accuracy for each round.

[0035] The method of the present invention is divided into two stages. The first stage is the structure-aware federated learning feature enhancement stage, which aims to improve the feature extraction ability of the client feature extractor without increasing the model parameters. The second stage is the personalized adaptive aggregation stage based on autoencoder optimization, which aims to improve the personalized ability of the client prediction head. By combining the structure-aware federated learning feature enhancement stage and the personalized adaptive aggregation stage based on autoencoder optimization, the present invention improves the accuracy of heterogeneous federated learning.

[0036] The second aspect of the present invention relates to a heterogeneous federated learning device based on structure-aware information optimization and adaptive aggregation, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation of the present invention.

[0037] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation of the present invention.

[0038] The working principle of the present invention is as follows: By combining structure-aware information optimization and adaptive aggregation technologies, the present invention realizes feature enhancement and personalized training of the client local model in the federated learning framework. Specifically, first, through multi-scale model construction and training, structure-aware information of the complex model is extracted, and a gradient adjustment factor is generated and incorporated into the local training process of the client, thereby significantly improving the feature extraction ability of the local model. In addition, more personalized feature representations are extracted through an autoencoder, and based on the weighted average of features and labels, the personalized information of the client is uploaded to the server side for optimizing the global prediction head. After the global prediction head is updated, each client further optimizes the local model through adaptive aggregation to better adapt to the local data distribution.

[0039] The innovation points of the present invention are as follows: By constructing a complex multi-scale model, extracting the structure-aware information of the model, and generating a gradient adjustment factor, the feature learning ability of the client local model is significantly enhanced, while avoiding the need to increase model parameters; An adaptive aggregation strategy is proposed: Using an autoencoder to extract personalized features and combining the dynamic weight aggregation method of the global prediction head and the local prediction head, personalized model optimization for heterogeneous data distributions is achieved, effectively improving the generalization ability and personalized performance of the client model.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. The present invention generates a gradient adjustment factor by introducing a complex multi-scale model and extracting its structure-aware information, effectively enhancing the feature extraction ability of the simple model on the client side. Without increasing the computational complexity, it enables the simple model to accurately capture complex features, thus improving the overall performance of federated learning.

[0042] 2. Through the personalized aggregation strategy based on autoencoders, the present invention enables each client to optimize its own feature representation according to the local data characteristics. The autoencoder can compress high-dimensional features to remove redundant information and restore key features through the decoder, thereby enhancing the expression ability and discrimination ability of the local model and significantly improving the personalized performance of the client. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a schematic flow diagram of the method of the present invention.

[0044] Figure 2 is a schematic diagram of constructing a multi-scale model of the method of the present invention.

[0045] Figure 3 is a schematic diagram of autoencoder feature compression and restoration of the method of the present invention.

[0046] Figure 4 is a schematic diagram of client prediction head aggregation of the method of the present invention.

[0047] Figure 5 is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0048] The following further describes the specific implementation manners of the present invention in conjunction with specific embodiments:

[0049] Embodiment 1

[0050] As Figure 1 shown, a heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation is proposed, including the following steps:

[0051] Structure-aware federated learning feature enhancement stage:

[0052] Step 1: Construct a complex multi-scale model and train it using local data;

[0053] Step 2: Obtain the structure-aware information of the multi-scale model in Step 1 and generate a gradient adjustment factor using the structure-aware information;

[0054] Step 3: Each client incorporates the gradient adjustment factor into the local model for training; after training is completed, each client extracts the features and labels of all data through the feature extractor;

[0055] Personalized Adaptive Aggregation Phase Optimized by Autoencoder:

[0056] Step 4: Each client uses the features and labels extracted in Step 3 as a dataset, and uses an autoencoder to extract more personalized feature representations and corresponding labels; after each client performs weighted averaging on the features of each label, it uploads them to the server together with the corresponding labels;

[0057] Step 5: The server receives the average features and corresponding labels of each client in Step 4, and uses this data to train the server's global prediction head;

[0058] Step 6: After the server's global prediction head is trained, the server broadcasts the updated global prediction head to the clients. Each client's prediction head adaptively aggregates the global prediction head based on local data;

[0059] Step 7: Repeat Steps 3 to 6 until the preset number of training rounds is reached, and calculate the test metrics for each round;

[0060] Step 8: Each client selects the optimal model parameters, imports them, and inputs the image to complete the classification task.

[0061] Step 1 specifically includes:

[0062] Step 1-1: In the structure-aware federated learning feature enhancement phase, such as Figure 2 , construct a complex multi-scale model. For the layers that need to improve the feature extraction ability locally, add an additional multi-scale feature enhancer MFE beside the original convolution, and add a scale regulator SC after the original convolution Conv and the newly added multi-scale feature enhancer MFE. The output formula after the construction is as follows:

[0063] (1)

[0064] Among them, represents the output of the j-th layer after passing through the constructed layer, X represents the input data, represents the scale regulator of the j-th layer's original convolution layer, represents the output of X passing through the j-th layer's original convolution, represents the output of X passing through the k-th multi-scale feature enhancer of the j-th layer, represents the k-th scale regulator of the j-th layer. Thus, the construction of the multi-scale model is completed. Subsequently, local data is used for training.

[0065] Step 1-2: After the construction of the multi-scale model is completed, use local data for training, calculate the cross-entropy loss between the prediction result of the multi-scale model and the true class label s. Then use the gradient descent method to update the parameters of the multi-scale model.

[0066] Step 2 specifically includes:

[0067] After the multi-scale model training is completed, extract the scale regulator parameter sd, that is, the structure-aware information, and generate the gradient adjustment factor GAF. For the multi-branch layer, we can calculate the gradient adjustment factor GAF through the following formula:

[0068] (2)

[0069] Where, represents the scale regulator parameter of the original convolution Conv, represents the scale regulator parameter of the multi-scale feature enhancer MFE.

[0070] For the layer with only a single convolutional branch, assuming its corresponding scale regulator parameter is , then the calculation formula of its GAF is:

[0071] (3)

[0072] Where, represents the scale regulator parameter of the convolution Conv of the layer with only a single convolutional branch.

[0073] Step 3 specifically includes:

[0074] Step 3-1: Incorporate the gradient adjustment factor into the local original model training, and adjust the gradient update process of the model to achieve a training process equivalent to that of the multi-scale model. The formula for the incorporated gradient update is:

[0075] (4)

[0076] Where is the model parameter after the th iteration, is the model parameter at the i-th iteration, is the learning rate, is the gradient adjustment factor, and L is the loss function. In the process of federated learning, each client has its own local gradient adjustment factor.

[0077] Step 3-2: After the training is completed, use the feature extractor to extract features from each training sample in the local dataset , and generate the feature .

[0078] Step 4 specifically includes:

[0079] Step 4-1: As Figure 3, each client uses an encoder to compress the local features generated by the local extractor to remove redundant information and noise. This compression process can be expressed as

[0080] (5)

[0081] where, represents the i-th compressed feature of the k-th client in the t-th round, represents the i-th unoptimized feature of the k-th client in the t-th round, represents the encoding operation of the k-th client in the t-th round.

[0082] Step 4-2: As Figure 3 , the compressed features are restored to the original feature dimension size through a decoder to ensure the dimensional consistency of the data uploaded to the client. The restoration process can be expressed as:

[0083] (6)

[0084] where, represents the i-th restored feature of the k-th client in the t-th round, represents the decoding operation of the k-th client in the t-th round, represents the i-th compressed feature of the k-th client in the t-th round.

[0085] Step 4-3: Each client calculates the average feature of all samples belonging to the same class, and uploads the average feature and the corresponding label to the server. The formula for calculating the average feature is as follows

[0086] (7)

[0087] where, represents the average feature, represents the class of the sample set. The client then uploads these optimized average features to the server.

[0088] Step 5 specifically includes:

[0089] The server receives the average features and the corresponding labels s from all participating clients, and uses these as a dataset to train the server's global prediction head. The number of training rounds for the server is 100 rounds. Calculate the cross-entropy loss between the prediction result of the server's global prediction head and the true class label s. Then use the gradient descent method to update the parameters of the global prediction head. The specific update formula is:

[0090] (8)

[0091] Among them, represents the global prediction head parameter, is the learning rate of the global prediction head, is the cross-entropy loss function, represents the local average representation uploaded by the k-th client in the t-th round of federated learning training, and s is the corresponding label. After training, the server's global prediction head has full-category knowledge obtained from datasets across multiple clients, and has stronger generalization ability compared to the client prediction heads that only contain local knowledge.

[0092] Step 6 specifically includes:

[0093] Each client receives the server's global prediction head after training, and according to the local data of each client, the local client prediction head adaptively aggregates the downloaded server's global prediction head. We initialize the local prediction head through the weighted adaptive aggregation of the corresponding elements of the global prediction head and the local prediction head . The aggregation calculation formula is:

[0094] (9)

[0095] Among them, represents the local prediction head parameter, represents the clustering weight of each element. Each client updates iteratively through feature . The update formula is as follows

[0096] (10)

[0097] Among them, represents the clustering weight, is the learning rate for learning the weight, is the local prediction head parameter, is the global prediction head parameter, is the cross-entropy loss function. After completing the update , the personalized parameters of the client prediction head can be calculated.

[0098] Step 7 specifically includes:

[0099] Repeat steps 3 to 6 until the preset number of training rounds (the preset number of training rounds is 1000 rounds) is reached, record the model parameters of each client, and calculate the global average accuracy rate for each round. The formula is as follows:

[0100] (11)

[0101] Among them, denotes the average accuracy, and N denotes the total number of images of all clients. denotes the number of images correctly predicted by the i-th client.

[0102] Step 8 specifically includes:

[0103] Select the model parameters of each client in the round with the optimal global average accuracy metric, import the model parameters into each client model, and complete image classification by inputting images.

[0104] Next, the specific data and scenario parameters of the present invention are described. The present invention constructs two data heterogeneity scenarios: pathological settings and practical settings. In the pathological settings, we allocate a dataset with non-overlapping categories and imbalanced data volumes to each client. Each client obtains data of 2 / 10 / 10 / 20 categories from the CIFAR-10, CIFAR-100, Flowers102, and Tiny-ImageNet datasets respectively. This setting simulates a heterogeneous training environment with significant differences in client data. In the practical settings, we use the Dirichlet distribution to allocate data to each client to simulate the data distribution in the real world. Specifically, a distribution ratio is drawn from the Dirichlet distribution, and the data points of the category in the dataset are allocated to client i according to the ratio.

[0105] To simulate the heterogeneous federated learning scenario, we set up an environment consisting of 20 clients with a client participation rate of 1. To simulate different federated learning scenarios, we adjusted the number of clients and the participation rate in subsequent experiments. Each client's dataset is divided into a training set and a test set in a 3:1 ratio. Table 1 shows the accuracy comparison between the present invention and other methods on the pathological settings dataset, and Table 2 shows the accuracy comparison between the present invention and other methods on the practical settings dataset. The results show that the accuracy of the present invention is better than other methods.

[0106] Table 1 Performance comparison under pathological settings

[0107]

[0108] Table 2 Performance comparison under practical settings

[0109]

[0110] Example 2

[0111] This example relates to a face recognition method applying the heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation of the present invention, as Figure 1As shown, the steps include:

[0112] Step 1: Construct a complex multi-scale model and train it using local data. Specifically:

[0113] Step 1-1: Each client prepares a face dataset and, in the structure-aware federated learning feature enhancement phase, such as Figure 2 , constructs a complex multi-scale model. For the layers that need to improve the feature extraction ability locally, an additional multi-scale feature enhancer MFE is added beside the original convolution, and a scale regulator SC is added after the original convolution Conv and the newly added multi-scale feature enhancer MFE. The output formula after construction is as follows:

[0114] (12)

[0115] Among them, represents the output of the j-th layer after the construction layer, X represents the input data, represents the scale regulator of the j-th layer's original convolution layer, represents the output of X passing through the j-th layer's original convolution, represents the output of X passing through the k-th multi-scale feature enhancer of the j-th layer, represents the k-th scale regulator of the j-th layer. Thus, the construction of the multi-scale model is completed. Subsequently, it is trained using local data.

[0116] Step 1-2: After the construction of the multi-scale model is completed, it is trained using local data, and the cross-entropy loss between the prediction result of the multi-scale model and the true class label s is calculated. Then, the parameters of the multi-scale model are updated using the gradient descent method.

[0117] Step 2: Obtain the structure-aware information of the multi-scale model in Step 1 and generate a gradient adjustment factor using the structure-aware information. Specifically:

[0118] After the multi-scale model is trained, extract the scale regulator parameters, that is, the structure-aware information, and generate a gradient adjustment factor GAF. For multi-branch layers, we can calculate the gradient adjustment factor GAF through the following formula:

[0119] (13)

[0120] Among them, represents the scale regulator parameters of the original convolution Conv, represents the scale regulator parameters of the multi-scale feature enhancer MFE.

[0121] For a layer with only a single convolution branch, assume its corresponding scale regulator parameter is , then the calculation formula of its GAF is as follows:

[0122] (14)

[0123] Among them, represents the scale regulator parameter of the convolution Conv with only a single convolutional branch layer.

[0124] Step 3: Each client incorporates the gradient adjustment factor into the local model for training. After training is completed, each client extracts the features and labels of all data through the feature extractor. Specifically:

[0125] Step 3-1: Incorporate the gradient adjustment factor into the local original model training, and adjust the gradient update process of the model to achieve a training process equivalent to that of a multi-scale model. The formula for the incorporated gradient update is:

[0126] (15)

[0127] Among them is the model parameter after the th iteration, the model parameter at the i-th iteration, is the learning rate, is the gradient adjustment factor, and L is the loss function. During the federated learning process, each client has its own local gradient adjustment factor.

[0128] Step 3-2: After training is completed, use the feature extractor to extract features from each training sample in the local dataset , generating features .

[0129] Step 4: Each client uses the features and labels extracted in Step 3 as the dataset, and uses the autoencoder to extract more personalized feature representations and corresponding labels; after each client performs weighted averaging on the features of each label, it uploads them to the server together with the corresponding labels. Specifically:

[0130] Step 4-1: As Figure 3 , each client uses the encoder to compress the local features generated from the local extractor to remove redundant information and noise. This compression process can be expressed as

[0131] (16)

[0132] Among them, represents the i-th compressed feature of the k-th client in the t-th round, represents the i-th unoptimized feature of the k-th client in the t-th round, Denotes the encoding operation of the k-th client in the t-th round.

[0133] Step 4-2: As Figure 3 , the compressed features are restored to the original feature dimension size through the decoder to ensure the dimensional consistency when uploading to the client. The restoration process can be expressed as:

[0134] (17)

[0135] Where, Denotes the i-th restored feature of the k-th client in the t-th round, Denotes the decoding operation of the k-th client in the t-th round, Denotes the i-th compressed feature of the k-th client in the t-th round.

[0136] Step 4-3: Each client calculates the average feature of all samples belonging to the same class, and uploads the average feature and the corresponding label to the server. The formula for calculating the average feature is as follows

[0137] (18)

[0138] Where, Denotes the average feature, Denotes the class Of the sample set. The client then uploads these optimized average features To the server.

[0139] Step 5: The server receives the average features and the corresponding labels of each client in Step 4, and uses this data to train the server's global prediction head. Specifically:

[0140] The server receives the average features From all participating clients and the corresponding label s, and uses this as a data set to train the server's global prediction head. The number of training rounds of the server is 100 rounds. Calculate the cross-entropy loss between the prediction result of the server's global prediction head and the true class label s. Then use the gradient descent method to update the parameters Of the global prediction head. The specific update formula is:

[0141] (19)

[0142] Where, Denotes the global prediction head parameters, Is the learning rate of the global prediction head, Is the cross-entropy loss function, Denote the local average representation uploaded by the k-th client in the t-th round of federated learning training, and s is the corresponding label. After the training is completed, the server's global prediction head has the full-class knowledge obtained from the datasets of multiple clients, and has stronger generalization ability compared with the client prediction head that only contains local knowledge.

[0143] Step 6: After the server's global prediction head training is completed, the server broadcasts the updated global prediction head to the clients. Each client's prediction head adaptively aggregates the global prediction head based on local data. Specifically:

[0144] Each client receives the server's global prediction head after the training is completed, such as Figure 4 , and according to the local data of each client, the local client prediction head adaptively aggregates the downloaded server's global prediction head. We initialize the local prediction head through the weighted adaptive aggregation of the corresponding elements of the global prediction head and the local prediction head . The aggregation calculation formula is:

[0145] (20)

[0146] where, represents the local prediction head parameter, represents the clustering weight of each element. Each client updates iteratively through the feature . The update formula is as follows

[0147] (21)

[0148] where, represents the clustering weight, is the learning rate for learning the weight, is the local prediction head parameter, is the global prediction head parameter, is the cross-entropy loss function. After completing the update , the personalized parameters of the client prediction head can be calculated.

[0149] Step 7: Repeat Step 3 to Step 6 until the preset number of training rounds is reached, and calculate the test metrics for each round. Specifically:

[0150] Repeat Steps 3 to 6 until the preset number of training rounds (the preset number of training rounds is 1000 rounds) is reached, record the model parameters of each client, and calculate the global average face recognition accuracy for each round. The formula is as follows:

[0151] (22)

[0152] Among them, represents the average accuracy rate, N represents the total number of images of all clients, represents the number of images correctly predicted by the i-th client.

[0153] Step 8: Each client selects the optimal model parameters, imports them, and inputs a face image to complete the face recognition task;

[0154] Select the model parameters of each client in the round with the optimal global average accuracy rate indicator, import the model parameters into each client model, and complete the face recognition task by inputting a face image.

[0155] The specific data and scenario parameters of this embodiment are described below. This embodiment constructs two data heterogeneity scenarios: pathological setting and practical setting. In the pathological setting, a dataset with non-overlapping categories and unbalanced data amounts is assigned to each client. This setting simulates a heterogeneous training environment with significant differences in client data. In the practical setting, the Dirichlet distribution is used to assign data to each client to simulate the data distribution in the real world. Specifically, a distribution ratio is drawn from the Dirichlet distribution, and the data points of the proportion of category c in the dataset are assigned to client i.

[0156] To simulate the heterogeneous federated learning scenario, this embodiment sets up an environment consisting of 10 clients, and the client participation rate is 1. The dataset of each client is divided into a training set and a test set according to a ratio of 3:1. Table 3 shows the comparison of face recognition accuracy rates between this embodiment and other methods on the MS-Celeb-1M face dataset in the pathological setting and the MS-Celeb-1M face dataset in the practical setting.

[0157] Table 3 Comparison of Face Recognition Accuracy Rates

[0158]

[0159] Embodiment 3

[0160] Referring to Figure 5 , this embodiment relates to a heterogeneous federated learning device based on structure-aware information optimization and adaptive aggregation, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation in Embodiment 1.

[0161] At the hardware level, the device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 method. Of course, in addition to the software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0162] Improvements to a technology can be clearly distinguished as hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system on a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL). There is not only one kind of hardware description language HDL, but many kinds, such as the Advanced Boolean Expression Language ABEL, the Altera Hardware Description Language AHDL, the Confluence Hardware Description Language, the Cornell University Programming Language CUPL, HDCal, the Java Hardware Description Language JHDL, the Hardware Description Language Lava, the Hardware Description Language Lola, the Hardware Description Language MyHDL, the Hardware Description Language PALASM, the Ruby Hardware Description Language RHDL, etc. The most commonly used ones currently are the Very-High-Speed Integrated Circuit Hardware Description Language VHDL and the Hardware Description Language Verilog.Those skilled in the art should also be clear that only by slightly logically programming the method flow with the above-mentioned several hardware description languages and programming it into an integrated circuit can a hardware circuit for implementing the logical method flow be easily obtained.

[0163] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: hardware description language ARC 625D, hardware description language Atmel AT91SAM, hardware description language Microchip PIC18F26K20, and hardware description language Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0164] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0165] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present invention, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0166] Embodiment 4

[0167] This embodiment relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation in Embodiment 1.

[0168] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0169] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0170] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0171] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0172] The embodiments of the present invention are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments.

[0173] The above description is only for the embodiments of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and changes can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation, characterized by: The following steps are involved: Step 1: Each client builds a complex multi-scale model and trains it using local data; Step 2: Obtain the structure-aware information of the multi-scale model in step 1, and use the structure-aware information to generate a gradient adjustment factor; specifically, including: After the multi-scale model training is completed, the scale regulator parameter sd, that is, the structural perception information, is extracted and the gradient adjustment factor GAF is generated; the gradient adjustment factor GAF is calculated by the following formula: ,in, represents the scaler parameters of the original convolution Conv, represents the scale regulator parameter of the multi-scale feature enhancer MFE; Step 3: Each client integrates the gradient adjustment factor into the local model for training. After the training is completed, each client extracts the features of all data through the feature extractor; Step 4: Each client inputs the features extracted in step 3 into the autoencoder to extract a more personalized feature representation and record the corresponding label. Each client performs a weighted average of the features of each label and uploads them to the server together with the corresponding label. Step 5: The server receives the average features and corresponding labels of each client in step 4, and uses these data to train the server global prediction head; Step 6: After the server-side global prediction header training is completed, the server broadcasts the updated global prediction header to the client; each client prediction header is based on local data and adaptively aggregates the global prediction header; Step 7: Repeat steps 3 to 6 until the preset training rounds are reached, and calculate the test indicators in each round; Step 8: Each client selects the optimal model parameters and imports the input image to complete the classification task.

2. The heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation according to claim 1 is characterized in that: Step 1 includes: Step 1-1: Build a complex multi-scale model. For the layers that need to improve the feature extraction capability locally, add an additional multi-scale feature enhancer MFE next to the original convolution, and add a scale regulator SC after the original convolution Conv and the newly added multi-scale feature enhancer MFE; Step 1-2: After the multi-scale model is built, it is trained using local data.

3. The heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation according to claim 1 is characterized in that: Step 3 includes: Step 3-1: Incorporate the gradient adjustment factor into the local original model training and adjust the model's gradient update process to achieve a training process equivalent to that of the multi-scale model; Step 3-2: After training, use its feature extractor For local datasets Extract features from each training sample and generate features .

4. The heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation according to claim 1 is characterized in that: Step 4 includes: Step 4-1: Each client uses the encoder to generate local features from the local extractor Perform feature compression to remove redundant information and noise; Step 4-2: Use the decoder to restore the compressed features to the original feature dimension size to ensure the dimensional consistency uploaded to the client; Step 4-3: Each client calculates the average features of all samples belonging to the same class and uploads the average feature representation and corresponding labels to the server.

5. The heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation according to claim 1, characterized in that: Step 5 includes: The server receives the average features and corresponding labels s from all participating clients and uses them as a dataset to train the server's global prediction head. It calculates the cross entropy loss between the prediction results of the server's global prediction head and the true category label s. It then uses the gradient descent method to adjust the parameters of the global prediction head. to update.

6. The heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation according to claim 1, characterized in that: Step 6 includes: The server broadcasts the server global prediction header after training. According to the local data of each client, the local client prediction header adaptively aggregates the downloaded server global prediction header.

7. The heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation according to claim 1, characterized in that: The test indicator described in step 7 is the global average accuracy.

8. A heterogeneous federated learning device based on structure-aware information optimization and adaptive aggregation, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that A program is stored thereon, and when the program is executed by a processor, the heterogeneous federated learning method based on structure-aware information optimization and adaptive aggregation as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Unbalanced data-oriented federal cross-modal retrieval method and system

    CN116244484A

  • Distributed traffic flow prediction method based on personalized federal map signal learning

    CN118675337A