An image classification method applicable to non-independent and identically distributed situations and with privacy protection
By introducing an equiangular tight frame (ETF) structure and memory vector into the image classification method, the problem of image classification performance degradation under non-independent homogeneous data is solved, efficient image classification under privacy protection is achieved, and the classification accuracy of the global model is improved.
Patent Information
- Application Number
- CN202310297826.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-03-23
AI Technical Summary
The existing federated learning method has deteriorated image classification performance under the conditions of non-independent and homogeneous data, making it difficult to effectively solve the classification problem of non-independent and homogeneous data among various clients. Especially in the scenario of privacy protection, the existing methods have failed to effectively improve the classification performance of the global model.
The global image classification system is initialized using an equiangular tightening framework (ETF) structure, and memory vectors are added to each category, and a consistent optimization goal is designed to make the client's local training learn in the same direction. At the same time, the imbalance problem within the class is corrected through memory vectors to achieve the limitations of feature learning.
The image classification performance of the global model is significantly improved, the classification accuracy under non-independent and same distribution conditions is improved, and the effectiveness and universality on multiple federated learning tasks are verified.
Smart Images

Figure CN116246117B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data security and privacy protection, and particularly relates to an image classification method applicable to the non-independent and identically distributed situation and with privacy protection. Background Art
[0002] Artificial intelligence needs to learn through a large amount of data. However, in real application scenarios, restricted by data privacy and security such as laws and regulations, policy supervision, trade secrets, and personal privacy, each client (data source) cannot directly exchange data. For example, information owned by different competitors in the same industry, personal private information of users, data of upstream and downstream industries, etc. cannot be shared. Privacy protection means protecting the data of the client so that it is not visible or shared externally, and transmitting the requirements or preferences of the client through a certain encryption method. To solve the problem of data cooperation under privacy protection and data security, the concept of federated learning was proposed. Its core idea is to conduct distributed model training among multiple clients with local data. Without the need to exchange local individual or sample data, only by exchanging model parameters or intermediate results, a global model based on virtual fusion data is constructed, so as to achieve the balance of data privacy protection and data sharing calculation, that is, the new application paradigm of "data can be used but not visible" and "data does not move but the model moves". For example, for an input method application, the user of the client hopes that language preferences (such as favorite things, language styles, etc.) and privacy data (such as names, work content, etc.) will not be shared with others, but at the same time can obtain accurate recommendations; while the server hopes to design a lightweight model that can be used on all clients in the case where the client data exists but cannot be directly used.
[0003] "Communication-efficient learning of deep networks from decentralized data", published in the international conference Artificial Intelligence and Statistics in 2017, proposed the FedAvg method. Numerous subsequent federated learning methods are based on this framework. Although it has good performance on independent and identically distributed data, its performance drops sharply when facing non-independent and identically distributed situations. Currently, methods for non-independent and identically distributed data mainly address the issue from three perspectives. The first perspective is to modify the local training process of each client. For example, the 2017 work "SCAFFOLD: stochastic controlled averaging for on-device federated learning" effectively corrects the updated gradient by adding a control variable during local updates; "Federated optimization in heterogeneous networks", published in the international conference Machine Learning and Systems in 2020, describes the distance between the local model and the global model through an approximation term and penalizes the situation where the local model is too far from the global model; "Fedrs: Federated learning with restricted softmax for label distribution non-iid data", published in the international top data mining conference ACM SIGKDD Conference on Knowledge Discovery and Data Mining in 2021, points out that the existing Softmax + cross-entropy method will cause the classification weights of missing classes to be updated in the wrong direction and proposes restricted Softmax. However, the results of the first two works on image datasets did not exceed FedAvg, and the last work only modified the last layer classifier and is not applicable to shallow networks.The second angle is to change the process of server-side fusion of various client models. For example, "Federated learning with matched averaging", published in the International Conference on Learning Representations in 2020, uses Bayesian non-parametric methods to set the fusion weights of each client in a hierarchical manner; "Tackling the objective inconsistency problem in heterogeneous federated optimization", published in the Annual Conference on Neural Information Processing Systems in 2020, normalizes the client before fusion. These works only prevent mutations in the learning process, but cannot fundamentally solve the problem of non-independent and identically distributed. The third angle is to increase restrictions on feature learning. For example, "Spherefed: Hyperspherical federated learning", published in the European Conference on Computer Vision in 2022, fixes the weight of the classifier to the unit sphere shared by the client.
[0004] The above methods do not fundamentally solve the problem of non-independent and identically distributed (IID), and in real scenarios, the data between clients are generally non-IID. Taking the input method as an example, for the same pinyin letters, different people prefer different characters. For example, for the pinyin letter "yu yan", for literary workers, they prefer the character "language"; for early childhood educators, they prefer the character "fable"; for computer vision workers, they prefer the character "fisheye". For each user, they hope that their commonly used character preferences can be displayed in the front part of the candidates. Similarly, in image classification tasks such as medical image diagnosis and bioinformatics discrimination, the data in each client is also non-independent and identically distributed. Summary of the invention
[0005] The present invention provides an image classification method applicable to non-IID situations and with privacy protection. From the two perspectives of local training of the client and increasing the restrictions on feature learning, a common optimization goal is designed for all clients, and they learn in the same direction. At the same time, the present invention adds a memory vector for each category in local training. This greatly improves the classification performance of the global model.
[0006] An image classification method applicable to non-independent and identically distributed situations with privacy protection, comprising the following steps:
[0007] (1) Randomly generate an equiangular tight frame (ETF) structure, where the dimension of the ETF matrix is p×K, p is the dimension of the feature, and K is the number of clients selected;
[0008] (2) Initialize the global image classification system with the ETF structure in step (1) so that each client has a consistent optimization objective;
[0009] (3) Randomly select K clients from all clients;
[0010] (4) For the clients selected in step (3), first initialize the local image classification system of the client to the global image classification system, and then update the feature extraction model with the private data of the client; use the memory vector of the same category in each round of update, and calculate the loss value with the objective function;
[0011] (5) Update the memory vector with the feature vectors of all clients in step (4);
[0012] (6) Update the global image classification system with the local image classification system of the client;
[0013] (7) Repeat steps (3) to (6) until the loss value of the objective function converges;
[0014] (8) Input the image to be classified into the trained global image classification system to obtain the image classification result.
[0015] Further, in step (1), the ETF structure needs to meet the following conditions:
[0016] (1-1)
[0017] (1-2) Each column is a unit norm,
[0018] (1-3) The columns are equiangular to each other.
[0019] In step (2), after initializing the global image classification system to the ETF structure, it remains fixed during training.
[0020] In step (3), in each round of iterative update, K clients are randomly selected again to ensure regularization and randomness.
[0021] In step (4), the specific process of updating the feature extraction model with its private data is as follows:
[0022] (4-1) Obtain the feature vector f using the feature extraction model i , where i is the sample sequence of private data;
[0023] (4-2) When the number of iterations is greater than R warm , correct the feature vector using the memory vector
[0024] (4-3) Calculate the loss value using the objective function and update the feature extraction model by gradient.
[0025] In step (5), for each category, update the memory vector μ using all the feature vector f values extracted by the client in step (4-1) i , where c represents the category number. c Compared with the prior art, the present invention has the following beneficial effects:
[0026] 1. Inspired by the Neural Collapse phenomenon, the present invention designs a consistent optimization objective shared by all clients, which solves to a certain extent the problem of performance degradation of image classification methods and systems in joint learning under non-independent and identically distributed conditions.
[0027] 2. In order to reduce the fluctuations caused by the intra-class inconsistency of conditional distributions between clients, a global memory vector is maintained for each category to correct the bias in backbone training.
[0028] 3. Experiments verify the generality and strong performance of the method of the present invention, and significantly improve the common classification methods on multiple joint learning classification task benchmarks.
[0029] BRIEF DESCRIPTION OF THE DRAWINGS The flowchart of the method of the present invention;
[0030] Figure 1 The schematic diagram of the motivation for adopting the consistent optimization objective of the present invention;
[0031] Figure 2 The schematic diagram after adopting the ETF as the consistent optimization objective of the present invention;
[0032] Figure 3 The schematic diagram after adding the memory vector of the present invention;
[0033] Figure 4 The schematic diagram of the motivation for adding the memory vector of the present invention;
[0034] Figure 5 The schematic diagram after adding the memory vector of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not impose any limitations on it.
[0036] As Figure 1 shown, an image classification method applicable to the non-independent and identically distributed situation and with privacy protection specifically includes the following steps:
[0037] S01. Inspired by Neural Collapse, randomly generate a simplex ETF structure. The dimension of the ETF matrix is p×K, where p is the dimension of the feature and K is the number of clients selected.
[0038] ETF refers to the geometric structure of Equiangular Tight Frame, which was proposed in the paper "Papyan, V., Han, X. Y., and Donoho, D. L. Prevalence of neural collapse during the terminal phase of deep learning training. CoRR, abs / 2008.08186, 2020. URL https: / / arxiv.org / abs / 2008.08186." The main feature of the ETF structure is the "equiangular" property, and its specific generation method is as follows:
[0039]
[0040] Among them, P∈R d×c ,(d≥c) is a partial orthogonal matrix. Here we set P T P = I c ; I c is a c×c identity matrix, and 1 c ∈I c×1 is a vector filled with 1.
[0041] S02. Initialize the classifier in the global image classification system with the simplex ETF structure in step S01 and fix it, which makes the optimization objectives of each client consistent.
[0042] S03. Randomly select K clients from all clients, and randomly select again in each round to ensure a certain degree of regularization and randomness.
[0043] S04. For the clients selected in step S03, first, and then update the feature extraction model with their private data, and consider the memory vector when updating.
[0044] (4-1) Initialize the local image classification system of the client as a global image classification system;
[0045] (4-2) Update the feature extraction model with its private data;
[0046] (4-2-1) Obtain the feature vector f using the feature extraction model F i (f i = F(x i ), where i is the sample sequence of the private data;
[0047] (4-2-2) When the number of iterations is less than or equal to R warm , it is not considered that the memory vector contains accurate information, that is, h i = f i ; when the number of iterations is greater than R warm , correct the feature vector with the memory vector
[0048] (4-2-3) Calculate the loss value using the objective function and update the feature extraction model by gradient.
[0049] S05. For each category, update the memory vector μ i with all the feature vector f c values extracted to the client in (4-2-1), where c represents the category serial number.
[0050] S06. Update the global image classification system with the client image classification system.
[0051] Repeat steps S03 - S06 for multiple rounds of training until convergence or the performance of the image classification system no longer improves.
[0052] As Figure 2 shown, when there is no consistent goal among clients, each client only achieves local optimality, making it impossible for the server to obtain global information after fusing the information of each client. Starting from the two aspects of local training of the client and increasing the restrictions on feature learning, the present invention designs a common optimization goal for all clients and enables them to learn in the same direction. As Figure 3 shown, after adding a common optimization goal, the optimization directions of each client are consistent with the optimization direction of the global server. Since there will also be an imbalance problem within the category, as Figure 4 shown. Therefore, the present invention adds a memory vector for each category in local training, which solves the problem of intra-class imbalance to a certain extent, as Figure 5 shown. This greatly improves the classification performance of the global model.
[0053] To verify the effectiveness of the present invention, three scenario experiments were conducted on the CIFAR10 and CIFAR100 test sets. CIFAR contains a total of 60,000 color images of 32*32, with no overlap of any type. The dataset is divided into five training batches and one test batch, with 10,000 images in each batch. The test batch contains 1,000 images randomly selected from each category. The training batches contain the remaining images in random order, but some training batches may contain more images from one category than another. Between these two batches, the training batches contain 5,000 pictures of each category. CIFAR10 divides CIFAR into 10 categories, namely airplane, car, bird, cat, deer, dog, frog, horse, ship, and truck. While CIFAR-100 divides the pictures into 100 categories more finely. Each category contains 600 images. Each category has 500 training images and 100 test images. The 100 categories in CIFAR-100 are divided into 20 supercategories. Each image comes with a "fine" label and a "coarse" label. And three scenarios were designed for these two datasets, namely CIFAR10-100-2, CIFAR10-100-5, and CIFAR100-100-20. Among them, the first part represents the dataset, the second part represents the number of clients, and the third part represents the average number of categories that each client can see.
[0054] In this embodiment, comparisons were made with the currently best-performing published methods on the test set, and the comparison results are shown in Table 1. Under three scenarios, the average precision of each method was compared. In Table 1, the top row is the currently published method; the row below is the effect verification of the present invention on the most basic FedAvg method. It can be seen that the present invention has achieved the best results in all indicators, and the present invention (FEDAVG_ETF) has a stronger detection accuracy compared to other methods.
[0055] Table 1
[0056]
[0057]
[0058] Table 2 shows the effects after the present invention is combined with other common federated learning methods in classification applications as a plugin. The experimental results show that the present invention is effective in multiple classification systems of federated learning.
[0059] Table 2
[0060]
[0061] The embodiments described above have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modification, supplement, and equivalent replacement made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image classification method applicable to the non - independent and identically - distributed case and with privacy protection, characterized in that, It includes the following steps: (1) Randomly generate an equiangular tight frame (ETF) structure, where the dimension of the ETF matrix is p×K, p is the dimension of the feature, and K is the number of selected clients; (2) Initialize the global image classification system with the ETF structure in step (1) so that each client has a consistent optimization objective; (3) Randomly select K clients from all clients; (4) For the clients selected in step (3), first initialize the local image classification system of the client as the global image classification system, and then update the feature extraction model with the private data of the client; use the memory vector for the same category in each round of update, and calculate the loss value with the objective function; (5) Update the memory vector with the feature vectors of all clients in step (4); (6) Update the global image classification system with the local image classification system of the client; (7) Repeat steps (3) to (6) until the loss value of the objective function converges; (8) Input the image to be classified into the trained global image classification system to obtain the image classification result.
2. The image classification method applicable to the non-independent and identically distributed situation and with privacy protection according to claim 1, wherein In step (1), the ETF structure needs to meet the following conditions: (1-1) (1-2) Each column is a unit norm, (1-3) The columns are equiangular to each other.
3. The image classification method applicable to the non-independent and identically distributed situation and with privacy protection according to claim 1, wherein In step (2), after initializing the global image classification system as the ETF structure, it remains fixed during training.
4. The image classification method applicable to the non - independent and identically - distributed situation and with privacy protection according to claim 1, wherein In step (3), in each round of iterative update, K clients are randomly selected again to ensure regularization and randomness.
5. The image classification method applicable to the non-independent and identically distributed situation and with privacy protection according to claim 1, characterized in that, In step (4), the specific process of updating the feature extraction model with its private data is as follows: (4-1) Obtain the feature vector f using the feature extraction model i , where i is the sample sequence of the private data; When the number of iteration rounds is greater than R warm the feature vector is corrected with the memory vector (4-3) Calculate the loss value with the objective function and update the feature extraction model by gradient.
6. The image classification method applicable to the non-independent and identically distributed situation and with privacy protection according to claim 5, wherein In step (5), for each category, use all the feature vectors f i values extracted at the client in step (4-1) to update the memory vector μ c , where c represents the category number.