Personalized federal learning method for image semantic segmentation

By introducing a two-layer parallel optimization personalized federated learning model in the federated learning environment, the personalized model is collaboratively trained and the global model is updated, the problem of promoting semantic segmentation data on different user devices is solved, and the generalization and personalization capabilities of the model are improved.

CN120163979APending Publication Date: 2025-06-17CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510311151.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In a federated learning environment, the classimetric and statistical heterogeneity of semantic segmentation data leads to a degradation of model performance and is difficult to promote on different user devices.

Method used

A personalized federated learning method for image semantic segmentation is proposed. By establishing a personalized federated learning model with two-layer parallel optimization, including inner layer optimization problems based on Moro envelope improvement and outer layer optimization problems used to generate local models, user nodes collaborate with edge servers to train personalized models, and update the global model through aggregation.

Benefits of technology

The generalization ability of the model is improved, the local personalization ability under the complex semantic segmentation task is maintained, and the classimonial and statistical heterogeneity of semantic segmentation data is effectively dealt with.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163979A_ABST
    Figure CN120163979A_ABST
Patent Text Reader

Abstract

The invention relates to an image semantic segmentation-oriented personalized federal learning method, which belongs to the technical field of wireless communication, and comprises the following steps: S1, establishing an edge network system consisting of a plurality of user nodes and an edge server, each user node being configured with a sensing device capable of collecting segmented image data, data among different users are not independently distributed; s2, establishing a personalized federated learning model of mutual decoupling double-layer parallel optimization, wherein the personalized federated learning model comprises an inner-layer optimization problem based on Moire envelope improvement and an outer-layer optimization problem used for generating a local model; s3, the user node and the edge server cooperatively train and update the personalized model, and then update the corresponding local model; s4, the edge server collects local models uploaded by the user nodes, aggregates and updates a global model, and then returns to the user nodes participating in training in next iteration; and S5, repeating the steps S3-S4, and obtaining a solution of the optimization problem after a convergence condition is reached.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning and relates to a personalized federated learning method for image semantic segmentation. Background Art

[0002] Semantic segmentation requires detailed classification of each pixel point in an image or video frame and assignment of accurate semantic labels. Compared with traditional classification tasks, semantic segmentation is more suitable for more complex task scenarios such as autonomous driving and medical image analysis. Therefore, a large amount of pixel-level labeled data is required for centralized training, which will bring the risk of privacy leakage.

[0003] Federated learning (FL) is an emerging distributed machine learning paradigm. Because it can train a suitable global model for all participating user devices while protecting user data privacy, it has now attracted extensive attention in the industrial and academic fields. However, the combination of FL and semantic segmentation faces many challenges. For example, when facing a large amount of heterogeneous device data, the model performance will be greatly reduced, resulting in the fact that the finally trained global model cannot be well generalized to each user device.

[0004] Data heterogeneity stems from differences in the attributes, preferences, and data collection modes of user devices. To address the challenges brought by data heterogeneity to federated learning, personalized federated learning (PFL) has been proposed. Its core lies in formulating a set of personalized solutions for each participating user device. Compared with traditional FL, PFL pays more attention to the data differences between different user devices, so as to improve the model performance as much as possible. At present, there are already a variety of PFLs that can effectively address the challenges of data heterogeneity, but most of their design solutions only focus on the statistical heterogeneity of data. However, there are more detailed semantic category differences in semantic segmentation data itself. Therefore, in the FL environment, there is also semantic category heterogeneity between different users. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a personalized federated learning method for image semantic segmentation.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A personalized federated learning method for image semantic segmentation includes the following steps:

[0008] S1: Establish an edge network system composed of a number of user nodes and an edge server. Among them, the user nodes are configured with sensing devices that can collect segmented image data, and the data between different users is non-independent and identically distributed (Non-IID);

[0009] S2: Establish a personalized federated learning model with a decoupled two-layer parallel optimization, including an inner-layer optimization problem improved based on the Moreau envelope and an outer-layer optimization problem for generating local models;

[0010] S3: The user nodes and the edge server cooperate to train and update the personalized model, and then update the corresponding local model through the obtained personalized model;

[0011] S4: The edge server collects the local models uploaded from the user nodes and updates the global model in an aggregated manner, and then returns to the user nodes participating in the training in the next iteration;

[0012] S5: Repeat steps S3 - S4. After several iterations, the convergence condition will be reached, and the solution to the optimization problem will be obtained.

[0013] Furthermore, in step S1, establish an edge network system composed of N user nodes and one edge server; use to represent the set of user nodes; the edge server is connected to each non-interfering user node through a wired backhaul link; the edge server communicates with the user device nodes through a wireless channel, and the dataset used by the user device is denoted as The total dataset is denoted as

[0014] The datasets of each user device are non-independent and identically distributed (Non-IID). According to the characteristics of semantic segmentation data, the data between each user has class heterogeneity and statistical heterogeneity.

[0015] Furthermore, step S2 specifically includes the following steps:

[0016] Establish the FL optimization problem P1 for the edge groups:

[0017]

[0018] where w is defined as the global model, and w * represents the optimal solution of the optimization problem P1; F n (w) is related to the Moreau envelope and is defined as the following optimization problem P2:

[0019]

[0020] where u n is defined as the personalized model of user n participating in the training; λ is a regularization parameter used to control the difference between u n and w; f n (u n) Represents the training losses of all users participating in the training. Represents the distillation loss used to enable the local personalized model to retain local features while benefiting from global knowledge. Specifically, f n (u n ) Represents the personalized model u n on the user device of the local dataset ;

[0021] In the two-layer parallel optimization process represented by P1 and P2, problem P1 is the outer optimization problem for solving the global model w, and problem P2 is the inner optimization problem for solving the personalized model u n ; these two optimization problems are decoupled from each other.

[0022] Define the optimal solution of the personalized model u n as The local optimal solution of problem P1 is where local optimal solution is calculated by the following formula:

[0023]

[0024] Use the gradient descent method to solve the local optimal solution where α is the gradient descent step size.

[0025] Furthermore, the specific steps of S3 are as follows:

[0026] S31: The edge server sends the latest global model to all user nodes, as follows:

[0027] Use to represent the global iteration index, where is the index set of T global iterations; use e ∈ ε to represent the index of the group iteration, where is the index set of E group iterations; the local model after the e-th local iteration in the t-th global iteration is represented as

[0028] Before each global iteration starts the group iteration, initialize the local model to w t-1 ;

[0029] The edge server broadcasts the latest global model w t-1,E to all user nodes;

[0030] S32: Knowledge distillation between the user node and the global, as follows:

[0031] The user node n utilizes the personalized model retained from the previous round of global iteration and the global model w t-1,E to achieve knowledge distillation through the following formula:

[0032]

[0033] where z p and z G represent the feature representations of the personalized model and the global model w t-1,E corresponding to the local input data x of user n after introducing the Projection Head layer PH in the semantic segmentation model; the PH layer consists of several fully connected mappings; the feature representation of x corresponding to the personalized model is the original logit output masked according to the local labels, and the mask is represented as follows:

[0034]

[0035] where c n is the semantic category contained in user n;

[0036] S33: Define the training loss f n (u n ) of the user node to address class heterogeneity, as follows:

[0037] l corresponding to the local input data x of user node n is the overall segmentation label, and q x (j, l j ) is the predicted probability of pixel j relative to the output logit; there is class-level heterogeneity in the local data of each user, and foreground class pixels need to be emphasized, while also considering the influence of background pixels in global training; the local objective loss function f n (u n ) is:

[0038]

[0039] where the predicted probability of pixel j is defined as:

[0040]

[0041] where represents the foreground annotation category of the local data of user , and represents the number of foreground categories owned by this user locally;

[0042] S34: The user node collaborates with the edge server to train the personalized model to address statistical heterogeneity, as follows:

[0043] The user device n calculates the user local model by performing K steps of stochastic gradient descent. The user local model is the model updated by the user device n based on the local dataset through the following formula:

[0044]

[0045] where η t represents the step size, and f n (·) represents the local loss function considering unbalanced semantic classes;

[0046] According to the above calculation process of the local optimal solution the update process of the local model in the e-th local iteration is as follows:

[0047]

[0048] After E local iterations, the user local model update task is completed, and the personalized model and the local model

[0049] Furthermore, the S4 specifically includes the following steps:

[0050] S41: The edge server collects the local models, as follows:

[0051] At the t-th global iteration of the model, the edge server obtains the local models from all participating training users n

[0052] S42: The edge server aggregates the collected user local models and updates the global model, as follows:

[0053] The edge server aggregates the user local models and updates the global model through the following formula:

[0054]

[0055] where the constant β > 0 is used to control the influence degree of the aggregation of each local model on the global model w t of.

[0056] Furthermore, the S5 specifically includes the following steps:

[0057] Repeat steps S3 and S4. Each time a global iteration is performed, the global iteration index t is updated to t + 1 until the specified global iteration number T is reached, and the iteration stops to obtain the solution w T of problem P1 and the solution of problem P2

[0058] On the other hand, an electronic device includes a memory and a processor;

[0059] The memory is used to store a computer program;

[0060] The processor is configured to, when executing the computer program, implement the personalized federated learning method for image semantic segmentation as described in any one of the above.

[0061] On yet another aspect, a computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the personalized federated learning method for image semantic segmentation as described in any one of the above is implemented.

[0062] The beneficial effects of the present invention are as follows:

[0063] 1. The present invention proposes a federated learning framework for semantic segmentation tasks, focusing on the real-world challenges of statistical heterogeneity and class heterogeneity. Compared with traditional centralized methods, federated learning can integrate multi-source data and improve the model generalization ability on the premise of protecting data privacy. This research has practical significance for large-scale, multi-institutional applications such as smart cities and telemedicine;

[0064] 2. The present invention promotes the collaborative training of personalized models between edge servers and user nodes through a double-layer parallel optimization mechanism, improves the generalization ability of the model, and maintains the local personalization ability under the complex semantic segmentation task of data.

[0065] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in preferred detail below in conjunction with the drawings, where:

[0067] Figure 1 is a flowchart of the personalized federated learning method for image semantic segmentation according to an embodiment of the present invention;

[0068] Figure 2 is a system architecture diagram of the personalized federated learning method for image semantic segmentation according to an embodiment of the present invention;

[0069] Figure 3When fixing the number of devices, considering two semantic segmentation datasets and the device data being non-independent and identically distributed (non-IID), the average test accuracy and average intersection over union of the method proposed in the present invention and the comparative methods; where (a) is the average test accuracy of each method for the Cityscapes dataset, (b) is the average test intersection over union of each method for the Cityscapes dataset, (c) is the average test accuracy of each method for the CamVid dataset, and (d) is the average test intersection over union of each method for the CamVid dataset;

[0070] Figure 4 The influence of the number of users participating in model training on the performance of the personalized model. Detailed implementation manners

[0071] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0072] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention schematically. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, numbers, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0073] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0074] The present invention provides a personalized federated learning method for semantic segmentation, which is applicable to non-IID datasets with richer semantic categories; the process of this method is as Figure 1 shown and includes the following steps:

[0075] Step S1, establish a wireless edge network system model as Figure 2 shown, and the specific content is as follows:

[0076] An edge network system consisting of N user nodes and an edge server is established. Among them, represents the set of user nodes. In the system model proposed by the present invention, the edge server connects each non-interfering user node through a wired backhaul link. Communication between the edge server and the user equipment node is carried out through a wireless channel, and the data set used by the user equipment is represented as Thus, the total data set can be represented as And the data sets of each user equipment are non-independent and identically distributed (Non-IID). Specifically, according to the characteristics of semantic segmentation data, the data between each user has class heterogeneity and statistical heterogeneity.

[0077] Step S2, establish a personalized federated learning optimization problem with double-layer parallel optimization, including an inner-layer optimization problem improved based on the Moreau envelope, and an outer-layer optimization problem for generating a local model. The specific content is as follows:

[0078] 1) Establish the FL optimization problem P1 for between-edge groups:

[0079]

[0080] Among them, w is defined as the global model, and w * represents the optimal solution of the optimization problem P1;

[0081] 2) F m (w) is related to the Moreau envelope and can be defined as the following optimization problem P2:

[0082]

[0083] Among them, u n is defined as the personalized model of the participating training user n; λ is a regularization parameter used to control the difference between u n and w; f n (u n ) represents the training loss of all participating training users, represents the distillation loss used to enable the local personalized model to retain local features and benefit from global knowledge at the same time. Specifically, f n (u n ) represents the loss function of the personalized model u n on the local data set of the user equipment ; In the double-layer parallel optimization process represented by P1 and P2, problem P1 is the outer-layer optimization problem for solving the global model w, and problem P2 is the outer-layer optimization problem for solving the personalized model u nThe inner layer optimization problem, and these two optimization problems are decoupled from each other; define the personalized model u n The optimal solution of The local optimal solution of problem P1 is Where Furthermore, the local optimal solution Can be calculated by the following formula:

[0084]

[0085] The above process of solving the local optimal solution Adopts the gradient descent method, where α is the gradient descent step size.

[0086] Step S3, the specific content includes the following sub-steps:

[0087] Sub-step S31, the edge server sends the latest global model to all user nodes and defines the model update parameters, and the specific content is as follows:

[0088] Use To represent the global iteration index, where Is the index set of T global iterations; use e ∈ ε to represent the index of the group iteration, where Is the index set of E group iterations; the local model after the e-th local iteration in the t-th global iteration is denoted as

[0089] Before each global iteration starts the group iteration, initialize the local model To w t-1 ;

[0090] The edge server broadcasts the latest global model w t-1,E To all user nodes;

[0091] Sub-step S32, the knowledge distillation between the user node and the global, and the specific content is as follows:

[0092] User node n uses the personalized model Retained in the previous global iteration and the global model w t-1,E To achieve knowledge distillation through the following formula:

[0093]

[0094] Where, z p And z G Represent after introducing the Projection Head (PH) layer in the semantic segmentation model, the local input data x of user n corresponds to the personalized model And the global model w t-1,EFeature representation; it should be noted that the PH layer consists of several fully connected mappings; x corresponds to the personalized model The feature representation is that the original logit output is masked according to the local label, and the mask representation is as follows:

[0095]

[0096] where c n is the semantic category included in user n;

[0097] Sub-step S33, the sub-server aggregates the personalized models uploaded by the user devices and updates the group model, and the specific content is as follows:

[0098] S33: Define the training loss f of the user node according to the category heterogeneity n (u n ) to cope with category heterogeneity, and the specific content is as follows:

[0099] For the overall segmentation label l corresponding to the local input data x of user node n, then q x (j, l j ) is the prediction probability of pixel j relative to the output logit; each user's local data has heterogeneity at the category level, and foreground category pixels need to be considered, and the influence of background pixels in global training also needs to be considered; the local objective loss function f n (u n ) is:

[0100]

[0101] Among them, the prediction probability of pixel j is defined as:

[0102]

[0103] where represents the foreground annotation category of user local data, then represents the number of foreground categories owned by this user locally;

[0104] S34: The user node collaborates with the edge server to train the personalized model to cope with statistical heterogeneity, and the specific content is as follows:

[0105] User device n calculates the user local model by performing K steps of stochastic gradient descent (the model updated by user device n based on the local dataset obtained through the following formula)

[0106]

[0107] Among them, η t represents the step size, and f n (·) represents the local loss function considering unbalanced semantic classes;

[0108] According to the above calculation process of the local optimal solution in the e-th local iteration, the update process of the local model is as follows:

[0109]

[0110] After E local iterations, the user local model update task is completed, and a personalized model and the local model

[0111] Step S4, the specific content includes the following sub-steps:

[0112] Sub-step S41, the edge server collects the local model, and the specific content is as follows:

[0113] At the t-th global iteration of the model, the edge server obtains the local model from all participating training users n

[0114] Sub-step S42, the edge server aggregates the collected user local models and updates the global model, and the specific content is as follows:

[0115] The edge server aggregates the user local models and updates the global model through the following formula:

[0116]

[0117] Among them, the constant β > 0 is used to control the influence degree of the aggregation of each local model on the global model w t of.

[0118] Step S5 can be regarded as judging whether the model reaches the convergence condition, and the specific content is as follows:

[0119] Repeat steps S3 and S4. Each time a global iteration is performed, the global iteration index t is updated to t + 1 until the specified global iteration number T is reached, and the iteration is stopped to obtain the solution w of problem P1 T and the solution of problem P2

[0120] Furthermore, the present invention conducts simulation experiments in the constructed wireless edge network system model through the proposed scheme, and the specific results are as Figure 3 and Figure 4 shown;

[0121] Furthermore, as Figure 3As shown in (a)-(d), the present invention uses the Cityscapes dataset and the CamVid dataset to give the relationship between the average test accuracy under the two datasets and the mean intersection over union (mIoU) of the local personalized model with respect to the number of model iterations;

[0122] Figure 3 It shows that when the number of fixed devices and user data are non-independent and identically distributed (non-IID), and the category richness varies according to different datasets, the personalized federated learning scheme for semantic segmentation proposed by the present invention is compared with Comparative Schemes 1, 2, and 3. Among them, Comparative Scheme 1 is a personalized federated learning scheme without introducing the knowledge distillation strategy, Comparative Scheme 2 is a federated learning scheme without introducing the knowledge distillation strategy and personalization, and Comparative Scheme 3 is traditional federated learning. Specifically, from Figure 3 the following content can be obtained:

[0123] 1) The performance of the personalized federated learning scheme is significantly better than that of traditional federated learning, proving the effectiveness of the personalized federated learning scheme in the case of non-IID data distribution and being able to alleviate and make up for the defects of traditional federated learning in the face of data statistical heterogeneity;

[0124] 2) For the personalized federated learning scheme considering the knowledge distillation strategy, compared with the personalized federated learning scheme without considering the knowledge distillation strategy, it has better effects on datasets with richer semantic categories such as Cityscapes. However, the category richness of CamVid is lower than that of the Cityscapes dataset, so the ability of the personalized model to obtain "knowledge" from the global model is limited;

[0125] 3) The performance of the personalized federated learning scheme for semantic segmentation proposed by the present invention is further superior to the federated learning method without considering the personalized scheme because the federated learning method without considering the personalized scheme only considers the category heterogeneity of semantic segmentation data in the local loss function;

[0126] Comprehensively Figure 3 Through analysis, it is obtained that the proposed scheme of the present invention can effectively improve the model performance in the realistic wireless edge network scenario where the data distribution at the group level is non-IID and the semantic categories are richer.

[0127] Figure 4 It shows the influence of the number of users participating in model training on the performance of the personalized model; as the number of users participating in training increases, the mIoU corresponding to the user's local personalized model increases, while the accuracy shows a downward trend after the number of users increases to a certain extent.

[0128] In the above embodiments, the reference in the specification to "this embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment are included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.

[0129] In the above embodiments, although the present invention has been described in connection with specific embodiments of the present invention, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed. Embodiments of the present invention are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims.

[0130] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which when executed by a processor implements any one of the methods in this embodiment.

[0131] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0132] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes any one of the methods in this embodiment.

[0133] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disk that can store program code.

[0134] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication therebetween. The memory is used to store a computer program, and the communication interface is used for communication. The processor and the transceiver are used to run the computer program to make the electronic terminal execute each step of the above method.

[0135] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0136] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it may also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0137] The present invention can be used in numerous general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0138] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A personalized federated learning method for image semantic segmentation, characterized by: The following steps are involved: S1: Establish an edge network system consisting of several user nodes and an edge server, where the user nodes are equipped with sensor devices that can collect segmented image data, and the data between different users are non-independent and identically distributed Non-IID; S2: Establish a personalized federated learning model with decoupled two-layer parallel optimization, including an inner optimization problem based on the improvement of the Morrow envelope and an outer optimization problem for generating a local model; S3: The user node and the edge server collaborate to train and update the personalized model, and then update the corresponding local model with the obtained personalized model; S4: The edge server collects local models uploaded from user nodes and updates the global model in an aggregated manner, and then returns to the user nodes participating in the training in the next iteration; S5: Repeat steps S3-S4 to obtain a solution to the optimization problem after reaching the convergence condition.

2. The personalized federated learning method for image semantic segmentation according to claim 1, characterized in that: In step S1, an edge network system consisting of N user nodes and an edge server is established; Represents a collection of user nodes; The edge server is connected to each user node without interfering with each other through a wired backhaul link; the edge server communicates with the user device node through a wireless channel. The dataset used is represented as The total dataset is represented as Dataset for each user device The distribution between them is Non-IID. According to the characteristics of semantic segmentation data, the data between each user has category heterogeneity and statistical heterogeneity.

3. The personalized federated learning method for image semantic segmentation according to claim 1, characterized in that: The S2 specifically includes the following steps: Establish the FL optimization problem P1 for edge groups: P1: Among them, w is defined as the global model, w * represents the optimal solution of the optimization problem P1; F n (w) is related to the Morrow envelope and is defined as the following optimization problem P2: P2: Among them, u n is defined as the personalized model participating in training user n; λ is a regularization parameter used to control u n The difference between w and f n (u n ) represents the training loss of all users participating in the training. represents the distillation loss used to make the local personalized model retain local features while benefiting from global knowledge. Specifically, f n (u n ) represents the personalized model u n On user device Local dataset The loss function on ; In the two-layer parallel optimization process represented by P1 and P2, problem P1 is to solve the outer optimization problem of the global model w, and problem P2 is to solve the personalized model u n The inner optimization problem of , these two optimization problems are decoupled from each other; Define the personalized model u n The optimal solution is The local optimal solution of problem P1 is in Local optimal solution Calculated by the following formula: Use gradient descent method to find the local optimal solution Where α is the gradient descent step size.

4. The personalized federated learning method for image semantic segmentation according to claim 1, characterized in that: The S3 specifically includes the following steps: S31: The edge server sends the latest global model to all user nodes. The content is as follows: use represents the global iteration index, where is the index set of T global iterations; e∈ε is used to represent the index of group iteration, where is the index set of E group iterations; the local model after the e-th local iteration of the t-th global iteration is expressed as Before each global iteration starts the group iteration, the local model Initialized to w t-1 ; The edge server sends the latest global model w t-1,E Broadcast to all user nodes; S32: Knowledge distillation between user nodes and the global world, the content is as follows: User node n uses the personalized model retained by the previous round of global iteration With the global model w t-1,E Knowledge distillation is achieved through the following formula: Among them, z p and z G It means that after the Projection Head layer PH is introduced into the semantic segmentation model, the local input data x of user n corresponds to the personalized model With the global model w t-1,E The feature representation of ; the PH layer is composed of several layers of fully connected mapping; x corresponds to the personalized model The feature representation of is the original logit output masked according to the local label, and the mask is represented as follows: where c n is the semantic category included by user n; S33: Define user node training loss f based on category heterogeneity n (u n ) to cope with category heterogeneity, as follows: The local input data x of user node n corresponds to l, which is the overall segmentation label, q x (j,l j ) is the predicted probability of pixel j relative to the output logit; the local data of each user is heterogeneous at the category level, and it is necessary to focus on the foreground category pixels while considering the impact of background pixels in global training; the local target loss function f n (u n )for: Among them, the predicted probability of pixel j is Defined as: in Indicates user Foreground annotation categories of local data, Indicates the number of foreground categories that the user has locally; S34: User nodes and edge servers collaborate to train personalized models to cope with statistical heterogeneity, as follows: User device n calculates the user local model by performing K steps of stochastic gradient descent The user local model Based on local data sets for user devices The model is updated by the following formula: in, η t represents the step length, f n (·) represents the local loss function considering the imbalanced semantic categories; According to the above local optimal solution The calculation process of the local model in the e-th local iteration is as follows: After E local iterations, the user local model update task is completed and a personalized model is obtained. and local models 5. The personalized federated learning method for image semantic segmentation according to claim 1, characterized in that: The S4 specifically comprises the following steps: S41: The edge server collects the local model, the content is as follows: At the tth global iteration of the model, the edge server obtains the local model from all participating training users n S42: The edge server aggregates the collected user local models and updates the global model, the content of which is as follows: The edge server aggregates the user's local model and updates the global model through the following formula: Among them, the constant β>0 is used to control the aggregation of each local model to the global model w t degree of impact.

6. The personalized federated learning method for image semantic segmentation according to claim 1, characterized in that: The S5 specifically includes the following steps: Repeat steps S3 and S4. Each time a global iteration is performed, the global iteration index t is updated to t+1 until the specified number of global iterations T is reached. Then the iteration is stopped and the solution w of problem P1 is obtained. T and the solution to problem P2 7. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the personalized federated learning method for image semantic segmentation as described in any one of claims 1 to 6 when executing the computer program.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the personalized federated learning method for image semantic segmentation as described in any one of claims 1 to 6 is implemented.