Federal learning-based weak supervision semantic segmentation method and device

By utilizing unbiased global prototype adaptive region expansion and refined training in federated learning to generate high-quality pseudo labels, the data heterogeneity and privacy protection issues of weakly supervised semantic segmentation in federated learning scenarios are resolved, and the segmentation accuracy and security of the model are improved.

CN120689614APending Publication Date: 2025-09-23SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510769070.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing weakly supervised semantic segmentation methods have difficulty coping with data distribution heterogeneity in federated learning scenarios, resulting in low model performance and insufficient privacy protection. In particular, there are data sharing security risks in fields such as medical imaging, remote sensing images, and autonomous driving.

Method used

By obtaining the first type of activation map and using the unbiased global prototype for adaptive region expansion, the second type of activation map is generated. The initial model is trained with local image data and pseudo labels, and the unbiased global prototype is updated on the server side until the preset training target is reached, thus forming a target local segmentation model.

Benefits of technology

It significantly improves the performance of weakly supervised semantic segmentation, alleviates the problem of data heterogeneity, improves the generalization ability and privacy protection of the model, and ensures data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689614A_ABST
    Figure CN120689614A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning-based weak supervision semantic segmentation method and device, and the method comprises the steps: obtaining a first-class activation graph according to local image data and an image-level label; through an unbiased global prototype, carrying out adaptive region expansion on the first-class activation graph to obtain a second-class activation graph; performing model training on an initial local segmentation model according to the local image data, the first type of activation graph and the second type of activation graph; and updating the unbiased global prototype, and returning to the step of obtaining the first-class activation graph according to the local image data and the image-level label until the model training reaches a preset training round number or a preset training target, thereby obtaining a target local segmentation model. The method can improve the performance of weak supervision semantic segmentation, and can be widely applied to the technical field of artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a weakly supervised semantic segmentation method and device based on federated learning. Background Art

[0002] In recent years, semantic segmentation technology has attracted much attention due to its outstanding performance in tasks such as understanding image content and achieving pixel-level fine-grained recognition. However, traditional semantic segmentation model training usually needs to be performed on centralized annotated data. This approach inevitably brings high data annotation costs and privacy leakage risks. Especially in data-sensitive fields such as medical imaging, remote sensing images, and autonomous driving, data sharing has significant security risks. Therefore, weakly supervised semantic segmentation technology is used to replace fully supervised semantic segmentation. However, when existing weakly supervised semantic segmentation methods are applied to federated learning scenarios, on the one hand, due to the heterogeneity of data distribution in federated learning environments, traditional weakly supervised methods are difficult to effectively deal with the impact of data distribution differences between different clients; on the other hand, under the federated learning framework, the performance of weakly supervised semantic segmentation models is low, and under the premise of protecting data privacy, the quality of weakly supervised labels is not high, which makes it difficult to improve model performance. Summary of the Invention

[0003] In view of this, the main purpose of the embodiments of the present invention is to provide a weakly supervised semantic segmentation method and device based on federated learning, in order to solve at least one of the problems of the existing technology. The present invention can improve the performance of weakly supervised semantic segmentation.

[0004] To achieve the above objectives, an embodiment of the present invention provides a weakly supervised semantic segmentation method based on federated learning, which includes the following steps:

[0005] Obtain the first type of activation map based on local image data and image-level labels;

[0006] The first type of activation map is adaptively expanded by using an unbiased global prototype to obtain the second type of activation map;

[0007] Performing model training on an initial local segmentation model according to the local image data, the first type activation map, and the second type activation map;

[0008] The unbiased global prototype is updated, and the step of obtaining the first type of activation map according to the local image data and the image-level label is returned until the model training reaches a preset number of training rounds or a preset training target, thereby obtaining a target local segmentation model.

[0009] In some embodiments, obtaining the first type of activation map based on the local image data and the image-level label includes the following steps:

[0010] Training a lightweight classification network using the local image data and the image-level labels;

[0011] According to the trained lightweight classification network, the local image data is processed through forward propagation to generate the first type activation map.

[0012] In some embodiments, the step of adaptively expanding the first activation map using an unbiased global prototype to obtain the second activation map comprises the following steps:

[0013] receiving the unbiased global prototype;

[0014] The first type of activation map is optimized by using the category prior knowledge and global information of the unbiased global prototype to obtain the second type of activation map.

[0015] In some embodiments, the performing model training on the initial local segmentation model based on the local image data, the first class activation map, and the second class activation map comprises the following steps:

[0016] Based on the first type of activation map, generating a first pixel-level pseudo label;

[0017] Based on the second type of activation map, generating a second pixel-level pseudo label;

[0018] An initial local segmentation model is trained according to the local image data, the first pixel-level pseudo-label, and the second pixel-level pseudo-label.

[0019] In some embodiments, the performing model training on the initial local segmentation model based on the local image data, the first pixel-level pseudo-label, and the second pixel-level pseudo-label comprises the following steps:

[0020] Presetting the training parameters and loss function of the initial local segmentation model;

[0021] Using the local image data, the first pixel-level pseudo-label, and the second pixel-level pseudo-label as training data, and inputting the training data into the initial local segmentation model;

[0022] The initial local segmentation model is trained according to the training parameters, the training data, and the loss function.

[0023] In some embodiments, updating the unbiased global prototype comprises the following steps:

[0024] Uploading the current unbiased global prototype to the server;

[0025] Performing aggregation calculation on the current unbiased global prototype through the server to obtain a trainable global prototype;

[0026] The trainable global prototype is debiased by a debiasing network to obtain the updated unbiased global prototype.

[0027] To achieve the above objectives, another aspect of an embodiment of the present invention provides a weakly supervised semantic segmentation device based on federated learning, the device comprising:

[0028] The first module is used to obtain a first-class activation map based on local image data and image-level labels;

[0029] The second module is used to perform adaptive region expansion on the first type of activation map through an unbiased global prototype to obtain a second type of activation map;

[0030] A third module is configured to perform model training on an initial local segmentation model based on the local image data, the first type activation map, and the second type activation map;

[0031] The fourth module is used to update the unbiased global prototype and return to the step of obtaining the first type of activation map based on the local image data and the image-level label until the model training reaches a preset number of training rounds or a preset training target to obtain the target local segmentation model.

[0032] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the weakly supervised semantic segmentation method based on federated learning described above.

[0033] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the weakly supervised semantic segmentation method based on federated learning described above.

[0034] To achieve the above objectives, another aspect of an embodiment of the present invention provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned federated learning-based weakly supervised semantic segmentation method.

[0035] The embodiments of the present invention include at least the following beneficial effects: the present invention provides a weakly supervised semantic segmentation method and device based on federated learning, which obtains a first type of activation map through local image data and image-level labels; adaptively expands the first type of activation map through an unbiased global prototype to obtain a second type of activation map; trains an initial local segmentation model based on the local image data, the first type of activation map and the second type of activation map; updates the unbiased global prototype and returns to the step of obtaining the first type of activation map based on the local image data and image-level labels, until the model training reaches a preset number of training rounds or a preset training target, and obtains a target local segmentation model, which can improve the performance of weakly supervised semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 Flowchart of a weakly supervised semantic segmentation method based on federated learning provided by an embodiment of the present invention;

[0038] Figure 2 Schematic diagram of a weakly supervised semantic segmentation process based on client-server federated learning provided by an embodiment of the present invention;

[0039] Figure 3 It is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the appended claims.

[0041] It should be noted that although the functional modules are divided in the system schematic and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first / S100" and "second / S200" in the specification and claims and the above-mentioned figures may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to a determination".

[0042] The terms "at least one", "plurality", "each", "any", etc. used in the present invention include at least one, two or more, multiple, two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. The terms used herein are for the purpose of describing embodiments of the present invention only and are not intended to limit the present invention.

[0044] Before describing the embodiments of the present invention in detail, some nouns and terms involved in the embodiments of the present invention are first described. The nouns and terms involved in the embodiments of the present invention are subject to the following explanations.

[0045] Federated learning is a distributed machine learning approach that allows multiple clients to jointly train a global model without sharing the original data. Each client trains the model locally using its own data. After several rounds of training, the updated model parameters are uploaded to a central server. Parameter aggregation enables knowledge sharing and model optimization, resulting in a globally optimized model.

[0046] Class activation maps are heatmap visualization techniques generated by classification networks, indicating which regions in an image contribute most to the network's prediction of a specific class. In weakly supervised semantic segmentation tasks, class activation maps are used as coarse segmentation masks that indicate the location and extent of objects. Class activation maps are typically derived by weighted summation of feature maps from the last convolutional layer of a classification network, reflecting the network's attention to different regions in the image.

[0047] Prototype learning is a machine learning method whose core idea is to represent the distribution of data by learning a set of "prototypes." Each prototype can be considered a representative sample or feature vector of a category in the dataset. In image classification tasks, prototypes can be feature representations of the image that capture its key characteristics.

[0048] Unbiased Global Prototype,In unsupervised or self-supervised learning, an unbiased global prototype,refers to a global and unbiased prototype representation that can fairly,represent the distribution of the entire dataset rather than being biased towards certain,specific categories or subsets.

[0049] In the federated learning scenario, due to differences in data distribution among clients, the training of weakly supervised semantic segmentation models is easily affected by data heterogeneity, which in turn causes a significant decline in model performance.

[0050] In view of this, if Figure 1 As shown, the embodiment of the present invention provides a weakly supervised semantic segmentation method based on federated learning, which may include but is not limited to steps S100 to S400:

[0051] Step S100, obtaining a first type of activation map based on local image data and image-level labels;

[0052] Step S200, performing adaptive region expansion on the first type of activation map using an unbiased global prototype to obtain a second type of activation map;

[0053] Step S300: performing model training on an initial local segmentation model according to the local image data, the first type activation map, and the second type activation map;

[0054] Step S400, updating the unbiased global prototype and returning to the step of obtaining the first type of activation map based on the local image data and the image-level label, until the model training reaches a preset number of training rounds or a preset training target, and obtaining the target local segmentation model.

[0055] In some embodiments, step S100 may include but is not limited to steps S110 to S120:

[0056] Step S110, training a lightweight classification network using the local image data and the image-level labels;

[0057] Step S120: Process the local image data according to the trained lightweight classification network through forward propagation to generate the first class activation map.

[0058] In step S110 of some embodiments, a lightweight classification network is trained using local image data and image-level labels. This lightweight classification network is designed to learn the classification features of the image and can generate preliminary class activation maps based on the image data and image-level labels, laying the foundation for subsequent pseudo-label generation. By training the lightweight classification network, classification accuracy can be improved and category-related regions in the image can be effectively located.

[0059] In step S120 of some embodiments, the trained lightweight classification network is used to process the local image data in a forward propagation manner to generate a preliminary class activation map, i.e., a first-class activation map. This first-class activation map can initially indicate the degree of activation of each region in the image for a specific class, and is used to indicate key information such as the location and range of objects in weakly supervised semantic segmentation. However, its boundaries may be relatively rough, and activated regions may be under-activated or over-activated.

[0060] In some embodiments, step S200 may include but is not limited to steps S210 to S220:

[0061] Step S210, receiving the unbiased global prototype;

[0062] Step S220 : Optimizing the first type of activation map by using the category priori knowledge of the unbiased global prototype and global information to obtain the second type of activation map.

[0063] In step S210 of some embodiments, an unbiased global prototype is received from the server. This unbiased global prototype is a key model component maintained by the server. It provides category prior knowledge and global category representation, guiding pseudo-label generation and local segmentation model training, and effectively alleviating data heterogeneity in federated learning scenarios. The unbiased global prototype is trained and updated using a server-side debiasing network and local prototypes uploaded by the client, ensuring that it provides stable and reliable category information and improving the generalization capabilities of the global model.

[0064] In step S220 of some embodiments, the unbiased global prototype provided by the server is received and combined with the generated first-class activation map to perform adaptive region expansion. The adaptive region expansion method uses the category prior knowledge and global information provided by the unbiased global prototype to optimize the first-class activation map, thereby expanding the underactivated areas in the class activation map and suppressing the overactivated areas, thereby improving the quality and accuracy of the class activation map and obtaining a more accurate refined class activation map, i.e., the second-class activation map. Exemplarily, when there are "underactivated" areas in the first-class activation map (i.e., the first-class activation map fails to fully cover the target object, resulting in the appearance of "underactivated" areas that belong to the target but have insufficient activation), the adaptive region expansion method can identify and enhance the activation values ​​of these areas based on the guidance of the global prototype, thereby effectively expanding its coverage range; at the same time, when the first-class activation map mistakenly activates non-target areas, or its activation range exceeds the actual boundary of the object, forming an "overactivated" area, the adaptive region expansion method can also use the constraints of the global prototype to identify and reduce the activation values ​​of these irrelevant or out-of-range areas to achieve a suppression effect. By specifically expanding under-activated areas and suppressing over-activated areas, the overall quality and positioning accuracy of the class activation map are significantly improved, ultimately obtaining a more accurate and refined second-class activation map.

[0065] In some embodiments, step S300 may include but is not limited to steps S310 to S330:

[0066] Step S310: generating a first pixel-level pseudo label based on the first activation map;

[0067] Step S320: generating a second pixel-level pseudo label based on the second activation map;

[0068] Step S330: Perform model training on an initial local segmentation model according to the local image data, the first pixel-level pseudo-label, and the second pixel-level pseudo-label.

[0069] In some embodiments, in steps S310 to S320, an initial pixel-level pseudo-label, i.e., a first pixel-level pseudo-label, is generated based on the generated first-class activation map. Simultaneously, a high-quality refined pixel-level pseudo-label, i.e., a second pixel-level pseudo-label, is generated based on the generated second-class activation map. Both the initial pixel-level pseudo-label and the refined pixel-level pseudo-label serve as supervisory information for subsequent local segmentation model training. The refined pixel-level pseudo-label, due to its higher quality, plays a more important role in local segmentation model training.

[0070] By combining the unbiased global prototype provided by the server, the generated preliminary class activation map is adaptively expanded, and high-quality refined pixel-level pseudo-labels are generated based on the refined class activation map. This serves as an accurate supervision signal for the local segmentation model. Through this refined training strategy, the quality of the pseudo-labels is effectively improved, thereby significantly improving the training efficiency and segmentation accuracy of the local segmentation model.

[0071] In some embodiments, step S330 may include but is not limited to steps S331 to S333:

[0072] Step S331, presetting the training parameters and loss function of the initial local segmentation model;

[0073] Step S332: using the local image data, the first pixel-level pseudo-label, and the second pixel-level pseudo-label as training data, and inputting the training data into the initial local segmentation model;

[0074] Step S333: training the initial local segmentation model according to the training parameters, the training data, and the loss function.

[0075] In steps S331 to S333 of some embodiments, the training parameters and loss function of the local segmentation model are pre-configured. The local image data and the generated initial pixel-level pseudo-labels and refined pixel-level pseudo-labels are loaded and used as training data input for the local segmentation model. The local segmentation model is iteratively optimized and trained based on the input training data using the configured training parameters. Optionally, in the local segmentation model training, different optimization algorithms and hyperparameter settings can be selected to train the local segmentation model. The model parameters are adjusted and updated according to the set loss function during the training process to minimize the prediction error. When the model training reaches a preset number of training rounds or convergence conditions, a trained local segmentation model can be obtained.

[0076] In some embodiments, step S400 may include but is not limited to steps S410 to S430:

[0077] Step S410, uploading the current unbiased global prototype to the server;

[0078] Step S420, performing aggregation calculation on the current unbiased global prototype through the server to obtain a trainable global prototype;

[0079] Step S430: Debiasing the trainable global prototype through a debiasing network to obtain the updated unbiased global prototype.

[0080] In step S410 of some embodiments, each client uploads the locally trained multiple prototypes to the server, wherein only the local prototype parameters are uploaded and other model parameters are retained locally.

[0081] In step S420 of some embodiments, the server receives local prototype parameters from all participating clients and performs adaptive weighted aggregation calculation to obtain a more representative and trainable global prototype.

[0082] In step S430 of some embodiments, the server uses a debiasing network to train and optimize the obtained trainable global prototype to achieve debiasing. Ultimately, the server obtains a more stable and reliable unbiased global prototype and uses the updated unbiased global prototype to replace the old version of the prototype.

[0083] In some embodiments, different prototype aggregation algorithms and debiased network structures can be selected and applied to the prototype aggregation calculation and update steps, and the server side sends the updated unbiased global prototype model parameters back to all participating clients. The client receives and updates the locally stored unbiased global prototype for subsequent client pseudo-label generation and model training steps. In each round of federated learning iteration, the client uses the updated unbiased global prototype to apply to the regional expansion step of local model training to perform a new round of local model training. During the iterative optimization process, the client or server determines whether the current federated learning training meets the preset termination conditions. The termination conditions can be reaching a preset training target (for example, the model performance index reaches a threshold) or reaching a set maximum number of training iterations. Optionally, the number of iterations can be adjusted or different model performance evaluation indicators can be set according to task requirements, such as selecting different segmentation indicators or setting different convergence thresholds to balance model performance and training efficiency. At the same time, the training time range of the debiased multi-prototype federated learning framework should also be adjusted accordingly. If the termination condition is met, the federated learning training ends and the target local segmentation model is obtained; otherwise, the next round of iterative training continues and returns to the step of obtaining the first type of activation map based on the local image data and image-level labels.

[0084] like Figure 2 As shown in the figure, based on the client and server, taking the training process on a certain dataset (such as the VOC dataset) as an example, the weakly supervised semantic segmentation processing flow of federated learning is as follows:

[0085] Step 1: The image dataset is divided into categories and distributed to multiple clients. The data of each client is non-independent and identically distributed (Non-IID). Each client initializes a lightweight classification network and a local prototype. The lightweight classification network is trained using local image data and image-level labels to generate a preliminary class activation map. The class activation map is then adaptively expanded with the unbiased global prototype provided by the server to obtain a refined class activation map. Initial pixel-level pseudo-labels are then generated based on the class activation map, and refined pixel-level pseudo-labels are generated based on the refined class activation map.

[0086] During this process, each client processes its own image data only on its own device. This image data and intermediate results are not uploaded or shared, protecting the instance-level privacy of user data. The debiased multi-prototype federated learning framework ensures that the generated pseudo-labels conform to the local data distribution and prevents the generation of pseudo-labels that include data from other clients, thus protecting the category-level privacy of user data.

[0087] Step 2: After a certain number of rounds of client-side pseudo-label generation, each client uses the initial pixel-level pseudo-labels and refined pixel-level pseudo-labels as precise supervisory signals to train a high-performance local segmentation model, enabling deep learning to accurately map images to pixel-level semantic segmentation. The client trains the local segmentation model only on its local device, and the local segmentation model parameters are not uploaded or shared, further protecting user data privacy.

[0088] Step 3: After a certain number of rounds of client-side local segmentation model training, each client uploads the locally trained multi-prototype parameters to the central server, while the local segmentation model parameters remain locally on the client. The central server then receives the local prototype parameters uploaded by each client and initiates the prototype aggregation and update process. During the aggregation phase, the central server uses a widely used federated averaging strategy to perform a weighted average based on the weight of the client's data contribution, thereby generating a more representative trainable global prototype. The aggregated trainable global prototype is further used to train the debiasing network, resulting in a more stable and reliable unbiased global prototype, which is then distributed to each client for the next round of pseudo-label generation. After receiving the global unbiased global prototype, the client combines it with its local data to further optimize the client-side pseudo-label generation process. In this way, the model achieves cross-client knowledge sharing while effectively alleviating data heterogeneity, improving the model's generalization capability.

[0089] Step 4: After receiving the global unbiased prototype, the client returns to Step 1 and repeats the federated learning training process. The client and the central server continuously optimize the model performance through repeated iterative training and aggregation. In each iteration, the client locally trains the model parameters and uploads them, while the central server continuously aggregates and updates the model parameters. After multiple cycles, the model's segmentation quality and data heterogeneity mitigation effect gradually improve until one of the following stopping conditions is met: the model's segmentation quality on the validation set (such as the mIoU metric) converges; or training reaches the preset maximum number of iterations.

[0090] As shown in Table 1, Table 1 shows the image generation quality improvement effect and category-level privacy protection effect of the weakly supervised semantic segmentation method based on federated learning on a certain dataset. The lower the three indicators, the better.

[0091] Table 1

[0092]

[0093] Among them, Lce in Table 1 refers to the loss between the initial pixel-level pseudo-label obtained from the first type of activation map and the true label, which is calculated by the cross-entropy function (Cross-Entropy Loss); Lrefined refers to the loss between the refined pixel-level pseudo-label obtained from the second type of activation map and the true label, which is also calculated by the cross-entropy function; Lsingle-proto means that the server maintains only one prototype for each category for training, while Lk-proto means that the server maintains k prototypes for each category for training. Optionally, the training method adopts the adversarial training mechanism method.

[0094] Experimental results demonstrate that the weakly supervised semantic segmentation method based on federated learning provided by the embodiments of the present invention has significant advantages and practical value. Experimental data show that the method of the embodiments of the present invention significantly outperforms existing methods in terms of image semantic segmentation quality, as shown in the following:

[0095] In terms of image semantic segmentation quality: In terms of the mIoU indicator, the method of the embodiment of the present invention (Server: Lk-proto, Client: Lce+Lrefined) has a result of 53.89% under the parameter setting of α=0.3, which is significantly better than other federated learning baseline methods. By introducing the debiased multi-prototype federated learning framework and the refined training strategy, the embodiment of the present invention effectively utilizes the advantages of federated learning and significantly improves the segmentation quality of the weakly supervised semantic segmentation model. In contrast, the segmentation quality of the baseline method based on federated averaging (FedAvg+GradCAM) is poor (mIoU is 50.22% under the parameter setting of α=0.3), indicating that it is difficult to effectively cope with the challenge of data heterogeneity in the federated learning scenario, and the model segmentation performance is limited.

[0096] In terms of alleviating data heterogeneity: By comparing the mIoU results of different federated learning methods under different α parameters, it can be seen that the method of the embodiment of the present invention (Server: Lk-proto, Client: Lce+Lrefined) has achieved relatively stable high performance under the two parameter settings of α=0.1 and α=0.3 (mIoU is 47.92% and 53.89% respectively), indicating that the method of the present invention can effectively alleviate the data heterogeneity problem in the federated learning scenario and ensure the robustness and stability of the model under different data distributions. This result shows that the present invention effectively improves the segmentation performance of the federated weakly supervised semantic segmentation model in data heterogeneity scenarios by introducing the debiased multi-prototype federated learning framework, and significantly enhances the practical value of the model. On this basis, the refined training strategy further improves the segmentation accuracy of the model, so that the method of the present invention exhibits excellent performance in the federated weakly supervised semantic segmentation task.

[0097] In summary, by introducing a debiased multi-prototype federated learning framework and a refined training strategy, this paper significantly improves the performance of weakly supervised semantic segmentation while effectively alleviating data heterogeneity, surpassing the existing federated weakly supervised semantic segmentation baseline method, and verifying its effectiveness and practical value in practical applications.

[0098] An embodiment of the present invention further provides a weakly supervised semantic segmentation device based on federated learning, which can implement the above-mentioned weakly supervised semantic segmentation method based on federated learning. The device includes:

[0099] The first module is used to obtain a first-class activation map based on local image data and image-level labels;

[0100] The second module is used to perform adaptive region expansion on the first type of activation map through an unbiased global prototype to obtain a second type of activation map;

[0101] A third module is configured to perform model training on an initial local segmentation model based on the local image data, the first type activation map, and the second type activation map;

[0102] The fourth module is used to update the unbiased global prototype and return to the step of obtaining the first type of activation map based on the local image data and the image-level label until the model training reaches a preset number of training rounds or a preset training target to obtain the target local segmentation model.

[0103] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0104] An embodiment of the present invention further provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the aforementioned federated learning-based weakly supervised semantic segmentation method. The electronic device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.

[0105] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0106] refer to Figure 3 , Figure 3 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0107] The processor 501 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.

[0108] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 502 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 502 and is called by the processor 501 to execute the weakly supervised semantic segmentation method based on federated learning in the embodiments of the present invention.

[0109] Input / output interface 503, used to implement information input and output;

[0110] Communication interface 504, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0111] Bus 505 , which transmits information between various components of the device (e.g., processor 501 , memory 502 , input / output interface 503 , and communication interface 504 );

[0112] The processor 501 , the memory 502 , the input / output interface 503 and the communication interface 504 are connected to each other in communication within the device via a bus 505 .

[0113] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned weakly supervised semantic segmentation method based on federated learning.

[0114] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0115] An embodiment of the present invention further provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned federated learning-based weakly supervised semantic segmentation method.

[0116] In summary, the weakly supervised semantic segmentation method and apparatus based on federated learning in the embodiments of the present invention have the following advantages:

[0117] 1. By combining the debiased multi-prototype federated learning framework with a refined training strategy, the client-side local model can effectively integrate global context information and local detail features, effectively addressing the data heterogeneity problem in the federated learning scenario while enhancing model performance, and achieving better performance in weakly supervised semantic segmentation tasks. Specifically, at the level of the debiased multi-prototype federated learning framework, the unbiased global prototype maintained by the server serves as the key information to guide the client model training during the federated learning training process, and combined with the multiple prototypes trained locally on the client, it effectively overcomes the negative impact of data heterogeneity on model training in the federated learning scenario, ensuring that the global model can learn more generalized category representations, thereby significantly improving the segmentation accuracy and generalization performance of the weakly supervised semantic segmentation model; at the level of refined training strategy, a refined training strategy is adopted in the client model training stage, and the unbiased global prototype provided by the server is used to adaptively expand the preliminary class activation map generated by the client to obtain a more accurate refined class activation map, and based on this, high-quality refined pixel-level pseudo-labels are generated, which effectively improves the problems of rough class activation map boundaries and regional underactivation in traditional weakly supervised semantic segmentation methods, significantly improves the quality of pseudo-labels, and thus greatly improves the training efficiency and segmentation accuracy of the local segmentation model.

[0118] 2. In the prototype aggregation and update steps, this embodiment of the present invention achieves cross-client knowledge sharing by uploading locally trained multiple prototypes to the client, followed by weighted aggregation and debiasing optimization on the server side. This effectively addresses data heterogeneity and improves the generalization capability of the global model. Repeated iterative optimization steps ensure continuous improvement in model segmentation performance and meet practical application requirements.

[0119] 3. In the embodiment of the present invention, the client trains the local segmentation model locally, and the parameters of the local segmentation model are only saved locally on the client and not uploaded to the server, ensuring that the model training process only uses information on local data distribution, while preventing other clients from obtaining the distribution information of local data, thereby enhancing data security.

[0120] 4. During the collaborative training process of each client using the federated unbiased global prototype and the client-side local model, the federated unbiased global prototype incorporates global knowledge from multiple parties, thereby improving the generalization performance of the client-side local model and overcoming the challenges brought by data heterogeneity.

[0121] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operation and logic flow presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.

[0122] Furthermore, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise indicated, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It will also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skill of an engineer. Therefore, a person skilled in the art using ordinary skill will be able to implement the present invention set forth in the claims without undue experimentation. It will also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0123] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0124] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0125] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0126] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0127] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0128] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

[0129] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A weakly supervised semantic segmentation method based on federated learning, characterized in that: The following steps are involved: Obtain the first type of activation map based on local image data and image-level labels; The first type of activation map is adaptively expanded by using an unbiased global prototype to obtain the second type of activation map; Performing model training on an initial local segmentation model according to the local image data, the first type activation map, and the second type activation map; The unbiased global prototype is updated, and the step of obtaining the first type of activation map according to the local image data and the image-level label is returned until the model training reaches a preset number of training rounds or a preset training target, thereby obtaining a target local segmentation model.

2. The weakly supervised semantic segmentation method based on federated learning according to claim 1, characterized in that The method of obtaining a first-class activation map based on local image data and image-level labels includes the following steps: Training a lightweight classification network using the local image data and the image-level labels; According to the trained lightweight classification network, the local image data is processed through forward propagation to generate the first type activation map.

3. The weakly supervised semantic segmentation method based on federated learning according to claim 1, characterized in that The method of performing adaptive region expansion on the first type of activation map by using an unbiased global prototype to obtain a second type of activation map comprises the following steps: receiving the unbiased global prototype; The first type of activation map is optimized by using the category prior knowledge and global information of the unbiased global prototype to obtain the second type of activation map.

4. The weakly supervised semantic segmentation method based on federated learning according to claim 1, characterized in that The performing model training on the initial local segmentation model according to the local image data, the first type activation map, and the second type activation map comprises the following steps: Based on the first type of activation map, generating a first pixel-level pseudo label; Based on the second type of activation map, generating a second pixel-level pseudo label; An initial local segmentation model is trained according to the local image data, the first pixel-level pseudo-label, and the second pixel-level pseudo-label.

5. The weakly supervised semantic segmentation method based on federated learning according to claim 4 is characterized in that The performing model training on the initial local segmentation model according to the local image data, the first pixel-level pseudo-label, and the second pixel-level pseudo-label comprises the following steps: Presetting the training parameters and loss function of the initial local segmentation model; Using the local image data, the first pixel-level pseudo-label, and the second pixel-level pseudo-label as training data, and inputting the training data into the initial local segmentation model; The initial local segmentation model is trained according to the training parameters, the training data, and the loss function.

6. The weakly supervised semantic segmentation method based on federated learning according to claim 1, characterized in that The updating of the unbiased global prototype comprises the following steps: Uploading the current unbiased global prototype to the server; Performing aggregation calculation on the current unbiased global prototype through the server to obtain a trainable global prototype; The trainable global prototype is debiased by a debiasing network to obtain the updated unbiased global prototype.

7. A weakly supervised semantic segmentation device based on federated learning, characterized in that: include: The first module is used to obtain a first-class activation map based on local image data and image-level labels; The second module is used to perform adaptive region expansion on the first type of activation map through an unbiased global prototype to obtain a second type of activation map; A third module is configured to perform model training on an initial local segmentation model based on the local image data, the first type activation map, and the second type activation map; The fourth module is used to update the unbiased global prototype and return to the step of obtaining the first type of activation map based on the local image data and the image-level label until the model training reaches a preset number of training rounds or a preset training target to obtain the target local segmentation model.

8. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.