Class activation graph-driven adaptive weight adjustment semantic segmentation method

Through the adaptive weight adjustment method driven by class activation graph, the problem of difficulty in realizing dynamic weight adjustment in the existing technology under category imbalance is solved, and the accuracy and robustness of semantic segmentation are achieved, and more flexible and personalized training strategies are provided.

CN120107574APending Publication Date: 2025-06-06YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510027901.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-06

Smart Images

  • Figure CN120107574A_ABST
    Figure CN120107574A_ABST
Patent Text Reader

Abstract

The invention discloses a class activation graph-driven adaptive weight adjustment semantic segmentation method, which comprises the following steps of: generating a pseudo mask: generating a thermodynamic diagram by utilizing a trained semantic segmentation model and a Grad-CAM method, and processing the thermodynamic diagram to obtain the pseudo mask; performance evaluation: utilizing ISTS to evaluate the segmentation performance of each category through comparison with a real mask; dynamic weight calculation: calculating the dynamic weight of each category, and adjusting the influence of different categories in the loss function; a training strategy: dynamically adjusting a loss function according to the calculated weight of each category; and iterative updating: carrying out the next round of network training according to the adjusted weight, and repeating the above steps to form a dynamically adjusted training process. The accuracy and robustness of the model in a semantic segmentation task can be remarkably improved, the limitation of the prior art is overcome, and the problems of complex scenes and unbalanced categories can be better solved. And more powerful support is provided for application in the fields of image understanding and computer vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for adaptive weight adjustment semantic segmentation driven by a class activation map. Background Art

[0002] Semantic segmentation is an important research field in computer vision, which aims to assign a category label to each pixel. With the development of deep learning, semantic segmentation methods based on neural networks have become mainstream, but in the case of class imbalance, the segmentation performance of the model often fails to meet expectations. In this case, Focal Loss was proposed and widely used due to its adaptive weighting characteristics. Focal Loss solves the problem of class imbalance by increasing the weight of samples that are difficult to classify. Although Focal Loss has obvious advantages in dealing with class imbalance problems, it has the following disadvantages and shortcomings similar to other existing loss functions:

[0003] (1) Static weight adjustment: Many existing loss functions (such as Focal Loss and other resampling-based loss functions) often use static parameters when calculating weights. This means that in each training epoch, the adjustment of class weights depends on the previously set fixed value and fails to be dynamically updated according to the current performance of the name, thus failing to achieve optimization for complex situations of individual classes.

[0004] (2) Ignoring the relative performance changes between categories: Many existing methods focus on one or two technical indicators or static objects (such as the number of samples or Mean Intersection Union), but fail to comprehensively consider the performance fluctuations of the model between different training stages, resulting in an inability to effectively focus on the types that are currently performing poorly or are ignored.

[0005] (3) Difficult to adapt to new categories: When dealing with multi-category problems, existing methods often need to readjust the loss function design when introducing new categories. It is difficult to achieve real-time adaptive adjustment, which limits the scalability and flexibility of the model.

[0006] (4) Lack of personalized optimization mechanism: Existing methods usually adopt global loss functions and weight strategies, and fail to provide personalized training strategies for the different characteristics of each category, resulting in some important categories not being fully trained. Summary of the invention

[0007] The purpose of the present invention is to provide a method for semantic segmentation driven by class activation maps with adaptive weight adjustment to address the problems existing in the current prior art, which can significantly improve the accuracy and robustness of the model in semantic segmentation tasks, overcome the limitations of the prior art, and better deal with complex scenes and unbalanced category problems. It provides more powerful support for applications in the fields of image understanding, computer vision, etc.

[0008] The technical solution of the present invention is as follows:

[0009] A method for adaptive weight adjustment semantic segmentation driven by class activation maps, comprising the following steps:

[0010] Generate pseudo-mask: Use the trained semantic segmentation model and Grad-CAM method to generate heat map and process the heat map to obtain pseudo-mask;

[0011] Performance evaluation: Use ISTS to evaluate the segmentation performance of each category by comparing with the true mask;

[0012] Calculate dynamic weights: Calculate the dynamic weights of each category and adjust the influence of different categories in the loss function;

[0013] Training strategy: dynamically adjust the loss function based on the calculated weights of each category;

[0014] Iterative update: Based on the adjusted weights, the next round of network training is carried out, and the above steps are repeated to form a dynamically adjusted training process.

[0015] Furthermore, the calculation process of ISTS includes:

[0016] Compare the pseudo-mask and the real mask and calculate the correct predictions, wrong predictions and missed detections for each category;

[0017] The ISTS value of each category is calculated using the following formula. The ISTS value reflects the segmentation effect of the category in the current epoch:

[0018]

[0019] Among them, m r and m f They represent the real mask and the pseudo mask respectively, and δ is a hyperparameter used to adjust the proportion of IoU and Pix to meet the requirements of different semantic segmentation tasks.

[0020] Furthermore, the calculation of the dynamic weight specifically includes the following steps:

[0021] Calculate Ratio_weight i: Generate Ratio_weight based on the comparison of the ISTS value of each category in the current epoch with that in the previous epoch i :

[0022]

[0023] Calculating Base_weight i : To measure the segmentation performance of each category relative to all categories:

[0024]

[0025] Calculate the final weight i :

[0026] weight i =Base_weight i ×Ratio_weight i .

[0027] Furthermore, the training strategy includes the following steps:

[0028] Introducing the calculated weights into the loss function so that the loss function reflects the relative importance and performance of the current category in each training batch;

[0029] The weighted loss function is used to ensure that the model gradually focuses on the categories with higher weights in the current epoch during training, thus achieving precise adjustment.

[0030] Furthermore, the generating of the pseudo mask comprises the following steps:

[0031] Grad-CAM back-propagates the model’s gradient information to the feature map of the input image, thereby highlighting the areas that have a greater impact on the final decision and obtaining important information about the model’s reasoning process;

[0032] The heat map is thresholded to extract salient areas and generate a pseudo mask.

[0033] Furthermore, the iterative update includes the following steps:

[0034] Feedback and performance evaluation after each epoch continuously update the model's learning strategy; the complete data feedback mechanism enables the model to focus not only on overall performance, but also pay special attention to the performance of subdivided categories, thereby achieving continuous optimization.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] 1. Dynamic weight adjustment mechanism: Provides a dynamic weight adjustment mechanism based on the actual training process. After each epoch, by introducing the comparison between the pseudo mask and the real mask, the weight of each category is calculated in real time to ensure that the model can adapt to the performance changes of different categories. This mechanism enables the model to dynamically adjust the focus according to the segmentation performance and maximize the segmentation accuracy.

[0037] 2. Comprehensive performance evaluation index: By introducing ISTS (a custom category segmentation performance evaluation index for this application) to comprehensively evaluate the segmentation performance of each category, the relative performance changes between different categories can be quantified; the evaluation method can effectively capture the performance changes of the model during training, thereby providing a basis for dynamic weight adjustment;

[0038] 3. Flexible scalability: It can easily adapt to the needs of adding new categories and allows dynamic updates in the existing training framework without redesigning the global loss function, making the model more scalable and flexible;

[0039] 4. Personalized training strategy: By dynamically adjusting weights, personalized training strategies are designed for different categories to ensure that important categories are fully trained, reduce missed and misclassified cases, and thus improve the overall segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 Flowchart of a method for adaptive weight adjustment for semantic segmentation driven by class activation maps. DETAILED DESCRIPTION

[0041] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0042] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.

[0043] See also Figure 1 , a method for semantic segmentation with adaptive weight adjustment driven by class activation maps,

[0044] A method for adaptive weight adjustment semantic segmentation driven by class activation maps, comprising the following steps:

[0045] Generate pseudo-mask: Use the trained semantic segmentation model and Grad-CAM method to generate heat map and process the heat map to obtain pseudo-mask;

[0046] Generating a pseudo mask consists of the following steps:

[0047] Grad-CAM back-propagates the model’s gradient information to the feature map of the input image, thereby highlighting the areas that have a greater impact on the final decision and obtaining important information about the model’s reasoning process;

[0048] Threshold the heatmap to extract the salient areas and generate a pseudo mask. Ensure that the strong feature areas are focused on to provide a reference for subsequent performance evaluation.

[0049] Performance evaluation: Use ISTS to evaluate the segmentation performance of each category by comparing with the true mask;

[0050] The calculation process of ISTS includes:

[0051] Compare the pseudo-mask and the real mask and calculate the correct predictions, wrong predictions and missed detections for each category;

[0052] The ISTS value of each category is calculated using the following formula. The ISTS value reflects the segmentation effect of the category in the current epoch:

[0053]

[0054] Among them, m r and m f Represent the real mask and pseudo mask respectively, and δ is a hyperparameter used to adjust the proportion of IoU and Pix to meet the requirements of different semantic segmentation tasks. Comprehensively analyze the performance of the model in the current epoch to identify poorly performing or ignored categories.

[0055] Calculate dynamic weights: Calculate the dynamic weights of each category and adjust the influence of different categories in the loss function;

[0056] The calculation of dynamic weight specifically includes the following steps:

[0057] Calculate Ratio_weight i : Generate Ratio_weight based on the comparison of the ISTS value of each category in the current epoch with that in the previous epoch i :

[0058]

[0059] It can effectively reflect the performance changes of the category in each round of training and help the model to be optimized more specifically.

[0060] Calculating Base_weight i : To measure the segmentation performance of each category relative to all categories:

[0061]

[0062] Ensure that the model can pay attention to the comparison between the current better performing categories and the overall performance, and protect the model from over-focusing on minority categories and ignoring the overall performance.

[0063] Calculate the final weight i :

[0064] weight i =Base_weight i ×Ratio_weight i .

[0065] The final weights will be used in the next round of training to adjust the influence of different categories in the loss function, giving priority to training and optimizing categories with poor performance.

[0066] Training strategy: dynamically adjust the loss function based on the calculated weights of each category;

[0067] The training strategy consists of the following steps:

[0068] Introducing the calculated weights into the loss function so that the loss function reflects the relative importance and performance of the current category in each training batch;

[0069] The weighted loss function is used to ensure that the model gradually focuses on the categories with higher weights in the current epoch during training, thus achieving precise adjustment.

[0070] Iterative update: Based on the adjusted weights, the next round of network training is carried out, and the above steps are repeated to form a dynamically adjusted training process.

[0071] The iterative update process includes the following steps:

[0072] Feedback and performance evaluation after each epoch continuously update the model's learning strategy; the complete data feedback mechanism enables the model to focus not only on overall performance, but also pay special attention to the performance of subdivided categories, thereby achieving continuous optimization.

[0073] It can realize dynamic and flexible training strategies, ensure that the model can fully pay attention to the characteristics of each category at different training stages, and improve the overall performance and accuracy of semantic segmentation. This dynamic weight adjustment mechanism can efficiently cope with complex real-world scenarios and highly unbalanced category distributions, providing new advantages and possibilities for computer vision applications.

[0074] This application has the following characteristics, which show obvious differences and advantages compared with the prior art:

[0075] 1. Dynamic weight adjustment mechanism:

[0076] Different from the traditional static weighting method (such as fixed category weights or weights based on the number of samples), this application proposes a dynamic weight adjustment mechanism that adjusts the weights of the above categories in real time through performance evaluation of each epoch. The focus of the model can be flexibly adjusted according to the current classification performance, ensuring that the model can respond to changes in the performance differences of each category during the training process in a timely manner.

[0077] 2. Feedback mechanism based on pseudo-mask:

[0078] This application introduces the generation and evaluation of pseudo masks, combining the Grad-CAM technology with the comparison of real masks to form a novel feedback mechanism. This mechanism allows the model to not only rely on the prediction results, but also obtain intuitive information about important feature areas, and conduct a deeper understanding and analysis of segmentation performance. Compared with traditional methods, this comprehensive mechanism can better represent the depth and accuracy of model learning.

[0079] 3. Comprehensive performance evaluation index ISTS:

[0080] This application uses the ISTS metric as a custom performance evaluation standard that can fully reflect the segmentation effect of each category. The ISTS metric comprehensively considers multiple factors such as segmentation accuracy, inter-category coordination, and error distribution, so it is more comprehensive than a single IoU, precision, or recall. Compared with many existing methods that only rely on a single metric, its applicability and flexibility are improved.

[0081] 4. Targeted training strategy:

[0082] This application provides a personalized training strategy that introduces the weights to be calculated in the loss function through dynamic weight adjustment of each category, so that the categories with poor performance can be given priority during the training process. This targeted training can effectively reduce missed and misclassified cases, and is particularly suitable for processing complex scenes and highly unbalanced segmentation tasks.

[0083] 5. Good scalability:

[0084] This method shows good scalability when dealing with the introduction of new categories. Traditional methods usually require resetting the loss function according to the category, but with the current method, after the introduction of new categories, the model can quickly adapt dynamically without redundant manual adjustments, thereby simplifying the model training process.

[0085] 6. Adaptive learning ability:

[0086] By dynamically adjusting weights and policy feedback, this application demonstrates stronger adaptive learning capabilities. The model can continuously learn and adjust during the training process to better adapt to the diversity and complexity of data, thereby improving the final segmentation accuracy.

[0087] 7. Effectively deal with category imbalance:

[0088] This application focuses on balancing the influence between categories through a dynamic weight mechanism, significantly improving the learning ability of minority categories. Compared with traditional resampling or fixed weight methods, the present invention can more efficiently optimize the learning effect in unbalanced data sets.

[0089] With the realization of the above features, this application not only improves the expressiveness of the semantic segmentation model, but also provides a new solution to deal with the complexity and uncertainty in practical applications, and has high practical application potential. These innovations give this method a clear advantage in semantic segmentation tasks and promote the development of related technologies.

[0090] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.

Claims

1. A method for semantic segmentation with adaptive weight adjustment driven by class activation maps, characterized in that The following steps are involved: Generate pseudo-mask: Use the trained semantic segmentation model and Grad-CAM method to generate heat map and process the heat map to obtain pseudo-mask; Performance evaluation: Use ISTS to evaluate the segmentation performance of each category by comparing with the true mask; Calculate dynamic weights: Calculate the dynamic weights of each category and adjust the influence of different categories in the loss function; Training strategy: dynamically adjust the loss function based on the calculated weights of each category; Iterative update: Based on the adjusted weights, the next round of network training is carried out, and the above steps are repeated to form a dynamically adjusted training process.

2. The method for semantic segmentation driven by class activation map adaptive weight adjustment according to claim 1, characterized in that: The calculation process of ISTS includes: Compare the pseudo-mask and the real mask and calculate the correct predictions, wrong predictions and missed detections for each category; The ISTS value of each category is calculated using the following formula. The ISTS value reflects the segmentation effect of the category in the current epoch: Among them, m r and m f They represent the real mask and the pseudo mask respectively, and δ is a hyperparameter used to adjust the proportion of IoU and Pix to meet the requirements of different semantic segmentation tasks.

3. The method for semantic segmentation driven by class activation map with adaptive weight adjustment according to claim 1 or 2, characterized in that: The calculation of the dynamic weight specifically includes the following steps: Calculate Ratio_weight i : Generate Ratio_weight based on the comparison of the ISTS value of each category in the current epoch with that in the previous epoch i : Calculating Base_weight i : To measure the segmentation performance of each category relative to all categories: Calculate the final weight i : weight i =Base_weight i ×Ratio_weight i 。 4. The method for semantic segmentation driven by class activation map adaptive weight adjustment according to claim 3, characterized in that: The training strategy includes the following steps: Introducing the calculated weights into the loss function so that the loss function reflects the relative importance and performance of the current category in each training batch; The weighted loss function is used to ensure that the model gradually focuses on the categories with higher weights in the current epoch during training, thus achieving precise adjustment.

5. The method for semantic segmentation driven by class activation map adaptive weight adjustment according to claim 1, characterized in that: The generating of the pseudo mask comprises the following steps: Grad-CAM back-propagates the model’s gradient information to the feature map of the input image, thereby highlighting the areas that have a greater impact on the final decision and obtaining important information about the model’s reasoning process; The heat map is thresholded to extract salient areas and generate a pseudo mask.

6. The method for semantic segmentation driven by class activation map adaptive weight adjustment according to claim 1, characterized in that: The iterative update comprises the following steps: Feedback and performance evaluation after each epoch continuously update the model's learning strategy; the complete data feedback mechanism enables the model to focus not only on overall performance, but also pay special attention to the performance of subdivided categories, thereby achieving continuous optimization.