Model training methods, devices, electronic equipment, and computer-readable storage media
By determining the preset number of positive samples and loss weights for the target segmentation model, and combining the class loss values for dense supervision and hard sample mining, the problem of low accuracy for sparse classes is solved, and the accuracy and overall performance of the model for sparse classes are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2026-03-06
AI Technical Summary
The training dataset for target segmentation models contains a small number of sparse class samples, resulting in low accuracy for sparse classes. Furthermore, existing loss functions fail to effectively account for the impact of missed detections on other classes.
By determining the number of positive samples in a preset category, calculating the loss weights, and determining the category loss value for each pixel based on the annotation information and segmentation results, the model is trained by combining the comprehensive loss value, thus achieving dense supervision and hard sample mining.
It improves the accuracy of the target segmentation model for sparse categories, reduces the impact of missed detections on sparse categories, optimizes the training direction of the model, and improves the overall performance.
Smart Images

Figure CN114821066B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to artificial intelligence technology, and in particular to a model training method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] The datasets used for training object segmentation models often have some categories with significantly fewer samples than other categories, or some categories have a very small percentage of samples in the entire dataset. These samples belong to sparse categories, and currently, object segmentation models have low accuracy for sparse categories. Summary of the Invention
[0003] To address the aforementioned technical problems, this disclosure is proposed. Embodiments of this disclosure provide a model training method, apparatus, and computer-readable storage medium.
[0004] According to one aspect of the present disclosure, a model training method is provided, comprising:
[0005] Based on the annotation information of the sample image, the number of positive samples corresponding to each of the N preset categories is determined, resulting in N positive sample counts. Any positive sample count is the number of pixels in the sample image that have the corresponding preset category.
[0006] Based on the number of N positive samples, the loss weights corresponding to the N preset categories are determined, resulting in N loss weights;
[0007] The sample image is segmented using the target segmentation model to be trained, and the segmentation result is obtained.
[0008] Based on the annotation information and the segmentation result, the category loss value corresponding to the N preset categories is determined for each pixel in the sample image, and the N category loss values of the pixel are obtained.
[0009] For each pixel in the sample image, a comprehensive loss value for that pixel is determined based on the N loss weights and the N category loss values for that pixel.
[0010] The target segmentation model is trained based on the comprehensive loss value of each pixel in the sample image.
[0011] According to another aspect of the present disclosure, a model training apparatus is provided, comprising:
[0012] The first acquisition module is used to determine the number of positive samples corresponding to N preset categories based on the annotation information of the sample image, and obtain N positive sample counts. The number of any positive sample is the number of pixels in the sample image that have the corresponding preset category.
[0013] The second acquisition module is used to determine the loss weights corresponding to the N preset categories based on the number of N positive samples obtained by the first acquisition module, and thus obtain N loss weights.
[0014] The third acquisition module is used to segment the sample image using the target segmentation model to be trained, and obtain the segmentation result.
[0015] The fourth acquisition module is used to determine the category loss value corresponding to the N preset categories for each pixel in the sample image based on the annotation information and the segmentation result obtained by the third acquisition module, so as to obtain the N category loss values of the pixel.
[0016] The determination module is used to determine the comprehensive loss value of each pixel in the sample image based on the N loss weights obtained by the second acquisition module and the N category loss values of the pixel obtained by the fourth acquisition module.
[0017] The training module is used to train the target segmentation model based on the comprehensive loss value of each pixel in the sample image determined by the determining module.
[0018] According to another aspect of the present disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program for performing the above-described model training method.
[0019] According to another aspect of the present disclosure, an electronic device is provided, the electronic device comprising:
[0020] processor;
[0021] Memory used to store the processor's executable instructions;
[0022] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-described model training method.
[0023] Based on the model training method provided in the above embodiments of this disclosure, the number of N positive samples corresponding to N preset categories can be determined based on the annotation information of the sample images. The number of positive samples of any preset category is closely related to the sparsity of the preset category. Based on the number of N positive samples, loss weights can be reasonably determined for each of the N preset categories according to their respective sparsity, thereby obtaining N loss weights. For each pixel in the sample image, based on the annotation information and the segmentation result obtained by the target segmentation model to be trained, N category loss values corresponding to N preset categories can be determined. These N category loss values reflect the loss of the pixel in each preset category. Subsequently, the N loss weights and N category loss values are combined to determine the comprehensive loss value of the pixel. The obtained comprehensive loss value is used to train the target segmentation model. On the one hand, it can ensure that all positive and negative samples corresponding to each preset category participate in the convergence supervision process of the model, so as to achieve dense supervision of samples. This takes into account the impact of missed detections in one preset category on other preset categories, which helps to reduce the impact of missed detections on the accuracy of sparse categories. On the other hand, the sparsity of each preset category is taken into account during the model training process, which helps to improve the contribution of sparse categories to the optimization direction of the model. Therefore, the embodiments of this disclosure can improve the accuracy of the finally trained target segmentation model for sparse categories.
[0024] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0025] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0026] Figure 1 This is a schematic flowchart of a model training method provided in an exemplary embodiment of this disclosure.
[0027] Figure 2 This is a flowchart illustrating a model training method provided in another exemplary embodiment of this disclosure.
[0028] Figure 3 This is a flowchart illustrating a model training method provided in yet another exemplary embodiment of this disclosure.
[0029] Figure 4 This is a flowchart illustrating a model training method provided in yet another exemplary embodiment of this disclosure.
[0030] Figure 5 This is a flowchart illustrating a model training method provided in yet another exemplary embodiment of this disclosure.
[0031] Figure 6 This is a schematic diagram of the model training principle in an exemplary embodiment of this disclosure.
[0032] Figure 7 This is a schematic diagram of the structure of a model training apparatus provided in an exemplary embodiment of the present disclosure.
[0033] Figure 8 This is a schematic diagram of the structure of a model training apparatus provided in another exemplary embodiment of this disclosure.
[0034] Figure 9 This is a schematic diagram of the structure of a model training apparatus provided in another exemplary embodiment of the present disclosure.
[0035] Figure 10 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation
[0036] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0037] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0038] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0039] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0040] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0041] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.
[0042] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0043] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0044] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0045] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0046] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0047] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.
[0048] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.
[0049] Application Overview
[0050] Object segmentation models can be used to perform object segmentation tasks and are often applied in fields such as autonomous driving. Before applying an object segmentation model, it needs to be trained.
[0051] In the process of realizing this disclosure, the inventors discovered that the datasets used for training object segmentation models generally have some categories with significantly fewer samples than other categories, or some categories have a small proportion of samples in the entire dataset. For example, for sample images in the dataset, most pixels are true categories 1 or 2, and only a small number of pixels are true categories 3.
[0052] It's important to note that during model convergence, the optimization direction is dominated by categories 1 and 2, with category 3 contributing very little. This leads to lower accuracy for category 3 in the final trained target segmentation model. Furthermore, even a small number of pixels that are actually in category 1 or 2 but are falsely detected as category 3 can significantly impact the accuracy of category 3, further reducing the accuracy of the final trained target segmentation model for category 3. Additionally, the loss function used during model training is typically the cross-entropy (CE) loss function. This function usually only calculates the loss value for positive sample regions of each category (the current training method can be considered sparse supervision), failing to consider the impact of a missed detection in one category on other categories. However, a missed detection in one category inevitably leads to false detections in other categories, and a certain number of false detections have a significantly greater impact on the accuracy of category 3 than on category 1 or 2.
[0053] Exemplary methods
[0054] Figure 1 This is a schematic flowchart of a model training method provided in an exemplary embodiment of this disclosure. Figure 1 The method shown includes steps 110, 120, 130, 140, 150 and 160, which are explained below.
[0055] Step 110: Based on the annotation information of the sample image, determine the number of positive samples corresponding to each of the N preset categories, and obtain the N positive sample counts. The number of any positive sample is: the number of pixels in the sample image that have the corresponding preset category.
[0056] It should be noted that training a target segmentation model generally requires a large number of sample images, and the processing method for each sample image is similar. Therefore, the embodiments of this disclosure mainly focus on the processing method for a single sample image. The sample image involved in step 110 refers to a single sample image.
[0057] In the embodiments of this disclosure, the sample images can be pre-annotated manually so that the sample images have annotation information, which may include the true category of each pixel in the sample image.
[0058] In this way, based on the annotation information of the sample images, for each of the N preset categories, the number of pixels in the sample image whose true category is that preset category (i.e., the number of pixels with that preset category) can be counted. The count obtained can be used as the number of positive samples corresponding to that preset category. In this way, the number of N positive samples corresponding one-to-one with the N preset categories can be obtained. Optionally, N can be 5, 10, 20, 30 or other values, which will not be listed here.
[0059] Step 120: Based on the number of N positive samples, determine the loss weights corresponding to the N preset categories to obtain N loss weights.
[0060] In step 120, based on the number of N positive samples, a loss weight corresponding to each of the N preset categories can be determined. This loss weight is the weight required for subsequent loss calculation, thus obtaining N loss weights that correspond one-to-one with the N preset categories.
[0061] Step 130: The sample image is segmented using the target segmentation model to be trained to obtain the segmentation result.
[0062] In step 130, the sample image can be provided as input to the target segmentation model to be trained. The target segmentation model can then perform calculations to achieve target segmentation of the sample image and obtain the segmentation result.
[0063] Step 140: Based on the annotation information and segmentation results, determine the category loss values corresponding to N preset categories for each pixel in the sample image, and obtain the N category loss values for that pixel.
[0064] In step 140, the annotation information can be compared with the segmentation results, and based on the comparison results, N category loss values corresponding to N preset categories can be determined for each pixel in the sample image.
[0065] Step 150: For each pixel in the sample image, determine the comprehensive loss value of the pixel based on N loss weights and N category loss values of the pixel.
[0066] Optionally, in step 150, for each pixel in the sample image, N loss weights can be directly used to weight the N category loss values of the pixel, and the processing result can be used as the comprehensive loss value of the pixel; wherein, the weighting processing can be weighted summation processing or weighted average processing.
[0067] In a specific example, N is 3, and the N preset categories are c1, c2, and c3. The loss weight corresponding to c1 is denoted as w1, the loss weight corresponding to c2 is denoted as w2, and the loss weight corresponding to c3 is denoted as w3. The category loss value of a pixel in the sample image corresponding to c1 is denoted as L1, the category loss value corresponding to c2 is denoted as L2, and the category loss value corresponding to c3 is denoted as L3. The comprehensive loss value is denoted as L... 总 Then we can have:
[0068] L 总 =w1*L1+w2*L2+w3*L3
[0069] It should be noted that for each pixel in the sample image, instead of directly using the N loss weights to weight the N class loss values of that pixel, you can first optimize the N loss weights and then weight the N class loss values of that pixel based on the optimization results.
[0070] It should be noted that after obtaining the processing result through weighted processing, the processing result may not be directly used as the comprehensive loss value of the pixel. Instead, the processing result can be optimized (for example, the processing result can be mapped to a specified numerical range), and the optimized result can be used as the comprehensive loss value of the pixel.
[0071] Step 160: Train the target segmentation model based on the comprehensive loss value of each pixel in the sample image.
[0072] Assuming the number of sample images is s1, and the total number of pixels in each sample image is s2, then the total number of pixels in s1 sample images is s1*s2. By executing step 150 above multiple times, s1*s2 comprehensive loss values corresponding one-to-one with s1*s2 pixels can be obtained.
[0073] In step 160, the ratio of the combined loss values s1*s2 to s1*s2 can be calculated, and the ratio can be used as the model loss value. Based on the model loss value, the model parameters of the target segmentation model can be adjusted according to the stochastic gradient descent method, thereby achieving the training of the target segmentation model.
[0074] Based on the model training method provided in the above embodiments of this disclosure, the number of N positive samples corresponding to N preset categories can be determined based on the annotation information of the sample images. The number of positive samples of any preset category is closely related to the sparsity of the preset category. Based on the number of N positive samples, loss weights can be reasonably determined for each of the N preset categories according to their respective sparsity, thereby obtaining N loss weights. For each pixel in the sample image, based on the annotation information and the segmentation result obtained by the target segmentation model to be trained, N category loss values corresponding to N preset categories can be determined. These N category loss values reflect the loss of the pixel in each preset category. Subsequently, the N loss weights and N category loss values are combined to determine the comprehensive loss value of the pixel. The obtained comprehensive loss value is used to train the target segmentation model. On the one hand, it can ensure that all positive and negative samples corresponding to each preset category participate in the convergence supervision process of the model, so as to achieve dense supervision of samples. This takes into account the impact of missed detections in one preset category on other preset categories, which helps to reduce the impact of missed detections on the accuracy of sparse categories. On the other hand, the sparsity of each preset category is taken into account during the model training process, which helps to improve the contribution of sparse categories to the optimization direction of the model. Therefore, the embodiments of this disclosure can improve the accuracy of the finally trained target segmentation model for sparse categories.
[0075] In one optional example, when determining the loss weights for each of the N preset categories based on the number of N positive samples, the loss weights are negatively correlated with the number of positive samples.
[0076] Here, the first objective function for obtaining the loss weights from the number of positive samples can be determined through multiple experiments. The first objective function can be a function that ensures a negative correlation between the independent and dependent variables. During model training, the number of N positive samples corresponding to each of the N preset categories can be counted online. By substituting the number of positive samples corresponding to each preset category into the first objective function for online calculation, the loss weight corresponding to that preset category can be obtained, thus obtaining the N loss weights corresponding to the N preset categories.
[0077] In the embodiments of this disclosure, the loss weight is negatively correlated with the number of positive samples. Therefore, the preset category with a smaller number of positive samples will be assigned a larger loss weight. That is, sparse categories will be assigned a larger weight than dense categories. This is beneficial to improve the contribution of sparse categories to the optimization direction of the model, thereby improving the accuracy of the target segmentation model trained in the end for sparse categories.
[0078] exist Figure 1 The illustrated embodiment, as Figure 2 As shown, step 120 includes steps 1201, 1203 and 1205.
[0079] Step 1201: Calculate the product of the first positive sample quantity and the preset coefficient. The first positive sample quantity is any positive sample quantity among N positive sample quantities.
[0080] Here, the preset coefficient can be represented as a, the number of the first positive samples can be represented as cur_value, and the product of the number of the first positive samples and the preset coefficient can be represented as a*cur_value.
[0081] Step 1203: Determine the relationship between the product and the first preset value.
[0082] Here, a comparator can be used to compare the product with a first preset value to obtain the relationship between the product and the first preset value. Optionally, the first preset value can be 1.
[0083] Step 1205: Based on the number and size relationship of the second positive samples, determine the loss weight corresponding to the first preset category. The second positive sample number is the positive sample number with the largest value among the N positive sample numbers, and the first preset category is the preset category corresponding to the first positive sample number.
[0084] Here, a comparator can be used to compare the sizes of the N positive sample counts to obtain their relative sizes, thereby determining the largest positive sample count (i.e., the second positive sample count). Optionally, the second positive sample count can be represented as max_value.
[0085] Next, the loss weight corresponding to the first preset category can be determined based on the second number of positive samples and the size relationship determined in step 1203. In one specific implementation, such as Figure 3 As shown, step 1205 includes steps 12051, 12053 and 12055.
[0086] Step 12051: Based on the size relationship, determine the larger value between the product and the first preset value.
[0087] Assuming the product is represented as a*cur_value and the first preset value is 1, the larger of the product and the first preset value can be represented as max(a*cur_value, 1).
[0088] Step 12053: Calculate the ratio of the number of the second positive sample to the larger value.
[0089] If the number of the second positive samples is denoted as max_value, then the ratio of the number of the second positive samples to the larger value can be expressed as max_value / max(a*cur_value, 1).
[0090] Step 12055: Based on the ratio, determine the loss weight corresponding to the first preset category. The loss weight corresponding to the first preset category is positively correlated with the ratio.
[0091] Optionally, max_value / max(a*cur_value, 1) can be directly used as the loss weight corresponding to the first preset category, or the product of max_value / max(a*cur_value, 1) and a set value greater than 0 can be used as the loss weight corresponding to the first preset category, so as to ensure that the loss weight corresponding to the first preset category and the ratio maintain a positive correlation.
[0092] In this implementation, by combining the relationship between the product and the first preset value, as well as simple division operations, the loss weight can be determined efficiently and accurately.
[0093] It should be noted that when determining the loss weights, a clamp function can be introduced to limit the loss weights to a given range. For example, let's assume the loss weight corresponding to the first preset category is represented as W. s Then we have:
[0094] W s =clamp[max_value / max(a*cur_value, 1), min=1, max=b]
[0095] Optionally, the values of a and b in the above formula can be adjusted according to the actual situation.
[0096] According to the above formula, if max_value / max(a*cur_value, 1) is greater than 1 and less than b, W s =max_value / max(a*cur_value, 1), if max_value / max(a*cur_value, 1) is less than or equal to 1, W s =1, if max_value / max(a*cur_value, 1) is greater than or equal to b, W s =b.
[0097] As can be seen, in the embodiments of this disclosure, the determination of loss weights can be achieved efficiently and accurately through the execution of simple operations such as multiplication and numerical comparison.
[0098] exist Figure 1 Based on the illustrated embodiments, as Figure 4 As shown, step 150 includes steps 1501, 1503, 1505 and 1507.
[0099] Step 1501: For each pixel in the sample image, based on the segmentation result, determine the predicted probability values corresponding to N preset categories for that pixel, and obtain N predicted probability values.
[0100] In step 130, after the sample image is provided as input to the target segmentation model to be trained, the target segmentation model can obtain the segmentation result by segmenting the sample image. The segmentation result can include N predicted probability values of N preset categories for each pixel in the sample image. There can be a one-to-one correspondence between the N preset categories and the N predicted probability values. The predicted probability value corresponding to any preset category is the probability value that the target segmentation model predicts that the category of the pixel is the preset category.
[0101] Thus, in step 1501, N predicted probability values corresponding to N preset categories can be obtained for each pixel from the segmentation results.
[0102] Step 1503: Based on N predicted probability values, determine N loss correction coefficients corresponding to each of the preset categories for the pixel point, and obtain N loss correction coefficients.
[0103] In step 1503, for each preset category, the correction loss coefficient corresponding to the preset category can be determined based on the predicted probability value corresponding to the preset category, thereby obtaining N correction loss coefficients that correspond one-to-one with the N preset categories.
[0104] Step 1505: For each of the N loss weights, the loss weight is corrected using the corresponding loss correction coefficient among the N loss correction coefficients to obtain the correction weight corresponding to the loss weight, thereby obtaining the N correction weights.
[0105] In step 1505, for each of the N loss weights, the loss weight can be multiplied by the corresponding loss correction coefficient, and the resulting product is used as the correction weight corresponding to the loss weight. In this way, N correction weights that correspond one-to-one with the N loss weights can be obtained, and there is also a one-to-one correspondence between the N correction weights and the N preset categories.
[0106] Step 1507: Based on N correction weights, the N category loss values of the pixel are weighted to obtain the comprehensive loss value of the pixel.
[0107] In a specific example, N is 3, and the N preset categories are c1, c2, and c3. The true category of a certain pixel in the sample image is c1. Then the information for that pixel in the annotation information of the sample image can be represented in the form of (1, 0, 0). In this form, the "1" in (1, 0, 0) corresponds to c1, the first "0" in (1, 0, 0) corresponds to c2, and the second "0" in (1, 0, 0) corresponds to c3.
[0108] Assuming the predicted probability value of the pixel obtained from the segmentation result is p1 for c1, p2 for c2, and p3 for c3, then based on the difference between p1 and 1 and the preset loss function, we can calculate the class loss value L1 for the pixel corresponding to c1; based on the difference between p2 and 0 and the preset loss function, we can calculate the class loss value L2 for the pixel corresponding to c2; and based on the difference between p3 and 0 and the preset loss function, we can calculate the class loss value L3 for the pixel corresponding to c3. Thus, we obtain three class loss values for this pixel, and there is a one-to-one correspondence between the three class loss values and the three preset classes c1, c2, and c3.
[0109] Assume the correction weight corresponding to c1 is denoted as z1, the correction weight corresponding to c2 is denoted as z2, the correction weight corresponding to c3 is denoted as z3, and the overall loss value is denoted as L. 总 Then we can have:
[0110] L 总 = z1*L1+z2*L2+z3*L3
[0111] In the embodiments of this disclosure, for each pixel in a sample image, based on the segmentation result, N predicted probability values corresponding to N preset categories can be determined for that pixel. These N predicted probability values reflect the difficulty of the target segmentation model in predicting the category for that pixel. Specifically, for a preset category that is the same as the true category of the pixel, if the predicted probability value is closer to 1 (i.e., the predicted value is closer to the true value), it means that for the target segmentation model, it is easier to obtain a prediction result that matches the true value when predicting for that preset category. In this case, the pixel can be considered a simple sample relative to the preset category. Conversely, if the predicted probability value is closer to 0 (i.e., the predicted value differs more from the true value), it means that for the target segmentation model, it is more difficult to obtain a prediction result that matches the true value when predicting for that preset category. In this case, the pixel can be considered a difficult sample relative to the preset category. Difficult samples: For a preset category different from the true category of a pixel, if the predicted probability value is closer to 0 (i.e., the predicted value is closer to the true value), then for the target segmentation model, it is easier to obtain a prediction result that matches the true value when predicting for that preset category. In this case, the pixel can be considered a simple sample relative to that preset category. Conversely, if the predicted probability value is closer to 1 (i.e., the predicted value differs more from the true value), then for the target segmentation model, it is more difficult to obtain a prediction result that matches the true value when predicting for that preset category. In this case, the pixel can be considered a difficult sample relative to that preset category. Based on N predicted probability values, loss correction coefficients can be reasonably determined for each of the N preset categories, taking into account the difficulty of category prediction by the target segmentation model. This results in N loss correction coefficients, which are then used to correct N loss weights. The corrected N correction weights are then used to weight the N category loss values of the pixel. This effectively utilizes the results of difficult sample mining, giving more attention to samples that have not converged well during model training, thereby improving the performance of the final trained target segmentation model.
[0112] In an optional example, when determining the loss correction coefficients for N preset categories for a pixel based on N predicted probability values, the following must be satisfied:
[0113] For a preset category that is the same as the true category of the pixel, the loss correction coefficient is negatively correlated with the predicted probability value;
[0114] For a preset category that is different from the true category of the pixel, the loss correction coefficient is positively correlated with the predicted probability value;
[0115] The true category of the pixel is determined based on the annotation information.
[0116] Here, the second objective function for obtaining the loss correction coefficient from the predicted probability value and the third objective function for obtaining the loss correction coefficient from the predicted probability value can be determined through multiple experiments. The second objective function can be a function that can ensure a negative correlation between the independent variable and the dependent variable, and the third objective function can be a function that can ensure a positive correlation between the independent variable and the dependent variable.
[0117] For a preset category that is the same as the true category of the pixel, the loss correction coefficient corresponding to that preset category can be obtained by substituting the predicted probability value of that preset category into the second objective function. For a preset category that is different from the true category of the pixel, the loss correction coefficient corresponding to that preset category can be obtained by substituting the predicted probability value of that preset category into the third objective function. In this way, N loss correction coefficients corresponding to N preset categories can be obtained.
[0118] In the embodiments of this disclosure, for a preset category that is the same as the true category of the pixel, the loss correction coefficient is negatively correlated with the predicted probability value. Therefore, the smaller the predicted probability value corresponding to the preset category (i.e., the closer it is to 0), the larger the loss correction coefficient will be assigned to that preset category; that is, more difficult samples will be assigned a larger loss correction coefficient. Conversely, for a preset category that is different from the true category of the pixel, the loss correction coefficient is positively correlated with the predicted probability value. Therefore, the larger the predicted probability value corresponding to the preset category (i.e., the closer it is to 1), the larger the loss correction coefficient will be assigned to that preset category; that is, more difficult samples will be assigned a larger loss correction coefficient. Thus, in the embodiments of this disclosure, difficult samples are assigned a larger loss correction coefficient than easy samples. This allows the model to give more attention to samples that have not converged well during training, thereby improving the performance of the final trained target segmentation model.
[0119] exist Figure 4 Based on the illustrated embodiments, as Figure 5 As shown, step 1503 includes steps 15031, 15033, and 15035.
[0120] Step 15031: Determine the true category of the pixel based on the annotation information.
[0121] Since the annotation information can include the true category of each pixel in the sample image, the true category of the pixel can be determined efficiently and accurately based on the annotation information.
[0122] Step 15033: In response to the second preset category being the same as the true category of the pixel, calculate the difference between the second preset value and the predicted probability value corresponding to the second preset category, and determine the loss correction coefficient corresponding to the second preset category based on the difference. The second preset category is any preset category among N preset categories, and the loss correction coefficient corresponding to the second preset category is positively correlated with the difference.
[0123] Here, the second preset value can be 1. Assuming the predicted probability value corresponding to the second preset category is represented as px, the difference between the second preset value and the predicted probability value corresponding to the second preset category can be represented as 1 - px.
[0124] In step 15033, 1-px can be directly used as the loss correction coefficient corresponding to the second preset category, or the product of 1-px and a set value greater than 0 can be used as the loss correction coefficient corresponding to the second preset category, or an exponentiation operation can be performed with 1-px as the base and a set value greater than 0 as the exponent, and the result can be used as the loss correction coefficient corresponding to the second preset category, so as to ensure that the loss correction coefficient corresponding to the second preset category and 1-px maintain a positive correlation.
[0125] Step 15035: In response to the fact that the second preset category is different from the true category of the pixel, a loss correction coefficient corresponding to the second preset category is determined based on the predicted probability value corresponding to the second preset category. The loss correction coefficient corresponding to the second preset category is positively correlated with the predicted probability value corresponding to the second preset category.
[0126] Assuming the predicted probability value corresponding to the second preset category is represented as px, in step 15035, px can be directly used as the loss correction coefficient corresponding to the second preset category, or the product of px and a set value greater than 0 can be used as the loss correction coefficient corresponding to the second preset category, or an exponential operation can be performed with px as the base and a set value greater than 0 as the exponent, and the result can be used as the loss correction coefficient corresponding to the second preset category, so as to ensure that the loss correction coefficient corresponding to the second preset category and px maintain a positive correlation.
[0127] In a specific example, N is 3, and the N preset categories are c1, c2, and c3. The true category of a certain pixel in the sample image is c1, and the predicted probability value of this pixel corresponding to c1 is p1, the predicted probability value corresponding to c2 is p2, and the predicted probability value corresponding to c3 is p3. It is easy to see that the preset category c1 is the same as the true category of the pixel, while the two preset categories c2 and c3 are different from the true category of the pixel. Therefore, it can be determined that the loss correction coefficient corresponding to c1 is 1-p1, the loss correction coefficient corresponding to c2 is p2, and the loss correction coefficient corresponding to c3 is p3.
[0128] In the embodiments of this disclosure, the loss correction coefficient can be determined efficiently and accurately through simple operations such as subtraction. In addition, when the second preset category is the same as the true category of the pixel, since the difference is negatively correlated with the predicted probability value corresponding to the second preset category, the loss correction coefficient corresponding to the second preset category is positively correlated with the difference. The superposition of a negative correlation and a positive correlation will result in the loss correction coefficient being negatively correlated with the predicted probability value. However, when the second preset category is different from the true category of the pixel, the loss correction coefficient is positively correlated with the predicted probability value. Thus, as described above, difficult samples will be assigned a larger loss correction coefficient than simple samples, which can give more attention to samples that have not converged well during model training, thereby improving the performance of the final trained target segmentation model.
[0129] In an optional example, the target segmentation model may include an image segmentation part and a loss calculation part. In specific implementations, such as... Figure 6 As shown, sample images can be provided as input to the target segmentation model. The image segmentation part can then perform target segmentation on the sample images (target segmentation includes a series of processes such as feature extraction) to obtain the segmentation results. The loss calculation part can perform supervised learning based on the segmentation results. The processing implemented by the loss calculation part includes dense supervision, hard sample mining, and sparse class loss weighting. Dense supervision means that loss calculation is performed on all positive and negative sample regions of each class in the target segmentation task, and all positive and negative samples participate in the convergence supervision of the model (corresponding to the scheme mentioned above that determines N class loss values corresponding to N preset classes for each pixel in the sample image). Hard sample mining refers to giving more attention to samples that are more difficult to learn from during model training, so that the model converges better to these hard samples (corresponding to the scheme mentioned above that corrects N loss weights with N loss correction coefficients); sparse class loss weighting refers to quantifying the different impacts of a certain number of false detections on each class, so that the model also supervises the degree of negative response of each class to false detections (corresponding to the scheme mentioned above that for each pixel in the sample image, based on N loss weights and N class loss values of that pixel, a comprehensive loss value for that pixel is determined, and the target segmentation model is trained based on the comprehensive loss value of each pixel in the sample image). Using Figure 6 The proposed solution can effectively improve the accuracy of the final trained target segmentation model for sparse categories and the average accuracy for each category.
[0130] Any model training method provided in this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any model training method provided in this disclosure can be executed by a processor, such as by a processor executing any model training method mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.
[0131] Exemplary device
[0132] Figure 7 This is a schematic diagram of the structure of a model training apparatus provided in an exemplary embodiment of the present disclosure. Figure 7 The apparatus shown includes a first acquisition module 710, a second acquisition module 720, a third acquisition module 730, a fourth acquisition module 740, a determination module 750, and a training module 760.
[0133] The first acquisition module 710 is used to determine the number of positive samples corresponding to N preset categories based on the annotation information of the sample image, and obtain N positive sample counts. Any positive sample count is the number of pixels in the sample image that have the corresponding preset category.
[0134] The second acquisition module 720 is used to determine the loss weights corresponding to the N preset categories based on the N positive sample counts obtained by the first acquisition module 710, and obtain the N loss weights.
[0135] The third acquisition module 730 is used to segment the sample image by the target segmentation model to be trained and obtain the segmentation result.
[0136] The fourth acquisition module 740 is used to determine the category loss values corresponding to N preset categories for each pixel in the sample image based on the annotation information and the segmentation results obtained by the third acquisition module 730, and obtain the N category loss values of the pixel.
[0137] The determination module 750 is used to determine the comprehensive loss value of each pixel in the sample image based on the N loss weights obtained by the second acquisition module 720 and the N category loss values of the pixel obtained by the fourth acquisition module 740.
[0138] Training module 760 is used to train the target segmentation model based on the comprehensive loss value of each pixel in the sample image determined by determination module 750.
[0139] In one optional example, when determining the loss weights for each of the N preset categories based on the number of N positive samples, the loss weights are negatively correlated with the number of positive samples.
[0140] In an optional example, such as Figure 8 As shown, the second acquisition module 720 includes:
[0141] The calculation submodule 7201 is used to calculate the product of the first positive sample quantity and the preset coefficient. The first positive sample quantity is any positive sample quantity among the N positive sample quantities obtained by the first acquisition module 710.
[0142] The first determining submodule 7203 is used to determine the relationship between the product calculated by the calculation submodule 7201 and the first preset value;
[0143] The second determining submodule 7205 is used to determine the loss weight corresponding to the first preset category based on the size relationship determined by the second positive sample quantity and the first determining submodule 7203. The second positive sample quantity is the positive sample quantity with the largest value among the N positive sample quantities, and the first preset category is the preset category corresponding to the first positive sample quantity.
[0144] In an optional example, such as Figure 8 As shown, the second determining submodule 7205 includes:
[0145] The first determining unit 72051 is used to determine the larger value between the product calculated by the calculation submodule 7201 and the first preset value based on the size relationship determined by the first determining submodule 7203.
[0146] The first calculation unit 72053 is used to calculate the ratio of the number of second positive samples to the larger value determined by the first determination unit 72051;
[0147] The second determining unit 72055 is used to determine the loss weight corresponding to the first preset category based on the ratio calculated by the first calculating unit 72053. The loss weight corresponding to the first preset category is positively correlated with the ratio.
[0148] In an optional example, such as Figure 9 As shown, module 750 includes:
[0149] The first acquisition submodule 7501 is used to determine the predicted probability values corresponding to N preset categories for each pixel in the sample image based on the segmentation result obtained by the third acquisition module 730, and obtain N predicted probability values.
[0150] The second acquisition submodule 7503 is used to determine the loss correction coefficients corresponding to the N preset categories for the pixel based on the N predicted probability values determined by the first acquisition submodule 7501, and obtain the N loss correction coefficients.
[0151] The third acquisition submodule 7505 is used to correct each of the N loss weights using the corresponding loss correction coefficient among the N loss correction coefficients obtained by the second acquisition submodule 7503, thereby obtaining the correction weight corresponding to the loss weight, and thus obtaining N correction weights.
[0152] The fourth acquisition submodule 7507 is used to weight the N category loss values of the pixel based on the N correction weights obtained by the third acquisition submodule 7505 to obtain the comprehensive loss value of the pixel.
[0153] In an optional example, when determining the loss correction coefficients for N preset categories for a pixel based on N predicted probability values, the following must be satisfied:
[0154] For a preset category that is the same as the true category of the pixel, the loss correction coefficient is negatively correlated with the predicted probability value;
[0155] For a preset category that is different from the true category of the pixel, the loss correction coefficient is positively correlated with the predicted probability value;
[0156] The true category of the pixel is determined based on the annotation information.
[0157] In an optional example, such as Figure 9 As shown, the second acquisition submodule 7503 includes:
[0158] The third determining unit 75031 is used to determine the true category of the pixel based on the annotation information;
[0159] The fourth determining unit 75033 is used to respond to the fact that the second preset category is the same as the true category of the pixel determined by the third determining unit 75031, calculate the difference between the second preset value and the predicted probability value corresponding to the second preset category, and determine the loss correction coefficient corresponding to the second preset category based on the difference. The second preset category is any preset category among N preset categories, and the loss correction coefficient corresponding to the second preset category is positively correlated with the difference.
[0160] The fifth determining unit 75035 is used to determine the loss correction coefficient corresponding to the second preset category based on the predicted probability value corresponding to the second preset category in response to the fact that the second preset category is different from the true category of the pixel determined by the third determining unit 75031. The loss correction coefficient corresponding to the second preset category is positively correlated with the predicted probability value corresponding to the second preset category.
[0161] Exemplary electronic devices
[0162] Below, for reference Figure 10This describes an electronic device according to embodiments of the present disclosure. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them.
[0163] Figure 10 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0164] like Figure 10 As shown, the electronic device 1000 includes one or more processors 1010 and memory 1020.
[0165] The processor 1010 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1000 to perform desired functions.
[0166] The memory 1020 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1010 may execute the program instructions to implement the model training methods of the various embodiments of this disclosure described above and / or other desired functions.
[0167] In one example, the electronic device 1000 may also include an input device 1030 and an output device 1040, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0168] For example, when the electronic device is a first device or a second device, the input device 1030 may be a microphone or a microphone array. When the electronic device is a standalone device, the input device 13 may be a communication network connector for receiving acquired input signals from the first device and the second device.
[0169] In addition, the input device 1030 may also include, for example, a keyboard, a mouse, etc. The output device 1040 can output various information to the outside. The output device 1040 may include, for example, a monitor, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0170] Of course, for the sake of simplicity, Figure 10Only some of the components of the electronic device 1000 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 1000 may include any other suitable components depending on the specific application.
[0171] Exemplary computer program products and computer-readable storage media
[0172] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps in the model training methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0173] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0174] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the model training methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0175] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0176] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0177] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0178] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0179] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0180] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0181] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0182] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A model training method, comprising: determining, based on annotation information of a sample image, a number of positive samples corresponding to each of N preset categories, to obtain N positive sample numbers, wherein any positive sample number is a number of pixel points with a corresponding preset category in the sample image; determining, based on the N positive sample numbers, a loss weight corresponding to each of the N preset categories, to obtain N loss weights; performing object segmentation on the sample image via a target segmentation model to be trained, to obtain a segmentation result; based on the annotation information and the segmentation result, determining, for each pixel point in the sample image, a category loss value corresponding to each of the N preset categories, to obtain N category loss values of the pixel point; for each pixel point in the sample image, determining, based on the N loss weights and the N category loss values of the pixel point, a comprehensive loss value of the pixel point; training the target segmentation model based on the comprehensive loss value of each pixel point in the sample image.
2. The method of claim 1, wherein, When the loss weight corresponding to each of the N preset categories is determined based on the N positive sample numbers, the loss weight is negatively correlated with the positive sample number.
3. The method of claim 1, wherein, The determination of the loss weight corresponding to each of the N preset categories based on the N positive sample numbers comprises: calculating a product of a first positive sample number and a preset coefficient, the first positive sample number being any positive sample number in the N positive sample numbers; determining a size relationship between the product and a first preset numerical value; determining, based on a second positive sample number and the size relationship, a loss weight corresponding to a first preset category, the second positive sample number being a positive sample number with the largest numerical value in the N positive sample numbers, and the first preset category being a preset category corresponding to the first positive sample number.
4. The method of claim 3, wherein, The determination of the loss weight corresponding to the first preset category based on the second positive sample number and the size relationship comprises: determining, based on the size relationship, a larger value between the product and the first preset numerical value; calculating a ratio of the second positive sample number to the larger value; determining, based on the ratio, the loss weight corresponding to the first preset category, the loss weight corresponding to the first preset category being positively correlated with the ratio.
5. The method of claim 1, wherein, The determination of the comprehensive loss value of each pixel point in the sample image based on the N loss weights and the N category loss values of the pixel point comprises: for each pixel point in the sample image, determining, based on the segmentation result, a predicted probability value corresponding to each of the N preset categories for the pixel point, to obtain N predicted probability values; determining, based on the N predicted probability values, a loss correction coefficient corresponding to each of the N preset categories for the pixel point, to obtain N loss correction coefficients; for each loss weight of the N loss weights, correcting the loss weight via a corresponding loss correction coefficient in the N loss correction coefficients, to obtain a correction weight corresponding to the loss weight, thereby obtaining N correction weights; weighting the N category loss values of the pixel point based on the N correction weights, to obtain the comprehensive loss value of the pixel point.
6. The method of claim 5, wherein, In determining the loss correction coefficients corresponding to the N preset categories for the pixel point based on the N prediction probability values, the following conditions need to be met: For a preset category that is the same as the true category of the pixel point, the loss correction coefficient is negatively correlated with the prediction probability value; For a preset category that is different from the true category of the pixel point, the loss correction coefficient is positively correlated with the prediction probability value; Wherein, the true category of the pixel point is determined based on the labeling information.
7. The method of claim 5, wherein, The determination of the loss correction coefficients corresponding to the N preset categories for the pixel point based on the N prediction probability values comprises: determining the true category of the pixel point based on the labeling information; in response to the second preset category being the same as the true category of the pixel point, calculating the difference between the second preset value and the prediction probability value corresponding to the second preset category, and determining the loss correction coefficient corresponding to the second preset category based on the difference, the second preset category being any of the N preset categories, and the loss correction coefficient corresponding to the second preset category being positively correlated with the difference; in response to the second preset category being different from the true category of the pixel point, determining the loss correction coefficient corresponding to the second preset category based on the prediction probability value corresponding to the second preset category, the loss correction coefficient corresponding to the second preset category being positively correlated with the prediction probability value corresponding to the second preset category.
8. A model training apparatus comprising: a first obtaining module configured to determine N positive sample quantities corresponding to N preset categories based on labeling information of a sample image, to obtain N positive sample quantities, and to determine the number of any positive sample quantity as the number of pixel points with a corresponding preset category in the sample image; a second obtaining module configured to determine N loss weights corresponding to the N preset categories based on the N positive sample quantities obtained by the first obtaining module, to obtain N loss weights; a third obtaining module configured to perform target segmentation on the sample image via a target segmentation model to be trained, to obtain a segmentation result; a fourth obtaining module configured to determine category loss values corresponding to the N preset categories for each pixel point in the sample image based on the labeling information and the segmentation result obtained by the third obtaining module, to obtain N category loss values for the pixel point; a determining module configured to determine, for each pixel point in the sample image, a comprehensive loss value for the pixel point based on the N loss weights obtained by the second obtaining module and the N category loss values for the pixel point obtained by the fourth obtaining module; a training module configured to train the target segmentation model based on the comprehensive loss value for each pixel point in the sample image determined by the determining module.
9. A computer-readable storage medium, the storage medium storing a computer program, the computer program being configured to execute the model training method of any one of claims 1-7.
10. An electronic device comprising: a processor; a memory configured to store instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the model training method according to any one of claims 1-7.
Citation Information
Patent Citations
Semantic segmentation method based on parameter importance incremental learning
CN112101364A