Semantic segmentation model training method, semantic segmentation method and device

By adopting the U-Net architecture, hollow convolution and adaptive learning rate scheduling methods in tree image segmentation, problems such as category imbalance and gradient disappearance in tree image segmentation are solved, and higher segmentation accuracy and model stability are achieved.

CN119579904BActive Publication Date: 2025-05-13北京爱宾果科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510138150.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-13
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

During the semantic segmentation of tree images, the number of pixels in the background category is much larger than that of the tree category, which leads to the model tending to predict background categories and it is difficult to capture the small and complex features of the edges of the tree, resulting in unsatisfactory segmentation effect. At the same time, deep neural networks are prone to encountering gradient explosion or gradient disappearance during training.

Method used

A semantic segmentation model based on U-Net architecture is adopted, and hollow convolution is introduced to capture tree structures of different scales. Combined with Dice coefficient loss and weighted cross entropy loss optimization model, the learning rate is adjusted with momentum, and the learning rate is smoothly adjusted through cosine annealing scheduling.

Benefits of technology

It effectively improves the segmentation accuracy of tree images and the generalization ability of the model, reduces the risk of gradient disappearance or explosion, and achieves a more stable training process and finer parameter adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119579904B_ABST
    Figure CN119579904B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and in particular to a training method, a semantic segmentation method and a device for a semantic segmentation model. The method includes obtaining an image containing trees and preprocessing it to construct an image set; performing image segmentation based on a U-Net architecture, wherein the U-Net architecture includes an encoder, a decoder and a jump connection; determining a total loss function based on a Dice coefficient loss and a weighted cross entropy loss to optimize the image background and tree category imbalance; adjusting the learning rate in each training step by combining the first-order moment estimation and the second-order moment estimation of the gradient with an Adam optimizer, and implementing the Adam optimizer learning rate attenuation adjustment in combination with a cosine annealing schedule; and implementing the training of a semantic segmentation model based on the configured Adam optimizer and the cosine annealing schedule until the semantic segmentation model is obtained by convergence. The semantic segmentation model thus obtained can facilitate the subsequent image segmentation processing application in an educational robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a training method for a semantic segmentation model, a semantic segmentation method and a semantic segmentation device. Background Art

[0002] In the field of deep learning and computer vision, image segmentation tasks, especially semantic segmentation tasks, are of great significance in multiple application scenarios, such as autonomous driving, medical image analysis, remote sensing image processing, etc. In these applications, the goal of the model is to classify each pixel in the input image to generate an accurate pixel-level segmentation map.

[0003] Educational robots can also use technologies such as convolutional neural networks (CNN) to analyze tree images, helping students learn and practice the basic principles of computer vision and machine learning. For example, by training models to identify different tree species, students can master skills such as image preprocessing, feature extraction, and model optimization.

[0004] However, in the process of analyzing tree images, such as the semantic segmentation of tree images included in educational robots, there are still the following problems:

[0005] In tree images, the number of pixels in the background category is usually much larger than that of the tree category (trunk, leaves, etc.). This category imbalance makes traditional loss functions (such as cross entropy loss) unable to train the model well because the model tends to predict the background category;

[0006] In addition, the edges of trees (such as leaves and branches) are usually complex and small, requiring the model to capture detailed features in detail. However, due to the large number of background pixels, the model tends to ignore these fine areas, resulting in unsatisfactory segmentation results.

[0007] Secondly, the current training of deep neural networks often faces problems such as gradient explosion and gradient vanishing. Especially when the network is deep and the training data is complex, the setting of the learning rate directly affects the convergence speed and final performance of the model. Summary of the invention

[0008] In view of the above-mentioned shortcomings of the prior art, the present invention provides a training method, a semantic segmentation method and a device for a semantic segmentation model, which can effectively solve the problem in the prior art that the semantic segmentation model for tree images cannot improve the segmentation accuracy and model generalization ability.

[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0010] The present invention provides a training method for a semantic segmentation model, a semantic segmentation method and a device, which at least include:

[0011] The steps include:

[0012] Obtain images containing trees and preprocess them to construct an image set;

[0013] Image segmentation is performed based on the U-Net architecture, which includes an encoder, a decoder, and skip connections;

[0014] The atrous convolution receptive field is introduced to capture tree structures of different scales. The expression of the atrous convolution is:

[0015] ;

[0016] in, Represents the value of the convolution result at position s, represents the value of the input image at position s, represents the i-th weight of the convolution kernel, represents the void ratio, Indicates the size of the convolution kernel;

[0017] The total loss function is determined based on Dice coefficient loss and weighted cross entropy loss to optimize the imbalance of image background and tree categories;

[0018] The Adam optimizer is combined with the first-order moment estimation and the second-order moment estimation of the gradient to adjust the learning rate in each training step, and the Adam optimizer learning rate attenuation adjustment is implemented in combination with the cosine annealing schedule. The expression of the learning rate attenuation adjustment is:

[0019] ;

[0020] in, Indicates the current number of steps The learning rate, represents the maximum learning rate threshold, represents the minimum learning rate, Indicates the maximum number of steps for training, and the learning rate is obtained by As a decay adjustment of the learning rate in the Adam optimizer;

[0021] The semantic segmentation model is trained based on the configured Adam optimizer and cosine annealing schedule until convergence to obtain a semantic segmentation model.

[0022] The encoder includes:

[0023] Convolution layer, performs 3x3 convolution on the input image to extract local features;

[0024] Pooling layer, performs 2x2 maximum pooling operation.

[0025] Furthermore, by introducing 1x1 convolution to optimize the feature fusion at different levels, we have:

[0026] Let the input image be the feature map , H and W represent height and width respectively, Indicates the number of input channels, and the output is a feature map , Represents the number of output channels, then the calculation formula for 1x1 convolution is:

[0027] ;

[0028] in, Represents the value of the input feature map at position (i, j) and channel c, represents the weight on the cth channel of the convolution kernel, Indicates offset top.

[0029] Furthermore, the total loss function is determined as follows:

[0030] For Dice coefficient :

[0031] ;

[0032] in, represents the i-th pixel value of the true label, Represents the i-th pixel value predicted by the model;

[0033] Defining Dice coefficient loss :

[0034] ;

[0035] Weighted Cross Entropy Loss :

[0036] ;

[0037] in, represents the total number of categories, represents the weight of the i-th category;

[0038] Combine the Dice coefficient and weighted cross entropy loss to define the total loss function for:

[0039] ;

[0040] represents the coefficient that balances the two losses.

[0041] Furthermore, the Adam optimizer combines the first-order moment estimate and the second-order moment estimate of the gradient to adjust the learning rate in each training step as follows:

[0042] First moment estimate:

[0043] ;

[0044] in, represents the attenuation coefficient of the first-order moment, represents the current gradient, represents the first moment estimate of the gradient;

[0045] Second moment estimate:

[0046] ;

[0047] in, represents the attenuation coefficient of the second-order moment, represents the second-order moment estimate of the gradient;

[0048] Modified first- and second-order moment estimates:

[0049] ;

[0050] in, and They represent the modified first-order moment estimate and second-order moment estimate respectively. Indicates the current number of steps;

[0051] Update the parameters:

[0052] ;

[0053] in, Indicates The parameters of the step, represents the learning rate, represents a constant, represents gradient noise.

[0054] Furthermore, the gradient noise Determined according to the following relationship:

[0055] ;

[0056] in, represents the noise scaling factor, It means that the mean is 0 and the variance is The noise drawn from the normal distribution of Represents the standard deviation of the noise.

[0057] Furthermore, the noise proportional factor is determined according to the following relationship:

[0058] ;

[0059] in, Indicates the current number of steps The noise scaling factor, Represents the initial noise intensity.

[0060] A semantic segmentation method, comprising:

[0061] Determine an image to be segmented;

[0062] The image to be segmented is input into the semantic segmentation model trained by any one of the above methods to obtain a semantic segmentation result of the image to be segmented.

[0063] A training device for a semantic segmentation model is implemented according to any one of the training methods for a semantic segmentation model described above.

[0064] A computer-readable storage medium stores a computer program, which implements the steps of any one of the above methods when executed by a processor.

[0065] Compared with the known prior art, the technical solution provided by the present invention has the following beneficial effects:

[0066] 1. By introducing atrous convolution, the model’s ability to capture features of trees of different scales is enhanced, and the ability to segment large trees, small trees, or trees of different shapes is improved. The Dice coefficient loss is used to help the model better handle the overlap of multi-scale features and optimize the overall segmentation effect of the model.

[0067] 2. By combining the Adam optimizer with momentum, the Adam optimizer can adaptively adjust the learning rate of each parameter, optimize the parameters more effectively during the training process, and smoothly adjust the learning rate in combination with the cosine annealing strategy to avoid the shock caused by excessive learning rate in the early stage of training, thus achieving a smooth training process;

[0068] 3. The learning rate is smoothly reduced through the cosine function of cosine annealing scheduling, so that in the training process, the learning rate is larger in the early stage to help the model converge quickly; and in the later stage of training, the learning rate gradually decreases, so that the model can perform more precise parameter adjustments to avoid overfitting, which can facilitate the subsequent image segmentation processing application in educational robots. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0070] Figure 1 It is a schematic diagram of the overall method of the present invention. DETAILED DESCRIPTION

[0071] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0072] In the image analysis application of educational robots, especially in tree images, the number of pixels in the background category is usually much larger than that in the tree category (trunk, leaves, etc.). This category imbalance makes it difficult for traditional loss functions (such as cross entropy loss) to train the model well, because the model tends to predict the background category.

[0073] In addition, the edges of trees (such as leaves and branches) are usually complex and small, requiring the model to capture detailed features in detail. However, due to the large number of background pixels, the model tends to ignore these fine areas, resulting in unsatisfactory segmentation results.

[0074] Secondly, the current training of deep neural networks often faces problems such as gradient explosion and gradient vanishing. Especially when the network is deep and the training data is complex, the setting of the learning rate directly affects the convergence speed and final performance of the model.

[0075] To this end, the present invention proposes a training method for a semantic segmentation model.

[0076] The present invention will be further described below in conjunction with the embodiments.

[0077] See also Figure 1 , the training method of the semantic segmentation model at least includes:

[0078] Collect images containing trees, obtain a variety of images of trees of different species, seasons, and lighting conditions, generate a label map for each image, and the value of each pixel in the label map represents the category to which the pixel belongs, such as background, trunk, leaves, etc., then construct an image set and perform preprocessing;

[0079] The tree image segmentation task based on U-Net architecture includes:

[0080] The encoder gradually extracts features from the image, using convolutional layers and pooling layers to reduce the spatial resolution while increasing the depth of the features. The specific structure is:

[0081] Convolutional layer: extracts local features by performing convolution operations on the input image. Each convolutional layer has multiple filters (convolution kernels), and each convolution kernel can extract specific patterns in the image (such as edges, textures, shapes, etc.);

[0082] Using 3x3 convolution can preserve spatial information while having low computational complexity;

[0083] After each convolution layer, the depth (number of channels) of the image increases, and the nonlinear characteristics are enhanced based on the activation function (such as ReLU);

[0084] Pooling layer: Pooling operation (maximum pooling or average pooling) is used to reduce the spatial resolution of the image, that is, to reduce the height and width of the image. The pooling layer can help reduce the amount of calculation and extract more semantic features;

[0085] Use 2x2 max pooling to reduce each region of the image to a maximum value, thereby reducing the spatial resolution;

[0086] The pooling operation helps capture the global information of the image while reducing the risk of overfitting;

[0087] Hierarchical structure: Through multiple convolution and pooling operations, the spatial resolution of the image data is gradually reduced, but each convolution operation increases the number of channels of the feature map, allowing the image to capture more semantic information;

[0088] The decoder gradually restores the spatial resolution through deconvolution (transposed convolution) to generate the final segmented image. The structure specifically includes:

[0089] Deconvolution: The convolution kernel is used to perform reverse operations in space, thereby increasing the spatial resolution of the feature map and expanding the spatial dimension through interpolation while maintaining the semantic information of the feature. Deconvolution inserts blank pixels and performs convolution calculations to gradually restore the spatial structure of the image.

[0090] Upsampling: Upsampling is performed using methods such as nearest neighbor interpolation or bilinear interpolation;

[0091] Convolutional layer: Use convolutional layers to refine the results, remove noise and unnecessary details, and improve the quality of the output. The decoder finally outputs a segmented image, and the spatial resolution has been restored to the same as the input image.

[0092] Among them, for tree images, when processing large-scale trees of different scales, there may be huge differences in tree morphology. Therefore, the hole convolution is introduced to increase the receptive field and better capture the tree structure of different scales.

[0093] ;

[0094] in, Represents the value of the convolution result at position s, represents the value of the input image at position s, represents the i-th weight of the convolution kernel, It represents the void rate, which describes how many positions are used to perform convolution. Indicates the size of the convolution kernel and the void rate Control the spacing between convolution kernel elements. The larger the dilation rate, the larger the receptive field, so that the dilated convolution can effectively capture a wide range of contextual information, thereby improving the perception of tree structures at different scales.

[0095] The jump connection connects the corresponding layers of the encoder and decoder, retains the low-level feature information, and ensures that no details are lost during image restoration. Among them, the introduction of 1x1 convolution further improves the feature fusion of different levels and enhances the interaction between the low layer (detail layer) and the high layer (semantic layer). The 1x1 convolution can compress or expand the dimension of the low-level features while maintaining the spatial information and enhancing the expression ability of the model. Then:

[0096] Let the input image be the feature map , H and W represent height and width respectively, Indicates the number of input channels, and the output is a feature map , Represents the number of output channels, then the calculation formula for 1x1 convolution is:

[0097] ;

[0098] in, Represents the value of the input feature map at position (i, j) and channel c, that is, the pixel value of the input image at that position and the feature on the cth channel. Represents the weight on the cth channel of the convolution kernel, which is used to perform convolution operation on each channel of the input feature map. Represents the bias top, which is used to adjust the result after convolution calculation;

[0099] Therefore, standard convolution and pooling layers are used in the encoder for feature extraction, and hole convolution is introduced to increase the receptive field and enhance the network's ability to capture features of trees at different scales. 1x1 convolution is introduced between the encoder and decoder to strengthen the transmission of low-level features, thereby improving the segmentation accuracy of fine areas (such as leaf edges).

[0100] It should be noted that the specific structures of the encoder, decoder and skip connection in the U-Net architecture are conventional structures and are not described one by one in this implementation.

[0101] Furthermore, the segmentation accuracy of tree images is ensured by optimizing the loss function, especially in dealing with class imbalance and detail segmentation, including:

[0102] Determine the total loss function based on Dice coefficient loss and weighted cross entropy loss;

[0103] For Dice coefficient :

[0104] ;

[0105] in, Represents the i-th pixel value of the true label, 1 represents the tree category, 0 represents the background category, Represents the i-th pixel value predicted by the model;

[0106] Defining Dice coefficient loss :

[0107] ;

[0108] Since the background category in the tree image usually occupies most pixels, it is necessary to give different weights to each category in the cross entropy loss to reduce the impact of the background on the model training. For the weighted cross entropy loss, there is:

[0109] ;

[0110] in, represents the total number of categories, represents the weight of the i-th category, the weight of the background is usually smaller, and the weight of the tree category is larger;

[0111] Combine the Dice coefficient and weighted cross entropy loss to define the total loss function for:

[0112] ;

[0113] represents the coefficient that balances the two losses;

[0114] In the image segmentation task of trees in the above scheme, the pixel difference between the background and the trees is large. The use of weighted cross entropy loss helps to reduce the dominant role of background pixels in the training process and increase the influence of the tree part. The Dice coefficient loss ensures more accurate segmentation by measuring the overlap between the true label and the predicted label. Combining these two loss functions can not only optimize the handling of class imbalance during training (through weighted cross entropy loss), but also achieve fine segmentation of minority categories (trees) (through Dice coefficient loss).

[0115] Furthermore, the Adam optimizer uses the first-order moment (momentum) and second-order moment (mean of the square of the gradient) of the gradient to adjust the learning rate to accelerate the training process, improve the convergence speed of the model, and effectively avoid the disappearance or explosion of the gradient. Then:

[0116] First moment estimate:

[0117] ;

[0118] in, represents the attenuation coefficient of the first-order moment, represents the current gradient, Represents the first-order moment estimate of the gradient, describing the exponential decay average of the gradient, and is used to estimate the direction of the gradient;

[0119] Second moment estimate:

[0120] ;

[0121] in, represents the attenuation coefficient of the second-order moment, Represents the second-order moment estimate of the gradient, which represents the exponentially decaying average of the square of the gradient and is used to estimate the magnitude of the gradient;

[0122] To avoid initialization phase bias, correct the first and second moment estimates:

[0123] ;

[0124] in, and They represent the modified first-order moment estimate and second-order moment estimate respectively. Indicates the current number of steps;

[0125] Update the parameters:

[0126] ;

[0127] in, Indicates The parameters of the step, represents the learning rate, represents a constant, Represents gradient noise (by adding it to the parameter update rule to promote random perturbations of the gradient, enhance the model's exploration ability, and avoid falling into local optimal solutions), Perform a square root operation and adjust the step size of each parameter to ensure that the learning process does not have gradient explosion or vanishing problems.

[0128] Among them, for the gradient noise :

[0129] ;

[0130] in, represents the noise scaling factor, It means that the mean is 0 and the variance is The noise drawn from the normal distribution of Represents the standard deviation of the noise and controls the intensity of the noise.

[0131] The Adam optimizer combines momentum and adaptive learning rate to help the model improve robustness in complex image segmentation tasks when faced with images of different resolutions, objects of different shapes, complex textures and backgrounds.

[0132] Furthermore, considering that the edge parts of tree images (such as leaves and branches) are usually small and complex, in order to ensure that the model can perform more refined optimization when segmenting detail areas, the learning rate in the Adam optimizer is further adjusted based on the cosine annealing schedule. Specifically, the cosine annealing schedule gradually decays the learning rate of the Adam optimizer during the training process, so that the model can be updated with a smaller step size in the later stage of training to ensure more refined convergence. Then:

[0133] ;

[0134] in, Indicates the current number of steps The learning rate, Represents the maximum learning rate threshold (by setting the maximum learning rate threshold, the learning rate can be trimmed to ensure that the learning rate during training is within a reasonable range and avoid instability caused by too large a step size). represents the minimum learning rate (the learning rate at the end of training), Indicates the maximum number of steps for training;

[0135] Furthermore, based on the above, the cosine annealing schedule will gradually reduce the learning rate , and the learning rate calculated by the cosine annealing schedule As in the Adam optimizer Dynamically adjusted learning rate helps the model to make more detailed parameter adjustments in the later stages of training;

[0136] In summary, the Adam optimizer can adaptively adjust the step size of parameter updates according to the gradient. However, considering that the learning rate is too large in the early stage of training, resulting in instability, gradually reducing the learning rate through cosine annealing scheduling helps the model to converge finely in the later stage of training. Therefore, the combination of the two can utilize the adaptive ability of the Adam optimizer and the progressive learning rate reduction strategy of cosine annealing to achieve the purpose of balancing training efficiency and stability, avoid unnecessary oscillations caused by excessive learning rate, and further improve model performance.

[0137] It should be noted that the above scheme can avoid falling into the local optimal solution and maintain a certain degree of exploration by introducing gradient noise. However, excessive noise will lead to instability in the training process or even failure to converge. Therefore, for the noise scaling factor Dynamically adjust as training progresses, that is, as the number of training steps increases As the learning rate increases, the noise intensity is gradually reduced, thereby ensuring that the model can converge more carefully in the later stage of training, and the noise intensity is adjusted by changing the learning rate. For example, as the learning rate decreases, the noise intensity gradually decreases, making the transition from exploration to fine-tuning in the training process smoother:

[0138] ;

[0139] in, Indicates the current number of steps The noise scaling factor, represents the initial noise intensity, through Substitute above to control gradient noise.

[0140] Therefore, based on the configured Adam optimizer and cosine annealing schedule, model training is performed, including:

[0141] Forward propagation: pass each input image into the network to get the model's predicted output;

[0142] Loss calculation: Calculate the loss (including Dice coefficient loss and weighted cross entropy loss) based on the true label and the model's prediction results;

[0143] Back propagation: calculate the gradient of the loss function with respect to the model parameters;

[0144] Parameter update: Update network weights and biases based on the calculated gradients using the Adam optimizer;

[0145] If the training reaches the expected convergence state, a semantic segmentation model applied to tree images is generated.

[0146] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. The training method of the semantic segmentation model includes the following steps: Obtain images containing trees and preprocess them to construct an image set; Image segmentation is performed based on the U-Net architecture, which includes an encoder, a decoder, and a jump connection, and is characterized in that: The atrous convolution receptive field is introduced to capture tree structures of different scales. The expression of the atrous convolution is: ; in, Represents the value of the convolution result at position s, represents the value of the input image at position s, represents the i-th weight of the convolution kernel, represents the void ratio, Indicates the size of the convolution kernel; The total loss function is determined based on Dice coefficient loss and weighted cross entropy loss to optimize the imbalance of image background and tree categories; The Adam optimizer is combined with the first-order moment estimation and the second-order moment estimation of the gradient to adjust the learning rate in each training step, and the Adam optimizer learning rate attenuation adjustment is implemented in combination with the cosine annealing schedule. The expression of the learning rate attenuation adjustment is: ; in, Indicates the current number of steps The learning rate, represents the maximum learning rate threshold, represents the minimum learning rate, Indicates the maximum number of steps for training, and the learning rate is obtained by As a decay adjustment of the learning rate in the Adam optimizer; The semantic segmentation model is trained based on the configured Adam optimizer and cosine annealing schedule until convergence to obtain a semantic segmentation model; In the U-Net architecture, 1x1 convolution is introduced to optimize the feature fusion of different levels, and then: Let the input image be the feature map , H and W represent height and width respectively, Indicates the number of input channels, and the output is a feature map , Represents the number of output channels, then the calculation formula for 1x1 convolution is: ; in, Represents the value of the input feature map at position (i, j) and channel c, represents the weight on the cth channel of the convolution kernel, Indicates bias top; Introduce gradient noise for parameter update; Gradient Noise Determined according to the following relationship: ; in, represents the noise scaling factor, It means that the mean is 0 and the variance is The noise drawn from the normal distribution of represents the standard deviation of the noise; The noise scaling factor is determined according to the following relationship: ; in, Indicates the current number of steps The noise scaling factor is Represents the initial noise intensity.

2. The method for training a semantic segmentation model according to claim 1, characterized in that: The encoder comprises: Convolution layer, performs 3x3 convolution on the input image to extract local features; Pooling layer, performs 2x2 maximum pooling operation.

3. The method for training a semantic segmentation model according to claim 1, characterized in that: The total loss function is determined as follows: For Dice coefficient : ; in, represents the i-th pixel value of the true label, Represents the i-th pixel value predicted by the model; Defining Dice coefficient loss : ; Weighted Cross Entropy Loss : ; in, represents the total number of categories, represents the weight of the i-th category; Combine the Dice coefficient and weighted cross entropy loss to define the total loss function : ; represents the coefficient that balances the two losses.

4. The method for training a semantic segmentation model according to claim 1, characterized in that: The Adam optimizer combines the first-order moment estimate and the second-order moment estimate of the gradient to adjust the learning rate in each training step as follows: First moment estimate: ; in, represents the attenuation coefficient of the first-order moment, represents the current gradient, represents the first moment estimate of the gradient; Second moment estimate: ; in, represents the attenuation coefficient of the second-order moment, represents the second moment estimate of the gradient; Modified first- and second-order moment estimates: ; in, and They represent the modified first-order moment estimate and second-order moment estimate respectively. Indicates the current number of steps; Update the parameters: ; in, Indicates The parameters of the step, represents the learning rate, represents a constant, represents gradient noise.

5. A semantic segmentation method, characterized in that: include: Determine an image to be segmented; The image to be segmented is input into the training method of the semantic segmentation model described in any one of claims 1 to 4 to obtain the semantic segmentation result of the image to be segmented.

6. A training device for a semantic segmentation model, characterized in that: This is achieved by the training method of the semantic segmentation model according to any one of claims 1-4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Remote sensing image culture pond detection method based on semantic segmentation

    CN110059758A

  • Remote sensing image vegetation extraction method and device

    CN118196629A