Semantic segmentation method and system based on improved PPLiteSeg
By introducing the Detail Head module of the STDC network model into the PPLiteSeg network model and setting it accordingly in the training and inference stages, the problem of taking into account performance and accuracy in the semantic segmentation of the big model is solved, and high precision and real-time performance are achieved.
Patent Information
- Application Number
- CN202510170205.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-13
AI Technical Summary
When semantic segmentation is performed based on large models, how to achieve the balance of semantic segmentation performance and accuracy, especially in real-time scenarios.
By introducing the Detail Head module of the STDC network model into the PPLiteSeg network model, Detail Head is turned on during the training phase to enhance segmentation accuracy, Detail Head is turned off during the inference phase to maintain real-time performance, and the model is optimized through data augmentation and joint loss functions.
The accuracy of semantic segmentation is improved, while maintaining real-time performance in the inference stage, and the adaptability of the model and the accuracy of the segmentation results are enhanced through data augmentation and loss function optimization.
Smart Images

Figure CN120147631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and specifically to a semantic segmentation method and system based on improved PPLiteSeg. Background Art
[0002] With the continuous development of deep learning technology and the continuous improvement of computer performance, semantic segmentation technology has made remarkable progress. However, traditional semantic segmentation algorithms have a relatively high computational complexity, resulting in limited real-time performance. Therefore, real-time semantic segmentation technology has emerged, aiming to improve the performance of semantic segmentation in real-time scenarios.
[0003] The main challenges faced by real-time semantic segmentation technology include high computational complexity and high real-time performance requirements.
[0004] To address these challenges, researchers have proposed various solutions:
[0005] 1. Lightweight model structure: Using lightweight model structures to reduce the computational amount of the model and improve real-time performance. These lightweight model structures reduce the complexity of the model through methods such as decomposed convolution and dilated convolution;
[0006] 2. Reducing the input image size: Reducing the size of the input image can significantly reduce the computational amount, but it will lead to information loss and accuracy degradation. Therefore, while ensuring real-time performance, it is necessary to maintain the size and information integrity of the input image as much as possible;
[0007] 3. Model pruning and compression: Reducing the number of model parameters and computational amount through model pruning and compression techniques to improve real-time performance. These techniques include weight pruning, quantization, knowledge distillation, etc.;
[0008] 4. Dual-path structure: Adopting a dual-path structure to process spatial information and semantic information separately. One path processes spatial information to retain rich detailed information; the other path processes semantic information to extract global context information, and then the two paths of information are fused to obtain the final segmentation result.
[0009] Practical applications have high requirements for semantic segmentation methods. Although semantic segmentation has made significant leaps with the help of deep learning, the performance of real-time methods is not satisfactory.
[0010] When performing semantic segmentation based on large models, how to balance semantic segmentation performance and accuracy is a technical problem that needs to be solved. Summary of the Invention
[0011] The technical task of the present invention is to address the above deficiencies and provide an improved semantic segmentation method based on improved PPLiteSeg to solve the technical problem of how to balance semantic segmentation performance and accuracy when performing semantic segmentation based on large models.
[0012] In a first aspect, an improved semantic segmentation method based on improved PPLiteSeg of the present invention includes the following steps:
[0013] Image acquisition: Collect segmentation images from the automatic wire arrangement project as sample images, and divide the sample images into a training set and a validation set;
[0014] Image processing: Perform data augmentation on the sample images to obtain enhanced sample images;
[0015] Model construction: Introduce the Detail Head module of the STDC network model into the PPliteSeg network model, and construct a semantic segmentation model based on the improved PPliteSeg network model. In the training phase, the Detail Head module in the semantic segmentation model is turned on, and a detail map is predicted and output through the Detail Head module, and a segmentation mask is predicted and output through the overall network of the semantic segmentation model. In the inference phase, the Detail Head module in the semantic segmentation model is turned off, and the segmentation mask is output through the overall network of the semantic segmentation model;
[0016] Model training: Construct a loss function based on the true annotation map corresponding to the sample images, the detail map predicted and output by the Detail Head module in the semantic segmentation model, and the segmentation mask predicted and output by the overall network of the semantic segmentation model. Using the enhanced sample images as input images, perform model training and model verification on the semantic segmentation model by minimizing the loss function to obtain the final trained semantic segmentation model;
[0017] Model deployment: Copy the trained semantic segmentation model from the PaddlePaddle framework to the Torch framework and apply it to the business scenario.
[0018] Preferably, when performing data augmentation on the sample images, one or more data augmentation operations are performed on the sample images, and the data augmentation operations include cropping, resizing, color distortion, grayscale, blurring, and sharpening.
[0019] Preferably, during model construction, the Detail Head module with good performance in the STDC network model is integrated into the 16x downsampling layer of the encoding layer of the PPliteSeg network model.
[0020] Preferably, during model training, the Detail True is obtained by calculating the sample image corresponding to the true annotation map through the Laplace operator. The binary cross-entropy loss Bce Loss between the detail map predicted and output by the Detail Head module and the Detail True is calculated. The Dice loss Dice Loss between the detail map predicted and output by the Detail Head module and the Detail True is calculated. The online hard example mining cross-entropy loss OhemCrossEntropy Loss between the segmentation mask predicted and output by the overall network of the semantic segmentation model and the true annotation map is calculated. The loss function of the semantic segmentation model is constructed based on the Bce Loss, Dice Loss, and OhemCrossEntropy Loss;
[0021] Among them, the true annotation map is an annotation map with the same resolution as the sample image, and the value of each pixel represents the semantic category to which the pixel belongs.
[0022] In a second aspect, a semantic segmentation system based on improved PPLiteSeg according to the present invention is used to implement semantic segmentation through a semantic segmentation method based on improved PPLiteSeg as described in any one of the first aspects. The system includes an image acquisition module, an image processing module, a model construction module, a model training module, and a model deployment module;
[0023] The image acquisition module is used to perform the following: collect segmentation images from the automatic wire arrangement project as sample images, and divide the sample images into a training set and a validation set;
[0024] The image processing module is used to perform the following: perform data augmentation on the sample images to obtain enhanced sample images;
[0025] The model construction module is used to perform the following: introduce the Detail Head module of the STDC network model into the PPliteSeg network model, and construct a semantic segmentation model based on the improved PPliteSeg network model. In the training stage, the module is used to perform the following. In the semantic segmentation model, the Detail Head module is turned on, and the detail map is predicted and output through the Detail Head module, and the segmentation mask is predicted and output through the overall network of the semantic segmentation model. The module is used to perform the following. In the inference stage, the Detail Head module in the semantic segmentation model is turned off, and the segmentation mask is output through the overall network of the semantic segmentation model;
[0026] The model training module is used to perform the following: Based on the ground truth annotation map corresponding to the sample image, the detail map predicted and output by the Detail Head module in the semantic segmentation model, and the segmentation mask predicted and output by the overall network of the semantic segmentation model, construct a loss function. Using the enhanced sample image as the input image, minimize the loss function to train and validate the semantic segmentation model, and obtain the final trained semantic segmentation model;
[0027] The model deployment module is used to perform the following: Copy the trained semantic segmentation model from the PaddlePaddle framework to the Torch framework and apply it to the business scenario.
[0028] Preferably, when performing data augmentation on the sample image, the image processing module is used to perform one or more data augmentation operations on the sample image. The data augmentation operations include cropping, resizing, color distortion, grayscale, blurring, and sharpening.
[0029] Preferably, the model construction module is used to integrate the well-performing Detail Head module in the STDC network model into the 16x downsampling layer of the encoding layer of the PPliteSeg network model.
[0030] Preferably, during model training, the model training module is used to perform the following: Calculate the ground truth annotation map corresponding to the sample image through the Laplace operator to obtain Detail True. Calculate the binary cross-entropy loss Bce Loss between the detail map predicted and output by the Detail Head module and Detail True. Calculate the Dice loss Dice Loss between the detail map predicted and output by the Detail Head module and Detail True. Calculate the online hard example mining cross-entropy loss OhemCrossEntropy Loss between the segmentation mask predicted and output by the overall network of the semantic segmentation model and the ground truth annotation map. Construct the loss function of the semantic segmentation model based on Bce Loss, Dice Loss, and OhemCrossEntropy Loss;
[0031] Among them, the ground truth annotation map is an annotation map with the same resolution as the sample image, and the value of each pixel represents the semantic category to which the pixel belongs.
[0032] The semantic segmentation method and system based on the improved PPLiteSeg of the present invention have the following advantages:
[0033] 1. Integrate the well-performing Detail Head in STDC into the PPLiteSeg model, use it during training, and turn it off during inference, which improves the network segmentation accuracy and does not incur additional costs during inference;
[0034] 2. Perform one or more (randomly determine the number of augmentation times) augmentation combinations on the collected sample images: cropping, resizing, color distortion, grayscale, blurring, and sharpening. Through one or more data augmentation combinations, the diversity of the dataset is greatly increased;
[0035] 3. Use Bce Loss, Dice Loss, and OhemCrossEntropyLoss to jointly optimize the improved PPLiteSeg model. Among them, Bce Loss and Dice Loss mainly constrain the Detail Head to assist the model in achieving higher segmentation result accuracy. OhemCrossEntropyLoss only selects difficult samples with higher loss values for gradient update during training, thus focusing on more difficult-to-train samples, which helps the model better adapt to these samples and improves the segmentation result accuracy of the model;
[0036] 4. Copy the improved PPLiteSeg from the PaddlePaddle framework to the Torch framework, which provides convenience for subsequent model acceleration. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] The present invention will be further described below with reference to the drawings.
[0039] Figure 1 It is a flowchart of an improved semantic segmentation method based on the improved PPLiteSeg for Embodiment 1;
[0040] Figure 2 It is a structural block diagram of the improved PPLiteSeg network model in an improved semantic segmentation method based on the improved PPLiteSeg for Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present invention will be further described below with reference to the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and implement it. However, the embodiments cited are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0042] The embodiments of the present invention provide an improved semantic segmentation method and system based on improved PPLiteSeg, which are used to solve the technical problem of how to balance semantic segmentation performance and accuracy when performing semantic segmentation based on a large model.
[0043] Embodiment 1:
[0044] An improved semantic segmentation method based on improved PPLiteSeg of the present invention includes five steps: image acquisition, image processing, model construction, model training, and model deployment.
[0045] Step S100 Image Acquisition: Collect segmentation images from the automatic wire arranging project as sample images, and divide the sample images into a training set and a validation set.
[0046] In this embodiment, the segmentation images come from the automatic wire arranging project. There are a total of 3000 RGB images with a resolution of 512*1024, including 2400 in the training set and 600 in the validation set, which are tested online.
[0047] Step S200 Image Processing: Perform data augmentation on the sample images to obtain enhanced sample images.
[0048] As a specific implementation of image processing, when performing data augmentation on the sample images, one or more data augmentation operations are performed on the sample images. The data augmentation operations include cropping, resizing, color distortion, grayscale, blurring, and sharpening.
[0049] In this embodiment, in the training data channel, one to multiple (randomly determine the number of augmentation times) augmentation combinations are performed: cropping, resizing, color distortion, grayscale, blurring, and sharpening.
[0050] Step S300 Model Construction: Introduce the Detail Head module of the STDC network model into the PPliteSeg network model, and construct a semantic segmentation model based on the improved PPliteSeg network model. In the training stage, the Detail Head module in the semantic segmentation model is turned on, and the detail map is predicted and output through the Detail Head module, and the segmentation mask is predicted and output through the overall network of the semantic segmentation model. In the inference stage, the Detail Head module in the semantic segmentation model is turned off, and the segmentation mask is output through the overall network of the semantic segmentation model.
[0051] As a specific implementation of model construction, the Detail Head module with good performance in the STDC network model is integrated into the downsampling 16-fold layer of the encoding layer of the PPliteSeg network model.
[0052] In this embodiment, the Detail Head with good performance in STDC is integrated into the 16x downsampling layer of the encoding layer of PPliteSeg. During training, the Detail Head is enabled to increase the segmentation accuracy using the detail parameters. During inference, the DetailHead layer is disabled to eliminate the additional inference cost introduced by the detail head.
[0053] For the improved PPliteSeg network model, the enhanced image is input into the improved PPliteSeg model. The training output has two parts, and the inference output is only the overall output of the network. The two parts of the training output are respectively:
[0054] One: The overall output of the network, the predicted mask;
[0055] Two: The output of the detail head, the predicted detail map.
[0056] Step S400 Model Training: Based on the ground truth map corresponding to the sample image, the detail map predicted by the DetailHead module in the semantic segmentation model, and the segmentation mask of the overall prediction output of the semantic segmentation model, a loss function is constructed. Using the enhanced sample image as the input image, the semantic segmentation model is trained and model-validated by minimizing the loss function to obtain the final trained semantic segmentation model.
[0057] As a specific implementation of model training, the ground truth map corresponding to the sample image is calculated using the Laplace operator to obtain Detail True. Calculate the binary cross-entropy loss Bce Loss between the detail map predicted by the Detail Head module and Detail True, calculate the Dice loss Dice Loss between the detail map predicted by the Detail Head module and Detail True, and calculate the online hard example mining cross-entropy loss OhemCrossEntropy Loss between the segmentation mask of the overall prediction output of the semantic segmentation model and the ground truth map. Based on Bce Loss, Dice Loss, and OhemCrossEntropy Loss, a loss function for the semantic segmentation model is constructed. Among them, the ground truth map is an annotation map with the same resolution as the sample image, and the value of each pixel represents the semantic category to which the pixel belongs.
[0058] In this embodiment, the Ground True is segmented to obtain the Detail True through the Laplacian operator. The improved PPliteSeg detail head outputs the detail inference result, and the Bce Loss and Dice Loss are calculated between the detail inference result and the Detail True. The overall output of the network and the Seg True are used to calculate the OhemCrossEntropyLoss between them. The three losses are combined, and during the backpropagation process, the model weights and biases are continuously updated to make the similarity between the network output and the Ground True as large as possible.
[0059] Bce Loss = -(y * log(p(x))) + (1 - y) * log(1 - p(x)),
[0060]
[0061] Among them, the Bce Loss and Dice Loss mainly constrain the Detail Head to assist the model to have a higher segmentation result accuracy. The OhemCrossEntropyLoss only selects difficult samples with higher loss values for gradient update during the training process, so as to focus on more difficult-to-train samples, which helps the model better adapt to these samples and thus improves the segmentation result accuracy of the model.
[0062] Step S500 Model Deployment: Copy the trained semantic segmentation model from the PaddlePaddle framework to the Torch framework and apply it to the business scenario.
[0063] In this embodiment, the improved PPLiteSeg is copied from the PaddlePaddle framework to the Torch framework, which provides convenience for subsequent model acceleration.
[0064] The method of this embodiment integrates the Detail Head that performs well in STDC into PPliteSeg, increases the detail constraint, improves the segmentation accuracy of the model, and turns off the Detail Head during inference, which improves the segmentation accuracy without increasing any inference cost. Copying the improved PPLiteSeg from the PaddlePaddle framework to the Torch framework provides convenience for subsequent Cambrian model acceleration and Tensorrt acceleration.
[0065] Embodiment 2:
[0066] A semantic segmentation system based on the improved PPLiteSeg of the present invention includes an image acquisition module, an image processing module, a model construction module, a model training module, and a model deployment module.
[0067] The image acquisition module is used to perform the following: collect segmented images from the automatic wire arranging project as sample images, and divide the sample images into a training set and a validation set.
[0068] In this embodiment, the segmented images come from the automatic wire arranging project. There are 3,000 RGB images with a resolution of 512*1024, including 2,400 in the training set and 600 in the validation set, which are tested online.
[0069] The image processing module is used to perform the following: perform data augmentation on the sample images to obtain enhanced sample images.
[0070] As a specific implementation of the image processing module, when performing data augmentation on the sample images, one or more data augmentation operations are performed on the sample images. The data augmentation operations include cropping, resizing, color distortion, grayscale, blurring, and sharpening.
[0071] In this embodiment, in the training data channel, one to multiple (the number of augmentation times is randomly determined) augmentation combinations are performed: cropping, resizing, color distortion, grayscale, blurring, sharpening.
[0072] The model construction module is used to perform the following: introduce the Detail Head module of the STDC network model into the PPliteSeg network model, and build a semantic segmentation model based on the improved PPliteSeg network model. In the training stage, the module is used to perform the following. In the semantic segmentation model, the Detail Head module is enabled, and the detail map is predicted and output through the Detail Head module, and the segmentation mask is predicted and output through the overall network of the semantic segmentation model. The module is used to perform the following. In the inference stage, the Detail Head module in the semantic segmentation model is turned off, and the segmentation mask is output through the overall network of the semantic segmentation model.
[0073] As a specific implementation of model construction, the Detail Head module with good performance in the STDC network model is integrated into the downsampling 16-fold layer of the encoding layer of the PPliteSeg network model.
[0074] In this embodiment, the Detail Head with good performance in STDC is integrated into the downsampling 16-fold layer of the PPliteSeg encoding layer. The Detail Head is enabled during training to increase the segmentation accuracy using the detail parameters, and the DetailHead layer is turned off during inference to eliminate the additional inference cost introduced by the detail head.
[0075] For the improved PPliteSeg network model, the enhanced images are input into the improved PPliteSeg model. The training output has two parts, and the inference output has only the overall network output. The two parts of the training output are respectively:
[0076] One: The overall network output, the predicted mask;
[0077] Two: The detail head output, the predicted detail map.
[0078] The model training module is used to perform the following: construct a loss function based on the true annotation map corresponding to the sample image, the detail map predicted by the Detail Head module in the semantic segmentation model, and the segmentation mask predicted by the overall network of the semantic segmentation model, use the enhanced post-sample image as the input image, and perform model training and model verification on the semantic segmentation model by minimizing the loss function to obtain the final trained semantic segmentation model.
[0079] As a specific implementation of model training, calculate the true annotation map corresponding to the sample image through the Laplace operator to obtain Detail True, calculate the binary cross-entropy loss Bce Loss between the detail map predicted by the Detail Head module and Detail True, calculate the Dice loss Dice Loss between the detail map predicted by the Detail Head module and Detail True, calculate the online hard example mining cross-entropy loss OhemCrossEntropy Loss between the segmentation mask predicted by the overall network of the semantic segmentation model and the true annotation map, and construct the loss function of the semantic segmentation model based on Bce Loss, Dice Loss, and OhemCrossEntropy Loss. Among them, the true annotation map is an annotation map with the same resolution as the sample image, and the value of each pixel represents the semantic category to which the pixel belongs.
[0080] In this embodiment, the segmentation Ground True is used to calculate Detail True through the Laplace operator, and the improved PPliteSeg detail head outputs the detail inference result and calculates the Bce Loss and Dice Loss between the result and Detail True. The overall network output and Seg True are used to calculate the OhemCrossEntropyLoss between them. The three losses are combined, and during the backpropagation process, the model weights and biases are continuously updated to make the similarity between the network output and Ground True as large as possible.
[0081] Bce Loss = -(y * log(p(x))) + (1 - y) * log(1 - p(x)),
[0082]
[0083] Among them, Bce Loss and Dice Loss mainly constrain the Detail Head to assist the model in achieving higher segmentation accuracy. OhemCrossEntropyLoss only selects difficult samples with higher loss values for gradient update during training, thereby focusing on more difficult-to-train samples, helping the model better adapt to these samples, and thus improving the segmentation accuracy of the model.
[0084] The model deployment module is used to perform the following: copy the trained semantic segmentation model from the PaddlePaddle framework to the Torch framework and apply it to the business scenario.
[0085] In this embodiment, the improved PPLiteSeg is copied from the PaddlePaddle framework to the Torch framework, which provides convenience for subsequent model acceleration.
[0086] The system of this embodiment can implement semantic segmentation by executing the method disclosed in Embodiment 1.
[0087] The improved semantic segmentation method and system based on the improved PPLiteSeg provided by the present invention have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An improved semantic segmentation method based on improved PPLiteSeg, characterized in that: The steps include: Image acquisition: Collect segmented images from the automatic wiring project as sample images, and divide the sample images into training sets and validation sets; Image processing: perform data enhancement on the sample image to obtain the enhanced sample image; Model construction: The DetailHead module of the STDC network model is introduced into the PPliteSeg network model, and a semantic segmentation model is constructed based on the improved PPliteSeg network model. In the training stage, the Detail Head module in the semantic segmentation model is turned on, and the detail map is predicted and output by the Detail Head module, and the segmentation mask is output by the overall prediction of the network of the semantic segmentation model. In the inference stage, the Detail Head module in the semantic segmentation model is turned off, and the segmentation mask is output by the overall network of the semantic segmentation model. Model training: A loss function is constructed based on the real annotation image corresponding to the sample image, the detail image predicted by the Detail Head module in the semantic segmentation model, and the segmentation mask predicted by the overall network of the semantic segmentation model. The enhanced sample image is used as the input image, and the semantic segmentation model is trained and verified by minimizing the loss function to obtain the final trained semantic segmentation model. Model deployment: Copy the trained semantic segmentation model from the PaddlePaddle framework to the Torch framework and apply it to business scenarios.
2. The semantic segmentation method based on improved PPLiteSeg according to claim 1, characterized in that: When data augmentation is performed on a sample image, one or more data augmentation operations are performed on the sample image, and the data augmentation operations include cropping, resizing, color distortion, grayscale, blurring, and sharpening.
3. The semantic segmentation method based on improved PPLiteSeg according to claim 1, characterized in that: When building the model, the Detail Head module with good performance in the STDC network model is integrated into the 16-fold downsampling layer of the encoding layer of the PPliteSeg network model.
4. The semantic segmentation method based on improved PPLiteSeg according to claim 1, characterized in that: During model training, the Laplacian operator is used to calculate the true annotation map corresponding to the sample image to obtain Detail True, and the binary cross entropy loss Bce Loss between the detail map predicted and output by the DetailHead module and Detail True is calculated. The Dice loss Dice Loss between the detail map predicted and output by the DetailHead module and Detail True is calculated, and the online hard example mining cross entropy loss OhemCrossEntropy Loss between the segmentation mask output by the overall network prediction of the semantic segmentation model and the true annotation map is calculated. The loss function of the semantic segmentation model is constructed based on Bce Loss, Dice Loss and OhemCrossEntropy Loss. Among them, the true annotation map is a annotation map with the same resolution as the sample image, in which the value of each pixel represents the semantic category to which the pixel belongs.
5. A semantic segmentation system based on improved PPLiteSeg, characterized in that: Used to implement semantic segmentation by a semantic segmentation method based on improved PPLiteSeg as described in any one of claims 1 to 4, the system comprising an image acquisition module, an image processing module, a model building module, a model training module and a model deployment module; The image acquisition module is used to perform the following: collect segmented images from the automatic wiring project as sample images, and divide the sample images into a training set and a validation set; The image processing module is used to perform the following: perform data enhancement on the sample image to obtain an enhanced sample image; The model construction module is used to perform the following: introduce the Detail Head module of the STDC network model into the PPliteSeg network model, and build a semantic segmentation model based on the improved PPliteSeg network model. In the training stage, the module is used to perform the following: the Detail Head module in the semantic segmentation model is turned on, the detail map is predicted and output through the Detail Head module, and the segmentation mask is output through the overall prediction of the network of the semantic segmentation model. The module is used to perform the following: in the inference stage, the Detail Head module in the semantic segmentation model is turned off, and the segmentation mask is output through the overall network of the semantic segmentation model; The model training module is used to perform the following: construct a loss function based on the real annotation image corresponding to the sample image, the detail image predicted by the Detail Head module in the semantic segmentation model, and the segmentation mask predicted by the overall network of the semantic segmentation model, and use the enhanced sample image as the input image to train and verify the semantic segmentation model by minimizing the loss function to obtain the final trained semantic segmentation model; The model deployment module is used to perform the following: copy the trained semantic segmentation model from the PaddlePaddle framework to the Torch framework and apply it to the business scenario.
6. The semantic segmentation system based on improved PPLiteSeg according to claim 5, characterized in that: When data enhancement is performed on a sample image, the image processing module is used to perform one or more data enhancement operations on the sample image, and the data enhancement operations include cropping, resizing, color distortion, grayscale, blurring, and sharpening.
7. The semantic segmentation system based on improved PPLiteSeg according to claim 5, characterized in that: The model building module is used to integrate the well-performing Detail Head module in the STDC network model into the 16-fold downsampling layer of the encoding layer of the PPliteSeg network model.
8. The semantic segmentation system based on improved PPLiteSeg according to claim 5, characterized in that: During model training, the model training module is used to perform the following: calculate the true annotation map corresponding to the sample image through the Laplacian operator to obtain Detail True, calculate the binary cross entropy loss Bce Loss between the detail map predicted and output by the Detail Head module and Detail True, calculate the Dice loss Dice Loss between the detail map predicted and output by the Detail Head module and Detail True, calculate the online hard example mining cross entropy loss OhemCrossEntropy Loss between the segmentation mask output by the overall network prediction of the semantic segmentation model and the true annotation map, and construct the loss function of the semantic segmentation model based on Bce Loss, Dice Loss and OhemCrossEntropy Loss; Among them, the true annotation map is a annotation map with the same resolution as the sample image, in which the value of each pixel represents the semantic category to which the pixel belongs.