Deployment Method, Device, Storage Medium and Equipment of Semantic Segmentation Model
By creating and training a semantic segmentation model for deep learning accelerator DLA and converting it into ONNX format deployment, the problem of multi-model computing resource preemption on the GPU is solved, and efficient deployment and performance improvement of semantic segmentation models are achieved.
Patent Information
- Application Number
- CN202510287660.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Computational resource preemption leads to performance degradation when deploying multiple semantic segmentation models on GPUs.
Create a semantic segmentation model suitable for deep learning accelerator DLA, train the model through the training set, perform INT8 quantization and convert it to ONNX format, and finally deploy it on the DLA.
Without significantly sacrificing accuracy, the inference speed of the semantic segmentation model is optimized and performance is improved.
Smart Images

Figure CN119810453B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and particularly to a method, device, storage medium, and equipment for deploying a semantic segmentation model. Background Art
[0002] A semantic segmentation model is an image segmentation technology in the field of computer vision. Its core task is to assign each pixel in an image to predefined semantic categories, which can be vehicles, pedestrians, buildings, etc.
[0003] The normal operation of the perception algorithm of autonomous driving vehicles relies on the support of embedded artificial intelligence chips, such as Jetson Orin, etc. Usually, these chips are centered around the Graphics Processing Unit (GPU) in the Nvidia Ampere architecture. To ensure that the semantic segmentation model can respond correctly and quickly, the perception algorithm containing the semantic segmentation model is usually deployed on the GPU.
[0004] However, as the number of models deployed on the GPU increases, it leads to competition for computing resources among multiple models, affecting the performance of the semantic segmentation model. Summary of the Invention
[0005] This application provides a method, device, storage medium, and equipment for deploying a semantic segmentation model, which is used to solve the problem that deploying the semantic segmentation model on the GPU causes multiple models to compete for computing resources and affects the performance of the semantic segmentation model. The technical solutions are as follows:
[0006] According to the first aspect of this application, a method for deploying a semantic segmentation model is provided. The method includes:
[0007] Create a semantic segmentation model suitable for a Deep Learning Accelerator (DLA). The semantic segmentation model includes a feature extraction module, a detail extraction branch, a semantic extraction branch, a boundary extraction branch, a fusion module, and an upsampling segmentation head. The detail prediction head in the detail extraction branch corresponds to a detail loss function, the two auxiliary heads and the optimized feature pyramid in the semantic extraction branch correspond to a semantic loss function respectively, the boundary prediction head in the boundary extraction branch corresponds to a boundary loss function, and the upsampling segmentation head corresponds to two segmentation loss functions with different sampling rates;
[0008] Obtain a training set. The training samples in the training set include sample images, true class labels, and true boundary labels. The true boundary labels are obtained by performing boundary extraction on the true class labels;
[0009] For each training sample, use the semantic segmentation model to process the sample image to obtain a prediction result, and use the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions to calculate the losses of the prediction result, the true class label, and the true boundary label, and train the model parameters of the semantic segmentation model according to the calculation results;
[0010] Perform INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format;
[0011] Deploy the semantic segmentation model in the ONNX format to the DLA.
[0012] In a possible implementation manner, the using the semantic segmentation model to process the sample image to obtain a prediction result includes:
[0013] Use the feature extraction module to extract features from the sample image to obtain an image feature vector;
[0014] Use the semantic extraction branch to perform semantic extraction on the image feature vector to obtain a semantic feature vector;
[0015] Use the detail extraction branch and the semantic extraction branch to perform detail extraction on the image feature vector to obtain a detail feature vector;
[0016] Use the boundary extraction branch and the semantic extraction branch to perform boundary extraction on the image feature vector to obtain a boundary feature vector;
[0017] Use the fusion module to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector;
[0018] Use the upsampling segmentation head to process the fused feature vector to obtain a prediction result.
[0019] In a possible implementation manner, the using the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions to calculate the losses of the prediction result, the true class label, and the true boundary label includes:
[0020] Use two auxiliary heads and the optimized feature pyramid to perform semantic prediction on the corresponding semantic feature vectors respectively, and use the semantic loss functions corresponding to the two auxiliary heads and the optimized feature pyramid to calculate the losses of the corresponding semantic prediction results and the true class label;
[0021] Use the detail prediction head to perform detail prediction on the detail feature vector, and use the detail loss function to calculate the loss between the detail prediction result and the true class label;
[0022] Use the boundary prediction head to perform boundary prediction on the boundary feature vector, and use the boundary loss function to calculate the loss between the boundary prediction result and the true boundary label;
[0023] Use the upsampling segmentation head to perform prediction on the fused feature vector, and use a segmentation loss function to calculate the loss between the prediction result and the true class label; Use the upsampling segmentation head to perform prediction after upsampling the fused feature vector, and use another segmentation loss function to calculate the loss between the prediction result and the true class label.
[0024] In a possible implementation, the step of using the fusion module to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector includes:
[0025] Create a fusion algorithm supported by DLA in the fusion module;
[0026] Use the fusion algorithm to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector.
[0027] In a possible implementation, the step of using the detail extraction branch and the semantic extraction branch to perform detail extraction on the image feature vector to obtain a detail feature vector includes:
[0028] When the detail extraction branch includes three detail modules and the semantic extraction branch includes three semantic modules, use the first detail module to perform detail extraction on the image feature vector to obtain a first detail feature, and use the first semantic module to perform semantic extraction on the image feature vector to obtain a first semantic feature;
[0029] Use the second fusion algorithm supported by DLA to fuse the first detail feature and the first semantic feature to obtain a first fused feature; Use the second detail module to perform detail extraction on the first fused feature to obtain a second detail feature, and use the second semantic module to perform semantic extraction on the first semantic feature to obtain a second semantic feature;
[0030] Use the second fusion algorithm supported by DLA to fuse the second detail feature and the second semantic feature to obtain a second fused feature; Use the third detail module to perform detail extraction on the second fused feature to obtain the detail feature vector.
[0031] In a possible implementation, the step of using the boundary extraction branch and the semantic extraction branch to perform boundary extraction on the image feature vector to obtain a boundary feature vector includes:
[0032] When the boundary extraction branch includes three boundary modules and the semantic extraction branch includes three semantic modules, use the first boundary module to perform boundary extraction on the image feature vector to obtain a first boundary feature, and use the first semantic module to perform semantic extraction on the image feature vector to obtain a first semantic feature;
[0033] Use the second boundary module to perform boundary extraction on the concatenation of the first boundary feature and the first semantic feature to obtain a second boundary feature, and use the second semantic module to perform semantic extraction on the first semantic feature to obtain a second semantic feature;
[0034] Use the third boundary module to perform boundary extraction on the concatenation of the second boundary feature and the second semantic feature to obtain the boundary feature vector.
[0035] In a possible implementation, the method further includes:
[0036] Obtain a perceptual image to be recognized;
[0037] Use the feature extraction module to perform feature extraction on the perceptual image to obtain an image feature vector;
[0038] Use the semantic extraction branch to perform semantic extraction on the image feature vector to obtain a semantic feature vector;
[0039] Use the detail extraction branch and the semantic extraction branch to perform detail extraction on the image feature vector to obtain a detail feature vector;
[0040] Use the boundary extraction branch and the semantic extraction branch to perform boundary extraction on the image feature vector to obtain a boundary feature vector;
[0041] Use the fusion module to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector;
[0042] Use the upsampling segmentation head to perform prediction on the upsampled fused feature vector to obtain an image segmentation result.
[0043] According to a second aspect of the present application, there is provided a deployment device for a semantic segmentation model, the device includes:
[0044] A creation module for creating a semantic segmentation model applicable to a deep learning accelerator DLA. The semantic segmentation model includes a feature extraction module, a detail extraction branch, a semantic extraction branch, a boundary extraction branch, a fusion module, and an upsampling segmentation head. The detail prediction head in the detail extraction branch corresponds to a detail loss function, the two auxiliary heads and the optimized feature pyramid in the semantic extraction branch respectively correspond to a semantic loss function, the boundary prediction head in the boundary extraction branch corresponds to a boundary loss function, and the upsampling segmentation head corresponds to two segmentation loss functions;
[0045] An acquisition module for acquiring a training set. The training samples in the training set include sample images, true class labels, and true boundary labels. The true boundary labels are obtained by performing boundary extraction on the true class labels;
[0046] A training module for, for each training sample, processing the sample image using the semantic segmentation model to obtain a prediction result, calculating losses for the prediction result, the true class label, and the true boundary label using the detail loss function, the three semantic loss functions, the boundary loss function, and the two segmentation loss functions, and training the model parameters of the semantic segmentation model according to the calculation results;
[0047] The training module is further configured to perform INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format;
[0048] A deployment module for deploying the semantic segmentation model in the ONNX format to the DLA.
[0049] According to the third aspect of the present application, there is provided a computer-readable storage medium in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the deployment method of the semantic segmentation model as described above.
[0050] According to the fourth aspect of the present application, there is provided a computer device, and the computer device includes the deployment device of the above semantic segmentation model.
[0051] The beneficial effects of the technical solution provided by the present application at least include:
[0052] By creating a semantic segmentation model suitable for DLA, training the semantic segmentation model using a training set, and then performing INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format. Finally, deploying the ONNX-format semantic segmentation model to DLA, thereby realizing the migration of the semantic segmentation model from GPU to DLA, optimizing the inference speed of the semantic segmentation model without significantly sacrificing accuracy, and improving the performance of the semantic segmentation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 It is a schematic structural diagram of a semantic segmentation model shown according to some exemplary embodiments;
[0055] Figure 2 It is a schematic structural diagram of an upsampling segmentation head shown according to some exemplary embodiments;
[0056] Figure 3 It is a flowchart of a method for deploying a semantic segmentation model provided by an embodiment of the present application;
[0057] Figure 4 It is a flowchart of a method for deploying a semantic segmentation model provided by an embodiment of the present application;
[0058] Figure 5 It is a flowchart of a method for using a semantic segmentation model provided by an embodiment of the present application;
[0059] Figure 6 It is a schematic diagram of an image segmentation result provided by an embodiment of the present application;
[0060] Figure 7 It is a schematic diagram of an image segmentation result provided by an embodiment of the present application;
[0061] Figure 8 It is a block diagram of the structure of a device for deploying a semantic segmentation model provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0063] Figure 1A structural schematic diagram of a semantic segmentation model applicable to a Deep Learning Accelerator (DLA) is shown. The semantic segmentation model includes a feature extraction module 110, a detail extraction branch 120, a semantic extraction branch 130, a boundary extraction branch 140, a fusion module 150, and an upsampling segmentation head 160. When the semantic segmentation model further includes a preprocessing module, the input end of the preprocessing module is the input end of the semantic segmentation model, and the output end of the preprocessing module is connected to the input end of the feature extraction module 110; when the semantic segmentation model does not include a preprocessing module, the input end of the feature extraction module 110 is the input end of the semantic segmentation model, the output end of the feature extraction module 110 is respectively connected to the input ends of the detail extraction branch 120, the semantic extraction branch 130, and the boundary extraction branch 140, the output ends of the detail extraction branch 120, the semantic extraction branch 130, and the boundary extraction branch 140 are respectively connected to the input end of the fusion module 150, the output end of the fusion module 150 is connected to the input end of the upsampling segmentation head 160, and the output end of the upsampling segmentation head 160 is the output end of the semantic segmentation model.
[0064] The preprocessing module is used to preprocess the input image. Specifically, after performing data augmentation such as scaling and rotation and normalization operations on the input image, the image is input into the feature extraction module at a resolution of 512×640.
[0065] The feature extraction module 110 is used to extract features from the image output by the preprocessing module to obtain an image feature vector.
[0066] The semantic extraction branch 130 includes a semantic module 1-3, an optimized feature pyramid, and three auxiliary heads. The semantic module 1-3 is connected in series to the optimized feature pyramid. The semantic module 1 is connected to the first auxiliary head, and the semantic loss function corresponding to this auxiliary head is Loss 16X ; the semantic module 2 is connected to the second auxiliary head, and the semantic loss function corresponding to this auxiliary head is Loss 32X ; the optimized feature pyramid is connected to the third auxiliary head, and the semantic loss function corresponding to this auxiliary head is Loss 64X , where 16X, 32X, and 64X represent the sampling multiples. The semantic extraction branch 130 is used to perform semantic extraction on the image feature vector output by the feature extraction module 110 to obtain a semantic feature vector.
[0067] The detail extraction branch 120 includes a detail module 1-3 and a detail prediction head. The detail module 1-3 is connected in series to the detail prediction head, and the detail prediction head corresponds to a detail loss function Loss pThe detail extraction branch 120 is used to perform detail extraction on the image feature vector output by the feature extraction module 110 and the semantic feature vector output by the semantic extraction branch 130 to obtain a detail feature vector.
[0068] The boundary extraction branch 140 includes boundary modules 1-3 and a boundary prediction head. The boundary modules 1-3 are connected in series and then connected to the boundary prediction head, and the boundary prediction head corresponds to a boundary loss function Loss d The boundary extraction branch 140 is used to perform boundary extraction on the image feature vector output by the feature extraction module 110 and the semantic feature vector output by the semantic extraction branch 130 to obtain a boundary feature vector.
[0069] The fusion module 150 is used to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector.
[0070] The upsampling segmentation head 160 corresponds to two segmentation loss functions with different sampling rates, namely Loss 8X and Loss 4X , where 8X and 4X represent the sampling multiples. The upsampling segmentation head 160 is used to perform prediction on the upsampled fused feature vector and output a prediction result, which is the image segmentation result.
[0071] The structure of the upsampling segmentation head 160 in this application is as Figure 2 shown. The upsampling segmentation head 160 first performs convolution (conv), normalization (norm), activation (act), and upsampling (upsample) operations on the input feature vector, and then performs normalization (norm), activation (act), convolution (conv), normalization (norm), activation (act), convolution (conv) operations, and finally outputs the image segmentation result.
[0072] As Figure 3 shown, it shows the method flow chart of the deployment method of the semantic segmentation model provided by an embodiment of this application. The deployment method of the semantic segmentation model can be applied to a computer device. The deployment method of the semantic segmentation model may include:
[0073] Step 301, create a semantic segmentation model suitable for DLA.
[0074] The semantic segmentation model created in this embodiment is the Figure 1 semantic segmentation model suitable for DLA shown.
[0075] Step 302, obtain a training set. The training samples in the training set include sample images, true class labels, and true boundary labels. The true boundary labels are obtained by performing boundary extraction on the true class labels.
[0076] The training set includes multiple training samples, and each training sample includes a sample image, a true class label, and a true boundary label. The true class label is a predetermined number of class labels, such as 16 class labels. For each class label, boundary extraction can be performed on it to obtain a boundary label.
[0077] Step 303, for each training sample, use the semantic segmentation model to process the sample image to obtain a prediction result. Use the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions to calculate the loss of the prediction result, the true class label, and the true boundary label, and train the model parameters of the semantic segmentation model according to the calculation results.
[0078] After obtaining the semantic segmentation model, the semantic segmentation model can be used to process each sample image, and multiple loss functions, the true class label, and the true boundary label are used to calculate the loss, and the model parameters of the semantic segmentation are trained through multiple iterations.
[0079] Step 304, perform INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format.
[0080] Post-Training Quantization (PTQ) is a technique for quantizing a model after model training. Its main purpose is to convert a trained floating-point model (usually with FP32 precision) into a low-precision fixed-point model (such as INT8) to reduce the size of the model and accelerate the inference speed. PTQ does not require retraining the model and only needs to perform quantization on the existing model, saving time and computing resources. The quantized model can be deployed faster in resource-constrained environments, and the low-precision representation reduces the storage space and computing overhead.
[0081] After the semantic segmentation model has been iteratively trained for 85 epochs, a semantic segmentation model can be obtained under the PyTorch framework. Performing PTQ INT8 quantization training on this semantic segmentation model can convert it into a semantic segmentation model in the ONNX format.
[0082] Step 305, deploy the semantic segmentation model in the ONNX format to the DLA.
[0083] The exported semantic segmentation model in the ONNX format can all run on the DLA with INT8 precision. Changing from originally running on the GPU to running on the DLA can not only release more computing resources for the GPU running of other models, but also optimize the inference speed of the semantic segmentation model and improve the performance of the semantic segmentation model.
[0084] In summary, the method for deploying a semantic segmentation model provided by the embodiments of the present application creates a semantic segmentation model suitable for DLA, trains the semantic segmentation model using a training set, and then performs INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format. Finally, the semantic segmentation model in the ONNX format is deployed to DLA, thereby realizing the migration of the semantic segmentation model from GPU to DLA. Without significantly sacrificing accuracy, the inference speed of the semantic segmentation model is optimized, and the performance of the semantic segmentation model is improved.
[0085] As Figure 4 shown, it shows a flowchart of a method for deploying a semantic segmentation model provided by an embodiment of the present application. The method for deploying the semantic segmentation model can be applied to a computer device. The method for deploying the semantic segmentation model may include:
[0086] Step 401, create a semantic segmentation model suitable for DLA.
[0087] The semantic segmentation model created in this embodiment is Figure 1 the semantic segmentation model suitable for DLA as shown.
[0088] Step 402, obtain a training set. The training samples in the training set include sample images, true class labels, and true boundary labels. The true boundary labels are obtained by performing boundary extraction on the true class labels.
[0089] The training set includes multiple training samples. Each training sample includes a sample image, a true class label, and a true boundary label. The true class labels are a predetermined number of class labels, such as 16 class labels. For each class label, boundary extraction can be performed on it to obtain a boundary label.
[0090] Step 403, for each training sample, use a feature extraction module to extract features from the sample image to obtain an image feature vector.
[0091] The feature extraction module extracts features from the sample image to obtain an 8X image feature vector, that is, an image feature vector with a resolution of 64×80.
[0092] Step 404, use a semantic extraction branch to perform semantic extraction on the image feature vector to obtain a semantic feature vector.
[0093] The semantic modules 1-3 are semantic feature extraction modules with different widths and depths of convolution. Through the semantic extraction branch, the image feature vector is downsampled all the way to 8X->16X->32X->64X. The 64X semantic context feature has a large receptive field and strong ability to perceive the global context. Finally, after the 64x semantic context feature is processed by the optimized feature pyramid, a semantic feature vector is obtained.
[0094] When extracting the semantic feature vector, the semantic module 1 outputs the extracted features to the first auxiliary head, the semantic module 2 outputs the extracted features to the second auxiliary head, and the optimized feature pyramid outputs the extracted semantic feature vector to the third auxiliary head.
[0095] The structures of the three auxiliary heads are the same, except for the different input scales. The processing formula of the auxiliary head for the input feature vector is:
[0096] x = CONV(ConCat(x, DCN(x))) (1)
[0097] where x represents the feature vector, CONV represents convolution, ConCat represents concatenation, and DCN represents deformable convolution.
[0098] In the prior art, the original feature pyramid uses adaptive average pooling (avgpooling) to aggregate the features to the 1×1 dimension and then performs the pyramid-shaped feature aggregation operation. However, DLA INT8 does not support the operation of this operator. Therefore, we directly use a 1×1 convolution to replace the feature pyramid module, and this module is called the optimized feature pyramid.
[0099] Step 405, use the detail extraction branch and the semantic extraction branch to perform detail extraction on the image feature vector to obtain a detail feature vector.
[0100] Specifically, using the detail extraction branch and the semantic extraction branch to perform detail extraction on the image feature vector to obtain a detail feature vector may include:
[0101] (1) When the detail extraction branch includes three detail modules and the semantic extraction branch includes three semantic modules, use the first detail module to perform detail extraction on the image feature vector to obtain the first detail feature, and use the first semantic module to perform semantic extraction on the image feature vector to obtain the first semantic feature.
[0102] (2) Use the second fusion algorithm supported by DLA to fuse the first detail feature and the first semantic feature to obtain the first fusion feature; use the second detail module to perform detail extraction on the first fusion feature to obtain the second detail feature, and use the second semantic module to perform semantic extraction on the first semantic feature to obtain the second semantic feature.
[0103] (3) Use the second fusion algorithm supported by DLA to fuse the second detailed feature and the second semantic feature to obtain a second fused feature; use the third detailed module to extract details from the second fused feature to obtain a detailed feature vector.
[0104] The first detailed module is Figure 1 the detailed module 1 in Figure 1 The second detailed module is Figure 1 the detailed module 2 in
[0105] The detailed modules 1-3 are respectively detailed feature extraction modules with different convolutional depths. In the detailed extraction branch, after the detailed feature of each layer is fused (F1) with the semantic feature from the corresponding layer, it is then passed to the next layer for detailed feature extraction. Among them, the fusion (F1) is a process of weighting the features from semantics and details, and the specific formula is as follows:
[0106] (2)
[0107] (3)
[0108] where l and l-1 respectively represent the l-th layer and the (l-1)-th layer, and respectively represent the semantic feature and the detailed feature from the (l-1)-th layer, W1 is the operation of 1×1 convolution -> upsample -> 1×1 convolution, W2 is the operation of 1×1 convolution -> 3×3 convolution -> 1×1 convolution, and b represents the offset. After the feature extraction of the detailed extraction branch, the detailed information of the sample image is enriched, and no downsampling is performed at this time, and finally an 8X detailed feature vector is obtained.
[0109] Step 406, use the boundary extraction branch and the semantic extraction branch to perform boundary extraction on the image feature vector to obtain a boundary feature vector.
[0110] Specifically, using the boundary extraction branch and the semantic extraction branch to perform boundary extraction on the image feature vector to obtain a boundary feature vector may include:
[0111] (1) When the boundary extraction branch includes three boundary modules and the semantic extraction branch includes three semantic modules, use the first boundary module to perform boundary extraction on the image feature vector to obtain a first boundary feature, and use the first semantic module to perform semantic extraction on the image feature vector to obtain a first semantic feature.
[0112] (2) Use the second boundary module to perform boundary extraction on the concatenation of the first boundary feature and the first semantic feature to obtain a second boundary feature, and use the second semantic module to perform semantic extraction on the first semantic feature to obtain a second semantic feature.
[0113] (3) Use the third boundary module to perform boundary extraction on the concatenation of the second boundary feature and the second semantic feature to obtain a boundary feature vector.
[0114] The first boundary module is Figure 1 the boundary module 1 in Figure 1 the second boundary module is Figure 1 the boundary module 2 in
[0115] The boundary extraction branch obtains an 8X boundary feature vector by ConCat{semantic, boundary} features.
[0116] Step 407: Use the fusion module to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector.
[0117] Specifically, using the fusion module to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a feature vector may include: creating a fusion algorithm supported by DLA in the fusion module; using the fusion algorithm to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector.
[0118] The formula of the fusion algorithm is:
[0119] f out = F2(f p , f i , f d ) (4)
[0120] F2(f p , f i , f d ) = W3 × f p + W4 × f i + W5 × f d + c (5)
[0121] Among them, f p represents an 8X detail feature vector, f i represents an 8X semantic feature vector, f d represents an 8X boundary feature vector, W3, W4, and W5 represent a series of operations, and c represents an offset.
[0122] Step 408: Use the upsampling segmentation head to process the fused feature vector to obtain a prediction result.
[0123] The original segmentation head outputs at the original scale without an upsampling process, and the final output result is an 8X feature vector. However, the 8X feature vector has insufficiently fine boundaries and unclear distant details during use. To optimize the usage effect, we added a layer of convolution and a 2x upsampling before the original segmentation head, upsampled the 8X feature vector to 4X, and then performed output supervision. The details in use have been optimized.
[0124] The fused feature vector obtained by the fusion module will be output to the upsampling segmentation head, and the upsampling segmentation head will output the final prediction result.
[0125] Step 409: Calculate the losses of the prediction result, the true class label, and the true boundary label using the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions, and train the model parameters of the semantic segmentation model according to the calculation results.
[0126] Specifically, calculating the losses of the prediction result, the true class label, and the true boundary label using the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions may include:
[0127] (1) Use two auxiliary heads and an optimized feature pyramid to perform semantic predictions on their respective corresponding semantic feature vectors, and calculate the losses of the corresponding semantic prediction results and the true class labels using the semantic loss functions corresponding to the two auxiliary heads and the optimized feature pyramid respectively.
[0128] (2) Use the detail prediction head to perform detail predictions on the detail feature vectors, and calculate the losses of the detail prediction results and the true class labels using the detail loss function.
[0129] (3) Use the boundary prediction head to perform boundary predictions on the boundary feature vectors, and calculate the losses of the boundary prediction results and the true boundary labels using the boundary loss function.
[0130] (4) Use the upsampling segmentation head to perform predictions on the fused feature vectors, and calculate the losses of the prediction results and the true class labels using one segmentation loss function; use the upsampling segmentation head to perform predictions after upsampling the fused feature vectors, and calculate the losses of the prediction results and the true class labels using another segmentation loss function.
[0131] These results need to calculate the per-pixel differences with the true class label and the true boundary label to update the model gradient and constrain the parameter optimization direction.
[0132] Among them, the segmentation loss function Loss corresponding to the upsampling segmentation head (including Loss 8X and Loss 4X)(OhemCrossEntropy), the detail loss function Loss corresponding to the detail prediction head p is CrossEntropyLoss, the boundary loss function Loss corresponding to the boundary prediction head d is BoundaryLoss, the semantic loss function Loss corresponding to the auxiliary head (including Loss 16X , Loss 32X , Loss 64X ) is CrossEntropyLoss.
[0133] Step 410: Perform INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format.
[0134] PTQ is a technique for quantizing a model after model training. Its main purpose is to convert a trained floating-point model (usually with FP32 precision) into a low-precision fixed-point model (such as INT8) to reduce the model size and accelerate the inference speed. PTQ does not require retraining the model and only needs to perform quantization on the existing model, saving time and computing resources. The quantized model can be deployed faster in resource-constrained environments, and the low-precision representation reduces storage space and computing overhead.
[0135] After the semantic segmentation model has been iteratively trained for 85 epochs, a semantic segmentation model can be obtained under the PyTorch framework. Performing PTQ INT8 quantization training on this semantic segmentation model can convert it into a semantic segmentation model in the ONNX format.
[0136] Step 411: Deploy the semantic segmentation model in the ONNX format to the DLA.
[0137] The exported semantic segmentation model in the ONNX format can all run on the DLA with INT8 precision. Changing from running on the GPU originally to running on the DLA can not only release more computing resources for the GPU running of other models, but also optimize the inference speed of the semantic segmentation model and improve the performance of the semantic segmentation model.
[0138] In summary, the method for deploying the semantic segmentation model provided by the embodiments of the present application creates a semantic segmentation model suitable for DLA, trains the semantic segmentation model using a training set, and then performs INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format. Finally, the semantic segmentation model in the ONNX format is deployed on DLA, thereby realizing the migration of the semantic segmentation model from GPU to DLA. Without significantly sacrificing accuracy, the inference speed of the semantic segmentation model is optimized, and the performance of the semantic segmentation model is improved.
[0139] After deploying the semantic segmentation model on DLA, the semantic segmentation model can be used to perform semantic segmentation on the perception images collected during the driving of the vehicle. The specific process is as Figure 5 shown:
[0140] Step 501, obtain the perception image to be recognized.
[0141] Step 502, use the feature extraction module to extract features from the perception image to obtain an image feature vector.
[0142] Step 503, use the semantic extraction branch to extract semantics from the image feature vector to obtain a semantic feature vector.
[0143] Step 504, use the detail extraction branch and the semantic extraction branch to extract details from the image feature vector to obtain a detail feature vector.
[0144] Step 505, use the boundary extraction branch and the semantic extraction branch to extract boundaries from the image feature vector to obtain a boundary feature vector.
[0145] Step 506, use the fusion module to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector.
[0146] Step 507, use the upsampling segmentation head to perform prediction after upsampling the fused feature vector to obtain an image segmentation result.
[0147] It should be noted that during the use of the semantic segmentation model, the detail prediction head, the three auxiliary heads, and the boundary prediction head do not work.
[0148] Figure 6 and Figure 7 shows the image segmentation result obtained using the semantic segmentation model. Finally, when the semantic segmentation model runs on the DLA platform with BatchSize = 5, the inference speed is about 75ms, achieving a balance between speed and accuracy.
[0149] As Figure 8As shown, it shows a structural block diagram of a deployment device for a semantic segmentation model provided by an embodiment of the present application. The deployment device for the semantic segmentation model can be applied to a computer device. The deployment device for the semantic segmentation model may include:
[0150] A model creation module 810, configured to create a semantic segmentation model applicable to DLA. The semantic segmentation model includes a feature extraction module, a detail extraction branch, a semantic extraction branch, a boundary extraction branch, a fusion module, and an upsampling segmentation head. The detail prediction head in the detail extraction branch corresponds to a detail loss function. The two auxiliary heads and the optimized feature pyramid in the semantic extraction branch respectively correspond to a semantic loss function. The boundary prediction head in the boundary extraction branch corresponds to a boundary loss function. The upsampling segmentation head corresponds to two segmentation loss functions;
[0151] A training set acquisition module 820, configured to acquire a training set. The training samples in the training set include sample images, true class labels, and true boundary labels. The true boundary labels are obtained by performing boundary extraction on the true class labels;
[0152] A model training module 830, configured to, for each training sample, process the sample image using the semantic segmentation model to obtain a prediction result, calculate losses for the prediction result, the true class label, and the true boundary label using the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions, and train the model parameters of the semantic segmentation model according to the calculation results;
[0153] The model training module 830 is further configured to perform INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format;
[0154] A model deployment module 840, configured to deploy the semantic segmentation model in the ONNX format to DLA.
[0155] In an optional embodiment, the model training module 830 is further configured to:
[0156] Extract features from the sample image using the feature extraction module to obtain an image feature vector;
[0157] Extract semantic features from the image feature vector using the semantic extraction branch to obtain a semantic feature vector;
[0158] Extract detail features from the image feature vector using the detail extraction branch and the semantic extraction branch to obtain a detail feature vector;
[0159] Extract boundary features from the image feature vector using the boundary extraction branch and the semantic extraction branch to obtain a boundary feature vector;
[0160] The fusion module is used to fuse the detail feature vector, the semantic feature vector and the boundary feature vector to obtain a fused feature vector;
[0161] The upsampling segmentation head is used to process the fused feature vector to obtain a prediction result.
[0162] In an optional embodiment, the model training module 830 is further configured to:
[0163] Use two auxiliary heads and an optimized feature pyramid to perform semantic prediction on the corresponding semantic feature vectors respectively, and use the corresponding semantic loss functions of the two auxiliary heads and the optimized feature pyramid to calculate the loss between the corresponding semantic prediction results and the true class labels;
[0164] Use the detail prediction head to perform detail prediction on the detail feature vector, and use the detail loss function to calculate the loss between the detail prediction result and the true class label;
[0165] Use the boundary prediction head to perform boundary prediction on the boundary feature vector, and use the boundary loss function to calculate the loss between the boundary prediction result and the true boundary label;
[0166] Use the upsampling segmentation head to perform prediction on the fused feature vector, and use a segmentation loss function to calculate the loss between the prediction result and the true class label; Use the upsampling segmentation head to perform prediction on the upsampled fused feature vector, and use another segmentation loss function to calculate the loss between the prediction result and the true class label.
[0167] In an optional embodiment, the model training module 830 is further configured to:
[0168] Create a fusion algorithm supported by DLA in the fusion module;
[0169] Use the fusion algorithm to fuse the detail feature vector, the semantic feature vector and the boundary feature vector to obtain a fused feature vector.
[0170] In an optional embodiment, the model training module 830 is further configured to:
[0171] When the detail extraction branch includes three detail modules and the semantic extraction branch includes three semantic modules, use the first detail module to perform detail extraction on the image feature vector to obtain a first detail feature, and use the first semantic module to perform semantic extraction on the image feature vector to obtain a first semantic feature;
[0172] Fuse the first detailed feature and the first semantic feature using the second fusion algorithm supported by DLA to obtain the first fused feature; use the second detailed module to perform detailed extraction on the first fused feature to obtain the second detailed feature, and use the second semantic module to perform semantic extraction on the first semantic feature to obtain the second semantic feature;
[0173] Fuse the second detailed feature and the second semantic feature using the second fusion algorithm supported by DLA to obtain the second fused feature; use the third detailed module to perform detailed extraction on the second fused feature to obtain the detailed feature vector.
[0174] In an optional embodiment, the model training module 830 is further configured to:
[0175] When the boundary extraction branch includes three boundary modules and the semantic extraction branch includes three semantic modules, use the first boundary module to perform boundary extraction on the image feature vector to obtain the first boundary feature, and use the first semantic module to perform semantic extraction on the image feature vector to obtain the first semantic feature;
[0176] Use the second boundary module to perform boundary extraction on the concatenation of the first boundary feature and the first semantic feature to obtain the second boundary feature, and use the second semantic module to perform semantic extraction on the first semantic feature to obtain the second semantic feature;
[0177] Use the third boundary module to perform boundary extraction on the concatenation of the second boundary feature and the second semantic feature to obtain the boundary feature vector.
[0178] In an optional embodiment, the apparatus further includes:
[0179] An image acquisition module, configured to acquire a perceptual image to be recognized;
[0180] A feature extraction module, configured to perform feature extraction on the perceptual image using the feature extraction module to obtain an image feature vector;
[0181] The feature extraction module is further configured to perform semantic extraction on the image feature vector using the semantic extraction branch to obtain a semantic feature vector;
[0182] The feature extraction module is further configured to perform detailed extraction on the image feature vector using the detailed extraction branch and the semantic extraction branch to obtain a detailed feature vector;
[0183] The feature extraction module is further configured to perform boundary extraction on the image feature vector using the boundary extraction branch and the semantic extraction branch to obtain a boundary feature vector;
[0184] The feature extraction module is further configured to fuse the detailed feature vector, the semantic feature vector, and the boundary feature vector using the fusion module to obtain a fused feature vector;
[0185] A result prediction module, configured to perform prediction on the upsampled fusion feature vector by using an upsampling segmentation head to obtain an image segmentation result.
[0186] In summary, the deployment device of the semantic segmentation model provided by the embodiments of the present application creates a semantic segmentation model suitable for DLA, trains the semantic segmentation model by using a training set, and then performs INT8 quantization training on the trained semantic segmentation model so as to convert the semantic segmentation model from the PyTorch format to the ONNX format. Finally, the ONNX format semantic segmentation model is deployed on the DLA, thereby realizing the migration of the semantic segmentation model from the GPU to the DLA, optimizing the inference speed of the semantic segmentation model without significantly sacrificing accuracy, and improving the performance of the semantic segmentation model.
[0187] An embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the deployment method of the semantic segmentation model as described above.
[0188] An embodiment of the present application provides a computer device, and the computer device includes the deployment device of any of the above semantic segmentation models.
[0189] It should be noted that when the deployment device of the semantic segmentation model provided in the above embodiment deploys the semantic segmentation model, only the above division of each functional module is used as an example for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the deployment device of the semantic segmentation model is divided into different functional modules to complete all or part of the functions described above. In addition, the deployment device of the semantic segmentation model provided in the above embodiment and the embodiment of the deployment method of the semantic segmentation model belong to the same concept, and the specific implementation process thereof can be seen in the method embodiment, which will not be elaborated here.
[0190] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0191] The above does not intend to limit the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.
Claims
1. A method for deploying a semantic segmentation model, characterized in that The method includes: Create a semantic segmentation model applicable to the deep learning accelerator DLA. The semantic segmentation model includes a feature extraction module, a detail extraction branch, a semantic extraction branch, a boundary extraction branch, a fusion module, and an upsampling segmentation head. The detail extraction branch includes detail modules 1-3 and a detail prediction head. The detail modules 1-3 are connected in series and then connected to the detail prediction head. The detail prediction head corresponds to a detail loss function Loss p , the semantic extraction branch includes semantic modules 1-3, an optimized feature pyramid, and three auxiliary heads. The semantic modules 1-3 are connected in series and then connected to the optimized feature pyramid. Semantic module 1 is connected to the first auxiliary head. The semantic loss function corresponding to the first auxiliary head is Loss 16X , semantic module 2 is connected to the second auxiliary head. The semantic loss function corresponding to the second auxiliary head is Loss 32X , the optimized feature pyramid is connected to the third auxiliary head. The semantic loss function corresponding to the third auxiliary head is Loss 64X , where 16X, 32X, and 64X represent the sampling multiples. The boundary extraction branch includes boundary modules 1-3 and a boundary prediction head. The boundary modules 1-3 are connected in series and then connected to the boundary prediction head. The boundary prediction head corresponds to a boundary loss function Loss d , the upsampling segmentation head corresponds to two segmentation loss functions with different sampling rates, namely Loss 8X and Loss 4X , where 8X and 4X represent the sampling multiples; Obtaining a training set, where the training samples in the training set include sample images, true class labels, and true boundary labels, and the true boundary labels are obtained by performing boundary extraction on the true class labels; For each training sample, using the semantic segmentation model to process the sample image to obtain a prediction result, and using the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions to calculate the losses of the prediction result, the true class label, and the true boundary label, and training the model parameters of the semantic segmentation model according to the calculation results; Performing INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format; Deploying the semantic segmentation model in the ONNX format to the DLA.
2. The deployment method of the semantic segmentation model according to claim 1, wherein The using the semantic segmentation model to process the sample image to obtain a prediction result includes: Using the feature extraction module to extract features from the sample image to obtain an image feature vector; Using the semantic extraction branch to perform semantic extraction on the image feature vector to obtain a semantic feature vector; Using the detail extraction branch and the semantic extraction branch to perform detail extraction on the image feature vector to obtain a detail feature vector; Using the boundary extraction branch and the semantic extraction branch to perform boundary extraction on the image feature vector to obtain a boundary feature vector; Using the fusion module to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector; Using the upsampling segmentation head to process the fused feature vector to obtain a prediction result.
3. The deployment method of the semantic segmentation model according to claim 2, wherein The using the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions to calculate the losses of the prediction result, the true class label, and the true boundary label includes: Using two auxiliary heads and the optimized feature pyramid to perform semantic prediction on the corresponding semantic feature vectors, and using the semantic loss functions corresponding to the two auxiliary heads and the optimized feature pyramid to calculate the losses of the corresponding semantic prediction results and the true class label; Using the detail prediction head to perform detail prediction on the detail feature vector, and using the detail loss function to calculate the losses of the detail prediction result and the true class label; Using the boundary prediction head to perform boundary prediction on the boundary feature vector, and using the boundary loss function to calculate the losses of the boundary prediction result and the true boundary label; Using the upsampling segmentation head to perform prediction on the fused feature vector, and using one segmentation loss function to calculate the losses of the prediction result and the true class label; using the upsampling segmentation head to perform prediction on the upsampled fused feature vector, and using another segmentation loss function to calculate the losses of the prediction result and the true class label.
4. The deployment method of the semantic segmentation model according to claim 2, characterized in that, The using the fusion module to fuse the detail feature vector, the semantic feature vector, and the boundary feature vector to obtain a fused feature vector includes: Create the fusion algorithm supported by DLA in the fusion module; Fuse the detail feature vector, the semantic feature vector, and the boundary feature vector by using the fusion algorithm to obtain a fused feature vector.
5. The deployment method of the semantic segmentation model according to claim 2, characterized in that The using the detail extraction branch and the semantic extraction branch to extract details from the image feature vector to obtain a detail feature vector includes: When the detail extraction branch includes three detail modules and the semantic extraction branch includes three semantic modules, use the first detail module to extract details from the image feature vector to obtain a first detail feature, and use the first semantic module to extract semantics from the image feature vector to obtain a first semantic feature; Fuse the first detail feature and the first semantic feature by using the second fusion algorithm supported by DLA to obtain a first fused feature; use the second detail module to extract details from the first fused feature to obtain a second detail feature, and use the second semantic module to extract semantics from the first semantic feature to obtain a second semantic feature; Fuse the second detail feature and the second semantic feature by using the second fusion algorithm supported by DLA to obtain a second fused feature; use the third detail module to extract details from the second fused feature to obtain the detail feature vector.
6. The deployment method of the semantic segmentation model according to claim 2, wherein The using the boundary extraction branch and the semantic extraction branch to extract boundaries from the image feature vector to obtain a boundary feature vector includes: When the boundary extraction branch includes three boundary modules and the semantic extraction branch includes three semantic modules, use the first boundary module to extract boundaries from the image feature vector to obtain a first boundary feature, and use the first semantic module to extract semantics from the image feature vector to obtain a first semantic feature; Use the second boundary module to extract boundaries from the concatenation of the first boundary feature and the first semantic feature to obtain a second boundary feature, and use the second semantic module to extract semantics from the first semantic feature to obtain a second semantic feature; Use the third boundary module to extract boundaries from the concatenation of the second boundary feature and the second semantic feature to obtain the boundary feature vector.
7. The deployment method of the semantic segmentation model according to claim 1, wherein The method further includes: Obtain a perceptual image to be recognized; Extract features from the perceptual image by using the feature extraction module to obtain an image feature vector; Extract semantics from the image feature vector by using the semantic extraction branch to obtain a semantic feature vector; Extract details from the image feature vector by using the detail extraction branch and the semantic extraction branch to obtain a detail feature vector; Extract boundaries from the image feature vector by using the boundary extraction branch and the semantic extraction branch to obtain a boundary feature vector; Fuse the detail feature vector, the semantic feature vector, and the boundary feature vector by using the fusion module to obtain a fused feature vector; Perform prediction on the upsampled fused feature vector by using the upsampling segmentation head to obtain an image segmentation result.
8. A deployment device for a semantic segmentation model, characterized in that, The device includes: A creation module is used to create a semantic segmentation model applicable to a deep learning accelerator DLA. The semantic segmentation model includes a feature extraction module, a detail extraction branch, a semantic extraction branch, a boundary extraction branch, a fusion module, and an upsampling segmentation head. The detail extraction branch includes detail modules 1-3 and a detail prediction head. The detail modules 1-3 are connected in series and then connected to the detail prediction head. The detail prediction head corresponds to a detail loss function Loss p , the semantic extraction branch includes semantic modules 1-3, an optimized feature pyramid, and three auxiliary heads. The semantic modules 1-3 are connected in series and then connected to the optimized feature pyramid. Semantic module 1 is connected to the first auxiliary head. The semantic loss function corresponding to the first auxiliary head is Loss 16X , semantic module 2 is connected to the second auxiliary head. The semantic loss function corresponding to the second auxiliary head is Loss 32X , the optimized feature pyramid is connected to the third auxiliary head. The semantic loss function corresponding to the third auxiliary head is Loss 64X , where 16X, 32X, and 64X represent the sampling multiples. The boundary extraction branch includes boundary modules 1-3 and a boundary prediction head. The boundary modules 1-3 are connected in series and then connected to the boundary prediction head. The boundary prediction head corresponds to a boundary loss function Loss d , the upsampling segmentation head corresponds to two segmentation loss functions, namely Loss 8X and Loss 4X , where 8X and 4X represent the sampling multiples; An acquisition module, configured to acquire a training set, where the training samples in the training set include sample images, true class labels, and true boundary labels, and the true boundary labels are obtained by performing boundary extraction on the true class labels; A training module, configured to, for each training sample, process the sample image by using the semantic segmentation model to obtain a prediction result, calculate losses of the prediction result, the true class label, and the true boundary label by using the detail loss function, three semantic loss functions, the boundary loss function, and two segmentation loss functions, and train model parameters of the semantic segmentation model according to the calculation results; The training module is further configured to perform INT8 quantization training on the trained semantic segmentation model to convert the semantic segmentation model from the PyTorch format to the ONNX format; A deployment module, configured to deploy the semantic segmentation model in the ONNX format to the DLA.
9. A computer-readable storage medium, characterized in that, At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the deployment method of the semantic segmentation model according to any one of claims 1 to 7.
10. A computer device, characterized in that, A computer device includes: a deployment device of the semantic segmentation model according to claim 8.
Citation Information
Patent Citations
Machine learning framework applied in semi-supervised environment to perform instance tracking in sequence of image frames
CN114792331A
Performing object detection, instance segmentation, and semantic correspondence from bounding box supervision using neural networks
CN114972742A