A method for training image semantic segmentation models for traffic road scenes

By constructing the basic semantic segmentation model of the combination of DeepLabV3plus and ResnetX networks and performing feature map stacking operations, the pixel information loss problem caused by the upsampling module of the image semantic segmentation model is solved, and the semantic segmentation results with high accuracy and edge continuity are achieved.

CN114419058BActive Publication Date: 2025-05-09BEIJING WENAN INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210103540.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2025-05-09
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

In the prior art, the upsampling module of the image semantic segmentation model uses the nearest interpolation mode to calculate, resulting in a large amount of pixel information loss compared to the original input image, reducing the semantic segmentation performance and poor accuracy of the segmentation result.

Method used

The semantic segmentation basic model of the combination of the DeepLabV3plus network and the ResnetX network is constructed, and the output of the DeepLabV3plus network is jumped through a convolutional layer group at the input end of the basic network, and the feature map is merged to perform feature map stacking operations to ensure that the final output segmented image does not lose the pixel information of the original input image.

Benefits of technology

Through this method, the stable deployment of the image semantic segmentation model in traffic road scenarios is ensured, and the accuracy and edge continuity of the segmentation results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114419058B_ABST
    Figure CN114419058B_ABST
Patent Text Reader

Abstract

The present invention provides a method for training an image semantic segmentation model for traffic road scenes, comprising: constructing a semantic segmentation basic model, adjusting the structure of a basic network to form a semantic segmentation initial model, and training the semantic segmentation initial model using a sample image training set of traffic road scenes to obtain an image semantic segmentation model. The present invention solves the problem in the prior art that the upsampling operator of the upsampling module of the image semantic segmentation model uses a nearest neighbor interpolation mode for calculation, which causes the feature map of the image semantic segmentation model output by the upsampling module to lose a large amount of pixel information compared to the original input image, thereby affecting the semantic segmentation performance of the image semantic segmentation model and causing the final image semantic segmentation result to have poor accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision image processing, and in particular to a method for training an image semantic segmentation model for traffic road scenes. Background Art

[0002] Image semantic segmentation is one of the core research issues in the field of computer vision. The goal of image semantic segmentation is to assign labels to each pixel of the input image, that is, to achieve pixel-level object classification tasks. The image semantic segmentation model is used to predict and classify the pixels of the input image, generate semantic labels, and finally segment the image into several pixel regions with certain specific semantic meanings.

[0003] Image semantic segmentation technology is widely used in traffic road scenes. It makes it possible to perceive information in traffic road scenes by accurately analyzing and distinguishing targets such as drivable areas, pedestrians, and vehicles. In the prior art, in order to ensure good deployment compatibility between the image semantic segmentation model and the traffic road scene analysis platform, the upsampling module of the image semantic segmentation model is usually calculated using the nearest neighbor interpolation mode; however, this early operator calculation method will cause the feature map output by the upsampling module of the image semantic segmentation model to lose a large amount of pixel information compared to the original input image, which will in turn reduce the semantic segmentation performance of the image semantic segmentation model, resulting in poor accuracy of the final image semantic segmentation result. Summary of the invention

[0004] The main purpose of the present invention is to provide an image semantic segmentation model training method for traffic road scenes, so as to solve the problem in the prior art that the upsampling operator of the upsampling module of the image semantic segmentation model uses the nearest neighbor interpolation mode for calculation, which causes the feature map of the image semantic segmentation model after the upsampling module outputs a large amount of pixel information compared with the original input image, affecting the semantic segmentation performance of the image semantic segmentation model and causing the final image semantic segmentation result to have poor accuracy.

[0005] In order to achieve the above-mentioned purpose, the present invention provides an image semantic segmentation model training method for traffic road scenes, comprising: step S1, constructing a semantic segmentation basic model whose basic network is a combination of a DeepLabV3plus network and a ResnetX network, wherein the upsampling operator of the DeepLabV3plus network is calculated using the nearest neighbor interpolation mode; step S2, adjusting the structure of the basic network to form an initial semantic segmentation model, wherein the adjustment process is: copying the convolution module of the DeepLabV3plus network to the bottom of the DeepLabV3plus network as an underlying convolution module, merging the input end of the basic network with the output end of the DeepLabV3plus network through a jump connection after passing through a convolution layer group, and using the merged end as the input end of the underlying convolution module, and the output of the underlying convolution module is the final semantic segmentation result; step S3, training the initial semantic segmentation model using a sample image training set of traffic road scenes to obtain an image semantic segmentation model.

[0006] Furthermore, the convolution layer group includes one or more convolution layers, the convolution kernel size of each convolution layer is 3×3, the convolution step size is 1, and the padding value is 0 or 1.

[0007] Furthermore, the convolutional layer group includes multiple convolutional layers, and the number of the multiple convolutional layers is greater than 1 layer and less than or equal to 3 layers.

[0008] Furthermore, the filling value includes a filling width value and a filling height value.

[0009] Furthermore, the structure of the convolution module and the underlying convolution module of the DeepLabV3plus network from top to bottom includes: convolution layer, BN layer, Relu layer and convolution layer.

[0010] Furthermore, the sample image training set includes a first training set and a second training set. The training images in the first training set are selected from the front traffic road scene images taken by the front vehicle-mounted imaging device along the road direction and the front traffic road scene images taken by the rear vehicle-mounted imaging device along the road direction in the Audi large-scale autonomous driving data set A2D2; the training images in the second training set are highway scene images.

[0011] Further, step S3 includes: step S31, using the first training set to pre-train the semantic segmentation initial model to obtain the semantic segmentation pre-training model; step S32, adjusting the learning rate of the model training, and using the second training set to continue training the semantic segmentation pre-training model to obtain the image semantic segmentation model.

[0012] Further, in step S32, the ratio of the learning rate when training the semantic segmentation pre-training model to the learning rate when training the semantic segmentation initial model is adjusted to be between [1 / 5, 1 / 2].

[0013] Further, the ratio of the number of training images in the second training set to the number of training images in the first training set is in the range of [1 / 10, 1].

[0014] Furthermore, the ResnetX network in the basic network is one of the Resnet18 network, the Resnet34 network, the Resnet50 network, the Resnet101 network and the Resnet152 network.

[0015] By applying the technical solution of the present invention, since in the basic network of the image semantic segmentation model, the upsampling operator of the DeepLabV3plus network is calculated using the nearest neighbor interpolation mode, in the DeepLabV3plus network, the feature map has undergone multiple upsamplings. During the interpolation process of the feature map, an interpolation operation is performed between two adjacent eigenvalues ​​in the feature map, and the eigenvalue of the point at the interpolation position is the eigenvalue of the nearest neighbor to the point. Therefore, the segmented image finally output by the model will suffer from pixel information loss and boundary jaggedness, thereby causing distortion of the segmented image finally output by the model. By using the technical solution of the present invention, the upsampling operator of the upsampling module can be kept in the nearest neighbor interpolation mode unchanged, thereby ensuring that the image semantic segmentation model can be stably, conveniently and completely deployed on the traffic road scene analysis platform; further, the input end of the basic network is merged with the output end of the DeepLabV3plus network through a jump connection after passing through the convolutional layer group, and the feature map containing the complete feature information of the input image and the feature map of the basic network that has lost part of the feature information after multiple upsampling calculations are stacked to obtain a merged feature map, and the underlying convolution module continues to perform feature learning on the feature information of the merged feature map and outputs it according to the output dimension size of the basic network, thereby ensuring that the final output segmented image does not lose the pixel information of the original input image during the image semantic segmentation process, and ensuring that the edges of the segmented image are continuous, and the entire segmented image is concrete and clear, which effectively improves the semantic segmentation processing effect of the image semantic segmentation model on the input image. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings constituting a part of the present application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0017] Figure 1 A flowchart of a method for training an image semantic segmentation model for a traffic road scene according to an optional embodiment of the present invention is shown;

[0018] Figure 2A schematic diagram of a network structure in which a basic network of a semantic segmentation basic model according to an optional embodiment of the present invention is DeepLabV3plus+Resnet50 is shown;

[0019] Figure 3 Shows the Figure 2 Schematic diagram of the network structure of the semantic segmentation initial model obtained after the structure of the basic network of the semantic segmentation basic model is adjusted. DETAILED DESCRIPTION

[0020] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0021] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so as to describe the embodiments of the present invention described herein. In addition, the terms "including", "and", "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] In order to solve the problem in the prior art that the upsampling operator of the upsampling module of the image semantic segmentation model uses the nearest neighbor interpolation mode for calculation, which causes the feature map of the image semantic segmentation model after the upsampling module loses a lot of pixel information compared with the original input image, affecting the semantic segmentation performance of the image semantic segmentation model and causing the final image semantic segmentation result to have poor accuracy, the present invention provides an image semantic segmentation model training method for traffic road scenes.

[0024] Figure 1 FIG. 1 is a flow chart of a method for training an image semantic segmentation model for a traffic road scene according to an embodiment of the present invention. Figure 1As shown, the method includes the following steps: step S1, constructing a basic semantic segmentation model whose basic network is a combination of a DeepLabV3plus network and a ResnetX network, wherein the upsampling operator of the DeepLabV3plus network is calculated using the nearest neighbor interpolation mode; step S2, adjusting the structure of the basic network to form an initial semantic segmentation model, wherein the adjustment process is: copying the convolution module of the DeepLabV3plus network to the bottom of the DeepLabV3plus network as the bottom convolution module, merging the input end of the basic network with the output end of the DeepLabV3plus network through a jump connection after passing through the convolution layer group, and using the merged end as the input end of the bottom convolution module, and the output of the bottom convolution module is the final semantic segmentation result; step S3, training the initial semantic segmentation model using a sample image training set of traffic road scenes to obtain an image semantic segmentation model.

[0025] like Figure 2 As shown, in the basic network of the image semantic segmentation model, the upsampling operator of the DeepLabV3plus network is calculated using the nearest neighbor interpolation mode. In the DeepLabV3plus network, the feature map has been upsampled multiple times. In the interpolation process of the feature map, the interpolation operation is performed between two adjacent eigenvalues ​​in the feature map. The eigenvalue of the interpolation position point is the eigenvalue of the nearest neighbor of the point. Therefore, the segmented image finally output by the model will have pixel information loss and boundary jagged edges, which will lead to the distortion of the segmented image finally output by the model. By using the technical solution of the present invention, the upsampling operator of the upsampling module can be kept unchanged in the nearest neighbor interpolation mode, thereby ensuring that the image semantic segmentation model can be stably, conveniently and completely deployed on the traffic road scene analysis platform; further, as Figure 3 As shown in the figure, the input end of the basic network is merged with the output end of the DeepLabV3plus network through a jump connection after passing through a convolutional layer group, and the feature map containing the complete feature information of the input image and the feature map of the basic network that has lost part of the feature information after multiple upsampling calculations are stacked to obtain a merged feature map. The underlying convolution module continues to learn the feature information of the merged feature map and outputs it according to the output dimension size of the basic network. Therefore, in the process of image semantic segmentation, it can ensure that the final output segmented image does not lose the pixel information of the original input image, and ensures that the edges of the segmented image are continuous, and the entire segmented image is clear and concrete, which effectively improves the semantic segmentation processing effect of the image semantic segmentation model on the input image.

[0026] It should be noted that in the illustrated embodiment of the present invention, Figure 2As shown in the figure, there are three positions in the basic network of the image semantic segmentation model where the upsampling operators are calculated using the nearest neighbor interpolation mode. One position is in the ASPP structure of the DeepLabV3plus network, and the other two positions are between the ASPP structure and the convolution module of the DeepLabV3plus network, and after the convolution module. This ensures that the image semantic segmentation model under the PyTorch deep learning framework can be smoothly converted to other deep learning frameworks.

[0027] In the basic network of the image semantic segmentation model of the embodiment of the present invention, the upsampling module of the DeepLabV3plus network includes an upsampling operator upsample to realize the upsampling operation. Of course, the upsampling module can also include multiple upsampling operators or other operators, which can also complete the upsampling operation. The basic networks of other structural forms are also within the protection scope of the present invention.

[0028] like Figure 3 As shown in the figure, the input of the basic network passes through the convolutional layer group and the output is merged with the DeepLabV3plus network through a jump connection. The feature map obtained after the input image enters the image semantic segmentation model without downsampling is connected with the feature map obtained after upsampling. Then, the underlying convolution module located at the bottom of the network structure of the image semantic segmentation model is used to learn the superimposed features, refine the output boundary information, and obtain a semantic segmentation result with a higher pixel accuracy, thereby improving the semantic segmentation performance of the image semantic segmentation model.

[0029] It should be noted that the feature map referred to above is in the form of a matrix. The feature map is a three-dimensional matrix, expressed as W×H×m, where m represents the number of segmentation elements in the semantic information contained in the feature map, and m is the same for the feature maps obtained by different basic semantic segmentation sub-models. In a preferred embodiment of the present invention, the number of segmentation elements in the semantic information of the output feature map is 9, and a semantic segmentation image is obtained after post-processing; the width and height of the feature map obtained by the final output of the basic network are the same as the width and height of the input image, but the number of segmentation elements in the semantic information is different.

[0030] In step S3, the initial semantic segmentation model is trained using a sample image training set of traffic road scenes to obtain an image semantic segmentation model. The parameters of the image semantic segmentation model are modified based on the predicted semantic segmentation results of the training images and the pre-labeled semantic segmentation information. Steps S1 to S2 are continuously iterated using a number of training images until the training results of the image semantic segmentation model meet the preset convergence conditions. The image semantic segmentation model is continuously iterated using different training images in the training set. When the value of the error calculated by the cross entropy loss function is less than a preset threshold, or the number of iterations reaches a predetermined value, the training result can be considered to have converged, the training is completed, and the trained image semantic segmentation model is obtained, which can be directly used for image semantic segmentation of the image to be processed (input image).

[0031] Optionally, the convolution layer group includes one or more convolution layers, the convolution kernel size of each convolution layer is 3×3, the convolution step is 1, and the padding value is 0 or 1. It should be noted that the padding value includes a padding width value and a padding height value. When the padding value (padding value) is 1, it includes pad-w and pad-h, that is, the width and height of the padding in the processing of the original input image are both 1, for example, the pixel size 2*2 is padded to 3*3.

[0032] Preferably, the convolutional layer group includes multiple convolutional layers, and the number of the multiple convolutional layers is greater than 1 and less than or equal to 3.

[0033] like Figure 2 and Figure 3 As shown in the figure, the structure of the convolution module and the underlying convolution module of the DeepLabV3plus network includes: convolution layer, BN layer, Relu layer and convolution layer from top to bottom.

[0034] In a preferred embodiment of the present invention, the sample image training set includes a first training set and a second training set. The training images in the first training set are selected from the traffic road scene images taken by the front vehicle-mounted imaging device along the road direction and the traffic road scene images taken by the rear vehicle-mounted imaging device along the road direction in the Audi large-scale autonomous driving data set A2D2; the training images in the second training set are highway scene images. In this way, it is ensured that the data source of the training images is more objective and more sufficient, and has the universality of general traffic road scenes, and can improve the performance of the image semantic segmentation model on limited data; the training images of the second training set are highway scene images, and the high-speed road scenes are superimposed, ensuring that the trained image semantic segmentation model has more targeted traffic environment scenes for semantic segmentation recognition and detection.

[0035] The A2D2 large-scale autonomous driving dataset contains a total of 41,277 training images. This embodiment selects the image parts of the front and rear perspectives (front-shot road scene images and rear-shot road scene images), a total of 33,142 images as training data, and uses the trained model as the initial model for semantic segmentation.

[0036] Furthermore, step S3 of an embodiment of the present invention includes: step S31, using the first training set to pre-train the semantic segmentation initial model to obtain a semantic segmentation pre-trained model; step S32, adjusting the learning rate of the model training, and using the second training set to continue training the semantic segmentation pre-trained model to obtain an image semantic segmentation model.

[0037] Optionally, in step S32, the ratio of the learning rate when training the semantic segmentation pre-training model to the learning rate when training the semantic segmentation initial model is adjusted to be between [1 / 5, 1 / 2]. In a preferred embodiment of the present invention, the ratio of the learning rate when training the semantic segmentation pre-training model to the learning rate when training the semantic segmentation initial model is adjusted to be 1 / 3.

[0038] Preferably, the ratio of the number of training images in the second training set to the number of training images in the first training set is in the range of [1 / 10, 1]. In a preferred embodiment of the present invention, the number of training images in the second training set is equal to the number of training images in the first training set.

[0039] Optionally, the ResnetX network in the basic network is one of a Resnet18 network, a Resnet34 network, a Resnet50 network, a Resnet101 network, and a Resnet152 network. In a preferred embodiment of the present invention, the ResnetX network in the basic network is a Resnet50 network.

[0040] Use Figure 2 and Figure 3 The semantic segmentation of the input image of the traffic road scene obtained by the image semantic segmentation model trained by the two network structures is compared with the semantic segmentation standard metric (MIOU) calculation as shown in the following table:

[0041]

[0042] It can be seen from the above table that the technical solution of the present invention Figure 2 After adjusting the network structure, we get Figure 3 The image semantic segmentation model trained with the network structure in the input image has a higher mIOU value than Figure 2 The image semantic segmentation model trained by the network structure in the input image has mIOU values ​​corresponding to each category, indicating that Figure 3The image semantic segmentation model trained by the network structure in the above method has better semantic segmentation performance and better semantic segmentation effect on the input image.

[0043] The present invention also provides an electronic device, which can be a tablet computer, a portable computer, a laptop computer, a desktop computer, etc. The electronic device includes a semantic segmentation network training device, an image semantic segmentation device, a memory, a storage controller and a processor. The memory, the storage controller and the processor are electrically connected to each other directly or indirectly to realize data transmission or interaction. Optionally, the above elements can be electrically connected to each other through one or more communication buses or signal lines. The semantic segmentation network training device and the image semantic segmentation device both include at least one software function module stored in the memory in the form of software or firmware or solidified in the operating system (OS) of the electronic device. The processor is used to execute the executable module stored in the memory, such as the software function module or computer program included in the semantic segmentation network training device, and to complete the image semantic segmentation model training method by implementing or executing the disclosed methods, steps and logic block diagrams in the embodiments of the present invention, and to obtain the image semantic segmentation model; the image semantic segmentation device includes a software function module or a computer program to realize the semantic segmentation of the image to be processed.

[0044] The memory may be selected as a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. Among them, the memory is used to store the program, and the processor executes the program after receiving the execution instruction. The processor may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a voice processor, and a video processor, etc.; it may also be a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0045] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0046] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0047] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0048] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0049] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0050] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0051] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for training an image semantic segmentation model for traffic road scenes, characterized in that: include: Constructing a semantic segmentation basic model whose basic network is a combination of DeepLabV3plus network and ResnetX network, wherein the upsampling operator of the DeepLabV3plus network is calculated using the nearest neighbor interpolation mode; Adjust the structure of the basic network to form an initial model for semantic segmentation, wherein the adjustment process is: copy the convolution module of the DeepLabV3plus network to the bottom of the DeepLabV3plus network as the bottom convolution module, merge the input end of the basic network with the output end of the DeepLabV3plus network through a jump connection after passing through the convolution layer group, and use the merged end as the input end of the bottom convolution module, and the output of the bottom convolution module is the final semantic segmentation result; the structure of the convolution module of the DeepLabV3plus network includes from top to bottom: convolution layer, BN layer, Relu layer and convolution layer, and the convolution layer group includes one or more convolution layers; The semantic segmentation initial model is trained using a sample image training set of traffic road scenes to obtain an image semantic segmentation model.

2. The image semantic segmentation model training method according to claim 1, characterized in that: The convolution kernel size of each convolution layer in the convolution layer group is 3×3, the convolution step size is 1, and the padding value is 0 or 1.

3. The image semantic segmentation model training method according to claim 2, characterized in that: The convolutional layer group includes multiple convolutional layers, and the number of the multiple convolutional layers is greater than 1 layer and less than or equal to 3 layers.

4. The image semantic segmentation model training method according to claim 2, characterized in that: The filling value includes a filling width value and a filling height value.

5. The image semantic segmentation model training method according to claim 1, characterized in that: The structure of the bottom convolution module includes: convolution layer, BN layer, Relu layer and convolution layer from top to bottom.

6. The image semantic segmentation model training method according to claim 1, characterized in that: The sample image training set includes a first training set and a second training set. The training images in the first training set are selected from the front traffic road scene images taken along the road direction by the front vehicle-mounted imaging device and the front traffic road scene images taken along the road direction by the rear vehicle-mounted imaging device in the Audi large-scale autonomous driving data set A2D2; the training images in the second training set are highway scene images.

7. The image semantic segmentation model training method according to claim 6, characterized in that: The method of training the initial semantic segmentation model using a sample image training set of traffic road scenes to obtain an image semantic segmentation model includes: Pre-training the semantic segmentation initial model using the first training set to obtain a semantic segmentation pre-training model; The learning rate of the model training is adjusted, and the semantic segmentation pre-training model is continued to be trained using the second training set to obtain the image semantic segmentation model.

8. The image semantic segmentation model training method according to claim 7, characterized in that: In the step of adjusting the learning rate of model training, the ratio of the learning rate when training the semantic segmentation pre-training model to the learning rate when training the semantic segmentation initial model is adjusted to be between [1 / 5, 1 / 2].

9. The image semantic segmentation model training method according to claim 6, characterized in that: The ratio of the number of training images in the second training set to the number of training images in the first training set is in the range of [1 / 10, 1].

10. The image semantic segmentation model training method according to claim 1, characterized in that: The ResnetX network in the basic network is one of the Resnet18 network, the Resnet34 network, the Resnet50 network, the Resnet101 network and the Resnet152 network.

Citation Information

Patent Citations

  • A real-time scene image semantic segmentation method based on lightweight network

    CN109145983A

  • Image semantic segmentation method and device, electronic device and computer-readable medium

    CN109447990A