Semantic segmentation method of nighttime street scenes based on difficult category perception mechanism
By introducing a method based on a difficult category perception mechanism in night street scene semantic segmentation, using implicitly perceived exposure texture and difficult category semantic segmentation network, the problem of unsatisfactory semantic segmentation effect at night street scene is solved, and more efficient semantic segmentation performance is achieved.
Patent Information
- Application Number
- CN202310958101.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-01
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-08-01
AI Technical Summary
The prior art is difficult to effectively perform semantic segmentation at night street scenes, especially under low image quality and interference from artificial light sources at night, resulting in loss of semantic information and unsatisfactory segmentation effect.
Using a method based on the difficult category perception mechanism, the auxiliary branches and difficult category semantic segmentation network of implicitly perceived exposure textures are constructed, and the encoder is used to learn the exposure features and texture features in the image by using gradient backpropagation, and the network's recognition and positioning ability of the difficult category through the difficult category perception module.
The performance of the night street scene semantic segmentation algorithm is significantly improved, and it can more accurately identify and locate difficult categories in night scenes, improving semantic segmentation results.
Smart Images

Figure CN116883667B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and computer vision, and in particular relates to a semantic segmentation method of nighttime street scenes based on a difficult category perception mechanism. Background Art
[0002] In today's society, artificial intelligence, as a representative of advanced science and technology, affects people's lives and social development in all aspects. The accuracy and timeliness of image processing technology are becoming increasingly important in the field of artificial intelligence. As autonomous driving and smart cities have been recognized by more people around the world. In terms of unmanned driving, given the high safety requirements of unmanned driving technology, the driving system needs to plan the route of the vehicle during driving and detect obstacles such as other vehicles and buildings in the ever-changing external environment. This requires high accuracy to complete this precise task. Semantic segmentation can be used to judge various marks on the road in real time. In these fields, understanding the semantic information of the surrounding environment is of great practical significance for avoiding obstacles and reducing collisions between vehicles or between vehicles and people. The main task of semantic segmentation is to mark the category information for pixels in the image. Specifically, similar to the classification task, a classification network plus a segmentation head is used to classify the pixels in the image.
[0003] In unmanned driving, the semantic segmentation of night images plays an equally important role as that of daytime images. However, due to the large and complex degradation of night images, under weak natural lighting conditions at night, the edge information, texture information and some semantic information in the image color will change dramatically, which may cause some semantic information in the image to be completely lost due to low image quality or interference from artificial light sources. In addition, it is more difficult to annotate night images, so the semantic segmentation of night images is more challenging. Most of the mainstream semantic segmentation methods today are not applicable to nighttime. This is because most models and frameworks are trained under daytime conditions, which will produce huge domain deviations, which leads to some traditional models being able to achieve satisfactory performance on daytime images, but the effect of night images is not very ideal. Therefore, semantic segmentation of night street scenes has a high application value in reality, which can assist smart cars in understanding night scenes, detecting obstacles and pedestrians at night, and thus preventing traffic accidents. In general, semantic segmentation of night street scenes is now a very difficult and meaningful topic because night street scene images have many complex semantic degradation phenomena and conventional traditional semantic segmentation models cannot obtain good segmentation processing results. Summary of the invention
[0004] In view of the defects and shortcomings of the prior art, the purpose of the present invention is to provide a nighttime street scene semantic segmentation method based on a difficult category perception mechanism. The method can use the exposure texture map and, through gradient back propagation, enable the encoder to implicitly learn the exposure features and texture features in the image. At the same time, the difficult category perception module is used to enhance the recognition and positioning capabilities of the difficult category semantic segmentation network, effectively improving the performance of the nighttime street scene semantic segmentation algorithm.
[0005] The main steps include:
[0006] Step S1: Divide the night street scene dataset into a training set and a test set, and perform data preprocessing on the night city landscape images and corresponding labels in the dataset, including data enhancement, normalization, etc.; Step S2: On the basis of the main semantic segmentation network, construct an auxiliary branch that implicitly perceives exposure texture, and through gradient back propagation, enable the encoder in the network to implicitly learn the exposure features and texture features in the image; Step S3: Construct a difficult category semantic segmentation network and a difficult category perception module, wherein the difficult category perception module uses the features of the main semantic segmentation network to enhance the recognition and positioning capabilities of the difficult category semantic segmentation network encoder, and finally fuse the segmentation results of the two networks; Step S4: Construct a training pipeline for night street scene semantic segmentation based on the difficult category perception mechanism, and use the pipeline and training set images to train a night street scene semantic segmentation model based on the difficult category perception mechanism; Step S5: Input the test image of the night street scene into the trained night street scene semantic segmentation model based on the difficult category perception mechanism, and output the corresponding semantic segmentation mask map.
[0007] The technical solution specifically adopted by the present invention to solve the technical problem is:
[0008] A nighttime street scene semantic segmentation method based on a difficult category perception mechanism is provided, wherein the nighttime street scene semantic segmentation is performed using a semantic segmentation mask map output by a nighttime street scene semantic segmentation model based on a difficult category perception mechanism, and the method comprises the following steps:
[0009] Step S1: Divide the nighttime street view dataset into a training set and a test set, and perform data preprocessing on the nighttime urban landscape images and corresponding labels in the dataset;
[0010] Step S2: Based on the main semantic segmentation network, an auxiliary branch for implicitly perceiving exposure and texture is constructed. Through gradient back propagation, the encoder in the network implicitly learns the exposure features and texture features in the image.
[0011] Step S3: construct a difficult category semantic segmentation network and a difficult category perception module, wherein the difficult category perception module uses the characteristics of the main semantic segmentation network to enhance the recognition and positioning capabilities of the difficult category semantic segmentation network encoder, and finally fuses the segmentation results of the two networks;
[0012] Step S4: constructing a training pipeline for nighttime street scene semantic segmentation based on a difficult category perception mechanism, and using the pipeline and training set images to train a nighttime street scene semantic segmentation model based on a difficult category perception mechanism;
[0013] Step S5: input the night street scene test image into the trained night street scene semantic segmentation model based on the difficult category perception mechanism, and output the corresponding semantic segmentation mask map.
[0014] Furthermore, step S1 specifically includes the following steps:
[0015] Step S11: Divide the data set into a training set and a test set according to a certain ratio;
[0016] Step S12: performing data augmentation on the images in the training set to increase the number of samples in the data set;
[0017] Step S13: Preprocess the image after data enhancement in step S12 and convert it into an input image for the night street scene semantic segmentation network: first, randomly crop the image, and then normalize the cropped image with uniform size to convert the image data into a standard normal distribution; in order to ensure that the size and position of the segmented area in the label correspond to the night city landscape image, the same operation is performed on the label while enhancing the image data and preprocessing the image in each step.
[0018] Furthermore, step S2 specifically includes the following steps:
[0019] Step S21: Input the night scene image into the encoder of the subject semantic segmentation network to obtain the dimensions of The four-layer feature map {F1, F2, F3, F4} is input into the decoder to obtain the main semantic segmentation map, where h, w and c represent the height, width and number of channels of the feature map respectively. The specific expression is:
[0020] S = ClsSeg(F1, F2, F3, F4)
[0021] Among them, ClsSeg(·) represents semantic segmentation of features to obtain the semantic segmentation mask map of night street scene;
[0022] Step S22: {F1, F2, F3, F4} obtained from step S21 is used as the input of the exposure texture auxiliary branch, where the feature maps {F1, F2, F3} are respectively input into the 1×1 convolution layer, and the number of channels of the corresponding features is adjusted. The specific expression is:
[0023] F i ′=w i (F i )+b i , τ=1,2,3
[0024] Among them, w i , b i are the weights and biases of the 1×1 convolutional layer; through the 1×1 convolutional layer, the dimensions of {F1′, F2′, F3′} are adjusted to
[0025] Step S23: The feature map F4 with the highest semantics is sent to the spatial pyramid pooling module. Pooling of different scales is used on the original feature map to obtain multiple feature maps of different sizes. These feature maps are then concatenated in the channel dimension, and finally a composite feature map that integrates multiple scales is output, thereby achieving an enhanced feature F4′ that takes into account both global semantic information and local detail information. Finally, the dimension of F4′ is adjusted to The specific expression is:
[0026] P i =AdaptiveAvgPool(F4,ε),i=1,2,3,4
[0027] P i ′=w i (P i )+b i , i=1,2,3,4
[0028] F4′=Concat(P1′, P2′, P3′, P4′)
[0029] Where AdaptiveAvgPool(·) represents the adaptive average pooling operation, ε is the pooling size, and w i , b i are the weights and biases of the 1×1 convolutional layer, and Concat(·) indicates that the features are concatenated in a new dimension;
[0030] Step S24: The four feature maps {F1′, F2′, F3′, F4′} obtained in steps S22 and S23 with the same number of channels are fused at adjacent levels. The operation is to upsample the more abstract and semantically stronger high-level feature maps and then fuse them with the feature maps of adjacent levels to further enhance the semantic information and position information. The specific expression is:
[0031] F4″=F4′
[0032] U i =Upsample(F i+1 ′)+F i ′, i = 1, 2, 3
[0033] U i =w i (U i )+b i , i=1,2,3
[0034] F i = Upsample(U i ′), i=1, 2, 3
[0035] E=Concat(F1″, F2″, F3″, F4″)
[0036] Among them, Upsample(·) represents the upsampling operation, U i ′ represents the output feature after the i-th deep convolutional layer, w i , b i is the weight and bias of the i-th depth convolutional layer, Concat(·) indicates that the features are concatenated in dimension;
[0037] Step S25: Send the exposure texture feature E obtained in step S24 to the exposure texture decoder to obtain an exposure texture map. Through gradient back propagation, the encoder implicitly learns the exposure features and texture features in the image. The specific expression is as follows:
[0038] E′=w(E)+b
[0039] I=tRGB(E′)
[0040] Where w, b are the weights and biases of the 3×3 convolutional layer, and tRGB(·) represents the conversion of features into a 3-channel image to obtain an exposure texture map.
[0041] Furthermore, step S3 specifically includes the following steps:
[0042] Step S31: The four-layer feature map in the encoder of the main semantic segmentation network Four-layer feature maps from the encoder of a semantic segmentation network for hard categories As the input of the difficult category perception module, They are the features extracted by the i-th layer encoder of the main semantic segmentation network and the difficult category semantic segmentation network respectively;
[0043] Step S32: Add a difficult category perception module before each layer of the encoder of the difficult category semantic segmentation network, calculate the affinity matrix for the features of the i-th layer to enhance the recognition and positioning of difficult categories, and the output of the module is the features of the enhanced perception of difficult categories The specific expression is as follows:
[0044]
[0045]
[0046] where ⊙ represents the element-by-element matrix multiplication, S i Represents the affinity matrix calculated for the i-th layer features.
[0047] Furthermore, step S4 specifically includes the following steps:
[0048] Step S41: Based on the SwinTransformer-UperNet semantic segmentation network, a main semantic segmentation network and a difficult category semantic segmentation network are constructed respectively: for the main semantic segmentation branch, SwinTransformer is used as an encoder to extract features from the night street scene image after preprocessing in step S1, and four feature maps {F1, F2, F3, F4} with different scales and different numbers of channels are obtained; the four features are sent as input to the UperNet decoder, the features are decoded into a semantic segmentation mask map, and the exposure texture auxiliary branch constructed in step S2 is inserted after the encoder of the network to obtain an exposure texture map, so that the encoder implicitly learns the exposure features and texture features in the image; the input of the network is the night street scene image I and the corresponding label L all , the output of the network is two images: night street scene semantic segmentation mask map P all and exposure texture segmentation map I;
[0049] Step S42: The network constructed in step S41 is referred to as M. In order to further enhance the semantic segmentation capability of the model for difficult categories, a difficult category semantic segmentation network H is added. The network structure of network H is based on network M. In order to enhance the perception capability of difficult categories, a difficult category perception module constructed in step S3 is added. The input of the difficult sample semantic segmentation network H is the night street scene image I and the corresponding label L. hard , where L hard The representative label only contains difficult categories, and non-difficult categories are all set as background categories; the output of the difficult category semantic segmentation network H is the night street scene semantic segmentation difficult category mask map P hard ;
[0050] Step S43: input a batch of images and corresponding labels in the training set of step S1 into the network of step S41 for training, predict and obtain a semantic segmentation mask map of night street scenes and an exposure texture map, and then freeze the network parameters of step S41 to jointly train the network in step S42 to obtain a semantic segmentation mask map of difficult categories of night street scenes;
[0051] Step S44: According to the loss function of the main semantic segmentation network M, the gradient of each parameter in the network is calculated using the back propagation method, and the parameters are updated using the stochastic gradient descent method; the loss function L of the network main as follows:
[0052] L main =loss ce +loss aux
[0053]
[0054] loss aux =loss exp +loss spa +loss sem
[0055]
[0056]
[0057]
[0058] Among them, L main is the loss function of network M, loss ce Represents the cross entropy loss function; loss aux Represents the exposure texture auxiliary loss function, which is composed of exposure loss loss exp , spatial consistency loss spa , semantic brightness consistency loss sem Composition; for exposure loss exp , M represents the number of non-overlapping local regions, Y k represents the average intensity value of the local area in the exposure texture segmentation map, E is the average intensity value of the assumed good exposure level; for the spatial consistency loss loss spa , K is the number of local regions, ω(i) is the four adjacent regions above, below, left and right centered on region i, Y and I represent the average intensity values of local regions in enhanced and ordinary images, respectively; for semantic brightness consistency loss loss sem , S represents the number of categories for semantic prediction, θ s Represented as a set of pixel indices belonging to category s, Represents the exposure texture segmentation image I H The intensity level at pixel i, B s represents the average intensity level of category s;
[0059] Step S45: According to the loss function of the difficult category semantic segmentation network H, the gradient of each parameter in the network is calculated using the back propagation method, and the parameters are updated using the stochastic gradient descent method; the loss function L of the network is as follows:
[0060] L hard =loss ce
[0061] Step S46: Repeat steps S43 to S45 in batches until the loss value calculated in step S35 converges and stabilizes, save the network parameters, and complete the training process of the night street scene semantic segmentation network with difficult category perception mechanism.
[0062] Furthermore, in step S42, when constructing the difficult category semantic segmentation network H, in order to save memory, the exposure texture auxiliary head branch is deleted.
[0063] Furthermore, step S5 specifically includes the following steps:
[0064] Step S51: Input the nighttime street scene images in the test set into the trained subject semantic segmentation network and the difficult category semantic segmentation network respectively, and output the corresponding nighttime street scene semantic segmentation mask images P respectively. all , Nighttime Street Scene Semantic Segmentation Difficult Category Mask Map P hard .
[0065] Furthermore, step S5 further includes the following steps:
[0066] Step S52: P obtained in step S51 all , P hard In order to improve the semantic segmentation effect of difficult categories, P is selected by the confidence level. hard The "high confidence region" in P all The confidence of the same position in is replaced to improve the model's ability to recognize difficult categories. The "high confidence area" is defined as the area where the difficult category semantic segmentation network predicts better than the main semantic segmentation network on the difficult category. The replacement calculation formula is as follows:
[0067]
[0068] Among them, i represents the current semantic category, C easy represents the simple category without the difficult category, C hard Indicates difficulty category.
[0069] And, a nighttime street scene semantic segmentation system based on a difficult category perception mechanism includes a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method described above can be implemented.
[0070] A computer-readable storage medium stores computer program instructions that can be executed by a processor. When the processor executes the computer program instructions, the method described above can be implemented.
[0071] Compared with the prior art, the present invention and its preferred solution utilize exposure texture maps and gradient back propagation to enable the encoder to implicitly learn the exposure features and texture features in the image. At the same time, the difficult category perception module is used to enhance the recognition and positioning capabilities of the difficult category semantic segmentation network, thereby effectively improving the performance of the night street scene semantic segmentation algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0073] Figure 1 It is a flow chart of the implementation method of the embodiment of the present invention.
[0074] Figure 2 It is a structural diagram of a network model in an embodiment of the present invention. DETAILED DESCRIPTION
[0075] In order to make the features and advantages of this patent more obvious and easy to understand, the following embodiments are specifically described in detail as follows:
[0076] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0077] The embodiment of the present invention provides a nighttime street scene semantic segmentation method based on a difficult category perception mechanism. Figure 1 , Figure 2 As shown in the figure, the core lies in the construction of a nighttime street scene semantic segmentation model based on the difficult category perception mechanism, which includes the following steps:
[0078] Step S1: Divide the nighttime street view dataset into a training set and a test set, and perform data preprocessing on the nighttime urban landscape images and corresponding labels in the dataset, including data enhancement and normalization processing;
[0079] Step S2: Based on the main semantic segmentation network, an auxiliary branch for implicitly perceiving exposure and texture is constructed. Through gradient back propagation, the encoder in the network implicitly learns the exposure features and texture features in the image.
[0080] Step S3: construct a difficult category semantic segmentation network and a difficult category perception module, wherein the difficult category perception module uses the characteristics of the main semantic segmentation network to enhance the recognition and positioning capabilities of the difficult category semantic segmentation network encoder, and finally fuses the segmentation results of the two networks;
[0081] Step S4: constructing a training pipeline for nighttime street scene semantic segmentation based on a difficult category perception mechanism, and using the pipeline and training set images to train a nighttime street scene semantic segmentation model based on a difficult category perception mechanism;
[0082] Step S5: input the night street scene test image into the trained night street scene semantic segmentation model based on the difficult category perception mechanism, and output the corresponding semantic segmentation mask map.
[0083] In this embodiment, step S1 specifically includes the following steps:
[0084] Step S11: using the nighttime urban landscape dataset NightCity, for the images in the dataset, the dataset is divided into a training set and a test set according to a certain ratio, the training set includes 2998 images, and the test set includes 1299 images;
[0085] Step S12: Perform data augmentation on the images in the training set to increase the number of samples in the data set, including randomly flipping images, randomly cropping images, photometric distortion, etc.;
[0086] Step S13: Preprocess the image after data enhancement in step S12 and convert it into the input image of the night street scene semantic segmentation network. First, the image size is cropped to 512×512 pixels, and then the cropped image with uniform size is normalized to convert the image data into a standard normal distribution; in order to ensure that the size and position of the segmented area in the label correspond to the night city landscape image, the same operation is performed on the label at each step of image data enhancement and image preprocessing.
[0087] In this embodiment, step S2 specifically includes the following steps:
[0088] Take c=128, h=512, w=512.
[0089] Step S21: Input the night scene image into the encoder of the subject semantic segmentation network to obtain the dimensions of The four-layer feature map {F1, F2, F3, F4} is input into the decoder to obtain the main semantic segmentation map, where h, w and c represent the height, width and number of channels of the feature map respectively. The specific expression is:
[0090] S = ClsSeg(F1, F2, F3, F4)
[0091] Among them, ClsSeg(·) represents semantic segmentation of features to obtain the semantic segmentation mask map of night street scene;
[0092] Step S22: {F1, F2, F3, F4} obtained from step S21 is used as the input of the exposure texture auxiliary branch, where the feature maps {F1, F2, F3} are respectively input into the 1×1 convolution layer, and the number of channels of the corresponding features is adjusted. The specific expression is:
[0093] F i ′=w i (F i )+b i , i=1,2,3
[0094] Among them, w i , b i are the weights and biases of the 1×1 convolutional layer. Through the 1×1 convolutional layer, the dimensions of {F1′, F2′, F3′} are adjusted to
[0095] Step S23: The feature map F4 with the highest semantics is sent to the spatial pyramid pooling module, that is, pooling of different scales is used on the original feature map to obtain multiple feature maps of different sizes, and then these feature maps are spliced in the channel dimension, and finally a composite feature map that integrates multiple scales is output, so as to achieve an enhanced feature F4′ that takes into account both global semantic information and local detail information. Finally, the dimension of F4′ is adjusted to The specific expression is:
[0096] P i =AdaptiveAvgPool(F4,ε),i=1,2,3,4
[0097] P i ′=w i (P i )+b i , i=1,2,3,4
[0098] F4′=Concat(P1′, P2′, P3′, P4′)
[0099] Where AdaptiveAvgPool(·) represents the adaptive average pooling operation, ε is the pooling size, and wi , b i are the weights and biases of the 1×1 convolutional layer, and Concat(·) indicates that the features are concatenated in a new dimension;
[0100] Step S24: The four feature maps {F1′, F2′, F3′, F4′} with the same number of channels obtained in steps S22 and S23 are fused at adjacent levels. The operation is to upsample the more abstract and semantically stronger high-level feature maps and then fuse them with the feature maps of adjacent levels, further enhancing the semantic information and position information. The specific expression is:
[0101] F4″=F4′
[0102] U i =Upsample(F i+1 ′)+F i ′, i = 1, 2, 3
[0103] U i ′=w i (U i )+b i , i=1,2,3
[0104] F i = Upsample(U i ′), i=1, 2, 3
[0105] E=Concat(F1″, F2″, F3″, F4″)
[0106] Among them, Upsample(·) represents the upsampling operation, U i ′ represents the output feature after the i-th deep convolutional layer, w i , b i is the weight and bias of the i-th depth convolutional layer, Concat(·) indicates that the features are concatenated in dimension;
[0107] Step S25: Send the exposure texture feature E obtained in step S24 to the exposure texture decoder to obtain an exposure texture map. Through gradient back propagation, the encoder implicitly learns the exposure features and texture features in the image. The specific expression is as follows:
[0108] E′=w(E)+b
[0109] I=tRGB(E′)
[0110] Where w, b are the weights and biases of the 3×3 convolutional layer, and tRGB(·) represents the conversion of features into a 3-channel image to obtain an exposure texture map.
[0111] In this embodiment, step S3 specifically includes the following steps:
[0112] Step S31: The four-layer feature map in the encoder of the main semantic segmentation network Four-layer feature maps from the encoder of a semantic segmentation network for hard categories As the input of the difficult category perception module, They are the features extracted by the i-th layer encoder of the main semantic segmentation network and the difficult category semantic segmentation network respectively;
[0113] Step S32: Add a difficult category perception module before each layer of the encoder of the difficult category semantic segmentation network, calculate the affinity matrix for the features of the i-th layer to enhance the recognition and positioning of difficult categories, and the output of the module is the features of the enhanced perception of difficult categories The specific expression is as follows:
[0114]
[0115]
[0116] where ⊙ represents the element-by-element matrix multiplication, S i Represents the affinity matrix calculated for the i-th layer features.
[0117] In this embodiment, step S4 specifically includes the following steps:
[0118] Step S41: Based on the SwinTransformer-UperNet semantic segmentation network, a main semantic segmentation network and a difficult category semantic segmentation network are constructed respectively. For the main semantic segmentation branch, SwinTransformer is used as an encoder, and it is used to extract features from the night street scene image after preprocessing in step S1, and four feature maps {F1, F2, F3, F4} with different scales and different numbers of channels are obtained. The four features are sent as input to the UperNet decoder, and the features are decoded into a semantic segmentation mask map. At the same time, the exposure texture auxiliary branch constructed in step S2 is inserted into the encoder of the network, and the exposure texture map is obtained through this branch, so that the encoder can implicitly learn the exposure features and texture features in the image. That is, the input of the network is the night street scene image I and the corresponding label L all , the output of the network is two images: night street scene semantic segmentation mask map P all and exposure texture segmentation map I;
[0119] Step S42: The network in step S41 is called M. To further enhance the semantic segmentation capability of the model for difficult categories, a difficult category semantic segmentation network H is added. The network structure of H is similar to that of M. To enhance the perception capability for difficult categories, the difficult category perception module constructed in step S3 is added. To save memory, the exposure texture auxiliary head branch is deleted. The input of the difficult sample semantic segmentation network H is changed to the night street scene image I and the corresponding label L hard , where L hard The representative label only contains difficult categories, and non-difficult categories are set as background categories. The output of the difficult category semantic segmentation network H is the night street scene semantic segmentation difficult category mask map P hard ;
[0120] Step S43: Input a batch of images and corresponding labels from the training set of step S1 into the network of step S41 for training, predict the nighttime street scene semantic segmentation mask map and exposure texture map, and save the network parameters after training for 90,000 iterations. Then freeze the network parameters of step S41 and jointly train the network in step S42, predict the nighttime street scene semantic segmentation difficult category mask map, and save the network parameters after training for 90,000 iterations;
[0121] Step S44: According to the loss function of the main semantic segmentation network M, the gradient of each parameter in the network is calculated using the back propagation method, and the parameters are updated using the stochastic gradient descent method. The loss function L of the network main as follows:
[0122] L main =loss ce +loss aux
[0123]
[0124] loss aux =loss exp +loss spa +loss sem
[0125]
[0126]
[0127]
[0128] Among them, L main is the loss function of network M, loss ce Represents the cross entropy loss function. loss auxRepresents the exposure texture auxiliary loss function, which is composed of exposure loss loss exp , spatial consistency loss spa , semantic brightness consistency loss sem Composition. For exposure loss exp , M represents the number of non-overlapping local regions of size 16×16, Y k represents the average intensity value of the local area in the exposure texture recovery map, E is the average intensity value of the assumed good exposure level, which is set to 0.6 here; for the spatial consistency loss loss spa , K is the number of local regions, the size of K is 4×4, ω(i) is the four adjacent regions (upper, lower, left, and right) centered on region i, Y and I represent the average intensity values of local regions in the enhanced image and the ordinary image, respectively; for semantic brightness consistency loss loss sem , S represents the number of categories for semantic prediction, which is set to 19 here, θ s Represented as a set of pixel indices belonging to category s, Represents the exposure texture restored image I H The intensity level at pixel i, B s represents the average strength level of category s.
[0129] Step S45: According to the loss function of the difficult category semantic segmentation network H, the gradient of each parameter in the network is calculated using the back propagation method, and the parameters are updated using the stochastic gradient descent method. The loss function L of the network is as follows:
[0130] L hard =loss ce
[0131] Step S46: Repeat the above steps S43 to S45 in batches until the loss value calculated in step S35 converges and stabilizes, save the network parameters, and complete the training process of the night street scene semantic segmentation network with difficult category perception mechanism.
[0132] In this embodiment, step S5 specifically includes the following steps:
[0133] Step S51: Input the nighttime street scene images in the test set into the trained subject semantic segmentation network and the difficult category semantic segmentation network respectively, and output the corresponding nighttime street scene semantic segmentation mask images P respectively. all , Nighttime Street Scene Semantic Segmentation Difficult Category Mask Map P hard ;
[0134] Step S52: P obtained in step S51 all , P hardIn order to improve the semantic segmentation effect of difficult categories, P is selected by the confidence level. hard The "high confidence region" in P all The confidence of the same position in is replaced to improve the model's ability to recognize difficult categories. The "high confidence area" is defined as the area where the difficult category semantic segmentation network predicts better than the main semantic segmentation network on the difficult category. The replacement calculation formula is as follows:
[0135]
[0136] Among them, i represents the current semantic category, C easy represents the simple category without the difficult category, C hard Indicates difficulty categories, which are set in a preferred embodiment as: pole, traffic light, zone, rider, motorcycle and bicycle.
[0137] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0138] The present invention is described with reference to flowcharts of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process in the flowchart, as well as the combination of processes in the flowchart, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A device that specifies functions in a process or multiple processes.
[0139] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple functions specified in a flowchart.
[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 The steps of a specified function in a process or multiple processes.
[0141] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
[0142] This patent is not limited to the above-mentioned optimal implementation mode. Anyone can derive other forms of nighttime street scene semantic segmentation methods based on difficult category perception mechanisms under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of the present invention should be covered by this patent.
Claims
1. A nighttime street scene semantic segmentation method based on difficult category perception mechanism, characterized in that: The semantic segmentation of night street scenes is performed using a semantic segmentation mask map output by a night street scene semantic segmentation model based on a difficult category perception mechanism, including the following steps: Step S1: Divide the nighttime street view dataset into a training set and a test set, and perform data preprocessing on the nighttime urban landscape images and corresponding labels in the dataset; Step S2: Based on the main semantic segmentation network, an auxiliary branch for implicitly perceiving exposure and texture is constructed. Through gradient back propagation, the encoder in the network implicitly learns the exposure features and texture features in the image. Step S3: construct a difficult category semantic segmentation network and a difficult category perception module, wherein the difficult category perception module uses the characteristics of the main semantic segmentation network to enhance the recognition and positioning capabilities of the difficult category semantic segmentation network encoder, and finally fuses the segmentation results of the two networks; Step S4: constructing a training pipeline for nighttime street scene semantic segmentation based on a difficult category perception mechanism, and using the pipeline and training set images to train a nighttime street scene semantic segmentation model based on a difficult category perception mechanism; Step S5: inputting the nighttime street scene test image into the trained nighttime street scene semantic segmentation model based on the difficult category perception mechanism, and outputting the corresponding semantic segmentation mask image; Step S3 specifically includes the following steps: Step S31: The four-layer feature map in the encoder of the main semantic segmentation network Four-layer feature maps from the encoder of a semantic segmentation network for hard categories As the input of the difficult category perception module, They are the features extracted by the i-th layer encoder of the main semantic segmentation network and the difficult category semantic segmentation network respectively; Step S32: Add a difficult category perception module before each layer of the encoder of the difficult category semantic segmentation network, calculate the affinity matrix for the features of the i-th layer to enhance the recognition and positioning of difficult categories, and the output of the module is the features of the enhanced perception of difficult categories The specific expression is as follows: where ⊙ represents the element-by-element matrix multiplication, S i represents the affinity matrix calculated for the i-th layer features; Step S4 specifically includes the following steps: Step S41: Based on the SwinTransformer-UperNet semantic segmentation network, a main semantic segmentation network and a difficult category semantic segmentation network are constructed respectively: for the main semantic segmentation branch, SwinTransformer is used as an encoder to extract features from the night street scene image after preprocessing in step S1, and four feature maps {F1, F2, F3, F4} with different scales and different numbers of channels are obtained; the four features are sent as input to the UperNet decoder, the features are decoded into a semantic segmentation mask map, and the exposure texture auxiliary branch constructed in step S2 is inserted after the encoder of the network to obtain an exposure texture map, so that the encoder implicitly learns the exposure features and texture features in the image; the input of the network is the night street scene image I and the corresponding label L all , the output of the network is two images: night street scene semantic segmentation mask map P all and exposure texture segmentation map I; Step S42: The network constructed in step S41 is referred to as M. In order to further enhance the semantic segmentation capability of the model for difficult categories, a difficult category semantic segmentation network H is added. The network structure of network H is based on network M. In order to enhance the perception capability of difficult categories, a difficult category perception module constructed in step S3 is added. The input of the difficult sample semantic segmentation network H is the night street scene image I and the corresponding label L. hard , where L hard The representative label only contains difficult categories, and non-difficult categories are all set as background categories; the output of the difficult category semantic segmentation network H is the night street scene semantic segmentation difficult category mask map P hard ; Step S43: input a batch of images and corresponding labels in the training set of step S1 into the network of step S41 for training, predict and obtain a semantic segmentation mask map of night street scenes and an exposure texture map, and then freeze the network parameters of step S41 to jointly train the network in step S42 to obtain a semantic segmentation mask map of difficult categories of night street scenes; Step S44: According to the loss function of the main semantic segmentation network M, the gradient of each parameter in the network is calculated using the back propagation method, and the parameters are updated using the stochastic gradient descent method; the loss function L of the network main as follows: L main =loss ce +loss aux loss aux =loss exp +loss spa +loss sem Among them, L main is the loss function of network M, loss ce Represents the cross entropy loss function; loss aux Represents the exposure texture auxiliary loss function, which is composed of exposure loss loss exp , spatial consistency loss spa , semantic brightness consistency loss sem Composition; for exposure loss exp , M represents the number of non-overlapping local regions, Y k represents the average intensity value of the local area in the exposure texture segmentation map, E is the average intensity value of the assumed good exposure level; for the spatial consistency loss loss spa , K is the number of local regions, ω(i) is the four adjacent regions above, below, left and right centered on region i, Y and I represent the average intensity values of local regions in enhanced and ordinary images, respectively; for semantic brightness consistency loss loss sem , S represents the number of categories for semantic prediction, θ s Represented as a set of pixel indices belonging to category s, Represents the exposure texture segmentation image I H The intensity level at pixel i, B s represents the average intensity level of category s; Step S45: According to the loss function of the difficult category semantic segmentation network H, the gradient of each parameter in the network is calculated using the back propagation method, and the parameters are updated using the stochastic gradient descent method; the loss function L of the network is as follows: L hard =loss ce Step S46: Repeat steps S43 to S45 in batches until the loss value calculated in step S35 converges and stabilizes, save the network parameters, and complete the training process of the night street scene semantic segmentation network with difficult category perception mechanism.
2. The method for nighttime street scene semantic segmentation based on difficult category perception mechanism according to claim 1, characterized in that: Step S1 specifically includes the following steps: Step S11: Divide the data set into a training set and a test set according to a certain ratio; Step S12: performing data augmentation on the images in the training set to increase the number of samples in the data set; Step S13: Preprocess the image after data enhancement in step S12 and convert it into an input image for the night street scene semantic segmentation network: first, randomly crop the image, and then normalize the cropped image with uniform size to convert the image data into a standard normal distribution; in order to ensure that the size and position of the segmented area in the label correspond to the night city landscape image, the same operation is performed on the label while enhancing the image data and preprocessing the image in each step.
3. The method for nighttime street scene semantic segmentation based on difficult category perception mechanism according to claim 1, characterized in that: Step S2 specifically includes the following steps: Step S21: Input the night scene image into the encoder of the subject semantic segmentation network to obtain the dimensions of The four-layer feature map {F1, F2, F3, F4} is input into the decoder to obtain the main semantic segmentation map, where h, w and c represent the height, width and number of channels of the feature map respectively. The specific expression is: S=ClsSeg(F1,F2,F3,F4) Among them, ClsSeg(·) represents semantic segmentation of features to obtain the semantic segmentation mask map of night street scene; Step S22: {F1, F2, F3, F4} obtained from step S21 is used as the input of the exposure texture auxiliary branch, where the feature maps {F1, F2, F3} are respectively input into the 1×1 convolution layer, and the number of channels of the corresponding features is adjusted. The specific expression is: F′ i =w i (F i )+b i ,i=1,2,3 Among them, w i ,b i are the weights and biases of the 1×1 convolutional layer; through the 1×1 convolutional layer, the dimensions of {F1', F2', F3'} are adjusted to Step S23: The feature map F4 with the highest semantics is sent to the spatial pyramid pooling module. Pooling of different scales is used on the original feature map to obtain multiple feature maps of different sizes. These feature maps are then concatenated in the channel dimension, and finally a composite feature map that integrates multiple scales is output, thereby achieving an enhanced feature F4' that takes into account both global semantic information and local detail information. Finally, the dimension of F4' is adjusted to The specific expression is: P i =AdaptiveAvgPool(F4,ε),i=1,2,3,4 P′ i =w i (P i )+b i ,i=1,2,3,4 F4′=Concat(P1′,P2′,P3′,P4′) Where AdaptiveAvgPool(·) represents the adaptive average pooling operation, ε is the pooling size, and w i ,b i are the weights and biases of the 1×1 convolutional layer, and Concat(·) indicates that the features are concatenated in a new dimension; Step S24: The four feature maps {F1', F2', F3', F4'} with the same number of channels obtained in steps S22 and S23 are fused at adjacent levels. The operation is to upsample the more abstract and semantically stronger high-level feature maps and then fuse them with the feature maps of adjacent levels to further enhance the semantic information and position information. The specific expression is: F4″=F4′ U i =Upsample(F i+1 ′)+F i ′,i=1,2,3 U i ′=w i (U i )+b i ,i=1,2,3 F i ″=Upsample(U i ′),i=1,2,3 E=Concat(F1″,F2″,F3″,F4″) Among them, Upsample(·) represents the upsampling operation, U i ' represents the output feature after the i-th deep convolutional layer, w i ,b i is the weight and bias of the i-th depth convolutional layer, Concat(·) indicates that the features are concatenated in dimension; Step S25: Send the exposure texture feature E obtained in step S24 to the exposure texture decoder to obtain an exposure texture map. Through gradient back propagation, the encoder implicitly learns the exposure features and texture features in the image. The specific expression is as follows: E′=w(E)+b I=tRGB(E′) Where w, b are the weights and biases of the 3×3 convolutional layer, and tRGB(·) represents the conversion of features into a 3-channel image to obtain an exposure texture map.
4. The method for nighttime street scene semantic segmentation based on difficult category perception mechanism according to claim 1, characterized in that: In step S42, when constructing the difficult category semantic segmentation network H, in order to save memory, the exposure texture auxiliary head branch is deleted.
5. The method for nighttime street scene semantic segmentation based on difficult category perception mechanism according to claim 1, characterized in that: Step S5 specifically includes the following steps: Step S51: Input the nighttime street scene images in the test set into the trained subject semantic segmentation network and the difficult category semantic segmentation network respectively, and output the corresponding nighttime street scene semantic segmentation mask images P respectively. all , Nighttime Street Scene Semantic Segmentation Difficult Category Mask Map P hard .
6. The method for nighttime street scene semantic segmentation based on difficult category perception mechanism according to claim 5, characterized in that: Step S5 also includes the following steps: Step S52: P obtained in step S51 all , P hard In order to improve the semantic segmentation effect of difficult categories, P is selected by the confidence level. hard The "high confidence region" in P all The confidence of the same position in the image is replaced to improve the model's ability to recognize difficult categories. The "high confidence area" is defined as the area where the difficult category semantic segmentation network predicts better than the main semantic segmentation network on the difficult category. The calculation formula for replacement is as follows: Among them, i represents the current semantic category, C easy represents the simple category without the difficult category, C hard Indicates difficulty category.
7. A nighttime street scene semantic segmentation system based on difficult category perception mechanism, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the method according to any one of claims 1 to 6 can be implemented.
8. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the method according to any one of claims 1 to 6 can be implemented.
Citation Information
Patent Citations
Image semantic segmentation method and system based on semantic propagation and foreground and background perception
CN114494699A
Automatic driving scene panorama segmentation method based on multi-modal fusion perception
CN116129233A