A Detection Method for the Waiting Area Based on Semantic Segmentation

Through the delayed area detection method based on semantic segmentation, features are extracted using feature pyramids and information transmission modules, and combined with clustering and curve fitting algorithms, the difficulty of delayed area detection in non-site traffic law enforcement scenarios is solved, and high accuracy detection is achieved in complex environments.

CN115223112BActive Publication Date: 2025-06-24HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210921648.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-06-24
Estimated Expiration
2042-08-02

AI Technical Summary

Technical Problem

In the non-site traffic law enforcement scenario, there are difficulties in testing the waiting area, such as vehicle occlusion, lack of obvious characteristics, and complex environmental interference, which makes it difficult for the existing technology to achieve accurate detection.

Method used

The delayed area detection method based on semantic segmentation is adopted. By obtaining the annotated training data set, high-quality images are screened, and a waiting area detection network is trained. The network consists of an encoder, multiple information transmission modules and decoders. Features are extracted using feature pyramids and information transmission modules, and combined with clustering and curve fitting algorithms, the delayed area is accurately detected.

Benefits of technology

In complex environments, the accuracy and efficiency of detection of the to-go area is improved, and the to-go area can be accurately identified under vehicle occlusion and complex lighting conditions, reducing the difficulty and workload of manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223112B_ABST
    Figure CN115223112B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting a waiting area based on semantic segmentation, which relates to the field of semantic segmentation in deep learning. The present invention can be used to detect a left-turn waiting area or a right-turn waiting area in traffic non-site law enforcement images. The method obtains pixel points representing the lane lines in the form of curves on both sides of the waiting area by performing pixel-level prediction on the image, then obtains a curve model through a clustering and fitting algorithm, and finally connects both ends of the curve to obtain the waiting area region. The present invention has a good effect on detecting the waiting area in a complex environment and has a high accuracy in detecting the waiting area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of semantic segmentation in deep learning, and specifically relates to a detection method for a waiting area in the scenario of traffic non-site law enforcement. Background Art

[0002] In the scenario of traffic non-site law enforcement, it is necessary to capture illegal images of vehicles through a camera installed above the intersection. When determining whether a vehicle has committed a red-light running violation, it is often necessary to detect the waiting area on the road. However, the traditional determination of the waiting area relies on manual identification, and its efficiency is extremely low. Currently, with the development of artificial intelligence-assisted determination technology, artificial intelligence algorithms such as neural networks are gradually introduced to identify the waiting area from traffic non-site law enforcement images, and then determine whether a vehicle has committed an illegal act such as running a red light in the waiting area. Most traffic accidents are caused by red-light running violations. There are a large number of law enforcement cameras for this illegal act, and the determination of this illegal act is one of the important links in realizing artificial intelligence-assisted determination. Specifically, for artificial intelligence to determine illegal red-light running, in addition to determining whether the vehicle stops before the stop line when the traffic light is red, it is also necessary to determine whether it is currently allowed to enter the left-turn waiting area or the straight-through waiting area. If it is allowed to enter the waiting area, it cannot be determined that the vehicle has committed a red-light running violation just because the vehicle crosses the stop line. At the same time, it is also necessary to determine whether the vehicle is within the specified area of the waiting area. If the vehicle moves forward beyond the waiting area during the red light but is allowed to enter the waiting area, it can be determined that the vehicle has committed a red-light running violation. Therefore, it is necessary to accurately and specifically detect the area where the waiting area is located. Currently, the detection of the waiting area is mainly manually marked, with extremely high marking difficulty, a large amount of marking work, and it is extremely easy to be blocked. There is an urgent need for a waiting area detection algorithm to assist manual marking to reduce the workload and improve efficiency. However, the current waiting area detection technology is extremely lacking. Only the approximate area of the waiting area can be detected through object detection technology based on deep learning, but it is still difficult to detect the specific area. Therefore, the research on the detection of the waiting area in the scenario of traffic non-site law enforcement is very necessary.

[0003] The waiting area is an area composed of two curved dotted lines and a solid line. Obviously, the target detection method cannot well detect the waiting area. On the one hand, the features of the waiting area are not obvious enough. On the other hand, even if the waiting area can be detected, what is required for illegal determination is an accurate and specific area of the waiting area, and the rectangular box obtained by target detection is difficult to give such an area. Therefore, the target detection method cannot be used to detect the waiting area.

[0004] Currently, most of the semantic segmentation methods for detecting curves are for the autonomous driving scenario. In actual road conditions, when it is difficult to detect many complex environments, such as shadows and blocked lane lines, the model cannot detect the lane lines. In this case, lane line tracking techniques, such as Kalman filtering, are required to supplement the detected missing lane lines based on the historical state and road geometry relationship, and make the lane lines more stable in terms of spatial position. Since the autonomous driving scenario is in the form of video and there is a connection between frames, this method can be adopted. However, there is no such connection in this non-traffic on-site law enforcement scenario, so the lane line tracking technique cannot be used. Therefore, these methods cannot be applied in the non-traffic on-site law enforcement scenario either.

[0005] Therefore, the detection of the waiting area in the non-traffic on-site law enforcement scenario has the following difficulties:

[0006] (1) Most of the waiting area will be blocked by passing vehicles, making it difficult to detect this area

[0007] (2) It is necessary to accurately obtain the specific area of the waiting area, rather than just a general range

[0008] (3) The features of the waiting area are not obvious enough and are easily confused with other targets, such as lane lines

[0009] (4) The actual traffic conditions are complex, and the wear and tear of the waiting area signs and night conditions will interfere with the detection Summary of the Invention

[0010] The purpose of the present invention is to solve the problems existing in the detection of the waiting area in the non-traffic on-site law enforcement scenario in the prior art, and provide a method for detecting the waiting area based on semantic segmentation.

[0011] The specific technical solution adopted by the present invention is as follows:

[0012] A method for detecting the waiting area based on semantic segmentation is used to detect the left-turn waiting area or the right-turn waiting area in the non-traffic on-site law enforcement image, and it includes:

[0013] S1. Obtain a labeled training data set, where each image sample contains an image taken from above by a law enforcement camera and including the waiting area. The waiting area lane lines in the form of curved dotted lines on both sides of the waiting area in the image are all marked with points; the image samples in the training data set belong to different intersection scenarios, and all image samples are divided into a daytime image subset taken during the day and a nighttime image subset taken at night according to the shooting time;

[0014] S2. For each image sample in the training dataset, comprehensively considering the two retention principles that the fewer the number of vehicles, the higher the priority, and daytime images are prior to nighttime images, filter all image samples under the same intersection scene by combining the grayscale value of the image and the number of vehicles in the image, and for each intersection scene, remove the image samples exceeding the threshold number;

[0015] S3. Aiming at minimizing the loss function, use the training dataset filtered by S2 to train the stop line detection network;

[0016] The stop line detection network consists of an encoder, a multi-information transfer module, and a decoder;

[0017] In the encoder, a feature pyramid based on the ResNet50 backbone network is used as the basic feature extraction network to extract 4 feature maps of different sizes from the original input image;

[0018] In the multi-information transfer module, multiple information transfer operations need to be iteratively performed on each feature map output by the encoder. Each information transfer operation requires slicing the feature map in 4 directions: from top to bottom, from left to right, from right to left, and from bottom to top. The information between the slices is mutually transferred, and the step size of the information transfer is controlled to increase during the iterative information transfer operation to ensure that each slice can receive the information of the entire feature map;

[0019] The decoder receives the 4 feature maps of different sizes output by the multi-information transfer module, performs upsampling on the feature maps in ascending order of size and fuses them with larger-sized feature maps until all 4 feature maps are fused together and then upsampled to restore to the size of the original input image;

[0020] The loss function is the weighted sum of the segmentation loss and the classification loss;

[0021] S4. Input the image to be detected containing the stop line into the trained stop line detection network to obtain all the pixel points recognized as the stop line in the image to be detected. Then, cluster these pixel points based on the point spacing, and the pixel points belonging to the same lane line are clustered into one category; then, perform curve fitting on each category of pixel points respectively to obtain the fitting curve segments of each stop line. Connect the endpoints of the fitting curve segments corresponding to the lane lines on both sides of the same stop line to obtain the stop line detection result.

[0022] Preferably, in the training dataset, the stop lines in each image sample are marked with dots using a marking tool, and the marked points on each stop line need to restore the curve segment corresponding to the lane line.

[0023] Preferably, the specific method of S2 is as follows:

[0024] S21. Convert each image sample in the training dataset from an RGB image to a grayscale image, then calculate the grayscale mean of all pixels in each image sample, and then calculate the average of the grayscale means of all image samples in each of the daytime image subset and the nighttime image subset respectively, which is used as the average brightness of the corresponding subset; use the average of the average brightness of the two subsets as the brightness discrimination threshold for distinguishing day and night;

[0025] S22. Use the trained object detection model to detect vehicles in each image sample in the training dataset, obtain the number of vehicles in each image sample, then calculate the average number of vehicles in all image samples in the training dataset, and finally calculate the vehicle weight of each image sample as the ratio of the number of vehicles in the image sample to the average number of vehicles multiplied by the average brightness of the daytime image subset;

[0026] S23. According to the brightness discrimination threshold and the vehicle weight, calculate the quality weight of each image sample in the training dataset = 255 + λ * α * gray - β * carWeight, where gray represents the grayscale mean of all pixels in the currently calculated image sample, carWeight represents the vehicle weight corresponding to the currently calculated image sample, α and β are two weights respectively, λ is the weight value determined by the brightness discrimination threshold bound and gray. If gray ≥ bound, then λ = λ1; if gray < bound, then λ = λ2, λ1 + λ2 = 1 and λ1 > λ2;

[0027] S24. For all image samples in each intersection scene in the training dataset, sort them according to their respective quality weights. If the number of image samples in an intersection scene exceeds the threshold number, then retain the image samples that meet the threshold number in descending order of quality weight. If the number of image samples in an intersection scene does not exceed the threshold number, then retain all image samples.

[0028] Preferably, the weights α and β are 1 and 2 respectively, and the weight values λ1 and λ2 are 0.6 and 0.4 respectively.

[0029] Preferably, in the multi - information transfer module, each feature output by the encoder Figure X is iteratively subjected to N information transfer operations, and in each information transfer operation, the feature map needs to be sliced horizontally or vertically in the four directions of top - to - bottom, left - to - right, right - to - left, and bottom - to - top respectively, and information is mutually transferred between the slices; where:

[0030] In the bottom - to - top direction, for the input feature Figure XPerform horizontal slicing and vertical information transfer between slices. The calculation formula for vertical information transfer between slices during any nth iteration is as follows:

[0031]

[0032] Perform vertical slicing on the input features in the right-to-left direction Figure X and horizontal information transfer between slices. The calculation formula for horizontal information transfer between slices during any nth iteration is as follows:

[0033]

[0034] In the formula: F p,l,q represents a set of convolution kernels. p, l, and q represent the number of input channels, the number of output channels, and the kernel width respectively; the symbol "·" is the convolution operator; f is the non-linear activation function ReLU; represents the feature Figure X at the nth iteration. k, i, and j represent the indices of the channel, row (H direction), and column (W direction) respectively; represents the after information transfer processing. n represents the current iteration number, and s n represents the step size of information transfer in the nth iteration. L is the width W and height H of the input feature respectively during vertical information transfer and horizontal information transfer. Figure X

[0035] Perform horizontal slicing on the input features in the top-to-bottom direction Figure X after vertically mirror-flipping along the horizontal symmetry plane, and perform the same vertical information transfer between slices as in the bottom-to-top direction;

[0036] Perform vertical slicing on the input features in the left-to-right direction Figure X after horizontally mirror-flipping along the vertical symmetry plane, and perform the same horizontal information transfer between slices as in the right-to-left direction.

[0037] Preferably, the decoder receives 4 feature maps of different sizes output by the multiple information transfer module, which are the first feature map, the second feature map, the third feature map, and the fourth feature map in order from largest to smallest size. First, use bilinear interpolation to upsample the fourth feature map to make its size the same as that of the third feature Figure 1 map, and at the same time reduce the number of channels by half, and then perform feature fusion with the third feature map to obtain the first fusion feature map; then use bilinear interpolation to upsample the first fusion feature map to make its size the same as that of the second feature Figure 1 ​Meanwhile, reduce the number of channels by half, then perform feature fusion with the second feature map to obtain a second fused feature map; then use bilinear interpolation to upsample the second fused feature map to make its size the same as the first feature Figure 1 Meanwhile, reduce the number of channels by half, then perform feature fusion with the first feature map to obtain a third fused feature map; upsample the third fused feature map and restore it to the size of the original input image to obtain an image to be classified, classify each pixel in the image to be classified to achieve semantic segmentation, and thus obtain a lane line recognition result.

[0038] Preferably, the calculation formula of the loss function is:

[0039] Loss = Loss CE + Loss BCE (3.5)

[0040]

[0041] Loss BCB = -αy c log(p c ) - (1 - α)(1 - y c )log(1 - p c ) (3.7)

[0042] where Loss BCE and Loss CE are the segmentation loss and the classification loss respectively; M represents the number of categories, c represents the category, ω c represents the weight of the loss; y c is a vector with values of 0 or 1, indicating whether the pixel category prediction is correct or not, 1 indicates correct, and 0 indicates wrong; p c represents the probability that the predicted pixel category is c; the segmentation loss is used to distinguish the background and the annotation, α represents the proportion of the background segmentation loss, and y c represents the true value corresponding to p c .

[0043] Preferably, the specific method for clustering all pixel points based on the point spacing is:

[0044] S41. Put all pixel points into the first set B initialized to be empty;

[0045] S42. Randomly take out a pixel point from the current first set B and add it to the second set A initialized to be empty;

[0046] S43. Traverse all the pixel points in the first set B, and determine whether there are pixel points in the second set A whose distance from the currently traversed pixel point is less than the maximum clustering distance. If so, add the currently traversed pixel point to the second set A; the maximum clustering distance is the maximum distance value allowed between adjacent pixel points on a lane line in the image.

[0047] S44. Keep repeating S43 until no new pixel points are added to the second set A. Then, take all the pixel points in the second set A as a clustering cluster, and the pixel points in this clustering cluster belong to the same lane line in the waiting area.

[0048] S45. Keep repeating S42 - S44 until all the pixel points in the first set B are divided into clustering clusters, and obtain the pixel points corresponding to each lane line in the waiting area.

[0049] Preferably, the curve fitting uses a cubic curve equation as the fitting equation.

[0050] Preferably, the curve fitting is implemented using the RANSAC algorithm.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] 1. The present invention has a good effect on detecting the waiting area in a complex environment. When constructing the data set, the present invention calculates the weights of the pictures in the same scene, and screens out the pictures with fewer vehicles (less occlusion of the waiting area position by vehicles) and good light environment (easier to identify in the daytime than in the nighttime) to participate in the training, which can well avoid the influence of the bad environment on the model.

[0053] 2. The present invention has a high accuracy in detecting the waiting area. Since the waiting area is easily occluded by passing vehicles, the present invention simultaneously detects the waiting area in multiple pictures in the same scene, and then clusters and fits the detection results of each picture, greatly improving the accuracy of detecting the waiting area in this scene. Description of the Drawings

[0054] Figure 1 is the annotation result of the curved dotted line;

[0055] Figure 2 is Figure 1 the enlarged view of the annotation in

[0056] Figure 3 is the effect after sorting a group of images according to the quality weight;

[0057] Figure 4 is the network architecture diagram for detecting the waiting area;

[0058] Figure 5It is the structural diagram of the encoder network;

[0059] Figure 6 It is the schematic diagram of a single information transfer operation;

[0060] Figure 7 It is the processing structure of information transfer operations in two directions;

[0061] Figure 8 It is the schematic diagram of N MP operations in the multiple information transfer module;

[0062] Figure 9 It is the schematic diagram of the upsampling process of the decoder;

[0063] Figure 10 It is an example of the model test result;

[0064] Figure 11 It is an example of the clustering result;

[0065] Figure 12 It is the RANSAC fitting result graph;

[0066] Figure 13 It is the detected graph of the area of the waiting area for turning. Specific embodiments

[0067] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.

[0068] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.

[0069] In a preferred embodiment of the present invention, a method for detecting the waiting area for turning based on semantic segmentation is provided, which is used to detect the left-turn waiting area or the right-turn waiting area in the traffic non-site law enforcement images. The main idea of the present invention is to obtain the pixel points representing the lane lines in the form of curves on both sides of the waiting area through pixel-level prediction of the image, then obtain the curve model through clustering and fitting algorithms, and finally connect the two ends of the curve to obtain the area of the waiting area. The method for detecting the waiting area for turning based on semantic segmentation specifically includes steps S1 to S4, which are described in detail as follows:

[0070] S1. Obtain the labeled training data set, where each image sample contains an image of a waiting area taken from above by a law enforcement camera. The waiting area lane lines in the form of curved dotted lines on both sides of the waiting area in the image are all marked with annotation points. The image samples in the training data set belong to different intersection scenarios, and all image samples are divided into a daytime image subset taken during the day and a nighttime image subset taken at night according to the shooting time.

[0071] In the present invention, in the above training data set, the waiting area lane lines in each image sample are marked with dots using an annotation tool, and the annotation points on each waiting area lane line need to restore the corresponding curve segment of the lane line.

[0072] S2. For each image sample in the training data set, combining the two retention principles that the fewer the number of vehicles, the higher the priority and daytime images are prior to nighttime images, filter all image samples in the same intersection scenario by combining the gray value of the image and the number of vehicles in the image, and respectively eliminate the image samples exceeding the threshold number for each intersection scenario.

[0073] In the present invention, the specific method of S2 above is as follows:

[0074] S21. Convert each image sample in the training data set from an RGB image to a grayscale image, then calculate the gray mean value of all pixels in each image sample, and then calculate the average value of the gray mean values of all image samples in each subset for the daytime image subset and the nighttime image subset respectively, as the average brightness of the corresponding subset. Use the average value of the average brightness of the two subsets as the brightness discrimination threshold for distinguishing day and night.

[0075] S22. Use the trained object detection model to detect vehicles in each image sample in the training data set to obtain the number of vehicles in each image sample, then calculate the average number of vehicles in all image samples in the training data set, and finally calculate the vehicle weight of each image sample as the ratio of the number of vehicles in the image sample to the average number of vehicles multiplied by the average brightness of the daytime image subset.

[0076] S23. According to the brightness discrimination threshold and the vehicle weight, calculate the quality weight of each image sample in the training data set = 255 + λ * α * gray - β * carWeight, where gray represents the gray mean value of all pixels in the currently calculated image sample, carWeight represents the vehicle weight corresponding to the currently calculated image sample, α and β are two weights respectively, λ is the weight value determined by the brightness discrimination threshold bound and gray. If gray ≥ bound, then λ = λ1; if gray < bound, then λ = λ2, λ1 + λ2 = 1 and λ1 > λ2.

[0077] In the present invention, the above-mentioned weights α and β are preferably 1 and 2 respectively, and the weights λ1 and λ2 are preferably 0.6 and 0.4 respectively.

[0078] S24. For all image samples in each intersection scene in the training dataset, sort them according to their respective quality weights. If the number of image samples in an intersection scene exceeds the threshold number, retain the image samples that meet the threshold number in descending order of quality weight. If the number of image samples in an intersection scene does not exceed the threshold number, retain all image samples.

[0079] S3. With the goal of minimizing the loss function, use the training dataset filtered by S2 to train the stop line detection network.

[0080] The above-mentioned stop line detection network is composed of an encoder, a multi-information transfer module, and a decoder, and the three are specifically as follows:

[0081] In the encoder, a feature pyramid based on the ResNet50 backbone network is used as the basic feature extraction network to extract 4 feature maps of different sizes from the original input image;

[0082] In the multi-information transfer module, it is necessary to perform multiple information transfer operations on each feature map output by the encoder iteratively. Each information transfer operation requires slicing the feature map in 4 directions: from top to bottom, from left to right, from right to left, and from bottom to top. The information between the slices is mutually transferred, and the step size of the information transfer is controlled to increase during the iterative information transfer operation to ensure that each slice can receive the information of the entire feature map;

[0083] The decoder receives the 4 feature maps of different sizes output by the multi-information transfer module, performs upsampling on the feature maps in ascending order of size and fuses them with larger-sized feature maps until all 4 feature maps are fused together and then upsampled to restore to the size of the original input image;

[0084] In the present invention, in the above-mentioned multi-information transfer module, it is necessary to perform N information transfer operations on each feature Figure X output by the encoder iteratively, and in each information transfer operation, it is necessary to perform slicing on the feature map in the horizontal or vertical direction in 4 directions: from top to bottom, from left to right, from right to left, and from bottom to top and transfer the information between the slices; where:

[0085] In the direction from bottom to top for the input feature Figure XPerform horizontal slicing and vertical information transfer between slices. The calculation formula for vertical information transfer between slices during any nth iteration is as follows:

[0086]

[0087] Perform vertical slicing on the input features in the right-to-left direction Figure X and horizontal information transfer between slices. The calculation formula for horizontal information transfer between slices during any nth iteration is as follows:

[0088]

[0089] In the formula: F p,l,q represents a set of convolution kernels, where p, l, and q represent the number of input channels, the number of output channels, and the kernel width respectively; the symbol "·" is the convolution operator; f is the non-linear activation function ReLU; represents the feature Figure X at the nth iteration, where k, i, and j represent the indices of the channel, row (H direction), and column (W direction) respectively; represents the one after information transfer processing n represents the current iteration number, and s n represents the step size of information transfer in the nth iteration, L is the width W and height H of the input feature respectively during vertical information transfer and horizontal information transfer Figure X ;

[0090] Perform horizontal slicing on the input features in the top-to-bottom direction Figure X after vertical mirror flipping along the horizontal symmetry plane, and perform the same vertical information transfer between slices as in the bottom-to-top direction;

[0091] Perform vertical slicing on the input features in the left-to-right direction Figure X after horizontal mirror flipping along the vertical symmetry plane, and perform the same horizontal information transfer between slices as in the right-to-left direction.

[0092] In the present invention, the above decoder receives the 4 feature maps with different sizes output by the multiple information transfer module. In order from largest to smallest in size, they are the first feature map, the second feature map, the third feature map, and the fourth feature map. First, use the bilinear interpolation method to upsample the fourth feature map to make its size the same as that of the third feature Figure 1 map, and at the same time reduce the number of channels by half, and then perform feature fusion with the third feature map to obtain the first fusion feature map; then use the bilinear interpolation method to upsample the first fusion feature map to make its size the same as that of the second feature Figure 1At the same time, reduce the number of channels by half, then perform feature fusion with the second feature map to obtain a second fused feature map; then use bilinear interpolation to upsample the second fused feature map to make its size the same as the first feature Figure 1 At the same time, reduce the number of channels by half, then perform feature fusion with the first feature map to obtain a third fused feature map; upsample the third fused feature map and restore it to the size of the original input image to obtain an image to be classified, and classify each pixel in the image to be classified to achieve semantic segmentation, thereby obtaining a lane line recognition result.

[0093] During the training process, the loss function adopted is the weighted sum of the segmentation loss and the classification loss.

[0094] In the present invention, the calculation formula of the above loss function is:

[0095] Loss=Loss CE +Loss BCE (3.5)

[0096]

[0097] Loss BCB =-αy c log(p c )-(1-α)(1-y c )log(1-p c ) (3.7)

[0098] Wherein, Loss BCE and Loss CE are the segmentation loss and the classification loss respectively; M represents the number of categories, c represents the category, ω c represents the weight of the loss; y c is a vector with values of 0 or 1, indicating whether the pixel category prediction is correct or not, 1 indicates correct, and 0 indicates wrong; p c represents the probability that the predicted pixel category is c; the segmentation loss is used to distinguish the background and the annotation, α represents the proportion of the background segmentation loss, and y c represents the true value corresponding to p c .

[0099] S4. Input the image to be detected containing the area to be traveled into the trained area-to-be-traveled detection network, obtain all the pixel points recognized as the lane lines in the area to be traveled in the image to be detected, and then cluster these pixel points based on the point spacing. The pixel points belonging to the same lane line are clustered into one category; then perform curve fitting on each category of pixel points respectively to obtain the fitted curve segments of each lane line in the area to be traveled, and connect the endpoints of the fitted curve segments corresponding to the lane lines on both sides of the same area to be traveled to obtain the detection result of the area to be traveled.

[0100] In the present invention, the specific method for clustering all pixels based on the point spacing is:

[0101] S41, putting all pixel points into a first set B which is initialized to be empty;

[0102] S42, randomly taking a pixel point from the current first set B and adding it to the second set A which is initialized to be empty;

[0103] S43, traversing all the pixel points in the first set B, determining whether there is a pixel point in the second set A whose distance to the currently traversed pixel point is less than the maximum clusterable distance, and if so, adding the currently traversed pixel point to the second set A; the maximum clusterable distance is the maximum distance value allowed between adjacent pixel points on a lane line in the image;

[0104] S44, continuously repeating S43 until no new pixel points are added to the second set A, taking all the pixel points in the second set A as a cluster, and the pixel points in the cluster belong to the same lane line of the waiting area;

[0105] S45, continuously repeating S42 to S44 until all the pixel points in the first set B are divided into clusters, and the pixel points corresponding to each lane line in the waiting area are obtained.

[0106] In the present invention, the curve fitting preferably adopts a cubic curve equation as the fitting equation, and the curve fitting method is preferably implemented by a RANSAC algorithm.

[0107] The waiting area detection method based on semantic segmentation shown in S1 to S4 above is applied to a specific example to demonstrate its specific implementation process and the technical effects that can be achieved.

[0108] Example

[0109] In this embodiment, the waiting area detection method based on semantic segmentation shown in S1 to S4 is specifically implemented through the following process:

[0110] Step 1. Create a dataset

[0111] The dataset used in this example is real data from non-site traffic enforcement scenarios. The image data comes from illegal captures captured by 678 law enforcement cameras, and contains a total of 3,498 image data, including various complex environments such as darkness, shadows and uneven lighting, rain on the road, stains and reflections, and vehicle occlusion. All data in this dataset are image data that includes the waiting area.

[0112] Since it is necessary to detect the waiting line area by detecting the curved dotted lines, it is necessary to label the dotted lines on both sides of the waiting line area. The labeling tool used is Labelme, and the labeling method of the public dataset TuSimple is adopted to label the self-made dataset in this chapter.

[0113] Considering that the position of the waiting line area is not fixed, some waiting line areas under law enforcement cameras are in a lower position, while some are in a higher position. Therefore, different from the TuSimple dataset that only labels 70% of the area below the image, the dataset of the present invention needs to label the entire image area. LineStrip is used to label the dotted lines, and more points need to be marked at places with a larger curvature to restore the curve as much as possible. The labeling result of one exemplary image sample is as Figure 1 and Figure 2 shown.

[0114] The above labeling result is converted into the existing TuSimple labeling file format through code. The result is as Figure 5 shown.

[0115] Step 2. Data preprocessing

[0116] The scenario targeted in this embodiment is the traffic non-site scenario, where each image sample is obtained by a law enforcement camera above an intersection taking a top-down photo of the intersection. Therefore, under the same law enforcement camera, the background of the image data is often the same, only the positions of people and vehicles are different. And each illegal vehicle corresponds to three pictures. That is to say, under the same device, we will obtain a large number of traffic image data with the same background, only the positions of people and vehicles are different. In this context, considering that various complex environments such as vehicle occlusion, night, shadow, uneven illumination, road surface rain, stains and reflections, and vehicle occlusion are likely to have a great impact on the detection of the waiting line area, so when making the dataset, it is necessary to try to select picture data in a relatively ideal environment with less vehicle occlusion and good weather conditions. Therefore, it is necessary to judge the weights of the dataset and make selections to improve the detection rate of the lane lines of the waiting line area.

[0117] To better evaluate the image quality, this embodiment combines the gray value of the image and the number of vehicles in the image, and adopts a quality weight formula to calculate the quality values of all traffic images under the same device through the quality weight formula. The higher the quality value, the higher the image quality. Then, the images are sorted according to the image quality. On the premise that the number of images is large enough, high-quality value images are preferentially used for lane line detection. By screening the picture quality through this quality weight, the detection rate of the lane lines of the waiting line area can be improved.

[0118] Step 2.1 Calculate the data thresholds for day and night

[0119] First, divide all the image data in the dataset into two batches of data. One batch is daytime image data (referred to as the daytime image subset), and the other batch is nighttime image data (referred to as the nighttime image subset). Calculate the average brightness of these two batches of image data respectively, and then obtain a brightness discrimination threshold that can distinguish between day and night by calculating the arithmetic mean of the daytime image data and the nighttime image data. The calculation formula is as follows:

[0120] Gray=R*0.299+G*0.587+B*0.114(2.1)

[0121]

[0122] Among them, R, G, and B in formula (2.1) represent the three channels of the image. The gray value of a single pixel is calculated through the values of these three channels. Gray represents the gray value of a single pixel. Formula (2.2) is used to calculate the average gray value of all pixels in an image. h and w are the height and width of the image respectively. Formula (2.3) represents calculating the average brightness of a batch of image data (that is, the average of the average gray values of all images). D represents a subset of data (which can be a batch of daytime data or a batch of nighttime data), and n represents the number of image samples in the subset. In formula (2.4), bright day represents the average brightness of the daytime image subset, and bright night represents the brightness of the nighttime image subset. Both can be calculated by formula (2.3), and bound represents the brightness discrimination threshold for distinguishing between day and night.

[0123] Step 2.2 Calculate the vehicle weight threshold according to the number of vehicles

[0124] Use the yolov3 pre-trained model for transfer learning in the scenario of the present invention to obtain an object detection model applicable to the scenario of the present invention. And use this model to detect the image data to obtain the number of vehicles and vehicle coordinates in the image data, and save the coordinates to a JSON file. In the JSON file format, the first two values in each list represent the upper left coordinates of the vehicle, and the last two numerical values represent the lower right coordinates of the vehicle.

[0125] Calculate the vehicle weight according to the number of vehicles in the image. The calculation formula is as follows:

[0126]

[0127] Among them, formula (2.5) calculates the average number of vehicles avg in all image data. n represents the number of images in the dataset, and S represents the dataset. Formula (2.6) calculates the vehicle weight of this image. carWeight represents the vehicle weight of the current image, carNum represents the number of vehicles in the current image, and bright dayrepresents the average brightness of the subset of daytime images. Here, the average gray value of the daytime images bright is used. day is to improve the influence of vehicle weight on image quality.

[0128] Step 2.3. Finally, according to the brightness discrimination threshold and vehicle weight, calculate the quality weight of each image sample in the training dataset. The calculation formula is as shown in the following formula (2.7):

[0129] weight = 255 + λ·α·gray - β·carWeight (2.7)

[0130]

[0131] Among them, α and β are the weights of brightness and the number of vehicles respectively. In this embodiment, these two values are set to 1 and 2 respectively. This is because the influence of the number of vehicles on lane lines is much greater than the influence of day and night. Therefore, the proportion of vehicle weight is increased. gray represents the aforementioned average gray value of the image, carWeight represents the vehicle weight of the image calculated by formula (2.6), and λ changes according to the value of gray. When the value of gray is greater than or equal to the boundary value between day and night, it is considered that this image is daytime, and the value of λ is 0.6; when the value of gray is less than the boundary value, it is considered that the image is night, and the value of λ is 0.4. This approach is mainly to increase the weight of daytime images. However, since the influence of day and night is not too large, only the daytime weight is increased to 0.6, and the night weight is set to 0.4.

[0132] In this embodiment, the above quality weights are used to sort the image data taken under the same intersection scene, that is, the same law enforcement camera. The reason for sorting instead of directly eliminating images with poor quality by setting a quality weight threshold is that sometimes there are only night images or images with a large number of vehicles. In this case, once the method of eliminating by this quality weight threshold is adopted, all images will be eliminated, and lane line detection cannot be performed. Therefore, the present invention adopts a sorting method to obtain high-quality images by selecting images with higher quality weights to improve the lane line detection effect.

[0133] In this embodiment, after calculating the quality weights of all images in the scene, sort the image samples according to the size of the quality weights, and then retain the first 6 image samples in ascending order of weights, and the remaining image samples are deleted from the dataset. The number threshold of the retained image samples is taken as 6, which is obtained through multiple experiments and can ensure the highest lane line recognition rate and the lowest error rate. If the number of image data is less than 6, there is no need to perform quality formula sorting, and all images are retained.

[0134] Figure 3Show a set of pictures to view the image quality sorting effect of this step. (a) The picture has few cars during the day, (b) the picture has few cars at night, (c) the picture has many cars during the day, (d) the picture has many cars at night. Since it is difficult to find pictures that meet the above 4 conditions under the same device, these 4 pictures come from 4 different devices. According to the weight formula, the quality weights of the pictures are calculated. It can be seen that the quality weight formula of the present invention emphasizes pictures with fewer vehicles.

[0135] Step 3. Construction of the SS-Net network model for the detection algorithm of the waiting area based on semantic segmentation

[0136] The network architecture is as Figure 4 shown. The model adopts a classic encoder-decoder structure. The encoder extracts features from the image and processes the feature map to obtain a feature map with rich semantics. Then, the decoder restores the image to the original image size and classifies each pixel to achieve the semantic segmentation effect.

[0137] Step 3.1 Construction of the encoder.

[0138] The encoder structure is as Figure 5 shown. The feature pyramid based on the ResNet50 backbone network is used as the basic feature extraction network.

[0139] The basic feature extraction network can initially extract features from the original image. The original image is scaled to the input size required by ResNet50, that is, 3×224×224. According to the characteristics of ResNet50 and FPN, 4 feature maps with different sizes (from small to large: 2048×7×7, 1024×14×14, 512×28×28, 256×56×56) will be output. In order to better capture the spatial relationship between row and column pixels, in this embodiment, a Message Passing (MP) module is used to transfer spatial information so that each pixel can obtain global information.

[0140] Step 3.2 Construction of the Message Passing (MP) module.

[0141] The MP module has 4 directions, namely from top to bottom, from left to right, from right to left, and from bottom to top. By slicing the feature map in the horizontal and vertical directions in these 4 directions and transferring the information between the slices, fine curve information can be extracted, and thus high-semantic features of the curve can be extracted.

[0142] Perform information transfer in these four directions (denoted as U, D, L, R respectively) on a feature map to form an MP module. The structure diagram of the MP module is as Figure 6As shown. However, this only conveys the information of adjacent slices, and the information of distant slices also needs to be transmitted. Therefore, a step size is added to the information transmission. The step size for information transmission between adjacent slices is 1. To ensure that the outermost slices can also receive the information of other slices, a cyclic shift method is adopted. When transmitting information from top to bottom, the information of the last slice will be transmitted to the first slice, and the same applies to information transmission in other directions. The step size for information transmission between every other slice is 2 until every slice has received the information of other slices, that is, the MP module is iterated N times with different step sizes for each iteration, so as to ensure that each slice can receive the information of the entire feature map.

[0143] As Figure 7 shown, the left figure is the processing structure for information transmission from bottom to top, and the right figure is the processing structure for information transmission from right to left. The calculation formula for slice information transmission is as follows:

[0144]

[0145] Among them, Equation (3.1) is the vertical information transmission formula, and F p,l,q represents a group of convolution kernels, and "·" is the convolution operator; p, l, and q represent the number of input channels, the number of output channels, and the kernel width respectively. Here, both p and l are 1. Equation (3.2) is the horizontal information transmission formula, and the specific information is the same as that of Equation (3.1). In Equation (3.3), f is the non-linear activation function ReLU, represents the value of the feature Figure X at the nth iteration, and k, i, and j represent the indices of the channel, row (in the H direction), and column (in the W direction) respectively, represents the processed In Equation (3.4), n represents the number of iterations, and s n represents the step size of information transmission in the nth iteration, L is the width W and height H of the input feature Figure X in Equations (3.1) and (3.2) respectively.

[0146] It should be noted that in the two directions from top to bottom and from bottom to top, the input feature Figure X is horizontally sliced after being vertically mirror-flipped along the horizontal symmetry plane, and vertical information transmission is performed between the slices. However, when transmitting information from bottom to top, the input feature Figure X is directly processed, while when transmitting information from top to bottom, the input feature Figure X needs to be horizontally sliced after being vertically mirror-flipped along the horizontal symmetry plane, and the vertical information transmission after slicing is all in accordance with Equation (3.1). Similarly, in the two directions from left to right and from right to left, the input feature Figure XAfter horizontal mirror flipping along the vertical symmetry plane, vertical slicing is performed, and horizontal information transfer is carried out between the slices. However, when transferring information from right to left, the input features Figure X are directly processed, while when transferring information from left to right, the input features Figure X After horizontal mirror flipping along the vertical symmetry plane, vertical slicing is performed, and the horizontal information transfer after slicing is all carried out according to Equation (3.2)

[0147] Through N times of MP operations, each slice has obtained the information on the entire feature map, and the semantic information is more complete. Then it is used for the subsequent decoder to upsample the feature map. The N times of MP operations are as Figure 8 shown.

[0148] Step 3.3 Upsampling process.

[0149] The 4 feature maps output by the encoder have all enriched the spatial semantic information through multiple iterative MP operations. These 4 feature maps are named F1, F2, F3, and F4 from large to small. The bilinear interpolation method is used to upsample the feature map F4, doubling the size of the feature map and reducing the number of channels to half at the same time. Then it is feature fused with the feature map F3 (implemented through the Concat operation, the same below). The above operations are continuously repeated until the 4 feature maps are all fused together and the image is restored to the size of the original input image. The decoder upsampling process is shown in Figure 9 shown. The image of the size of the original input image obtained after the decoder upsampling can further classify each pixel (for example, it can be binarized through a threshold of 0.5) to determine whether it belongs to the background or the lane line to achieve the semantic segmentation effect.

[0150] Step 3.4 Loss function.

[0151] The loss function formula used to train the above network is as follows:

[0152] Loss = Loss CE + Loss BCE (3.5)

[0153]

[0154] Loss BCE = -αy c log(p c ) - (1 - α)(1 - y c )log(1 - p c ) (3.7)

[0155] Among them, Equation (3.5) indicates that the loss function consists of two parts, namely the segmentation loss BCE and the classification loss CE. Equation (3.6) represents the classification loss, M represents the number of categories, c represents the category, and ω c represents the weight of the loss. In the scenario of this embodiment, the sample numbers of the background and the lane lines are extremely different, and the negative samples far exceed the positive samples. To prevent label imbalance from resulting in poor training effects, it is necessary to reduce the weight of the background category loss. Therefore, in this embodiment, the background category loss weight α is set to 0.3, and the curve annotation loss weight ω c is 1, y c is a vector with values of 0 or 1, indicating whether the pixel category prediction is correct. 1 indicates correct, and 0 indicates incorrect. p c represents the probability that the predicted pixel category is c. Equation (3.7) represents the segmentation loss used to distinguish the background and the annotation. α represents the proportion of the background segmentation loss, and p c represents the predicted category, and y c represents the corresponding true category label.

[0156] Step 3.5 Model Training and Prediction

[0157] The data set is divided into a training set, a validation set, and a test set according to 7:2:1, and the model of the present invention is trained based on the training set. The training degree of the model is judged according to the loss changes of the training set and the validation set. When the training loss drops to the lowest and the loss of the validation set begins to rise, it indicates that the model has been basically trained, and continued training may lead to overfitting, so the training is stopped and the model is saved. In the experiment of the present invention, when the training loss drops to 0.17 and the validation loss is 0.23, the model reaches the optimal state. After the model training is completed, the model of the present invention is tested on the test set, and an example of the test result is shown in Figure 10 .

[0158] Step 4. Lane Line Clustering and Fitting in the Waiting Area

[0159] Considering that the waiting area is extremely easy to be blocked by vehicles, multiple images are detected to complete the blocked part of the waiting area. First, all the images under the same device are preprocessed as above, and then the model is called to detect the curves of the preprocessed multiple images, and all the detected results are merged into the same set. The points in the set are clustered and fitted to obtain a curve model, and finally the accurate and specific waiting area is obtained.

[0160] Step 4.1 Curve Point Clustering

[0161] Since the curves in the scenario of the present invention are the dotted lines on both sides of the waiting area, they will not intersect in the distance, and even the two closest curves still maintain a certain distance. The PDcluster clustering algorithm based on point spacing is used to cluster the model detection results, so that the points belonging to the same curve can be clustered into one category, and the points of different curves will not be clustered into one category. Using this clustering method, different curves can be well distinguished.

[0162] The PDcluster algorithm for clustering based on point spacing is as follows:

[0163] 1) Put all pixel points into the first set A initialized to be empty;

[0164] 2) Randomly take out a pixel point from the current first set A and add it to the second set B initialized to be empty;

[0165] 3) Traverse all pixel points in the first set A, and judge whether there is a pixel point in the second set B whose distance from the currently traversed pixel point is less than the maximum clustering distance. If so, add the currently traversed pixel point to the second set B; the maximum clustering distance is the maximum distance value allowed between adjacent pixel points on a lane line in the image;

[0166] 4) Keep repeating 3) until no new pixel points are added to the second set B, and take all pixel points in the second set B as a clustering cluster, and the pixel points in this clustering cluster belong to the same waiting area lane line;

[0167] 5) Keep repeating 2) to 4) until all pixel points in the first set A are divided into clustering clusters, and the pixel points corresponding to each waiting area lane line are obtained.

[0168] In this embodiment, the clustering result in an example is as Figure 11 shown, and the pixel points corresponding to three curves are obtained by clustering, which are distinguished by different labels 1, 2, and 3 respectively.

[0169] Step 4.2 Curve fitting

[0170] After obtaining the clustering result, each category needs to be fitted to obtain the final curve result. Since there are certain detection errors in semantic segmentation, there may be abnormal points. Using polynomial fitting based on the least squares method may have deviations, and a small number of abnormal points will affect the fitted curve, resulting in the fitted result not fitting the real curve well and reducing the algorithm effect. The RANSAC algorithm is more robust than the least squares method. The RANSAC algorithm can filter out abnormal points in the sample and will not be affected by abnormal points or outliers, and can ensure that the fitted result is extremely close to the real curve. Therefore, the RANSAC algorithm is used in the present invention to fit the clustering points.

[0171] The present invention uses a cubic curve equation to fit a curve, and the curve equation formula is as follows:

[0172] y = x0 + x1x + w2x 2 + w3x 3 (3.8)

[0173] Among them, w i represents the i-th coefficient, and x and y represent the abscissa and ordinate respectively. The RANSAC algorithm performs fitting according to the formula (3.8), and the fitting result is as Figure 12 shown.

[0174] (3) Obtain the area of the waiting lane

[0175] After clustering and fitting, multiple curve equations can be obtained. Each curve equation represents a curve. According to the points in the cluster, a section of the curve that fits the dotted line of the waiting lane can be intercepted. In order to judge the spatial position relationship of each curve, a horizontal line is drawn at a vertical coordinate, and intersections can be taken for multiple curves. Each curve can obtain an intersection. According to the abscissas of these intersections, the curves can be sorted from left to right.

[0176] Since both the upper and lower sides of the waiting lane are straight lines, it is not necessary to separately connect multiple regions. By directly connecting the endpoints of the leftmost and rightmost curves, the accurate and specific area of the waiting lane can be obtained. The connection result diagram of the waiting lane area is as Figure 13 shown.

[0177] The area of the waiting lane is mainly composed of three curves. However, since it is difficult to correspond the modification of the curve to the function equation, the present invention divides the curve into segments every 10 pixels, and then saves the coordinates after the curve segmentation into a JSON file. During visualization, the curve is modified in the form of a B-spline, and some deviated coordinates on the curve are manually modified, and then the modified coordinates of each point are saved.

[0178] The present invention has a good effect on the detection of the waiting lane in a complex environment and has a high accuracy in the detection of the waiting lane. As shown in Table 1, Table 2, and Table 3, in this embodiment, the detection accuracies in two datasets TuSimple, CULane and mIoU evaluation metrics widely used for lane line evaluation can reach 96.78%, 96.89%, and 95.47% respectively.

[0179] Table 1 Comparison table of the performance of the curve detection model under the TuSimple evaluation metric

[0180]

[0181] Table 2 Comparison table of the performance of the curve detection model under the CULane evaluation metric

[0182]

[0183] Table 3 Comparison table of the performance of the curve detection model under the mIoU evaluation index

[0184]

[0185] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A detection method for the waiting area to be traveled based on semantic segmentation, which is used to detect the left-turn waiting area or the right-turn waiting area in traffic non-site law enforcement images, and is characterized in that Including: S1. Obtain a labeled training data set, where each image sample contains an image of a stop bar area captured from above by a law enforcement camera. The stop bar area lane lines in the form of curved dotted lines on both sides of the stop bar area in the image are all marked with points; the image samples in the training data set belong to different intersection scenarios, and all image samples are divided into a daytime image subset captured during the day and a nighttime image subset captured at night according to the shooting time; S2. For each image sample in the training data set, comprehensively considering the two retention principles of the fewer the number of vehicles, the higher the priority and daytime images being prior to nighttime images, and combining the grayscale value of the image and the number of vehicles in the image to filter and screen all image samples in the same intersection scenario. For each intersection scenario, image samples exceeding the threshold number are removed respectively; S3. With the goal of minimizing the loss function, use the training data set filtered by S2 to train the stop bar area detection network; The stop bar area detection network consists of an encoder, a multi-information transfer module and a decoder; In the encoder, a feature pyramid based on the ResNet50 backbone network is used as the basic feature extraction network to extract 4 feature maps of different sizes from the original input image; In the multi-information transfer module, it is necessary to iteratively perform multiple information transfer operations on each feature map output by the encoder. Each information transfer operation requires slicing the feature map in 4 directions: from top to bottom, from left to right, from right to left, and from bottom to top. The information between the slices is transferred to each other, and the step size of the information transfer is controlled to increase during the iterative information transfer operation to ensure that each slice can receive the information of the entire feature map; The decoder receives the 4 feature maps of different sizes output by the multi-information transfer module, and performs upsampling on the feature maps in order from the smallest size to the largest size and fuses them with the feature maps of larger sizes until all 4 feature maps are fused together and then upsampled to restore to the size of the original input image; The loss function is the weighted sum of the segmentation loss and the classification loss; S4. Input the image to be detected containing the stop bar area into the trained stop bar area detection network to obtain all the pixel points recognized as stop bar area lane lines in the image to be detected. Then, cluster these pixel points based on the point spacing, and the pixel points belonging to the same lane line are clustered into one category; then, perform curve fitting on each category of pixel points respectively to obtain the fitted curve segments of each stop bar area lane line, and connect the endpoints of the fitted curve segments corresponding to the lane lines on both sides of the same stop bar area to obtain the stop bar area detection result.

2. The method for detecting the waiting area based on semantic segmentation according to claim 1, wherein, In the training data set, the stop bar area lane lines in each image sample are marked with points using a marking tool, and the marked points on each stop bar area lane line need to restore the curve segment corresponding to the lane line.

3. The method for detecting the waiting area based on semantic segmentation according to claim 1, wherein The specific method of S2 is as follows: S21. Convert each image sample in the training dataset from an RGB image to a grayscale image, then calculate the grayscale mean of all pixels in each image sample, and then calculate the average of the grayscale means of all image samples in each of the daytime image subset and the nighttime image subset respectively as the average brightness of the corresponding subset; use the average of the average brightnesses of the two subsets as the brightness discrimination threshold for distinguishing day and night. S22. Use the trained object detection model to detect vehicles in each image sample in the training dataset to obtain the number of vehicles in each image sample, then calculate the average number of vehicles in all image samples in the training dataset, and finally calculate the vehicle weight of each image sample as the ratio of the number of vehicles in the image sample to the average number of vehicles multiplied by the average brightness of the daytime image subset. S23. According to the brightness discrimination threshold and the vehicle weight, calculate the quality weight of each image sample in the training dataset as 255 + λ * α * gray - β * carWeight, where gray represents the grayscale mean of all pixels in the currently calculated image sample, carWeight represents the vehicle weight corresponding to the currently calculated image sample, α and β are two weights respectively, λ is a weight value determined by the brightness discrimination threshold bound and gray. If gray ≥ bound, then λ = λ1; if gray < bound, then λ = λ2, λ1 + λ2 = 1 and λ1 > λ2. S24. For all image samples in each intersection scene in the training dataset, sort them according to their respective quality weights. If the number of image samples in an intersection scene exceeds the threshold number, then retain the image samples that meet the threshold number in descending order of quality weight. If the number of image samples in an intersection scene does not exceed the threshold number, then retain all image samples.

4. The method for detecting a waiting area based on semantic segmentation according to claim 3, wherein The weights α and β are 1 and 2 respectively, and the weight values λ1 and λ2 are 0.6 and 0.4 respectively.

5. The method for detecting a waiting area based on semantic segmentation according to claim 1, wherein In the multi - information transfer module, it is necessary to perform N information transfer operations iteratively on each feature map X output by the encoder, and in each information transfer operation, it is necessary to perform slicing on the feature map in the horizontal or vertical direction in 4 directions: from top to bottom, from left to right, from right to left, and from bottom to top respectively, and perform mutual information transfer between the slices. Among them: Perform horizontal slicing on the input feature map X in the from - bottom - to - top direction and perform vertical information transfer between the slices. The calculation formula for vertical information transfer between the slices in any n - th round of iteration is as follows: Perform vertical slicing on the input feature map X in the from - right - to - left direction and perform horizontal information transfer between the slices. The calculation formula for horizontal information transfer between the slices in any n - th round of iteration is as follows: Where: F p,l,q represents a set of convolutional kernels, where p, l, and q represent the number of input channels, the number of output channels, and the kernel width respectively; the symbol "·" is the convolution operator; f is the non-linear activation function ReLU; represents the value of the feature map X at the n-th iteration, where k, i, and j represent the indices of the channel, row (H direction), and column (W direction) respectively; represents the after information passing processing, n n represents the current iteration number, and s represents the step size of information passing in the n-th iteration. L is the width W and height H of the input feature map X during vertical and horizontal information passing respectively; Perform a vertical mirror flip on the input feature map X along the horizontal symmetry plane in the from - top - to - bottom direction and then perform horizontal slicing, and perform the same vertical information transfer as in the from - bottom - to - top direction between the slices. After horizontally mirror-flipping the input feature map X along the vertical symmetry plane in the left-to-right direction, vertical slicing is performed, and the same horizontal information transfer as in the right-to-left direction is performed between the slices.

6. The method for detecting a waiting area based on semantic segmentation according to claim 1, wherein The decoder receives 4 feature maps of different sizes output by the multiple information transfer module. In order from largest to smallest size, they are the first feature map, the second feature map, the third feature map, and the fourth feature map. First, bilinear interpolation is used to upsample the fourth feature map to make its size the same as that of the third feature map, and at the same time, the number of channels is reduced by half. Then, feature fusion is performed with the third feature map to obtain the first fusion feature map; bilinear interpolation is used again to upsample the first fusion feature map to make its size the same as that of the second feature map, and at the same time, the number of channels is reduced by half. Then, feature fusion is performed with the second feature map to obtain the second fusion feature map; bilinear interpolation is used again to upsample the second fusion feature map to make its size the same as that of the first feature map, and at the same time, the number of channels is reduced by half. Then, feature fusion is performed with the first feature map to obtain the third fusion feature map; the third fusion feature map is upsampled and restored to the size of the original input image to obtain the image to be classified. Each pixel in the image to be classified is classified to achieve semantic segmentation, thereby obtaining the lane line recognition result.

7. The method for detecting the waiting area based on semantic segmentation according to claim 1, wherein, The calculation formula of the loss function is: Loss=Loss CE +Loss BCE (3.5) Loss BCE =-αy c log(p c )-(1-α)(1-y c )log(1-p c ) (3.7) Among them, Loss BCE and Loss CE are the segmentation loss and the classification loss respectively; M represents the number of categories, c represents the category, and ω c represents the weight of the loss; y c is a vector with values of 0 or 1, indicating whether the pixel category prediction is correct or not, 1 means correct, and 0 means wrong; p c represents the probability that the predicted pixel category is c; the segmentation loss is used to distinguish the background and the annotation, α represents the proportion of the background segmentation loss, and y c represents the true value corresponding to p c .

8. The method for detecting a waiting area based on semantic segmentation according to claim 1, wherein The specific method for clustering all pixel points based on the point spacing is: S41: Put all pixel points into the first set initialized as empty; S42: Randomly take out a pixel point from the current first set and add it to the second set initialized as empty; S43: Traverse all pixel points in the first set, and judge whether there is a pixel point in the second set whose distance from the currently traversed pixel point is less than the maximum clustering distance. If so, add the currently traversed pixel point to the second set; the maximum clustering distance is the maximum distance value allowed between adjacent pixel points on a lane line in the image; S44: Keep repeating S43 until no new pixel points are added to the second set. All pixel points in the second set are used as a clustering cluster, and the pixel points in this clustering cluster belong to the same lane line in the area to be traveled; S45: Keep repeating S42 - S44 until all pixel points in the first set are divided into clustering clusters, and the pixel points corresponding to each lane line in the area to be traveled are obtained.

9. The method for detecting a waiting area based on semantic segmentation according to claim 1, wherein The curve fitting uses a cubic curve equation as the fitting equation.

10. The method for detecting a waiting area based on semantic segmentation according to claim 1, wherein, The curve fitting is implemented using the RANSAC algorithm.

Citation Information

Patent Citations

  • Lane line detection method, device and system based on semantic segmentation and storage medium

    CN112613392A

  • Intelligent vehicle lane line detection method based on ResNeSt and self-attention distillation

    CN113158768A