Marking line segmentation method based on shape template constraint
Through the subdivision methods of background reconstruction, color clustering and shape template constraints, combined with the depth Boltzmann machine to optimize the marking profile, the noise interference and occlusion problems of traffic marking detection under complex backgrounds are solved, and the detection accuracy and recall rate are improved.
Patent Information
- Application Number
- CN202510559620.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art has noise interference in traffic marking detection and segmentation under complex background conditions, it is difficult to distinguish between small target intermittent marking lines and occlusions, and lacks effective data marking methods, resulting in insufficient detection accuracy.
The reticle segmentation method based on shape template constraints is adopted, including background reconstruction, color clustering coarse segmentation and subdivision of shape template constraints, combined with a depth Boltzmann machine (DBM) for contour optimization, and precise segmentation is performed through motion abnormality detection, color clustering and shape template constraints.
It improves the accuracy and recall of traffic marking detection, reduces the impact of pixel changes in background areas on modeling, and ensures that markings can still be effectively segmented under complex conditions to meet the requirements of subsequent tasks.
Smart Images

Figure CN120471949A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection and recognition and object segmentation in road scenes, and in particular to a marking line segmentation method based on shape template constraints. Background Art
[0002] In the vehicle violation detection task, road traffic marking detection and segmentation is also a basic task with pre-condition semantics. Although the research on target detection and segmentation in road scenes has achieved many very exciting results, directly transferring these methods to the task of traffic marking detection and segmentation under complex background conditions still has certain limitations. The main limitations include: (1) the high similarity of traffic marking appearance. (2) the appearance feature information of traffic markings is easily lost.
[0003] Semantic segmentation methods tailored to traffic scenarios have been proposed to address these issues. Existing solutions, such as deep network models like EL-GAN and H-Net, attempt to extract more comprehensive object features and perform semantic and instance segmentation. While these methods can help the model generate more realistic textures and details, they fail to distinguish between foreground and background objects, causing foreground moving objects to become noise interference in traffic marking segmentation, making them difficult to directly transfer to existing traffic marking segmentation tasks.
[0004] In the existing technology, a spatial convolutional CNN (SCNN) network structure is proposed for continuous linear structure targets such as lanes and telephone poles. This structure realizes message transmission between pixels across rows and columns through layer-by-layer convolution in the feature map, fully exploiting the spatial relationship between pixels in rows and columns. Even under weak coherence conditions such as target occlusion and missing, it can still accurately locate and infer the filling of occluded parts. Although this method takes into account the guidance of shape priors on target retrieval, it ignores the influence of pixel-level semantics, resulting in the loss of pixel continuity information and ineffectiveness in distinguishing small targets with discontinuous markings and occlusions.
[0005] In addition, existing traffic scene datasets lack data samples with labeled traffic line key points, and manual labeling methods often increase the complexity of system implementation. Summary of the Invention
[0006] The technical solution of the present invention addresses the technical problem that the existing technical solutions are too single, and provides a solution that is significantly different from the existing technology. It mainly provides a road marking segmentation method based on shape template constraints to solve the technical problems raised in the above background technology, such as the noise interference of foreground moving targets on traffic marking segmentation and the inability to effectively distinguish between small target intermittent markings and occlusions.
[0007] The technical solution adopted by the present invention to solve the above technical problems is:
[0008] A method for marking line segmentation based on shape template constraints includes the following steps:
[0009] S1. Reconstruct the background by using a background reconstruction method based on motion anomaly factors, remove the moving foreground object, and extract a static background reconstructed image from the video set;
[0010] S2, performing image segmentation on the static background reconstructed image in step S1 by using a color clustering coarse segmentation method;
[0011] S3, performing fine segmentation on the coarse segmentation result in step S2 by using a fine segmentation method constrained by a shape template;
[0012] S4. Optimize the outline of the fine segmentation result in step S3 by using an association-based outline optimization method.
[0013] Specifically, step S1 includes:
[0014] S1-1. The random process of video frame pixel values changing over time is described by a linear weighted combination of k Gaussian distribution functions as follows:
[0015]
[0016] where w t,i is the weight of the i-th Gaussian distribution function in the mixed Gaussian model at time t, η is the probability density function of the multivariate Gaussian distribution, where the parameter c is the covariance matrix of each Gaussian component, E is the identity matrix, μ is the mean, and δ is the standard deviation;
[0017] S1-2, the pixel value x of pixel P t Perform threshold matching with the mean of each Gaussian component in turn and continuously iteratively update the parameters and weights of the Gaussian model; the threshold matching condition is:
[0018] |x t -μ t,i |<2.5σ t,i
[0019] When the threshold matching condition of the i-th Gaussian component is met, the current pixel P is determined to be a background pixel.
[0020] Furthermore, the Gaussian component parameter update strategy is:
[0021] w t,i =(1-λ)·w t,i +λ
[0022] μ t,i =(1-ρ)·μ t,i +ρ·xt
[0023]
[0024] λ is the learning rate, ρ is the parameter update factor, It is a movement abnormality factor.
[0025] Furthermore, the motor abnormality factor The update strategy is:
[0026]
[0027] in, is the pixel value of the pixel p(i, j) with coordinates (i, j) at time t, M and N are the video image resolution sizes respectively;
[0028] After the parameter update is completed, the first N components with larger weight w and smaller standard deviation σ are taken as the background distribution description B, which is expressed as:
[0029]
[0030] Among them, w i is the proportion of background pixels of each component, which is greater than the threshold T.
[0031] Specifically, in step S2, the image is first reconstructed based on the static background, the color components of the color background image in each color space are extracted, and pixel-level feature vectors are formed according to the color components; then, clustering is performed using the K-means value algorithm based on Euclidean distance in units of pixels; finally, image segmentation is performed according to the clustering results.
[0032] Furthermore, the K-means value algorithm is:
[0033]
[0034] Among them, c i Represents pixel x i The cluster to which it belongs, μ c Represents the cluster center of the cluster, I is the total number of pixels, and K is the total number of cluster centers.
[0035] Furthermore, for the color cluster images covering most pixels, the SIFT local feature descriptor is extracted to obtain strong robust image features, and the bag-of-words Bow model is used to complete the ROI area retrieval of the color cluster images.
[0036] Specifically, in step S3, shape template modeling is performed and shape loss is constructed to constrain the region classification task; and the shape modeling task is implemented by training using a deep Boltzmann machine (DBM).
[0037] Furthermore, based on the pre-extracted color and texture features, color and texture features are extracted again on the image region of interest, and a two-classification feedforward neural network is used for secondary classification of the region of interest. The classification network structure of the secondary classification is expressed as follows:
[0038]
[0039] Where W0, W1, and W2 are the linear transformation matrices to be learned, and b1 and b2 are the corresponding bias values; X i is the color and texture feature of the current region, and I is the scene marking mask matrix. By introducing the scene marking mask matrix, the classification model will only output the focus marking area that matches the recognized scene; is the classification output of the model;
[0040] Based on the principle of minimizing the DBM energy function, the shape constraint term is added to obtain the loss expression of the classification network:
[0041]
[0042] Where θ={W 1 ,W 2 ,a 1 ,a 2 ,b} are the network parameters obtained by DBM model training, v is the visible unit input of DBM model; τ is the weight of the shape constraint regularization term.
[0043] Furthermore, in step S4, after the Gibbs iterative sampling converges, a shape that approximately obeys the edge probability distribution of the marking template is obtained.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] (1) In the static background reconstruction step, the present invention proposes a mixed Gaussian model static background reconstruction method based on motion anomaly detection. It is based on the assumption that the center motion coefficient of the road monitoring video lens is the largest, makes full use of the overall pixel distribution structure and the pixel difference between the previous and next frames, and uses the moving target state mixed Gaussian model to measure the update strategy of the model learning rate, thereby reducing the influence of the complex changes of the background area pixels on the background modeling, ensuring accurate background modeling when the foreground target moves slowly; and removes the moving foreground target, effectively solving the problem of moving target occlusion in a single video image frame.
[0046] (2) In the image segmentation step, the present invention first performs color clustering coarse segmentation and then performs fine segmentation based on shape template constraints. In the color clustering coarse segmentation, SIFT local feature descriptors are extracted to obtain strong robust image features, and the ROI area retrieval of the color clustering image is completed by combining the Bag-of-Words model with a sliding window in a progressive manner, providing a more refined semantic precondition to the target segmentation algorithm and eliminating interference for subsequent fine segmentation; in the fine segmentation step, the matching degree of the training data reinforcement template is described by shape expression parameters, and the form of shape regularization is added to the classification loss function to avoid the interference of scattered pixels, ensure the smooth edge of the degraded marking segmentation, and still achieve good target detection in the case of partial pixel loss due to occlusion, wear, and small targets, thereby improving the accuracy of target area segmentation.
[0047] (3) In the contour optimization step based on association, after the DBM model training is completed in the previous step to determine its parameters, the shape that approximately obeys the edge probability distribution of the marking template is obtained after Gibbs iterative sampling convergence, and the marking contour of the segmented area is optimized, which solves the problem that the morphological opening and closing operations have no directionality in the operation direction and are prone to shape loss, thereby ensuring the reliability of the image segmentation results.
[0048] In summary, the present invention proposes a progressive road marking segmentation method consisting of background reconstruction, color clustering coarse segmentation, fine segmentation and contour optimization tasks, which improves the accuracy of target area segmentation and reduces the impact of complex changes in background area pixels on background modeling. Experiments show that the method of the present invention has achieved good results in both precision and recall performance comparisons, and its overall performance in road traffic marking detection tasks meets the requirements of subsequent tasks.
[0049] The present invention will be explained in detail below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flow chart of the steps in the present invention;
[0051] Figure 2 Reconstructing an image of a static background in an embodiment of the present invention;
[0052] Figure 3 This is an image segmentation diagram based on color clustering in an embodiment of the present invention; (a)-(f) are the clustering effects of each partition (the number of clusters is 6) in the color texture feature space.
[0053] Figure 4 This is a result diagram of extracting a region of interest from a background image in an embodiment of the present invention;
[0054] Figure 5 A schematic diagram of the DBM model structure for shape modeling in an embodiment of the present invention;
[0055] Figure 6 : Detection and segmentation result diagram based on shape constraints in an embodiment of the present invention; (a) is a zebra crossing; (b) is a dotted line; (c) is a double yellow line;
[0056] Figure 7 Schematic diagram of reconstruction of the double yellow line information loss area in an embodiment of the present invention; (a) shows the target details; (b) shows the morphological operation; (c) shows the sampling reconstruction; and (d) shows the final shape of the double yellow line.
[0057] Figure 8 A line graph of the sum of squared errors under the number of clusters in an embodiment of the present invention;
[0058] Figure 9 This is a cross-validation graph of prediction accuracy and computational cost under different numbers of words in an embodiment of the present invention;
[0059] Figure 10 This is a mean square error verification diagram for different hidden layer nodes in an embodiment of the present invention;
[0060] Figure 11 : Comparison diagram of background reconstruction effects of road monitoring videos in the embodiment; wherein (a) is a background reconstruction diagram of a traditional method, and (b) is a background reconstruction diagram of the method provided by the present invention;
[0061] Figure 12 Schematic diagram of road traffic line segmentation using three different methods in the embodiment;
[0062] Figure 13 1 and 2 are the line marking segmentation and reconstruction effects under different road scenes in the embodiment; (a) is the original image; (b) is the result image after segmentation of the present invention. DETAILED DESCRIPTION
[0063] To facilitate understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in different forms and is not limited to the embodiments described in the text. On the contrary, these embodiments are provided to make the content disclosed in the present invention more thorough and comprehensive.
[0064] Please refer to the attached Figure 1 A line segmentation method based on shape template constraints fully utilizes color and texture information under shape semantic constraints based on traditional image segmentation algorithms, and sequentially completes the steps of static background reconstruction, color clustering coarse segmentation, target fine segmentation, and associative contour optimization. Specifically:
[0065] A road marking mask matrix is pre-generated using scene recognition methods. Each row vector represents the type of road marking detected for each road scenario. This matrix is a crucial component of the model training input for the fine segmentation phase. Table 1 shows the correspondence between road scenes and detected road markings. This method utilizes defined road marking templates for supervised learning and training, using samples cut from various road markings for each model's training. These templates are uniformly sized at 32×32.
[0066] Table 1 Correspondence between road scenes and traffic markings to be segmented
[0067]
[0068] 1. Traffic Marking Target Feature Extraction
[0069] Road markings like double yellow lines and zebra crossings are susceptible to complex conditions such as shadows, water accumulation, and contamination, resulting in significant loss of target edge information. Direct edge extraction using operators like Candy and line detection using the Hough transform can easily produce irregular, discontinuous, and noisy results. Therefore, for traffic markings with rich color information and regular, continuous shapes, this paper uses color and texture features and the SIFT (Scale Invariant Feature Transform) local feature descriptor for image segmentation.
[0070] 1.1 Color and texture feature extraction
[0071] Image color texture reflects the spatial relationship between different components of the image's color space. To accurately extract color and texture information from linear objects in traffic scenes, a multi-channel color co-occurrence matrix was chosen to integrate these two features. The color co-occurrence matrix is calculated by computing the joint probability density of pairs of pixels at a fixed distance in the horizontal, vertical, and diagonal (bidirectional) scanning directions. This matrix reflects the distribution characteristics of the quantized values of each component across each pixel in different color spaces. The matrix features a color histogram structure.
[0072] Assuming that there are color components α and β in a certain color space, and the coordinates of the image primitive pixel are (x, y), the calculation formula of the co-occurrence matrix CCM is as follows:
[0073]
[0074] Among them, Δx and Δy represent the increments of the primitive pixel coordinates on the x-axis and y-axis, and the co-occurrence matrix CCM i,jIt reflects the number of times the two component values of the pixel at a distance (Δx, Δy) from the primitive pixel are i and j, respectively. In the direction of the matrix sliding window movement, the CCM is still calculated in the four directions of 0°, 45°, 90°, and 135° and the average is taken to obtain the final color co-occurrence matrix.
[0075] Based on the statistical results of CCM, we further calculate the five classic non-correlation eigenvectors derived from it, namely inverse moment (Idm), contrast (Con), energy (Asm), correlation (Corr), and entropy (Ent). Inverse moment (Idm) expresses the degree of local texture variation, contrast (Con) expresses the depth of texture grooves, energy (Asm) expresses the uniformity of color component value distribution, and correlation (Corr) expresses the consistency of texture in the row and column directions. The calculation method is as follows:
[0076]
[0077] where μ i 、μ j and σ i , σ j Denote the mean and variance in the row and column directions, respectively, and D is the quantization level of each color component. To reduce computational cost, the quantization level can be defined as 16 or 32, resulting in co-occurrence matrices with dimensions of 16×16 and 32×32, respectively.
[0078] At the same time, in the selection of color space components, according to the illumination changes of road monitoring video images and the color saliency characteristics of traffic markings, seven color components, R, G, B in RGB space, a*, b* in Lab space, and H and S in HSI space, are selected to generate color texture features for traffic marking segmentation, which is defined as F = [F rr ,F rg ,F gg ,F rb ,F gb ,F bb ,F a*a* ,F a*b* ,F b*b* ,F hh ,F hs ,F ss Based on the feature vector category and color component combination dimension, a 60-dimensional color texture feature is finally obtained.
[0079] 1.2 SIFT local feature description
[0080] As a classic local feature description operator, SIFT features have good rotation invariance and scale invariance. They are robust and stable to interference such as perspective changes, illumination, and affine transformations. Therefore, they are very suitable for extracting features from road traffic monitoring images with complex shooting angles and illumination. The algorithm steps are as follows:
[0081] (1) Scale space generation: To establish the image scale space, it is first necessary to generate a two-dimensional Gaussian function G with varying scales, and then generate the scale space L by performing a convolution operation with the image I.
[0082]
[0083] L(x,y,σ)=G(x,y,σ)*I(x,y) (8)
[0084] In the above formula, (x, y) is the pixel coordinate and σ is the scale space factor.
[0085] (2) Scale space extreme value detection: The Gaussian difference pyramid function Q is established to detect the location of the feature key points. The calculation formula is as follows:
[0086] Q(x,y,σ)=L(x,y,kσ)-L(x,y,σ) (9)
[0087] In the above formula, k is the scale factor of the scale space. By comparing with the 26 pixels in the pixel neighborhood, it is determined whether the Gaussian difference factor can obtain an extreme value and then determine whether it is a key point.
[0088] (3) Key point screening: Use the Taylor function expansion of the Gaussian difference function D obtained in (2) to perform curve fitting, delete pseudo key points and obtain the final stable feature points.
[0089] (4) Feature point direction assignment: The gradient direction statistical histogram is established based on the local gradient direction of the pixels in the neighborhood of the feature point, and the direction corresponding to the histogram peak is used as the feature point direction;
[0090] (5) Feature point description: A descriptor vector is created for each feature point. The present invention divides the feature point neighborhood into 4×4 subregions. Each subregion performs gradient calculations in eight directions: horizontal, vertical, and diagonal. This results in a 128-dimensional descriptor vector that is then normalized.
[0091] 2. The specific steps of the line segmentation method constrained by shape template are as follows:
[0092] 2.1 Background reconstruction based on motion anomaly factors
[0093] The video background reconstruction method removes the moving foreground target and can convert the video analysis problem into a static image segmentation problem. Situations such as road congestion and waiting at traffic lights can cause the foreground target in the video to move slowly. In order to achieve accurate background modeling when the foreground target moves slowly, the present invention designs a mixed Gaussian model static background reconstruction method based on motion anomaly detection. It is based on the assumption that the center motion coefficient of the road monitoring video lens is the largest, and makes full use of the overall pixel distribution structure and the pixel difference between the previous and next frames. The pixel values in the continuous video frame image have Gaussian distribution characteristics on the time axis, and the pixel values are around the mean μ with a variance σ 2 Distributed within a distance range. Assume that the pixel value of pixel point p at time t in the scene coordinate system is x t , the random process of the video frame pixel value changing over time can be described by a linear weighted combination of k Gaussian distribution functions as follows:
[0094]
[0095] where w t,i is the weight of the i-th Gaussian distribution function in the mixed Gaussian model at time t, η is the probability density function of the multivariate Gaussian distribution, the parameter c is the covariance matrix of each Gaussian component, and E is the identity matrix. In the formula, the initial value of the parameters is set to the pixel grayscale value of the first video frame, the initial value of the standard deviation δ is 15, and the number of Gaussian components and the initial value of the weight are set to 1 and 0.001 respectively. t The threshold is matched with the mean of each Gaussian component in turn and the parameters and weights of the Gaussian model are updated iteratively.
[0096] |x t -μ t,i |<2.5σ t,i (13)
[0097] If the threshold matching condition of the i-th Gaussian component is met (i.e., formula (13)), the current pixel P is determined to be a background pixel. The Gaussian component parameter update strategy is:
[0098] w t,i =(1-λ)·w t,i +λ (14)
[0099] μ t,i =(1-ρ)·μ t,i +ρ·x t (15)
[0100]
[0101] λ is the learning rate, and ρ is the parameter update factor. Assuming that in a road surveillance video scene, the moving target close to the center of the lens needs more attention, the present invention evaluates the driving state of the moving target in the surveillance video and adds a motion abnormality factor when there is vehicle congestion or slow-moving pedestrians. Its update strategy is:
[0102]
[0103] in, is the pixel value of pixel p(i, j) at coordinates (i, j) at time t, where M and N are the video image resolution dimensions. The principle is that when a moving foreground pixel is close to the center of the image coordinate system and the overall distribution area of the differentially changing pixels between the previous and next frames is small, the parameter update rate should be reduced to prevent slowly changing pixels from being misclassified as background. After the parameter update is completed, the first n components with the largest weights w and the smallest standard deviations σ are selected as the background distribution description B. Furthermore, the proportion of background pixels in each component should be greater than a threshold T, which is set to 0.75.
[0104]
[0105] The static background extracted from the video set samples is used to reconstruct the image. Figure 2 shown.
[0106] 2.2 Color Clustering Rough Segmentation
[0107] The video background reconstruction method effectively solves the problem of moving object occlusion in a single video image frame. The present invention extracts the color components of the video background image in various color spaces and forms a pixel-level feature vector based on the seven color components: R, G, B in the RGB space, a*, b* in the Lab space, and H and S in the HSI space. Under 1080p surveillance video conditions, the color space feature vector dimension is 1920×1080×7. The present invention uses the K-means value algorithm based on Euclidean distance to perform clustering based on pixel points:
[0108]
[0109] Among them, c i Represents pixel x i The cluster to which it belongs, μ c Represents the cluster center of the cluster, I is the total number of pixels, and K is the total number of cluster centers (clusters).
[0110] According to the calculation results, the feature point x i Assign cluster c i And recalculate the cluster center μ c ,After multiple iterative clustering converges, image segmentation can be ,implemented according to the clustering results.
[0111] Through color clustering, the goal is to divide the target traffic marking pixels into 1-2 cluster spaces as much as possible. Choosing an appropriate K value is a key issue in unsupervised K-means color clustering. This paper uses the "elbow method" to select the number of clusters in K-means clustering and verifies this in subsequent experimental implementations. Figure 3 The segmentation results when the number of clusters is 6 are shown.
[0112] from Figure 3 As can be seen from the image, the double yellow line in the center of the road has been segmented into one category, as have the majority of pixels in the main road's dotted lines and zebra crossing areas, meeting the requirements for a Region of Interest (ROI) search. In complex road traffic scenes, after color clustering segmentation, a large number of non-target pixels remain in the color cluster image in addition to the double yellow line targets. Because the SIFT descriptor primarily describes spatially local features and does not rely on the spatial relationships between local features, it has a low missed detection rate for targets with information loss in complex backgrounds and can effectively eliminate outlier pixels from non-marked line targets.
[0113] In order to provide more refined semantic preconditions for the target segmentation algorithm, the present invention is designed to extract SIFT local feature descriptors to obtain strong robust image features for color cluster images covering most pixels, and complete the ROI area retrieval of color cluster images through the Bag-of-Words model. In view of the regularity of marking features such as double yellow lines, the present invention adopts a sliding window strategy in SIFT feature extraction and search. When the window slides over the target area containing segmented pixels, the SIFT local feature description operator of the target area is extracted and a dictionary is constructed. The index of visual words to images is established based on the word frequency histogram, and whether the sliding window area is an ROI area is determined based on the index result. Figure 4 shown.
[0114] Specifically, the Bow model requires an appropriate number of words and min-max normalization of the word frequency histogram. For a reconstructed background image of 1920×1080 pixels, the ROI sliding operation uses a window size of 128×64 with a step size of 64 pixels. The window traverses the entire image from top to bottom and left to right, starting from the upper left corner.
[0115] In order to improve the computational efficiency of the sliding window retrieval method, SIFT features will only be calculated when the number of pixels in the window exceeds a certain threshold. The present invention sets the threshold to the maximum side length of the sliding window size, that is, 128 pixels. During the training phase, the grayscale images of traffic marking templates of various forms will provide training data set samples for Bow model learning. By using the ROI area obtained by setting a sliding window classification of a reasonable size, a large number of pixels of distant non-road targets can be effectively removed, and targets of similar shapes but not traffic markings can be excluded, and associative reconstruction under natural discontinuities can be avoided. The coarse segmentation area obtained based on the present invention eliminates interference for the next step of fine segmentation of marking detection. However, it was also found that due to the influence of light, weathering and wear, some pixels in the target area after color clustering segmentation are in a scattered state.
[0116] 2.3 Fine Segmentation with Shape Template Constraints
[0117] Although the coarse segmentation results are obtained after matching the SIFT+Bow model, due to the complex background conditions filled with noise targets with the same local features, there are still some misclassifications of the areas of interest in the coarse segmentation results. Compared with local invariant feature points, traffic marking targets contain more color texture and regional shape information. Although most road markings have straight line characteristics in local areas, the Hough transform algorithm can be used to quickly complete the detection task under certain conditions. However, not all traffic markings have only simple shape features, such as zebra crossings, guide lines, no-parking lines, etc. At the same time, due to the problem of the shooting angle of the monitoring equipment, some regular shapes will have certain geometric deformations. In order to expand the versatility of the proposed algorithm, unlike the traditional learning and training method that only uses color texture, the present invention constrains the regional classification task by modeling the shape template and constructing the shape loss, so as to achieve better target detection when some pixels are lost due to occlusion, wear, and small targets. The present invention uses the deep Boltzmann machine DBM for training to realize the shape modeling task, such as Figure 5 shown.
[0118] By training the DBM network composed of two stacked Boltzmann machines RBM, its model parameters can be expressed as θ = {W 1 ,W 2 ,a 1 ,a 2 ,b}, where W1 and W2 are the visible layer v and the hidden layer h1, hidden layer h 1 With h 2 The connection weight between b and a 1 、a 2 They are the visible layer v and the hidden layer h respectively 1 , hidden layer h 2The training process is carried out in accordance with the proposed layer-by-layer pre-training combined with fine-tuning. The training sample is a binary image of the marking template with a pixel size of 32×32. Before formal training, the data samples need to be translated, rotated, and scaled for data enhancement operations. The final parameter expression of each training sample is ω = {s (i) ,r (i) ,z (i)}. Among them, s=0,1,…,16 represents the translation mode of the original i-th template in the horizontal, vertical and diagonal directions of 4, 8, 12 and 16 pixels respectively, r=0,1,…,7 represents the clockwise rotation mode in units of 45 degrees, and z=0,1,…,6 represents the scaling mode of enlarging or reducing the size by 2 times, 3 times and 4 times. The parameter value of 0 means that the current parameter state of the original template remains unchanged. Assume that the binary pixel matrix of the local area to be retrieved is U, and the pixel matrix of each sample template after data enhancement is U ω , the shape parameter expression corresponding to the input v of the local area to be retrieved can be obtained through the following shape matching.
[0119]
[0120] Here, edis is a similarity comparison function based on the Euclidean distance metric. The resulting shape expression parameter ω can provide data priors for image processing applications such as region merging. Based on the color and texture features described in 1.1 above, this paper extracts color and texture features again using a 32×32 sliding window size on the region of interest. A binary feedforward neural network is designed for secondary classification of the region of interest, completing fine segmentation for line marking detection. The classification network structure is represented as follows:
[0121]
[0122] Where W0, W1, and W2 are the linear transformation matrices to be learned, and b1 and b2 are the corresponding bias values. i is the color and texture feature of the current region, and I is the scene marking mask matrix. By introducing the scene marking mask matrix, the classification model will only output the focus marking area that matches the recognized scene. is the classification output of the model.
[0123] Based on the principle of minimizing the DBM energy function, the shape constraint term is added to obtain the loss expression of the classification network:
[0124]
[0125] Where θ={W 1 ,W 2 ,a 1 ,a 2,b} are the network parameters obtained from DBM model training, and v is the visible unit input of the DBM model. τ is the weight of the shape constraint regularization term. After learning and training the color and texture features based on shape template constraints, the overfitting problem of the classification model caused by partial pixel scattering can be effectively eliminated. Figure 6 It can be seen from the figure that in the case of local truncation due to wear and loss of small target information, the algorithm proposed in this invention can still ensure the segmentation quality of traffic markings such as double yellow lines, zebra crossings, and solid white lines.
[0126] 2.4 Contour Optimization Based on Association
[0127] After color clustering segmentation based on shape constraints, the impact of local wear on target detection can be eliminated. However, the segmentation result still retains the scattered distribution of actual pixels after target wear, which is different from the intuitive perception of traffic markings by vehicle drivers. Although morphological opening and closing operations can be simply used to close scattered pixels and expand connected areas, it is found in practice that such a filling process is subject to certain uncertainties. Since morphological opening and closing operations are both performed in pixel units, there is no directionality in the operation direction, which can easily lead to shape loss. The present invention optimizes the contour of the markings in the segmented area by completing the training of DBM.
[0128] After completing the DBM model training and determining its parameters θ, the shape that approximately obeys the edge probability distribution of the marking template can be obtained after the Gibbs iterative sampling convergence. Assume that the i-th sample template vector is Sampling can be performed using known vector components and the vector components in the previous sample. The specific process is as follows:
[0129]
[0130] After multiple iterations of convergence, the above sampling process can complete the probability distribution In the contour shape modeling process, the number of sampling iterations in this embodiment is set to 90 rounds.
[0131] Figure 7 (a), (b), and (c) take the missing part of the double yellow line as an example, showing the comparison effects of the original image, simple opening and closing operations, and sampling reconstruction respectively. (d) shows the final effect of line segmentation in a traffic scene.
[0132] 3. Experimental results and discussion
[0133] 1. Test Dataset
[0134] The test samples used in the experiment of this embodiment also come from the public security video network, with a resolution of 1080p (1980×1020 pixels). The dataset is a subset of the road scene video set in the previous chapter. Since the scenes covered by road surveillance videos are different, urban roads and intersections contain all the markings to be identified. Therefore, this embodiment selects daytime surveillance video samples from these two types of road scenes as test samples. According to the morphology of traffic markings, the experiment classifies traffic markings into linear targets (double yellow lines, dashed lines, single solid lines, stop lines) and regional targets (zebra crossings, no-parking lines, guide lines), and counts the number of marking templates used for experimental training, as shown in Table 2.
[0135] Table 2 Training samples of different traffic markings
[0136]
[0137] Each video sample contained at least one of the aforementioned traffic markings in its reconstructed background. During the preprocessing phase, the videos were uniformly trimmed to 30 seconds to form the experimental video samples used in this experiment. Finally, the actual pixel regions of each traffic marking were annotated using manual association.
[0138] 2. Parameter setting
[0139] (1) Number of K-means color cluster centers
[0140] The use of unsupervised K-means color clustering achieves a rough segmentation of the video background, providing a basis for subsequent tasks. Among them, the size of the color cluster K value determines the actual effect of the rough segmentation. This embodiment will select the appropriate K value based on the elbow rule and the calculation result of the sum of squared errors (SSE) between the particles of each class and the sample points within the class. Multiple K-means models are trained under the conditions of different numbers of color clusters, and the SSE results obtained on the random data set of 1000 pixels are as follows: Figure 8 shown.
[0141] The SSE value reflects the degree of class distortion after K-means clustering. It can be seen that increasing the cluster K value effectively improves the average distortion, making the sample set within the cluster more dense. When K = 6, the SSE value shows a significant plateau, which can be considered a critical point. Therefore, the number of K-means color cluster centers in the algorithm of this invention is set to 6.
[0142] (2) Number of words in the BOW model
[0143] The Bow model based on SIFT features completes the classification of the ROI area of interest in the coarse segmentation and eliminates the outlier pixels of some non-marked targets. In the process of setting the parameters of the Bow model, it is necessary to reasonably set the number of visual words according to the prediction accuracy and computational cost. To this end, the present invention has carried out relevant cross-validation experiments on the classification accuracy, recall rate and computational time of the ROI sliding window in the range of 2 to 8 words. The verification results are as follows Figure 9 shown.
[0144] Increasing the number of visual words can improve the recognition accuracy and recall rate of ROI region classification, but it also significantly increases the computational cost. Based on the above cross-validation results, and balancing prediction accuracy and computational cost, the present invention sets the number of visual words in the Bow model in the algorithm to 5.
[0145] (3) DBM model training parameters
[0146] This paper uses a DBM model to learn the pixel distribution of the marking template and establish prior knowledge of the target's shape in the search area. Because the pixel size of the marking template training samples is fixed, the number of visible units in the DBM model training parameters is set to 1024. However, the number of hidden unit nodes may have a certain impact on the performance of the DBM-reconstructed data. To select the optimal parameter settings, the paper verifies the mean squared error of the DBM-reconstructed image. Figure 10 The mean square error (MSE) obtained with different numbers of network nodes is shown when other training parameters are the same.
[0147] The DBM network with hidden layer 512-256 in the figure converges slowly, and the minimum mean square error is achieved with hidden layer 1024-512. However, when the learning rate is fixed, the training jitter occurs after 40 iterations, while the hidden layer 512-512 can achieve an ideal MSE score in a relatively stable convergence process. Based on the above verification conclusions, the present invention sets the hidden layer h 1 , hidden layer h 2 The number of units is set to 512, the number of training iterations is set to 50, and the learning rate is set to 0.005.
[0148] 3. Algorithm Comparative Analysis
[0149] (1) Evaluation method
[0150] To verify the effectiveness of the proposed algorithm in identifying and segmenting traffic markings in real-world scenarios, we used precision and recall to evaluate the algorithm's performance. We also calculated the mean average precision (mAP), mean average recall (mAR), and harmonic mean (F1-Measure) to measure the difference in segmentation performance between the proposed method and the baseline method. The specific formulas include:
[0151]
[0152] Where G0 and G1 represent the pixel sets labeled as background and road markings in the binary image of the test set samples, respectively; T0 and T1 represent the pixel sets of background and road markings detected by the algorithm, respectively; and N is the total number of road markings of all types in the test data samples. Precision indicates the proportion of pixels predicted as traffic markings that are actually target road markings. Recall indicates the proportion of pixels within the labeled road marking area that are correctly segmented.
[0153] (2) Comparison results
[0154] This example migrated some lane semantic segmentation algorithms to road marking segmentation and compared their performance with the algorithm presented in this paper. The benchmark algorithms included: CANNY-SURF, PBIM-MSAC, EL-GAN, H-Net, UNet-ConvLSTM, and SCNN. The experiment used road markings in urban highway and intersection scenarios as test objects, and the evaluation results shown in Table 3 were obtained based on the evaluation criteria.
[0155] Table 3 Comparison of segmentation performance between the method of the present invention and related methods
[0156]
[0157] Through performance comparison, the algorithm proposed in the present invention achieved the best experimental results in the average segmentation precision mAP and harmonic mean F1 under the test data set. Part of the performance improvement is due to the shape association generated by the algorithm proposed in the present invention under the shape semantic constraint. Although the SCNN method also adopts the method of continuous region association and achieves the best results in the experimental test results of the segmentation recall rate mAR, it is easy to produce excessive association when there is no shape constraint during regional reconstruction, resulting in the precision and harmonic mean being slightly lower than the method of the present invention. In addition, due to the strategy of shape semantic constraint, the segmentation performance of the algorithm of the present invention for linear or regional road marking targets does not fluctuate greatly. However, the CANNY-SURF method and the PBIM+MSAC method are unable to better identify and process regional marking targets, resulting in poor overall segmentation performance.
[0158] (3) Comparison of segmentation effects
[0159] In the algorithm for extracting static background from videos, in the case of vehicle congestion and slow-moving pedestrians, the present invention improves the accuracy of video background pixel classification by evaluating the driving status of moving targets.
[0160] Figure 11 (a) and (b) show the background images of road monitoring videos obtained by the traditional background reconstruction method based on mixed Gaussian model and the method of the present invention respectively. Figure 11 As can be seen in (a), in the traditional method of reconstructing the background image, the background model cannot overcome the complex background conditions of slow moving targets, resulting in some moving target pixels being misclassified as background. Due to the learning rate update strategy based on the motion anomaly factor, it can effectively ensure that the visible double yellow line area can be displayed in the reconstructed background image ( Figure 11 (b) Box marks the area).
[0161] In order to verify the strong robustness of the algorithm of the present invention under the condition of line marking information loss, the present invention shows the experimental results of two types of benchmark algorithms UNet-ConvLSTM and SCNN with similar overall performance comparison scores and conducts subjective comparison, such as Figure 12 shown.
[0162] Depend on Figure 12 As can be seen, the UNet-ConvLSTM method transforms the object detection problem into an instance segmentation problem through end-to-end training. However, it cannot effectively associate objects with some information loss in image scenes. The SCNN method also suffers from insufficient association ability in some connected areas and excessive association in some unconnected areas, resulting in incorrect segmentation of structurally similar marked objects such as dotted lines and missing solid lines.
[0163] In addition, the segmentation effects of various types of road markings defined in the experiment were fully tested. Figure 13 The results of road marking segmentation and reconstruction for intersection and urban road scenarios are demonstrated in the following examples. These experimental results demonstrate that the proposed algorithm can effectively combine color and texture features to detect various types of traffic marking areas, avoid interference from noisy pixels, and produce associative segmentation results based on shape priors. The experimental results are more consistent with subjective judgment.
[0164] The above description of the present invention is exemplified in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as such non-substantial improvements are made using the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the scope of protection of the present invention.
Claims
1. A marking line segmentation method based on shape template constraints, characterized by: The steps include: S1. Reconstruct the background by using a background reconstruction method based on motion anomaly factors, remove the moving foreground object, and extract a static background reconstructed image from the video set; S2, performing image segmentation on the static background reconstructed image in step S1 by using a color clustering coarse segmentation method; S3, performing fine segmentation on the coarse segmentation result in step S2 by using a fine segmentation method constrained by a shape template; S4. Optimize the outline of the fine segmentation result in step S3 by using an association-based outline optimization method.
2. The method for marking line segmentation based on shape template constraint according to claim 1, characterized in that: Step S1 includes: S1-1. The random process of video frame pixel values changing over time is described by a linear weighted combination of k Gaussian distribution functions as follows: where w t,i is the weight of the i-th Gaussian distribution function in the mixed Gaussian model at time t, η is the probability density function of the multivariate Gaussian distribution, where the parameter c is the covariance matrix of each Gaussian component, E is the identity matrix, μ is the mean, and δ is the standard deviation; S1-2, the pixel value x of pixel P t Perform threshold matching with the mean of each Gaussian component in turn and continuously iteratively update the parameters and weights of the Gaussian model; the threshold matching condition is: |x t -m t,i |<2.5σ t,i When the threshold matching condition of the i-th Gaussian component is met, the current pixel P is determined to be a background pixel.
3. The method for marking line segmentation based on shape template constraint according to claim 2, characterized in that: The Gaussian component parameter update strategy is: w t,i =(1-λ)·w t,i +λ m t,i =(1-ρ)·μ t,i +p·x t λ is the learning rate, ρ is the parameter update factor, It is a movement abnormality factor.
4. The method for marking line segmentation based on shape template constraint according to claim 3, characterized in that: Movement abnormality factor The update strategy is: in, is the pixel value of the pixel p(i, j) with coordinates (i, j) at time t, M and N are the video image resolution sizes respectively; After the parameter update is completed, the first N components with larger weight w and smaller standard deviation σ are taken as the background distribution description B, which is expressed as: Among them, w i is the proportion of background pixels of each component, which is greater than the threshold T.
5. The method for marking line segmentation based on shape template constraint according to claim 1, characterized in that: In step S2, the image is first reconstructed based on the static background, the color components of the color background image in each color space are extracted, and pixel-level feature vectors are formed according to the color components; then, clustering is performed using the K-means value algorithm based on Euclidean distance in units of pixels; finally, image segmentation is performed according to the clustering results.
6. The method for marking line segmentation based on shape template constraint according to claim 5, characterized in that: The K-means value algorithm is: Among them, c i Represents pixel x i The cluster to which it belongs, μ c Represents the cluster center of the cluster, I is the total number of pixels, and K is the total number of cluster centers.
7. The method for marking line segmentation based on shape template constraint according to claim 5, characterized in that: For color cluster images covering most pixels, the SIFT local feature descriptor is extracted to obtain strong robust image features, and the bag-of-words Bow model is used to complete the ROI area retrieval of the color cluster image.
8. The method for marking line segmentation based on shape template constraint according to claim 1, characterized in that: In step S3, shape template modeling is performed and shape loss is constructed to constrain the region classification task; the shape modeling task is implemented by training using a deep Boltzmann machine (DBM).
9. The method for marking line segmentation based on shape template constraint according to claim 8, characterized in that: Based on the pre-extracted color and texture features, color and texture features are extracted again on the image region of interest, and a two-class feedforward neural network is used for secondary classification of the region of interest. The classification network structure of the secondary classification is expressed as follows: Where W0, W1, and W2 are the linear transformation matrices to be learned, and b1 and b2 are the corresponding bias values; X i is the color and texture feature of the current region, and I is the scene marking mask matrix. By introducing the scene marking mask matrix, the classification model will only output the focus marking area that matches the recognized scene; is the classification output of the model; Based on the principle of minimizing the DBM energy function, the shape constraint term is added to obtain the loss expression of the classification network: Where θ={W 1 ,W 2 ,a 1 ,a 2 ,b} are the network parameters obtained by DBM model training, v is the visible unit input of DBM model; τ is the weight of the shape constraint regularization term.
10. The method for marking line segmentation based on shape template constraint according to claim 1, characterized in that: In step S4, after the Gibbs iterative sampling converges, a shape that approximately obeys the edge probability distribution of the marking template is obtained.