A Vehicle Target Segmentation Method Based on Dual-Branch Unet Noise Suppression

By introducing a dual-branch module and asymmetric index loss function in the Unet network, the problems of low vehicle detection efficiency and poor effect in complex environments are solved, and more efficient and accurate vehicle target segmentation is achieved.

CN114463205BActive Publication Date: 2025-06-27ARMY ENG UNIV OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210066965.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-20
Publication Date
2025-06-27
Estimated Expiration
2042-01-20

AI Technical Summary

Technical Problem

In complex environments, the vehicle detection method has problems such as large calculation amount, cumbersome processing process, low detection efficiency and poor effect. Especially in the vehicle pictures taken by drones, the vehicle target size is small, the details are blurred, and the ground environment is complex, and the light and imaging angles change greatly, resulting in difficulty in segmenting the vehicle.

Method used

The vehicle target segmentation method based on dual-branch Unet noise suppression is adopted. By embedding the prediction branch module and the noise suppression branch module in the Unet network, and introducing an asymmetric exponential loss function, the model's ability to identify difficult samples is improved.

Benefits of technology

It improves the accuracy and efficiency of vehicle segmentation, enhances the model's ability to identify vehicle targets in complex environments, and improves the training effect under small sample conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463205B_ABST
    Figure CN114463205B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle target segmentation method based on dual-branch UNet noise suppression. First, a vehicle dataset is collected and these datasets are sorted out. Then, a part of the UNet network is selected as the backbone network, and on this basis, a prediction branch module and a noise suppression branch module are embedded. The prediction branch module mainly fine-tunes the obtained feature information and performs pixel classification based on this. The noise suppression branch module mainly suppresses the noise interference in the data through the loss function to achieve the accuracy of feature acquisition. Finally, the obtained vehicle dataset is imported into the model, and then the image feature information extracted from the backbone network is transmitted to the prediction branch module and the noise suppression branch module. These two branch modules alternately optimize the model parameters using the binary cross-entropy loss function and the asymmetric exponential loss function respectively, so as to improve the discriminative ability of the model for difficult samples and further enhance the overall performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision semantic segmentation, and to the technical field of a vehicle target segmentation method with a multi-branch hybrid network structure. Background Art

[0002] With the development of science and technology, vehicle detection has developed rapidly in fields such as aerial drone traffic law enforcement and intelligent transportation systems. The flight altitude of drones is relatively high, resulting in relatively small vehicle target sizes and blurred detail information in the captured images. Secondly, due to the complex ground environment, large variations in lighting and imaging angles, and significant differences in the appearance of ground vehicles, there are many interferences in the vehicle pictures taken by drones, which poses new challenges for vehicle detection. How to accurately segment the background and vehicle regions is of great significance to the traffic field.

[0003] Currently, vehicle detection methods are mainly divided into two categories: traditional methods and deep learning methods. Wen Xuezhi et al. obtained the image edges and texture features of the ROI using the Haar wavelet feature extraction algorithm in an article on a vehicle detection algorithm based on low-contrast images, and used a support vector machine to detect vehicles in the ROI; Wu Junwen et al. extracted features from vehicle pictures using wavelet transform and designed a classifier through principal component analysis to complete the vehicle detection task in an article on a static image detection method for vehicles using wavelet transform and PCA; Li Ying et al. used a Gaussian pyramid to estimate the background of the source image and then performed OTSU threshold segmentation on the source image in an article on an infrared image vehicle detection method based on road auxiliary information and saliency detection; Su Ang et al. used a cascaded boosting classifier and a gradient histogram feature extraction method based on a circular filter to detect vehicles in an article on fast calculation of circular filter HOG features in aerial image vehicle detection; Guo Lei et al. proposed a method for vehicle detection using monocular vision in an article on a feature-based vehicle detection method, and adopted an adaptive double threshold in image preprocessing to meet the usage requirements under different lighting conditions; the accuracy of vehicle vertical boundary recognition was improved by using energy density verification; Wu Xinsheng et al. proposed an algorithm for vehicle segmentation that combines the optimal segmentation double threshold method and a conditional random field model, etc. in an article on multi-vehicle segmentation based on the optimal threshold and random labeling method.

[0004] Zhou Kangming et al. adjusted the importance of each category in the model training task by configuring the weights corresponding to the categories in the article "Training Generation Method of Semantic Segmentation Model, Vehicle Appearance Detection Method, and Device", improving the accuracy of vehicle appearance detection; Zhang Yongfei et al. proposed a unique vehicle semantic segmentation model trained by the pyramid scene parsing network and the post-segmentation processing module in the article "A Vehicle Re-identification Method Based on Background Segmentation", using the deep residual network combined with the triplet loss function to achieve accurate vehicle re-identification; Lichao Mou et al. proposed a semantic boundary-aware unified multi-task learning FCN for vehicle instance segmentation in the article "Vehicle Instance Segmentation From Aerial Image and Video Using a Multitask Learning Residual Fully Convolutional Network" to address the problem that most vehicle semantic segmentation methods are difficult to separate objects individually in remote sensing images; Zhang Le et al. proposed a fully convolutional neural network to segment vehicles in images in the article "Research on Vehicle Segmentation Based on Complex Scenes of Full Convolutional Neural Network" to address the problem of insufficient vehicle segmentation accuracy in existing complex traffic scenes; D Wu et al. proposed a training sample iterative selection strategy based on convolutional neural networks in the article "Vehicle Detection in High-Resolution Images Using Superpixel Segmentation and CNN Iteration Stratey" to address the problem of unstable detection performance caused by randomly selecting samples, extracting representative features with high vehicle and background discrimination ability from specific samples for detection; Hao Liying et al. proposed a vehicle segmentation branch and a background segmentation branch in the article "A Vehicle Image Segmentation Method in Complex Traffic Scenes Based on Joint Corner Pooling" to address the problem that existing technologies perform poorly in detecting vehicles in complex traffic scenes; Wang Xue et al. proposed a vehicle detection algorithm based on region convolutional neural network for large-scale and fast vehicle detection and counting using high-resolution satellite image data in the article "Region Convolutional Neural Network for Vehicle Detection in Remote Sensing Images"; Deng Jianhua et al. adopted a neural network built for vehicle detection and obtained a detection box obtained by the detection network, an object detection network improved based on YOLOV3tiny in the article "A Vehicle Detection and Landing Point Location Method Based on Convolutional Neural Network".

[0005] Although these traditional methods have solved the problem of vehicle detection to a certain extent, there are still problems such as large computational complexity, cumbersome processing process, low detection efficiency and poor effect in complex environments. Deep learning technology has been widely used in the field of vehicle detection. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a vehicle target segmentation method based on double-branch Unet noise suppression. On the basis of the Unet network, a prediction branch module and a noise suppression branch module are embedded, and an asymmetric exponential loss function is introduced, providing an efficient solution to the vehicle segmentation problem.

[0007] The present invention adopts the following technical solutions to solve the above technical problems:

[0008] A vehicle target segmentation method based on double-branch Unet noise suppression includes the following steps:

[0009] Step S1, obtain a data set, including real pictures of vehicle scenes and corresponding label pictures, and use image data augmentation technology to expand the data set;

[0010] Step S2, construct a backbone network model;

[0011] Step S3, construct the network model structure of the prediction branch module, and this module uses the binary cross-entropy loss function;

[0012] Step S4, construct the network model structure of the noise suppression branch module, and design an asymmetric exponential loss function for the noise suppression branch module;

[0013] Step S5, import the training set of vehicle data into the backbone network. The backbone network transmits the extracted image features to the prediction branch module and the noise suppression branch module. The prediction branch module updates the network parameters using the binary cross-entropy loss function, and the noise suppression branch module further optimizes the network parameters through the asymmetric exponential loss function. The two modules alternately optimize the network parameters until the training ends to obtain the model parameters corresponding to this data set;

[0014] Step S6, load the saved model parameters, import the test set of vehicle data into the corresponding model, and thus obtain the corresponding test results.

[0015] As a preferred solution of the present invention, the image data augmentation technology in step S1 includes rotation, translation, projective transformation, scaling, flipping and pixel filling.

[0016] As a preferred solution of the present invention, for the backbone network model structure in step S2, the backbone network model includes two parts: a contracting path and an expanding path. Among them,

[0017] The contraction path mainly consists of convolution operations and pooling operations, specifically as follows: For the input image, two convolution operations are used in the first layer. After performing a pooling operation on the feature map output by the first layer, it enters the second layer. In the second layer, two convolution operations are used. After performing a pooling operation on the feature map output by the second layer, it enters the third layer. In the third layer, two convolution operations are used. After performing a pooling operation on the feature map output by the third layer, it enters the fourth layer. In the fourth layer, two convolution operations are used. After performing a pooling operation on the feature map output by the fourth layer, it enters the fifth layer. In the fifth layer, two convolution operations are used;

[0018] The expansion path mainly consists of deconvolution, concatenation, and convolution operations, specifically as follows: In the sixth layer, a deconvolution operation is performed on the feature map output by the fifth layer. The result is concatenated with the feature map output by the fourth layer by channel. Finally, two convolution operations are performed to enter the seventh layer. In the seventh layer, a deconvolution operation is performed on the feature map output by the sixth layer. The deconvolution result is concatenated with the feature map output by the third layer by channel. Finally, two convolution operations are performed and then enter the eighth layer. In the eighth layer, a deconvolution operation is performed on the feature map output by the seventh layer. The deconvolution result is concatenated with the feature map output by the second layer by channel. Finally, two convolution operations are performed to enter the ninth layer. In the ninth layer, a deconvolution operation is performed on the feature map output by the eighth layer. The deconvolution result is concatenated with the feature map output by the first layer by channel. Finally, two convolution operations are performed to obtain the output result;

[0019] Among them, for the convolution operations used in the first to ninth layers, the selected convolution kernel size is 3*3 for all, and the stride is 1 for all; for pooling, the selected convolution kernel size is 2*2 for all. Upsampling uses deconvolution operations, and the selected deconvolution kernel size is 2*2 for all. The number of filters used in the first to ninth layers is 64, 128, 256, 512, 1024, 512, 256, 128, 64 in sequence.

[0020] As a preferred solution of the present invention, the specific structure of the prediction branch module described in step S3 is as follows: The feature map output by the backbone network is input into the prediction head module, and four convolution operations are performed. Among them, the selected convolution kernel size for the first to fourth times is 3*3 for all, and the number of filters used is 64, 64, 64, 2 in sequence. The binary cross-entropy loss function is as follows:

[0021]

[0022] where y i is the pixel value of pixel point i in the ground truth, is the pixel value of pixel point i in the prediction result.

[0023] As a preferred embodiment of the present invention, the specific structure of the noise suppression branch module described in step S4 is as follows: The feature map output by the backbone network is input into the noise suppression module, and two convolution operations are performed. The selected convolution kernel sizes are both 3*3, and the numbers of filters used are 64 and 2 respectively. The asymmetric exponential loss function has the following formula:

[0024]

[0025] where α, β, and γ are hyperparameters. It is found that when α = 1, β = 1, and γ = 0.07, the experimental results are the best. α and β control the severity of the penalty for incorrect predictions, while γ specifies the degree of asymmetry. x is the difference between the prediction result and the ground truth. When x > 0, the ground truth is the background class; when x ≤ 0, the ground truth is the target class.

[0026] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0027] 1. The present invention uses the Unet network as the backbone structure, which can better identify vehicles and effectively improve the inconvenience caused by a small number of samples to training.

[0028] 2. The present invention designs a new asymmetric exponential loss function for the noise suppression branch module, which effectively improves the model's ability to identify difficult samples, thereby improving the overall performance of the model.

[0029] 3. The present invention uses a dual-branch structure with two loss functions. The role of the prediction branch is to continuously make the prediction result approach the label, and the role of the noise suppression branch is to improve the model's ability to identify samples with noise interference. Moreover, the two branch structures alternately optimize the network parameters, making the result of model training more accurate. Description of the Drawings

[0030] Figure 1 is the flowchart of the vehicle target segmentation method based on dual-branch Unet noise suppression of the present invention.

[0031] Figure 2 is the image of the asymmetric exponential loss function in the vehicle target segmentation method based on dual-branch Unet noise suppression of the present invention.

[0032] Figure 3 is the network model structure diagram in the vehicle target segmentation method based on dual-branch Unet noise suppression of the present invention. Detailed Embodiments

[0033] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0034] Due to the relatively complex ground environment, large variations in lighting and imaging angles, and significant differences in the appearance of ground vehicles, vehicle segmentation is subject to more interference. We collected and sorted vehicle data sets under different environmental conditions and perspectives, made corresponding labeled pictures, selected part of Unet as the backbone network. To improve the accuracy of experimental results, a prediction branch module and a noise suppression branch module are added after the Unet network. At the same time, to improve the model's ability to suppress interference, the prediction branch module uses the binary cross-entropy loss function, and the noise suppression branch module uses a new asymmetric exponential loss function to improve the model's recognition ability for difficult samples. The two branch modules alternately optimize the network parameters. Based on this idea, the present invention proposes a vehicle target segmentation method based on dual-branch Unet noise suppression.

[0035] As Figure 1 shown, a vehicle target segmentation method based on dual-branch Unet noise suppression includes the following steps:

[0036] Step S1, obtain a data set, including real pictures of vehicle scenes and corresponding labeled pictures, and use image data augmentation technology to expand the data set;

[0037] Step S2, construct a backbone network model;

[0038] Step S3, construct the network model structure of the prediction branch module, and this module uses the binary cross-entropy loss function;

[0039] Step S4, construct the network model structure of the noise suppression branch module, and design an asymmetric exponential loss function for the noise suppression branch module;

[0040] Step S5, import the training set of vehicle data into the backbone network respectively. The backbone network transmits the extracted image features to the prediction branch module and the noise suppression branch module. The prediction branch module updates the network parameters using the binary cross-entropy loss function, and the noise suppression branch module further optimizes the network parameters through the asymmetric exponential loss function. The two modules alternately optimize the network parameters until the training ends to obtain the model parameters corresponding to this data set;

[0041] Step 6, load the saved model parameters, import the test set of vehicle data into the corresponding model, and thus obtain the corresponding test results.

[0042] As Figure 2As shown in the figure, in step S1, a data set is collected, including vehicle images and their corresponding ground truth label images under various environments and perspectives; and the data set is augmented by means of rotation, translation, projective transformation, scaling, flipping, and pixel filling, etc., so as to increase the total number of its samples and improve the accuracy of the Unet network model.

[0043] Rotation means that the original image can be rotated by different angles to increase samples; translation includes moving the image along the X-axis or Y-axis or both directions simultaneously; projective transformation means that the x-coordinate (or y-coordinate) of all points remains unchanged, while the corresponding y-coordinate (or x-coordinate) is translated proportionally, and the size of the translation is proportional to the perpendicular distance from the point to the x-axis (or y-axis); scaling is to enlarge or reduce the image according to a certain ratio, then crop the enlarged image, and make assumptions about the boundary content of the reduced image to ensure that the size of the scaled image is the same as the original image; flipping is to perform horizontal or vertical flipping operations on the image; pixel filling means that when operations such as translation, scaling, and projection are performed on the image, some missing parts in the image are filled with pixels to keep its size the same as the original image.

[0044] In step S2, a part of the Unet network is selected to construct a backbone network model, and the model structure is as Figure 3 shown, and the specific construction steps are as follows:

[0045] (1) Construct a Unet model;

[0046] Unet is built on the network architecture of FCNs. Unet is a U-shaped structure, consisting of a contracting path and an expanding path. Among them, the contracting path is used to obtain context information, and the expanding path is used for accurate localization, and the two paths are symmetrical to each other. In addition, Unet uses the splicing method to fuse deep features and shallow features, splicing the features together in the channel dimension, and this network does not use fully connected layers.

[0047] (2) In the contracting path, for each layer of the network, the image is first subjected to two convolution operations. Each convolution sets multiple convolution kernels, and the number of convolution kernels used in each layer is different, which are 64, 128, 256, 512, and 1024 respectively. The convolution kernel with a size of 3*3 moves along the height and width of the input feature map with a stride of 1, and the number of channels of the output feature map is the same as the number of convolution kernels. The size of the output feature map is calculated using the following formula:

[0048]

[0049] where, w out is the output image size, w in is the size of the input image, k is the size of the convolution kernel, t is the number of filled pixels, and s is the stride.

[0050] After that, a pooling operation is performed on the feature map group, generally including two types: max pooling and average pooling. A pooling size of 2*2 is selected to downsample the feature map, so that the size of the feature map becomes half of the original. Five feature map groups with different sizes will be obtained in the contraction path.

[0051] As the number of network layers continues to deepen, the feature information of the image can be extracted from shallow to deep.

[0052] (3) In the expansion path, there are upsampling, concatenation, and convolution operations. There are many ways of upsampling. In the Unet model, deconvolution is used for upsampling. The size of the deconvolution kernel is 2*2. The size of the image after deconvolution is calculated using the following formula:

[0053] w in =(w out -1)×s + k - 2×t

[0054] Therefore, after upsampling, the size of the image doubles. To make up for the information lost during the downsampling process in the contraction path, Unet concatenates the upsampled feature map with the feature map with the same number of channels in the contraction path, utilizes the rich feature information in the contraction path, and the number of channels of the concatenated feature map increases. Finally, a secondary convolution operation is performed using a 3*3 convolution kernel, and the number of convolution kernels used in each layer is 512, 256, 128, and 64 respectively. The final output image has the same size as the original input image.

[0055] Step S3: Based on the backbone network model constructed in step S2, construct a prediction branch module. The main function of this module is to classify the samples in the image. It consists of 4 identical convolutional layers, the size of the convolution kernel is 3*3, the stride is 1, and the number of convolution kernels is 64. Finally, the Sigmoid activation function is used to map the output value to the range (0,1). To make the obtained result more accurate, a loss function is used to optimize the network parameters. There are only two categories in the image, namely vehicles and the background, which belongs to a binary classification problem. The binary cross-entropy loss function is used as the loss function of this module, and its formula is as follows:

[0056]

[0057] where y i is the pixel value of pixel point i in the groundtruth, is the pixel value of pixel point i in the prediction result.

[0058] Step S4: Construct a noise suppression branch module. The main function of this module is to suppress noise and improve the model's recognition ability for difficult samples. It consists of two identical convolutional layers with a kernel size of 3*3, a stride of 1, and 64 kernels. Finally, the Sigmoid activation function is used to map the output value to the range (0, 1). To achieve the aggregation of features among similar samples and the separation of features among different samples, and improve the model's discrimination ability for difficult samples, a new loss function, namely the asymmetric exponential loss function, is designed to optimize the features. The formula is as follows:

[0059]

[0060] Among them, α, β, and γ are hyperparameters. It is found that when α = 1, β = 1, and γ = 0.07, the experimental results are the best. α and β control the severity of the penalty for mispredictions, while γ specifies the degree of asymmetry. x is the difference between the prediction result and the ground truth. When x > 0, the ground truth is the background class; when x ≤ 0, the ground truth is the target class.

[0061] The schematic diagram of the asymmetric exponential loss function is as shown in Figure 1 and can be divided into three parts. When -0.3 < x < 0.3, at this time the model's prediction of the sample and the ground truth belong to the same class, and the difference between them is small. At this time, the sample belongs to an easy sample, so a small loss function value and gradient are assigned to this sample, so that the model pays less attention to it during training and pays more attention to difficult samples. When -0.5 < x < -0.3 and 0.3 < x < 0.5, at this time the model's prediction of the sample and the ground truth belong to the same class, but the difference between them is large. At this time, the sample belongs to a relatively difficult sample, so a slightly larger loss function value is assigned to this sample, and at the same time its gradient is increased, so that the model pays more attention to this sample during training and moves it towards a more accurate prediction result. When x < -0.5 and x > 0.5, at this time the model's prediction of the sample and the ground truth belong to different classes, and there is an obvious difference between them. At this time, the sample is a difficult sample, so a larger loss function value and gradient are assigned to this sample, so that the model pays more attention to it during training and moves it towards the correct prediction direction with a larger gradient. Through this loss function, the sample features extracted by the model can be optimized, and the separability of vehicle class and background class features can be achieved, which helps to improve the model's discrimination ability for difficult samples and enhance the model performance.

[0062] Step S5: Based on the augmented vehicle dataset in Step S1, perform model training and testing. Set the number of batches, the number of iterations, and the number of times per iteration during the training process, train the data, use the output feature map of the backbone network as the input of the two branch modules, and the two branch modules optimize the network parameters through their own structures and loss function optimization loops, and finally save the model parameters. Load the saved model parameters to obtain the test results of the model on the test set.

[0063] Finally, the model needs to be evaluated. Therefore, first calculate the accuracy and recall rate of the model. The accuracy represents the proportion of the results where the prediction is a positive example and the actual situation is also a positive example among the positive prediction results. The recall rate represents the proportion of the results where the prediction is a positive example and the actual situation is also a positive example among the actual positive results. These two indicators cannot comprehensively evaluate the detection results. The comprehensive evaluation indicator f β The formula is as follows:

[0064]

[0065] where β 2 takes a value of 0.3 to make the weight of the accuracy higher than that of the recall rate; p is the accuracy, and r is the recall rate.

[0066] The vehicle target segmentation method based on Unet noise suppression of the present invention improves the model's recognition ability for difficult samples and further enhances the segmentation accuracy by improving the structure of the network model and designing a new loss function, and using two loss functions to alternately optimize the network parameters.

[0067] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modifications made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the present invention.

Claims

1. A vehicle target segmentation method based on dual-branch Unet noise suppression, characterized in that It includes the following steps: Step S1: Obtain a dataset, including real pictures of vehicle scenarios and corresponding labeled pictures, and use image data augmentation technology to expand the dataset; Step S2: Construct a backbone network model; The backbone network model includes two parts: a contraction path and an expansion path. Among them, The contraction path includes convolution and pooling operations. Specifically: for the input image, two convolution operations are used in the first layer. After pooling the feature map output by the first layer, it enters the second layer. Two convolution operations are used in the second layer. After pooling the feature map output by the second layer, it enters the third layer. Two convolution operations are used in the third layer. After pooling the feature map output by the third layer, it enters the fourth layer. Two convolution operations are used in the fourth layer. After pooling the feature map output by the fourth layer, it enters the fifth layer. Two convolution operations are used in the fifth layer; The expansion path includes deconvolution, concatenation, and convolution operations. Specifically: in the sixth layer, deconvolution is performed on the feature map output by the fifth layer, and the result is concatenated with the feature map output by the fourth layer by channel. Finally, two convolution operations are performed to enter the seventh layer. In the seventh layer, deconvolution is performed on the feature map output by the sixth layer, and the deconvolution result is concatenated with the feature map output by the third layer by channel. Finally, two convolution operations are performed and then enter the eighth layer. In the eighth layer, deconvolution is performed on the feature map output by the seventh layer, and the deconvolution result is concatenated with the feature map output by the second layer by channel. Finally, two convolution operations are performed to enter the ninth layer. In the ninth layer, deconvolution is performed on the feature map output by the eighth layer, and the deconvolution result is concatenated with the feature map output by the first layer by channel. Finally, after two convolution operations, the output result is obtained; Among them, for the convolution operations used in the first to ninth layers, the selected convolution kernel size is 3*3 for all, and the stride is 1 for all; the selected pooling convolution kernel size is 2*2 for all. Upsampling uses deconvolution operations, and the selected deconvolution kernel size is 2*2 for all. The number of filters used in the first to ninth layers is 64, 128, 256, 512, 1024, 512, 256, 128, 64 in sequence; Step S3: Construct the network model structure of the prediction branch module, and this module uses the binary cross-entropy loss function; The specific structure of the prediction branch module is as follows: the feature map output by the backbone network is input into the prediction branch module, and four convolution operations are performed. Among them, the selected convolution kernel size for the first to fourth times is 3*3 for all, and the number of filters used is 64, 64, 64, 2 in sequence. The binary cross-entropy loss function, the formula is as follows: where y i is the pixel value of pixel point i in the ground truth, and is the pixel value of pixel point i in the prediction result; Step S4: Construct the network model structure of the noise suppression branch module, and design an asymmetric exponential loss function for the noise suppression branch module; The specific structure of the noise suppression branch module is as follows: the feature map output by the backbone network is input into the noise suppression module, and two convolution operations are performed. The selected convolution kernel size is 3*3 for both, and the number of filters used is 64 and 2 respectively. The asymmetric exponential loss function, the formula is as follows: Among them, α, β, and γ are hyperparameters, α = 1, β = 1, γ = 0.07; α and β control the severity of the penalty for incorrect predictions, while γ specifies the degree of asymmetry. x is the difference between the prediction result and the ground truth. When x > 0, the ground truth is the background class; when x ≤ 0, the ground truth is the target class. In step S5, the training set of vehicle data is imported into the backbone network respectively. The backbone network transfers the extracted image features to the prediction branch module and the noise suppression branch module. The prediction branch module updates the network parameters using the binary cross-entropy loss function, and the noise suppression branch module further optimizes the network parameters through the asymmetric exponential loss function. The two modules alternately optimize the network parameters until the training ends, obtaining the model parameters corresponding to this data set. In step S6, the saved model parameters are loaded, and the test set of vehicle data is imported into the corresponding model to obtain the corresponding test results.

2. The vehicle target segmentation method based on dual-branch Unet noise suppression according to claim 1, wherein The image data augmentation technique described in step S1 includes rotation, translation, projective transformation, scaling, flipping, and pixel filling.

Citation Information

Patent Citations

  • Target detection method and system for intelligent construction site

    CN112488015A

  • Fundus image classification method and imaging method under data deviation

    CN112560948A