Semantic Segmentation Method, Terminal and Storage Medium
By determining the position weights of key areas and non-key areas in the semantic segmentation model, the problem of low segmentation accuracy in key areas and categories of existing models is solved, and higher segmentation accuracy and improved safety and comfort of BSD system are achieved.
Patent Information
- Application Number
- CN202111489134.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-07
AI Technical Summary
The existing semantic segmentation model has low segmentation accuracy in the categories and regions that are of focus and does not meet the actual usage requirements.
By obtaining the training sample set and the pre-set target category, the key area and non-key area of the target image are determined according to the occurrence frequency of the pixel points in the target category in the training sample set, and the first position weight of each pixel point in the key area and the second position weight of each pixel point in the non-key area are determined, so that the first position weight is greater than the second position weight, thereby semantically segmenting the target image.
The segmentation accuracy of the semantic segmentation model for the key categories and regions is improved, and it is more in line with actual usage needs, and the safety and comfort of the BSD system are improved.
Smart Images

Figure CN114387434B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic segmentation, and in particular to a semantic segmentation method, a terminal, and a storage medium. Background Art
[0002] A Blind Spot Detection (BSD) system can perform semantic segmentation and target detection on the scene in the blind spot based on a camera, and measure the distance to the target based on the internal and external parameters calibrated by the camera. When the distance to the target is lower than a set safety distance, an alarm message is sent.
[0003] The semantic segmentation model in the BSD system can classify each pixel of an image. The classified categories usually include pedestrians, vehicles, road surfaces, traffic signs, backgrounds, etc. Currently, when training the semantic segmentation model, the cross-entropy is usually used as the loss function. The cross-entropy loss function is L = -∑y q log(p q )), where y q is the label value of the q-th pixel point in the image, and p q is the predicted probability of the q-th pixel point in the image.
[0004] The above cross-entropy loss function treats the regions and categories in the image equally, resulting in a higher segmentation accuracy for pixel categories with a larger proportion in the semantic segmentation model, a lower segmentation accuracy for pixel categories with a smaller proportion, a better learning effect for pixel regions that are easy to learn, and a worse learning effect for pixel regions with a greater learning difficulty. However, pixel categories with a larger proportion are not necessarily the categories that the BSD system needs to focus on, and pixel regions that are easy to learn are not necessarily the regions that the BSD system focuses on, resulting in a lower segmentation accuracy of the semantic segmentation model for the categories and regions that are focused on, which does not meet the actual usage requirements. Summary of the Invention
[0005] Embodiments of the present invention provide a semantic segmentation method, a terminal, and a storage medium to solve the problem that the existing technology has a lower segmentation accuracy for the categories and regions that are focused on and does not meet the actual usage requirements.
[0006] In a first aspect, an embodiment of the present invention provides a semantic segmentation method, including:
[0007] Obtain a training sample set; the training sample set includes multiple training images taken by the same imaging device at a preset angle, and each pixel point of each training image is labeled with a corresponding category;
[0008] Obtain a preset target category;
[0009] Determine the key area and non-key area of the target image according to the occurrence frequency of the pixel points of the target category in the training sample set; the target image is any image captured when the camera device is at a preset angle;
[0010] Determine the first position weight of each pixel point in the key area and the second position weight of each pixel point in the non-key area, where the first position weight is greater than the second position weight;
[0011] Perform semantic segmentation on the target image according to the first position weight and the second position weight.
[0012] In a possible implementation, determining the second position weight of each pixel point in the non-key area includes:
[0013] According to the training sample set, determine the set of road surface pixel points in the non-key area that are closest to the vehicle area;
[0014] Determine the dividing line according to the set;
[0015] According to the dividing line, determine the vehicle area and non-vehicle area of the non-key area;
[0016] Set the second position weight of each pixel point in the vehicle area to the first preset weight, and set the second position weight of each pixel point in the non-vehicle area to the second preset weight; the first preset weight is less than the second preset weight.
[0017] In a possible implementation, determining the dividing line according to the set includes:
[0018] Perform linear fitting on each pixel point in the set to obtain a fitting line;
[0019] Obtain the abscissas of the intersection points of the fitting line and the upper and lower boundaries of the target image in the preset coordinate system, and denote them as x1 and x2 respectively;
[0020] Determine the line where y = (x1 + x2) / 2 as the dividing line.
[0021] In a possible implementation, the set of road surface pixel points in the non-key area that are closest to the vehicle area is the set of road surface pixel points with the smallest abscissa in each row of pixel points in the non-key area in the preset coordinate system;
[0022] According to the dividing line, determining the vehicle area of the non-key area includes:
[0023] Determine the area within the non-key area that is within the range of the dividing line and the coordinate axes of the preset coordinate system as the vehicle area;
[0024] Wherein, when the imaging device is located on the right side of the vehicle, the lower left corner of the target image is taken as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower left corner of the target image to the right, and the positive y-axis direction of the preset coordinate system is from the lower left corner of the target image upwards;
[0025] When the imaging device is located on the left side of the vehicle, the lower right corner of the target image is taken as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower right corner of the target image to the left, and the positive y-axis direction of the preset coordinate system is from the lower right corner of the target image upwards.
[0026] In a possible implementation manner, determining the first position weight of each pixel point in the key area includes:
[0027] Determining the target segmentation error degree of each pixel point in the key area according to the training sample set;
[0028] Determining the weight of each pixel point in the key area in the first direction;
[0029] Determining the weight of each pixel point in the key area in the second direction;
[0030] Determining the first position weight of each pixel point in the key area according to the target segmentation error degree of each pixel point in the key area, the weight of each pixel point in the key area in the first direction, and the weight of each pixel point in the key area in the second direction.
[0031] In a possible implementation manner, the first direction and the second direction are different directions;
[0032] Determining the first position weight of each pixel point in the key area according to the target segmentation error degree of each pixel point in the key area, the weight of each pixel point in the key area in the first direction, and the weight of each pixel point in the key area in the second direction includes:
[0033] For each pixel point in the key area, multiplying the target segmentation error degree of the pixel point, the weight of the pixel point in the first direction, and the weight of the pixel point in the second direction to obtain the first position weight of the pixel point.
[0034] In a possible implementation manner, determining the target segmentation error degree of each pixel point in the key area according to the training sample set includes:
[0035] Dividing the training sample set into K training sample subsets on average; each training sample subset contains N training images;
[0036] Selecting one of the K training sample subsets as the test set in turn from the K training sample subsets, and using the remaining K - 1 training sample subsets as the training set to obtain K combinations of test sets and training sets;
[0037] For each combination of the test set and the training set, train a preset semantic segmentation model according to the training set to obtain a trained semantic segmentation model; use the trained semantic segmentation model to perform semantic segmentation on N training images in the test set respectively to obtain semantic segmentation results corresponding to the N training images in the test set respectively; count the number of times of segmentation errors of each pixel point in the key area according to the semantic segmentation results corresponding to the N training images in the test set respectively; divide the number of times of segmentation errors of each pixel point in the key area by N respectively to obtain the test segmentation error degree of each pixel point in the key area corresponding to the combination of the test set and the training set.
[0038] For each pixel point in the key area, take the average value of the test segmentation error degrees of this pixel point corresponding to K combinations of the test set and the training set respectively to obtain the average segmentation error degree of this pixel point.
[0039] Obtain the target segmentation error degree of each pixel point in the key area according to the average segmentation error degree of each pixel point in the key area.
[0040] In a possible implementation manner, obtaining the target segmentation error degree of each pixel point in the key area according to the average segmentation error degree of each pixel point in the key area includes:
[0041] Select the maximum average segmentation error degree value max and the minimum average segmentation error degree value min from the average segmentation error degrees of each pixel point in the key area;
[0042] According to Calculate the target segmentation error degree of each pixel point in the key area;
[0043] Among them, R i is the target segmentation error degree of the i-th pixel point in the key area; s i is the average segmentation error degree of the i-th pixel point in the key area; e is the third preset weight, which is the weight corresponding to the minimum average segmentation error degree value min; 1 ≤ i ≤ F, and F is the number of pixel points in the key area.
[0044] In a possible implementation manner, the first direction is the x-axis direction of the preset coordinate system;
[0045] When the imaging device is located on the right side of the vehicle, take the lower left corner of the target image as the origin of the preset coordinate system, the right direction from the lower left corner of the target image is the positive x-axis direction of the preset coordinate system, and the upward direction from the lower left corner of the target image is the positive y-axis direction of the preset coordinate system;
[0046] When the imaging device is located on the left side of the vehicle, the lower right corner of the target image is taken as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower right corner of the target image to the left, and the positive y-axis direction of the preset coordinate system is from the lower right corner of the target image upwards;
[0047] Determine the weights of each pixel point in the key area in the first direction, including:
[0048] For each row of pixel points in the key area, according to Determine the weights of each pixel point in this row of pixel points in the first direction; among them, the pixel points in this row of the key area are sorted in ascending order of abscissa, T j is the weight of the jth pixel point in this row of pixel points in the key area in the first direction; m j is the abscissa of the jth pixel point in this row of pixel points in the key area; is the fourth preset weight, and is the weight of the pixel point with abscissa m1 in this row of pixel points in the first direction; b is the fifth preset weight, and is the weight of the pixel point with abscissa m G in this row of pixel points in the first direction; a > b; G is the number of pixel points in this row of pixel points in the key area.
[0049] In a possible implementation, the second direction is the y-axis direction of the preset coordinate system;
[0050] Determine the weights of each pixel point in the key area in the second direction, including:
[0051] For each column of pixel points in the key area, according to Determine the weights of each pixel point in this column of pixel points in the second direction; among them, the pixel points in this column of the key area are sorted in descending order of ordinate, U l is the weight of the lth pixel point in this column of pixel points in the key area in the second direction; n l is the ordinate of the lth pixel point in this column of pixel points in the key area; c is the sixth preset weight, and is the weight of the pixel point with ordinate n1 in this column of pixel points in the second direction; d is the sixth preset weight, and is the weight of the pixel point with ordinate n H in this column of pixel points in the second direction; c > d; H is the number of pixel points in this column of pixel points in the key area.
[0052] In a possible implementation, according to the occurrence frequency of the target category pixel points in the training sample set, determine the key area and non-key area of the target image, including:
[0053] For each pixel point of the target image, if the occurrence frequency of the target category pixel point corresponding to this pixel point is greater than the preset frequency threshold, then determine the position where this pixel point is located as the key area, otherwise, determine the position where this pixel point is located as the non-key area.
[0054] In a second aspect, an embodiment of the present invention provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the semantic segmentation method described in the first aspect above or any possible implementation manner of the first aspect are implemented.
[0055] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the semantic segmentation method described in the first aspect above or any possible implementation manner of the first aspect are implemented.
[0056] An embodiment of the present invention provides a semantic segmentation method, a terminal, and a storage medium. By obtaining a training sample set and a preset target category, determining a key area and a non-key area of a target image according to the occurrence frequency of the target category pixel points in the training sample set, and determining a first position weight of each pixel point in the key area and a second position weight of each pixel point in the non-key area, with the first position weight being greater than the second position weight, the key area to be focused on and the non-key area to be focused on can be distinguished. According to the first position weight and the second position weight, semantic segmentation is performed on the target image, which can improve the segmentation accuracy of the semantic segmentation model for the category and area to be focused on, and better meet the actual usage requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 is a flowchart of the implementation of the semantic segmentation method provided by an embodiment of the present invention;
[0059] Figure 2 is a schematic diagram of an image captured by a camera device installed on the right side of a vehicle at a preset angle;
[0060] Figure 3 is when the target category is a pedestrian in an embodiment of the present invention Figure 2 schematic diagram of the frequency distribution image;
[0061] Figure 4 is provided by an embodiment of the present invention for Figure 3 schematic diagram of the image after dilation processing;
[0062] Figure 5 It is a schematic structural diagram of a semantic segmentation device provided by an embodiment of the present invention;
[0063] Figure 6 It is a schematic diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners
[0064] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present invention.
[0065] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will be described through specific embodiments with reference to the accompanying drawings.
[0066] Refer to Figure 1 , which shows a flowchart of the implementation of a semantic segmentation method provided by an embodiment of the present invention. Among them, the execution subject of the semantic segmentation method can be a terminal. This terminal can be a data processing device or a controller of a BSD system, etc.
[0067] Refer to Figure 1 , and the semantic segmentation method is described in detail as follows:
[0068] In S101, a training sample set is obtained; the training sample set includes multiple training images taken by the same imaging device at a preset angle, and each pixel point of each training image has been labeled with a corresponding category.
[0069] Since large vehicles have large blind spots, the BSD system is usually applied to some large vehicles, such as buses, trains, etc. Of course, the BSD system can also be applied to small cars, and no specific limitation is made here.
[0070] Among them, the above imaging device can be installed on the left side and / or the right side of the vehicle. The installation positions on the left and right sides can be symmetrical, and it is usually installed at a position closer to the tail. For example, for a large vehicle, it can be installed at a position 1 meter from the tail and a height of 2 meters. Usually, the blind spot on the right side of the vehicle is larger, so the imaging device installed on the right side of the vehicle has a greater effect. Figure 2 is an image taken when the imaging device installed on the right side of the vehicle is at a preset angle. In this image, the right rear wheel arch, the right side mirror, the right side ground, and the right side pedestrians of the vehicle can be seen.
[0071] When the imaging device is at a preset angle, the shooting range is the optimal required range, and the preset angle can be calibrated according to actual requirements. The imaging device can be a device such as a camera.
[0072] In this embodiment, multiple training images captured by the same imaging device in the BSD system of the vehicle when it is at a preset angle can be obtained to form a training sample set. The multiple training images can be images captured at different times, and each training image is labeled with the category of each pixel using an existing labeling method. The categories can be pedestrians, vehicles, road surfaces, lane lines, traffic signs, backgrounds, etc. Each training image has the same size and the same number of pixel points.
[0073] To improve the training accuracy, the training sample set contains as many and as comprehensive training images as possible.
[0074] In S102, a preset target category is obtained.
[0075] In this embodiment, the target category can be preset. The target category is the category that is the key focus, so that the key area can be determined according to the category that is the key focus. The target category can be one category or multiple categories. Exemplarily, the target category can be pedestrians, vehicles, etc.
[0076] In S103, according to the occurrence frequency of the pixel points of the target category in the training sample set, the key area and the non-key area of the target image are determined; the target image is any image captured by the imaging device when it is at a preset angle.
[0077] In this embodiment, statistical analysis can be performed on the training sample set according to the target category to determine the key area and the non-key area of the target image. Among them, the key area is the area that is the key focus and can be the area where the target category has appeared. The non-key area is the area other than the key area and can be the area where the target category has not appeared.
[0078] The pixel points of the target category refer to the pixel points that belong to the target category.
[0079] This embodiment can count the occurrence frequency of the pixel points of the target category corresponding to the positions of each pixel point of the target image according to the training sample set. The specific process is as follows:
[0080] For each pixel point of the target image, count the number of categories labeled as the target category in the areas at the same position as this pixel point in each training image in the training sample set, and divide this number by the number of training images in the training sample set as the occurrence frequency of the pixel points of the target category corresponding to this pixel point.
[0081] Among them, since both the target image and the training image are images captured by the same imaging device at a preset angle, their sizes are the same, and each pixel point in one image has a corresponding pixel point in the other image. Pixel points in the same row and the same column of different images can be pixel points at the same position in these different images. That is to say, the area in the above-mentioned training image that is at the same position as the pixel point of the target image represents a certain pixel point in the training image, and the position (the row and column where it is located) of this pixel point in the training image is the same as the position of the pixel point of the target image in the target image.
[0082] In some embodiments, S103 may include:
[0083] For each pixel point of the target image, if the occurrence frequency of the target category pixel point corresponding to this pixel point is greater than a preset frequency threshold, then determine the position where this pixel point is located as a key area; otherwise, determine the position where this pixel point is located as a non-key area.
[0084] In this embodiment, by determining whether the occurrence frequency of the target category pixel point corresponding to each pixel point is greater than the preset frequency threshold, to determine whether the position where each pixel point is located is a key area. If the occurrence frequency of the target category pixel point corresponding to a certain pixel point is greater than the preset frequency threshold, then determine the position where this pixel point is located as a key area; if the occurrence frequency of the target category pixel point corresponding to a certain pixel point is not greater than the preset frequency threshold, then determine the position where this pixel point is located as a non-key area.
[0085] Among them, the preset frequency threshold can be set according to actual needs. In one possible implementation, the preset frequency threshold can be 0.
[0086] Figure 3 is the frequency distribution image when the target category is a pedestrian. The greater the brightness, the greater the frequency of pedestrians appearing. A brightness of 0 indicates that no pedestrians have appeared in this area. Set the gray value of the image area with a frequency greater than 0 to 255, and after Figure 3 performing dilation processing, obtain Figure 4 . Figure 4 The key area and the non-key area can be clearly distinguished. The black part is the non-key area, and the white part is the key area.
[0087] In S104, determine the first position weight of each pixel point in the key area and the second position weight of each pixel point in the non-key area, and the first position weight is greater than the second position weight.
[0088] In this embodiment, the first position weight of each pixel point in the key area is greater than the second position weight of each pixel point in the non-key area, so that during semantic segmentation, key areas can be focused on.
[0089] The first position weights of the pixel points in the key area can be the same or different; the second position weights of the pixel points in the non-key area can be the same or different, and all can be determined according to actual requirements.
[0090] In some embodiments, "determining the second position weights of the pixel points in the non-key area" in the above S104 may include:
[0091] According to the training sample set, determine the set of road surface pixel points closest to the vehicle area within the non-key area;
[0092] Determine the dividing line according to the set;
[0093] According to the dividing line, determine the vehicle area and the non-vehicle area of the non-key area;
[0094] Set the second position weights of the pixel points in the vehicle area to the first preset weight, and set the second position weights of the pixel points in the non-vehicle area to the second preset weight; the first preset weight is less than the second preset weight.
[0095] In this embodiment, the non-key area can be divided into a vehicle area and a non-vehicle area. Since the camera device usually captures the vehicle when shooting at a preset angle, and the vehicle area is not the area we need to focus on, the vehicle area can be separated and a relatively small position weight can be set. For the non-vehicle area in the non-key area, compared with the vehicle area, the target category may appear in the non-vehicle area. Therefore, the position weight of the non-vehicle area can be set slightly larger than the position weight of the vehicle area, that is, the second preset weight is greater than the first preset weight.
[0096] The second preset weight and the first preset weight can be set according to actual requirements, and no specific limitation is made here.
[0097] In this embodiment, the non-key area is divided into a vehicle area and a non-vehicle area by a dividing line, and the dividing line can be determined by the set of road surface pixel points closest to the vehicle area within the non-key area.
[0098] In some embodiments, the above determining the dividing line according to the set includes:
[0099] Perform linear fitting on each pixel point in the set to obtain a fitting line;
[0100] Obtain the abscissas of the intersection points of the fitting line and the upper and lower boundaries of the target image in the preset coordinate system, and denote them as x1 and x2 respectively;
[0101] Determine the line where y = (x1 + x2) / 2 as the dividing line.
[0102] Perform linear fitting on each pixel point in the set to obtain a fitted line. Among them, the least squares method can be used for linear fitting. The fitted line has an intersection with both the upper and lower boundaries of the image. Denote the abscissas of the two intersections in the preset coordinate system as x1 and x2 respectively. According to x1 and x2, a line y = (x1 + x2) / 2 is obtained, and this line is the dividing line.
[0103] In some embodiments, the set of pavement pixel points closest to the vehicle area within the above non-key area is the set of pavement pixel points with the smallest abscissa in the preset coordinate system among the pixel points in each row within the non-key area;
[0104] The above-mentioned determination of the vehicle area of the non-key area according to the dividing line includes:
[0105] Determine the area within the non-key area that is within the range of the dividing line and the coordinate axes of the preset coordinate system as the vehicle area;
[0106] Among them, when the imaging device is located on the right side of the vehicle, take the lower left corner of the target image as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower left corner of the target image to the right, and the positive y-axis direction of the preset coordinate system is from the lower left corner of the target image upwards;
[0107] When the imaging device is located on the left side of the vehicle, take the lower right corner of the target image as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower right corner of the target image to the left, and the positive y-axis direction of the preset coordinate system is from the lower right corner of the target image upwards.
[0108] When the imaging device is located on the right side of the vehicle, the vehicle area of the captured image is located on the left side of the image. When the imaging device is located on the left side of the vehicle, the vehicle area of the captured image is located on the right side of the image. Therefore, when the imaging device is located at different positions of the vehicle, different origins and coordinate axes are selected for the preset coordinate system.
[0109] Among them, in the non-key area, the area other than the vehicle area is the non-vehicle area.
[0110] In this embodiment, the above-mentioned determination of the set of pavement pixel points closest to the vehicle area within the non-key area according to the training sample set may include:
[0111] For each training image in the training sample set, obtain the coordinates of the pavement pixel points with the smallest abscissa in the preset coordinate system among the pixel points in each row of the training image;
[0112] For each row of pixel points, take the average of the coordinates of the road surface pixel points with the smallest abscissa in the pixel points of this row of each training image, and obtain the average coordinate corresponding to this row of pixel points; among them, the set of average coordinates corresponding to each row of pixel points is the set of road surface pixel points closest to the vehicle area within the non-key area.
[0113] For each training image in the training sample set, obtain the coordinates of the road surface pixel points with the smallest abscissa in the pixel points of each row of this training image in the preset coordinate system. That is to say, if the training image is taken by a camera device on the right side of the vehicle, obtain the coordinates of the leftmost road surface pixel point of this training image; if the training image is taken by a camera device on the left side of the vehicle, obtain the coordinates of the rightmost road surface pixel point of this training image.
[0114] Take the average value of the coordinates of the road surface pixel points in the same row to obtain the average coordinate corresponding to each row of pixel points. Among them, taking the average value of the coordinates can be to calculate the average value of the abscissa and ordinate respectively. Since the ordinates of the pixel points in each row are the same, therefore, only the abscissa can be averaged.
[0115] Perform linear fitting on the average coordinates corresponding to each row of pixel points, and a fitting line can be obtained. Furthermore, the dividing line y = (x1 + x2) / 2 can be obtained. The dividing line, the coordinate axes, and the upper boundary of the image enclose a range. The non-key area within this range is the vehicle area, and the non-key area outside this range is the non-vehicle area. That is to say, if the camera device is on the right side of the vehicle, the non-key area on the left side of the dividing line is the vehicle area, and the non-key area on the right side of the dividing line is the non-vehicle area; if the camera device is on the left side of the vehicle, the non-key area on the right side of the dividing line is the vehicle area, and the non-key area on the left side of the dividing line is the non-vehicle area.
[0116] In some embodiments, the "determining the first position weight of each pixel point in the key area" in the above S104 may include:
[0117] According to the training sample set, determine the target segmentation error degree of each pixel point in the key area;
[0118] Determine the weight of each pixel point in the key area in the first direction;
[0119] Determine the weight of each pixel point in the key area in the second direction;
[0120] According to the target segmentation error degree of each pixel point in the key area, the weight of each pixel point in the key area in the first direction, and the weight of each pixel point in the key area in the second direction, determine the first position weight of each pixel point in the key area.
[0121] Among them, the target segmentation error degree of a pixel can be used to characterize the error degree of classifying the pixel using a semantic segmentation model.
[0122] The weight of a pixel in the first direction can be used to characterize the importance of the pixel in the first direction; the weight of a pixel in the second direction can be used to characterize the importance of the pixel in the second direction. Among them, the first direction and the second direction can be the row direction of the target image and the column direction of the target image respectively. That is to say, the weight of a pixel in the first direction and the weight of a pixel in the second direction can be the weight of the pixel in its corresponding row and the weight of the pixel in its corresponding column respectively.
[0123] Since a driver usually pays attention to the areas that affect their driving during driving and pays less attention to the areas that do not affect their driving, even if a target category appears in the areas that do not affect their driving, they will not pay too much attention to the areas that do not affect their driving. Therefore, even if the pixels are all in key areas, their position weights may be different. Usually, the areas closer to the vehicle may be focused on, and the areas farther from the vehicle may not be paid too much attention to. Therefore, the weights of each pixel in the first direction in the key areas may be different, and the weights of each pixel in the second direction in the key areas may be different.
[0124] In this embodiment, by determining the target segmentation error degree, the weight in the first direction, and the weight in the second direction of the pixels in the key area, the position weight of the pixel in the key area can be obtained.
[0125] In some embodiments, the above-mentioned first direction and second direction are in different directions;
[0126] The above-mentioned method for determining the first position weight of each pixel in the key area according to the target segmentation error degree of each pixel in the key area, the weight of each pixel in the first direction in the key area, and the weight of each pixel in the second direction in the key area includes:
[0127] For each pixel in the key area, multiply the target segmentation error degree of the pixel, the weight of the pixel in the first direction, and the weight of the pixel in the second direction to obtain the first position weight of the pixel.
[0128] In some embodiments, according to the training sample set, determining the target segmentation error degree of each pixel in the key area includes:
[0129] Divide the training sample set into K training sample subsets on average; each training sample subset contains N training images;
[0130] Take one of the K training sample subsets in turn as the test set, and use the remaining K - 1 training sample subsets as the training set to obtain K combinations of test sets and training sets;
[0131] For each combination of test set and training set, train a preset semantic segmentation model according to the training set to obtain a trained semantic segmentation model; use the trained semantic segmentation model to perform semantic segmentation on the N training images in the test set respectively to obtain the semantic segmentation results corresponding to the N training images in the test set respectively; count the number of times of segmentation errors of each pixel point in the key area according to the semantic segmentation results corresponding to the N training images in the test set respectively; divide the number of times of segmentation errors of each pixel point in the key area by N to obtain the test segmentation error degree of each pixel point in the key area corresponding to the combination of the test set and the training set;
[0132] For each pixel point in the key area, take the average value of the test segmentation error degrees of this pixel point corresponding to the K combinations of test sets and training sets to obtain the average segmentation error degree of this pixel point;
[0133] Obtain the target segmentation error degree of each pixel point in the key area according to the average segmentation error degree of each pixel point in the key area.
[0134] Among them, K can be a positive integer greater than or equal to 2, and N is a positive integer.
[0135] In this embodiment, the K-fold cross-validation method is adopted to evenly divide the training sample set into K training sample subsets. Each time, one of the training sample subsets is selected as the test set and selected in turn, that is, each time a different training sample subset is selected as the test set, and the remaining K - 1 training sample subsets are used as the training set. Perform K times of training and testing. Use the trained semantic segmentation model for each time to perform semantic segmentation on the test set of that time, and the semantic segmentation results corresponding to the N training images in the test set can be obtained, that is, the predicted categories of each pixel point of each training image in the test set.
[0136] Compare the predicted category with the labeled category. If they are different, the segmentation is incorrect. It can be statistically obtained that in this test, the number of times of category segmentation errors of the pixel points in the key area at the same position, and this number divided by N gives the test segmentation error degree of this pixel point in this test.
[0137] For the pixel points at the same position in the key area, take the average value of the test segmentation error degrees of the pixel points at this position obtained K times to obtain the average segmentation error degree of the pixel points at this position. According to the average segmentation error degree of the pixel points at this position, the target segmentation error degree of the pixel points at this position can be obtained.
[0138] In some embodiments, obtaining the target segmentation error degree of each pixel point in the key area based on the average segmentation error degree of each pixel point in the key area may include:
[0139] Selecting the maximum average segmentation error degree value max and the minimum average segmentation error degree value min from the average segmentation error degrees of each pixel point in the key area;
[0140] According to Calculating the target segmentation error degree of each pixel point in the key area;
[0141] Where R i is the target segmentation error degree of the i-th pixel point in the key area; s i is the average segmentation error degree of the i-th pixel point in the key area; e is the third preset weight, which is the weight corresponding to the minimum average segmentation error degree value min; 1 ≤ i ≤ F, and F is the number of pixel points in the key area.
[0142] The value of e can be set according to actual needs. In one possible implementation, e = 0.3.
[0143] In some embodiments, the first direction is the x-axis direction of the preset coordinate system;
[0144] When the imaging device is on the right side of the vehicle, taking the lower left corner of the target image as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower left corner of the target image to the right, and the positive y-axis direction of the preset coordinate system is from the lower left corner of the target image upward;
[0145] When the imaging device is on the left side of the vehicle, taking the lower right corner of the target image as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower right corner of the target image to the left, and the positive y-axis direction of the preset coordinate system is from the lower right corner of the target image upward;
[0146] Determining the weight of each pixel point in the key area in the first direction may include:
[0147] For each row of pixel points in the key area, according to Determining the weight of each pixel point in this row of pixel points in the first direction; among them, the pixel points in this row of the key area are sorted in ascending order of abscissa, and T j is the weight of the j-th pixel point in this row of pixel points in the key area in the first direction; m j is the abscissa of the j-th pixel point in this row of pixel points in the key area; is the fourth preset weight, which is the weight of the pixel point with abscissa m1 in this row of pixel points in the first direction; b is the fifth preset weight, which is the weight of the pixel point with abscissa m in this row of pixel points GThe weight of the pixel points in the first direction; a > b; G is the number of pixel points in the row of the key area.
[0148] In this embodiment, for each row of pixel points in the key area of the target image, the pixel points in this row can be scanned in the order of increasing abscissa (that is, when the imaging device is on the right side of the vehicle, in the order from left to right; when the imaging device is on the left side of the vehicle, in the order from right to left). The abscissa of the first pixel point in the key area scanned is m1, and the abscissa of the last pixel point in the key area is m G , and the weight of the pixel points in this row in the first direction satisfies that it gradually decreases from the first pixel point in the key area of this row to the last pixel point in the key area of this row.
[0149] In a possible implementation, the weight of the pixel point with abscissa m1 in this row in the first direction can be set to a, and the weight of the pixel point with abscissa m G in this row in the first direction is set to b. According to the formula calculate the weight of the j-th pixel point in the first direction of the pixel points in this row of the key area, where 1 ≤ j ≤ G.
[0150] It should be noted that the number of pixel points included in each row of pixel points in the key area may be different. Therefore, when calculating the weights of the pixel points in different rows in the first direction, the G value may be different.
[0151] Among them, the specific values of a and b can be set according to actual needs. In a possible implementation, a = 1 and b = 0.3.
[0152] The weight of the pixel points in the first direction can also be referred to as the weight of the pixel points in the x-axis direction.
[0153] In some embodiments, the second direction is the y-axis direction of the preset coordinate system;
[0154] The above determination of the weights of the pixel points in the key area in the second direction may include:
[0155] For each column of pixel points in the key area, according to determine the weights of the individual pixel points in this column in the second direction; among them, the pixel points in this column of the key area are sorted in the order of decreasing ordinate, and U l is the weight of the l-th pixel point in this column of the key area in the second direction; n l is the ordinate of the l-th pixel point in this column of the key area; c is the sixth preset weight, which is the weight of the pixel point with ordinate n1 in this column in the second direction; d is the sixth preset weight, which is the weight of the pixel point with ordinate n in this column HThe weight of the pixel points in the second direction; c > d; H is the number of pixel points in this column of pixel points in the key area.
[0156] In this embodiment, for each column of pixel points in the key area of the target image, the pixel points in this row can be scanned in the order from the largest to the smallest ordinate (i.e., in the order from top to bottom). The ordinate of the first pixel point in the key area scanned is n1, and the ordinate of the last pixel point in the key area is n H , and the weight of this column of pixel points in the second direction satisfies that it gradually decreases from the pixel point of the first key area in this column to the pixel point of the last key area in this column.
[0157] In a possible implementation manner, the weight of the pixel point with the ordinate of n1 in this column in the second direction can be set as c, and the weight of the pixel point with the ordinate of n H in the second direction is set as d. According to the formula Calculate the weight of the l-th pixel point in this column of pixel points in the key area in the second direction, 1 ≤ l ≤ H.
[0158] It should be noted that the number of pixel points included in each column of pixel points in the key area may be different. Therefore, when calculating the weights of pixel points in different columns in the second direction, the H value may be different.
[0159] Among them, the specific values of c and d can be set according to actual needs. In a possible implementation manner, c = a = 1, d = b = 0.3.
[0160] The weight of the pixel point in the second direction can also be referred to as the weight of the pixel point in the y-axis direction.
[0161] In a possible implementation manner, in order to ensure that the position weights of all pixel points in the key area are greater than the second preset weight, it is necessary to ensure that e × b × d > the second preset weight.
[0162] In S105, semantic segmentation is performed on the target image according to the first position weight and the second position weight.
[0163] In this embodiment, the existing cross-entropy loss function is improved through the first position weight and the second position weight to obtain an improved cross-entropy loss function. The pre-constructed semantic segmentation model is trained through the improved cross-entropy loss function, and the trained semantic segmentation model is used to perform semantic segmentation on the target image.
[0164] In a possible implementation manner, the improved cross-entropy loss function is L = -∑w q y q log(p q ), where w qis the position weight of the q-th pixel of the target image.
[0165] In a possible implementation, the improved cross-entropy loss function is L = -∑w q log(p q ), where w q is the position weight of the q-th pixel of the target image.
[0166] In a possible implementation, training the pre-constructed semantic segmentation model with the improved cross-entropy loss function and performing semantic segmentation on the target image using the trained semantic segmentation model may include:
[0167] Training the pre-constructed semantic segmentation model with training images to obtain the semantic segmentation results corresponding to the training images, calculating the error between the semantic segmentation results corresponding to the training images and the categories of each pixel pre-annotated in the training images using the improved cross-entropy loss function, and using the backpropagation algorithm to update the parameters of the pre-constructed semantic segmentation model according to the error until the calculated error is less than a certain value or the number of iterative training reaches a certain number of times, obtaining the trained semantic segmentation model, and performing semantic segmentation on the target image using the trained semantic segmentation model to obtain the semantic segmentation results of the target image. In this embodiment, using the improved cross-entropy loss function to train the semantic segmentation model can make the training effect of the semantic segmentation model better, so that semantic segmentation of the target image using the trained semantic segmentation model can achieve a better segmentation effect. Among them, the semantic segmentation model can be an existing model for semantic segmentation, for example, it can be a neural network model, etc.
[0168] In the application specific to the BSD system, when the vehicle is going straight, the BSD system is more concerned about the area closer to the vehicle itself and less concerned about the area far from the vehicle itself; when the vehicle turns right, it is more concerned about the area in the front right of the vehicle head and less concerned about the area near the vehicle tail. In addition, the BSD system pays more attention to categories such as people and vehicles and less attention to the background category, that is, the proportion of the loss of the attention area and category in the total loss should be increased. For difficult-to-segment pixels (usually manifested as false positives and false negatives), the learning of these pixels needs to consider their occurrence frequency, regional importance, etc. Not all difficult-to-segment pixels have an equal impact on the application.
[0169] Based on the above considerations, this embodiment proposes a semantic segmentation method. By obtaining a training sample set and a preset target category, according to the occurrence frequency of the target category pixel points in the training sample set, the key area and non-key area of the target image are determined, and the first position weight of each pixel point in the key area and the second position weight of each pixel point in the non-key area are determined, with the first position weight being greater than the second position weight. Thus, the areas that need to be focused on and the non-focused areas can be distinguished. According to the first position weight and the second position weight, semantic segmentation is performed on the target image, which can improve the segmentation accuracy of the semantic segmentation model for the categories and areas that need to be focused on, and better meet the actual usage requirements.
[0170] In this embodiment, according to the specific characteristics of BSD applications, the loss weights of specific regions and categories are adjusted, enabling the semantic segmentation model to increase the learning of the regions and categories that need to be focused on. For the learning of pixels that are difficult to segment, the probability of their occurrence, category frequency, etc. are comprehensively considered to adjust their losses, which can improve the segmentation accuracy of the semantic segmentation model for the regions, categories that need to be focused on, and pixels that are difficult to segment. It can overcome the problem that the general loss function causes the semantic segmentation model to equally treat image regions and categories, resulting in low segmentation accuracy for the regions, categories that need to be focused on, and difficult-to-segment samples, and can improve the accuracy and effectiveness of the BSD system using semantic segmentation, thereby enhancing the security and comfort of the BSD system.
[0171] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0172] The following is the device embodiment of the present invention. For the details not described in detail herein, reference can be made to the corresponding method embodiment above.
[0173] Figure 5 The structural schematic diagram of the semantic segmentation device provided by the embodiment of the present invention is shown. For the sake of convenience of description, only the parts related to the embodiment of the present invention are shown and are described in detail as follows:
[0174] As Figure 5 shown, the semantic segmentation device 100 includes: a first acquisition module 101, a second acquisition module 102, a division module 103, a position weight determination module 104, and a semantic segmentation module 105.
[0175] The first acquisition module 101 is used to acquire a training sample set; the training sample set includes multiple training images taken by the same imaging device at a preset angle, and each pixel point of each training image has been labeled with a corresponding category;
[0176] The second acquisition module 102 is configured to acquire a preset target category;
[0177] The division module 103 is configured to determine a key area and a non-key area of the target image according to the occurrence frequency of the target category pixel points in the training sample set; the target image is any image captured by the imaging device at a preset angle;
[0178] The position weight determination module 104 is configured to determine a first position weight of each pixel point in the key area and a second position weight of each pixel point in the non-key area, and the first position weight is greater than the second position weight;
[0179] The semantic segmentation module 105 is configured to perform semantic segmentation on the target image according to the first position weight and the second position weight.
[0180] In a possible implementation manner, the position weight determination module 104 is specifically configured to:
[0181] According to the training sample set, determine a set of road surface pixel points closest to the vehicle area in the non-key area;
[0182] Determine a dividing line according to the set;
[0183] According to the dividing line, determine the vehicle area and the non-vehicle area of the non-key area;
[0184] Set the second position weight of each pixel point in the vehicle area to a first preset weight, and set the second position weight of each pixel point in the non-vehicle area to a second preset weight; the first preset weight is less than the second preset weight.
[0185] In a possible implementation manner, the position weight determination module 104 is specifically configured to:
[0186] Perform linear fitting on each pixel point in the set to obtain a fitting line;
[0187] Obtain the abscissas of the intersection points of the fitting line and the upper and lower boundaries of the target image in a preset coordinate system, and denote them as x1 and x2 respectively;
[0188] Determine the line where y = (x1 + x2) / 2 is located as the dividing line.
[0189] In a possible implementation manner, the set of road surface pixel points closest to the vehicle area in the non-key area is the set of road surface pixel points with the smallest abscissa in the preset coordinate system among the pixel points in each row in the non-key area;
[0190] The position weight determination module 104 is specifically configured to:
[0191] Determine the vehicle area as the area within the range of the dividing line and the coordinate axes of the preset coordinate system in the non-key area;
[0192] Among them, when the imaging device is located on the right side of the vehicle, the lower left corner of the target image is used as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower left corner of the target image to the right, and the positive y-axis direction of the preset coordinate system is from the lower left corner of the target image upwards;
[0193] When the imaging device is located on the left side of the vehicle, the lower right corner of the target image is used as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower right corner of the target image to the left, and the positive y-axis direction of the preset coordinate system is from the lower right corner of the target image upwards.
[0194] In a possible implementation manner, the position weight determination module 104 is specifically configured to:
[0195] Determine the target segmentation error degree of each pixel point in the key area according to the training sample set;
[0196] Determine the weight of each pixel point in the key area in the first direction;
[0197] Determine the weight of each pixel point in the key area in the second direction;
[0198] Determine the first position weight of each pixel point in the key area according to the target segmentation error degree of each pixel point in the key area, the weight of each pixel point in the key area in the first direction, and the weight of each pixel point in the key area in the second direction.
[0199] In a possible implementation manner, the first direction and the second direction are different directions;
[0200] The position weight determination module 104 is specifically configured to:
[0201] For each pixel point in the key area, multiply the target segmentation error degree of the pixel point, the weight of the pixel point in the first direction, and the weight of the pixel point in the second direction to obtain the first position weight of the pixel point.
[0202] In a possible implementation manner, the position weight determination module 104 is specifically configured to:
[0203] Divide the training sample set into K training sample subsets on average; each training sample subset contains N training images;
[0204] Select one of the K training sample subsets as the test set in turn from the K training sample subsets, and use the remaining K-1 training sample subsets as the training set to obtain K combinations of test sets and training sets;
[0205] For each combination of the test set and the training set, train a preset semantic segmentation model according to the training set to obtain a trained semantic segmentation model; use the trained semantic segmentation model to perform semantic segmentation on N training images in the test set respectively to obtain semantic segmentation results corresponding to the N training images in the test set respectively; count the number of times of segmentation errors of each pixel point in the key area according to the semantic segmentation results corresponding to the N training images in the test set respectively; divide the number of times of segmentation errors of each pixel point in the key area by N respectively to obtain the test segmentation error degree of each pixel point in the key area corresponding to the combination of the test set and the training set.
[0206] For each pixel point in the key area, take the average value of the test segmentation error degrees of this pixel point corresponding to K combinations of the test set and the training set respectively to obtain the average segmentation error degree of this pixel point.
[0207] Obtain the target segmentation error degree of each pixel point in the key area according to the average segmentation error degree of each pixel point in the key area.
[0208] In a possible implementation manner, the position weight determination module 104 is specifically configured to:
[0209] Select the maximum average segmentation error degree value max and the minimum average segmentation error degree value min from the average segmentation error degrees of each pixel point in the key area;
[0210] According to Calculate the target segmentation error degree of each pixel point in the key area;
[0211] Among them, R i is the target segmentation error degree of the i-th pixel point in the key area; s i is the average segmentation error degree of the i-th pixel point in the key area; e is the third preset weight, which is the weight corresponding to the minimum average segmentation error degree value min; 1 ≤ i ≤ F, and F is the number of pixel points in the key area.
[0212] In a possible implementation manner, the first direction is the x-axis direction of the preset coordinate system;
[0213] When the imaging device is located on the right side of the vehicle, take the lower left corner of the target image as the origin of the preset coordinate system, the right direction from the lower left corner of the target image is the positive x-axis direction of the preset coordinate system, and the upward direction from the lower left corner of the target image is the positive y-axis direction of the preset coordinate system;
[0214] When the imaging device is located on the left side of the vehicle, take the lower right corner of the target image as the origin of the preset coordinate system, the left direction from the lower right corner of the target image is the positive x-axis direction of the preset coordinate system, and the upward direction from the lower right corner of the target image is the positive y-axis direction of the preset coordinate system;
[0215] The position weight determination module 104 is specifically configured to:
[0216] For each row of pixel points in the key area, according to Determine the weights of each pixel point in this row of pixel points in the first direction; wherein, the pixel points in this row of the key area are sorted in ascending order of abscissa, T j is the weight of the jth pixel point in the first direction among the pixel points in this row of the key area; m j is the abscissa of the jth pixel point in this row of pixel points in the key area; is the fourth preset weight, and is the weight of the pixel point with abscissa m1 in this row of pixel points in the first direction; b is the fifth preset weight, and is the weight of the pixel point with abscissa m G in this row of pixel points in the first direction; a > b; G is the number of pixel points in this row of pixel points in the key area.
[0217] In a possible implementation, the second direction is the y-axis direction of the preset coordinate system;
[0218] The position weight determination module 104 is specifically configured to:
[0219] For each column of pixel points in the key area, according to Determine the weights of each pixel point in this column of pixel points in the second direction; wherein, the pixel points in this column of the key area are sorted in descending order of ordinate, U l is the weight of the lth pixel point in the second direction among the pixel points in this column of the key area; n l is the ordinate of the lth pixel point in this column of pixel points in the key area; c is the sixth preset weight, and is the weight of the pixel point with ordinate n1 in this column of pixel points in the second direction; d is the sixth preset weight, and is the weight of the pixel point with ordinate n H in this column of pixel points in the second direction; c > d; H is the number of pixel points in this column of pixel points in the key area.
[0220] In a possible implementation, the partitioning module 103 is specifically configured to:
[0221] For each pixel point of the target image, if the occurrence frequency of the target category pixel point corresponding to this pixel point is greater than the preset frequency threshold, then determine the position where this pixel point is located as the key area, otherwise, determine the position where this pixel point is located as the non-key area.
[0222] Figure 6 is a schematic diagram of the terminal provided by the embodiments of the present invention. As Figure 6As shown, the terminal 11 of this embodiment includes: a processor 110, a memory 111, and a computer program 112 stored in the memory 111 and executable on the processor 110. When the processor 110 executes the computer program 112, it implements the steps in the above-mentioned embodiments of each semantic segmentation method, such as Figure 1 S101 to S105 shown. Alternatively, when the processor 110 executes the computer program 112, it implements the functions of each module / unit in the above-mentioned device embodiments, such as Figure 5 the functions of the modules / units 101 to 105 shown.
[0223] Exemplarily, the computer program 112 may be divided into one or more modules / units. The one or more modules / units are stored in the memory 111 and executed by the processor 110 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 112 in the terminal 11. For example, the computer program 112 may be divided into Figure 5 the modules / units 101 to 105 shown.
[0224] The terminal 11 may be a data processing device or controller of a BSD system, and may also be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal 11 may include, but is not limited to, a processor 110 and a memory 111. Those skilled in the art can understand that Figure 6 this is only an example of the terminal 11, and does not constitute a limitation on the terminal 11. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal may also include input / output devices, network access devices, a bus, etc.
[0225] The so-called processor 110 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0226] The memory 111 may be an internal storage unit of the terminal 11, such as the hard disk or memory of the terminal 11. The memory 111 may also be an external storage device of the terminal 11, such as a plug-in hard disk equipped on the terminal 11, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 111 may also include both the internal storage unit of the terminal 11 and external storage devices. The memory 111 is used to store the computer program and other programs and data required by the terminal. The memory 111 may also be used to temporarily store data that has been output or is to be output.
[0227] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0228] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0229] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0230] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0231] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0232] In addition, each functional unit in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0233] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned semantic segmentation method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0234] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A semantic segmentation method, characterized in that, Including: Obtain a training sample set; the training sample set includes multiple training images taken by the same imaging device at a preset angle, and each pixel point of each training image has been labeled with a corresponding category; Obtain a preset target category; Determine the key area and non-key area of the target image according to the occurrence frequency of the pixel points of the target category in the training sample set; the target image is any image taken by the imaging device at the preset angle; Determine the first position weight of each pixel point in the key area and the second position weight of each pixel point in the non-key area, and the first position weight is greater than the second position weight; Perform semantic segmentation on the target image according to the first position weight and the second position weight.
2. The semantic segmentation method according to claim 1, wherein Determining the second position weight of each pixel point in the non-key area includes: According to the training sample set, determine the set of road surface pixel points closest to the vehicle area in the non-key area; Determine the dividing line according to the set; According to the dividing line, determine the vehicle area and non-vehicle area of the non-key area; Set the second position weight of each pixel point in the vehicle area to a first preset weight, and set the second position weight of each pixel point in the non-vehicle area to a second preset weight; the first preset weight is less than the second preset weight.
3. The semantic segmentation method according to claim 2, wherein The determining the dividing line according to the set includes: Perform linear fitting on each pixel point in the set to obtain a fitting line; Obtain the abscissas of the intersection points of the fitting line and the upper and lower boundaries of the target image in a preset coordinate system, and denote them as x1 and x2 respectively; Determine the line where y = (x1 + x2) / 2 as the dividing line.
4. The semantic segmentation method according to claim 2, characterized in that The set of road surface pixel points closest to the vehicle area in the non-key area is the set of road surface pixel points with the smallest abscissa in the preset coordinate system among the pixel points in each row in the non-key area; The determining the vehicle area of the non-key area according to the dividing line includes: Determine the area within the range of the dividing line and the coordinate axes of the preset coordinate system in the non-key area as the vehicle area; Wherein, when the imaging device is located on the right side of the vehicle, take the lower left corner of the target image as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower left corner of the target image to the right, and the positive y-axis direction of the preset coordinate system is from the lower left corner of the target image upward; When the imaging device is located on the left side of the vehicle, take the lower right corner of the target image as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower right corner of the target image to the left, and the positive y-axis direction of the preset coordinate system is from the lower right corner of the target image upward.
5. The semantic segmentation method according to claim 1, characterized in that Determining the first position weight of each pixel point in the key area includes: According to the training sample set, determine the target segmentation error degree of each pixel point in the key area; Determine the weight of each pixel point in the key area in the first direction; Determine the weight of each pixel point in the key area in the second direction; Determine the first position weight of each pixel point in the key area according to the target segmentation error degree of each pixel point in the key area, the weight of each pixel point in the key area in the first direction, and the weight of each pixel point in the key area in the second direction.
6. The semantic segmentation method according to claim 5, wherein The first direction and the second direction are different directions; The determining the first position weight of each pixel point in the key area according to the target segmentation error degree of each pixel point in the key area, the weight of each pixel point in the key area in the first direction, and the weight of each pixel point in the key area in the second direction includes: For each pixel point in the key area, multiply the target segmentation error degree of this pixel point, the weight of this pixel point in the first direction, and the weight of this pixel point in the second direction to obtain the first position weight of this pixel point.
7. The semantic segmentation method according to claim 5, characterized in that, The determining the target segmentation error degree of each pixel point in the key area according to the training sample set includes: Evenly divide the training sample set into K training sample subsets; each training sample subset contains N training images; Select one of the K training sample subsets as the test set in turn from the K training sample subsets, and use the remaining K - 1 training sample subsets as the training set to obtain K combinations of test sets and training sets; For each combination of test set and training set, train a preset semantic segmentation model according to this training set to obtain a trained semantic segmentation model; use the trained semantic segmentation model to perform semantic segmentation on the N training images in this test set respectively to obtain the semantic segmentation results corresponding to the N training images in this test set respectively; count the number of times of segmentation errors of each pixel point in the key area according to the semantic segmentation results corresponding to the N training images in this test set respectively; divide the number of times of segmentation errors of each pixel point in the key area by N respectively to obtain the test segmentation error degree of each pixel point in the key area corresponding to this combination of test set and training set; For each pixel point in the key area, take the average value of the test segmentation error degrees of this pixel point corresponding to the K combinations of test sets and training sets to obtain the average segmentation error degree of this pixel point; Obtain the target segmentation error degree of each pixel point in the key area according to the average segmentation error degree of each pixel point in the key area.
8. The semantic segmentation method according to claim 7, wherein The obtaining the target segmentation error degree of each pixel point in the key area according to the average segmentation error degree of each pixel point in the key area includes: Select the maximum average segmentation error degree value max and the minimum average segmentation error degree value min from the average segmentation error degrees of each pixel point in the key area; According to Calculate the target segmentation error degree of each pixel point in the key area; Among them, R i is the target segmentation error of the i-th pixel in the key area; s i is the average segmentation error of the i-th pixel in the key area; e is the third preset weight, which is the weight corresponding to the minimum average segmentation error value min; 1≤i≤F, F is the number of pixels in the key area.
9. The semantic segmentation method according to claim 5, characterized in that, The first direction is the x - axis direction of the preset coordinate system; When the imaging device is located on the right side of the vehicle, take the lower - left corner of the target image as the origin of the preset coordinate system, the positive x - axis direction of the preset coordinate system is from the lower - left corner of the target image to the right, and the positive y - axis direction of the preset coordinate system is from the lower - left corner of the target image upwards; When the imaging device is located on the left side of the vehicle, the lower right corner of the target image is taken as the origin of the preset coordinate system, the positive x-axis direction of the preset coordinate system is from the lower right corner of the target image to the left, and the positive y-axis direction of the preset coordinate system is from the lower right corner of the target image upwards; The determining the weights of the pixel points in the first direction of the key area includes: For each row of pixel points in the key area, according to determine the weights of each pixel point in the row of pixel points in the first direction; wherein, the pixel points in the row of the key area are sorted in ascending order of abscissa, T j is the weight of the jth pixel point in the row of pixel points in the key area in the first direction; m j is the abscissa of the jth pixel point in the row of pixel points in the key area; a is the fourth preset weight, which is the weight of the pixel point with abscissa m1 in the row of pixel points in the first direction; b is the fifth preset weight, which is the weight of the pixel point with abscissa m G in the row of pixel points in the first direction; a > b; G is the number of pixel points in the row of pixel points in the key area.
10. The semantic segmentation method according to claim 9, wherein The second direction is the y-axis direction of the preset coordinate system; The determining the weights of the pixel points in the second direction of the key area includes: For each column of pixel points in the key area, according to determine the weights of each pixel point in the second direction among the pixel points in this column; wherein, the pixel points in this column of the key area are sorted in descending order of the ordinate, U l is the weight of the l-th pixel point in the second direction among the pixel points in this column of the key area; n l is the ordinate of the l-th pixel point among the pixel points in this column of the key area; c is the sixth preset weight, which is the weight of the pixel point with the ordinate of n1 in the second direction among the pixel points in this column; d is the sixth preset weight, which is the weight of the pixel point with the ordinate of n H in the second direction among the pixel points in this column; c > d; H is the number of pixel points among the pixel points in this column of the key area.
11. The semantic segmentation method according to any one of claims 1 to 10, characterized in that, The determining the key area and the non-key area of the target image according to the occurrence frequency of the target category pixel points in the training sample set includes: For each pixel point of the target image, if the occurrence frequency of the target category pixel point corresponding to this pixel point is greater than the preset frequency threshold, it is determined that the position where this pixel point is located is the key area, otherwise, it is determined that the position where this pixel point is located is the non-key area.
12. A terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the semantic segmentation method described in any one of claims 1 to 11 above.
13. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the semantic segmentation method described in any one of claims 1 to 11 above.
Citation Information
Patent Citations
A road scene image semantic segmentation method based on importance weighting
CN109740451A