Parking space detection method and device and vehicle
The parking space image is processed through a full convolutional segmentation network, and the multi-modal feature map is obtained for corner point matching, which solves the problems of insufficient corner point detection accuracy and occlusion cut-off in the existing technology, and achieves higher parking space detection accuracy and robustness.
Patent Information
- Application Number
- CN202311656741.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-06
AI Technical Summary
The existing parking space detection method is based on the target detection neural network, and there are problems such as insufficient corner detection accuracy and inaccurate detection caused by corner occlusion or truncation, which affects the robustness of parking space detection.
The parking space image is processed using a full convolutional segmentation network to obtain multi-modal feature maps, including corner point prediction information, occlusion attribute information and truncated attribute information, and corner point matching to improve the accuracy and robustness of parking space detection.
The pixel-level corner point prediction information is obtained through a fully convolutional segmentation network, which improves the accuracy of corner point prediction, and avoids matching failures caused by occlusion or truncation, and significantly improves the robustness of parking space detection.
Smart Images

Figure CN120107922A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of assisted driving technology, and more specifically to a parking space detection method, device and vehicle. Background Art
[0002] Automatic parking is an advanced driver assistance technology that can help drivers park more safely and efficiently in complex or narrow environments. Parking space detection is an important part of automatic parking technology.
[0003] The current parking space detection method is generally based on the target detection neural network, but the target detection neural network only uses the detection branch to regress the corner point information of the parking space, which has the problem of insufficient corner point detection accuracy. In addition, the current parking space detection method does not consider the situation where the corner point position is inaccurate due to truncation or occlusion in the picture, which will lead to unsuccessful corner point matching, thus affecting the robustness of parking space detection.
[0004] Therefore, a new parking space detection solution is needed. Summary of the invention
[0005] According to one aspect of the present application, a parking space detection method is provided, the method comprising: acquiring a parking space image; processing the parking space image using a trained fully convolutional segmentation network to obtain a multimodal feature map; obtaining parking space information based on the multimodal feature map, the parking space information comprising corner point prediction information, the parking space information further comprising corner point occlusion attribute information and / or corner point truncation attribute information; performing corner point matching based on the parking space information to obtain a parking space detection result.
[0006] In one embodiment of the present application, the corner point prediction information includes a candidate point set for each of the four corner points of a parking space, and the corner point matching based on the parking space information includes: based on a preset corner point matching threshold, matching the points in the candidate point set for each of the four corner points with the points in the candidate point sets of other corner points; wherein the points in the candidate point set are divided into non-special corner points and special corner points due to the corner point occlusion attribute information and / or the corner point truncation attribute information, and the special corner points include occluded corner points and / or truncated corner points; wherein the corner point matching threshold preset for the special corner point is greater than the corner point matching threshold preset for the non-special corner point.
[0007] In one embodiment of the present application, the parking space information also includes parking space line position information and parking space line direction information, the parking space line direction information includes the direction from a corner point at the parking space line position to another corner point and the direction from a non-corner point to the another corner point, and the corner point matching based on the parking space information includes: taking the one corner point as the starting point, translating to a non-corner point based on the direction from the one corner point to the other corner point, and corner point matching the non-corner point with the other corner point; when the matching fails, translating to the next non-corner point based on the direction from the non-corner point to the other corner point, and corner point matching the next non-corner point with the other corner point, until the matching is successful.
[0008] In one embodiment of the present application, when the matching fails, the corner point matching based on the parking space information further includes: determining whether the non-corner point is located on the parking space line; when the non-corner point is located on the parking space line, translating to the next non-corner point based on the direction in which the non-corner point points to the other corner point; when the non-corner point is not located on the parking space line, deleting the corner point serving as the starting point as a false alarm point.
[0009] In one embodiment of the present application, the multimodal feature map includes eight parking line direction segmentation maps, and the parking line direction information is obtained based on the parking line direction segmentation maps, wherein every two of the parking line direction segmentation maps respectively record the sine value and cosine value of each point on a parking line, and the direction of each point on the parking line pointing to the other corner point is obtained based on the sine value and the cosine value.
[0010] In one embodiment of the present application, the multimodal feature map includes four parking space line segmentation probability maps, and the parking space line position information is obtained based on binarization processing of the parking space line segmentation probability maps.
[0011] In one embodiment of the present application, obtaining the direction of each point on the parking space line pointing to the other corner point based on the sine value and the cosine value includes: for each point among the points on the parking space line, obtaining the sine value and the cosine value corresponding to the point, calculating the ratio of the sine value and the cosine value, and calculating the inverse tangent function value of the ratio to obtain an angle value, wherein the angle value represents the direction of the point pointing to the other corner point.
[0012] In one embodiment of the present application, the multimodal feature map includes four corner heat maps, and the corner prediction information is obtained based on threshold filtering and maximum pooling of the corner heat maps.
[0013] In one embodiment of the present application, the parking line position information is obtained based on binarizing the parking line segmentation probability map, including: for each point in the parking line segmentation probability map, obtaining the pixel value of the point; when the pixel value is greater than a preset threshold, binarizing the pixel value into a first value; when the pixel value is not greater than the preset threshold, binarizing the pixel value into a second value; obtaining a set of points whose pixel values are binarized into the first value, and obtaining the parking line position information based on the position of each point in the set of points.
[0014] In one embodiment of the present application, the parking space information further includes parking space entrance line occupancy attribute information, and the method further includes: determining a parking space occupancy status based on the parking space entrance line occupancy attribute information, and the parking space occupancy status is used as a part of the parking space detection result.
[0015] In one embodiment of the present application, the multimodal feature map includes an occupancy probability map, and the parking space entrance line occupancy attribute information is obtained based on the parking space line position information and the occupancy probability map.
[0016] In one embodiment of the present application, the parking space entrance line occupancy attribute information is obtained based on the parking space line position information and the occupancy probability map, including: obtaining a set of points corresponding to the parking space entrance line in the occupancy probability map based on the parking space line position information; calculating the average value of the pixel values of all points in the point set; when the average value is greater than the preset threshold, binarizing the average value into a first value; when the average value is not greater than the preset threshold, binarizing the average value into a second value; wherein the first value and the second value are both the parking space entrance line occupancy attribute information, the first value indicates that the parking space is occupied, and the second value indicates that the parking space is not occupied.
[0017] In one embodiment of the present application, the multimodal feature map also includes a corner occlusion probability map and / or a corner truncation probability map, wherein the corner occlusion attribute information is obtained based on the corner occlusion probability map, and the corner truncation attribute information is obtained based on the corner truncation probability map.
[0018] In one embodiment of the present application, acquiring the parking space image includes: acquiring multiple images captured by a vehicle-mounted camera, stitching the multiple images to obtain a stitched image, and transforming the stitched image into a bird's-eye view to obtain the parking space image.
[0019] According to another aspect of the present application, a parking space detection device is provided, the device comprising a memory and a processor, wherein the memory stores a computer executable program executed by the processor, and when the computer executable program is executed by the processor, the processor executes the above-mentioned parking space detection method.
[0020] According to yet another aspect of the present application, a vehicle is provided, the vehicle comprising the above-mentioned parking space detection device.
[0021] According to another aspect of the present application, a storage medium is provided, on which a computer program executed by a processor is stored. When the computer program is executed by the processor, the processor executes the above-mentioned parking space detection method.
[0022] According to another aspect of the present application, a computer program product is provided. When the computer program product is executed by a processor, the processor is enabled to execute the above parking space detection method.
[0023] The parking space detection method, device and vehicle of the present application can obtain pixel-level corner point prediction information through a fully convolutional segmentation network, improve the accuracy of corner point prediction, and can also obtain corner point occlusion attribute information and / or corner point truncation attribute information. For parking space detection scenarios where corner points are occluded or truncated, the problem of failure to successfully achieve matching due to occlusion or truncation of corner points can be avoided, thereby improving the robustness of parking space detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0025] Figure 1 A schematic flow chart of a parking space detection method according to an embodiment of the present application is shown.
[0026] Figure 2 An example diagram showing the structure of a fully convolutional segmentation network used in a parking space detection method according to an embodiment of the present application.
[0027] Figure 3 An exemplary flow chart of a corner point matching process in a parking space detection method according to an embodiment of the present application is shown.
[0028] Figure 4 A schematic diagram showing the translation of matching points based on dense parking space line direction information in a parking space detection method according to an embodiment of the present application is shown.
[0029] Figure 5 An exemplary schematic diagram showing a parking space instance generation result obtained by the parking space detection method according to an embodiment of the present application.
[0030] Figure 6 A schematic structural block diagram of a parking space detection device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of the present application more obvious, the example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application described in the present application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present application.
[0032] Figure 1 FIG. 1 is a schematic flow chart of a parking space detection method 100 according to an embodiment of the present application. Figure 1 As shown, the parking space detection method 100 according to the embodiment of the present application may include the following steps:
[0033] In step S110 , a parking space image is acquired.
[0034] In step S120, the parking space image is processed using the trained fully convolutional segmentation network to obtain a multimodal feature map.
[0035] In step S130, parking space information is obtained based on the multimodal feature map, where the parking space information includes corner point prediction information and corner point occlusion attribute information and / or corner point truncation attribute information.
[0036] In step S140, corner point matching is performed based on the parking space information to obtain a parking space detection result.
[0037] In an embodiment of the present application, after acquiring a parking space image, it is input into a trained full convolution segmentation network, and a multimodal feature map (such as a feature map that can extract corner point prediction information and a feature map that can extract corner point occlusion and / or truncation attributes) can be obtained. Information extraction from the feature map can obtain parking space information. Since the full convolution segmentation network is an application network for instance segmentation, the feature map it outputs can present pixel-level categories. For parking space detection, the feature map it outputs can present the probability of whether each pixel is a parking space corner point (i.e., the four corner points corresponding to the four parking space lines of the parking space), which makes the detection of parking space corner points more accurate. In contrast, when performing parking space detection, the target detection network can only identify the predicted corner point position through the detection box, and cannot achieve the above-mentioned pixel-level classification. Therefore, the present application obtains more accurate corner point prediction information based on the full convolution segmentation network, thereby improving the parking space detection results.
[0038] In addition, in an embodiment of the present application, the feature map output by the fully convolutional segmentation network can be used to extract corner point occlusion attribute information and / or corner point truncation attribute information. Among them, the corner point occlusion attribute information can reflect whether the corner point is occluded; the corner point truncation attribute information can reflect whether the corner point is truncated (corner point truncation means that due to the image shooting angle, a parking space only presents part of the corner points, but not all four corner points, and a corner point may be "truncated" by the boundary line of the image so that it is not presented in the image). Based on such corner point occlusion attribute information and corner point truncation attribute information, the special features of corner points being occluded or truncated are taken into account in the parking space detection process. Therefore, for parking space detection scenarios where corner points are occluded or truncated, it is possible to avoid the failure of corner point matching due to occlusion or truncation, thereby improving the robustness of parking space detection.
[0039] In an embodiment of the present application, the acquisition of the parking space image described in step S110 may include: acquiring multiple images captured by a vehicle-mounted camera, stitching the multiple images to obtain a stitched image, and transforming the stitched image into a bird's-eye view to obtain the parking space image. Specifically, images can be captured by multiple cameras installed in multiple directions of the vehicle, and the captured multiple images are stitched, and the stitched images are transformed into a panoramic stitched image from a bird's-eye view through the Inverse Perspective Mapping (IPM) technology. In an example, the size of the panoramic stitched image is 640x640 pixels, and the multiple directions may include the front, rear, left, and right directions of the vehicle. Stitching the images taken by multiple cameras and performing perspective conversion is conducive to reducing image distortion, thereby improving the accuracy of parking space detection.
[0040] In an embodiment of the present application, the fully convolutional segmentation network used in step S120 may be a bottom-up fully convolutional segmentation network. In one example, the fully convolutional segmentation network includes an encoder-decoder structure. After the parking space image is input as an input image to the encoder of the fully convolutional segmentation network, the encoder can extract the semantic features of the input image through multiple convolution and downsampling operations, and gradually obtain feature maps of different semantic levels. Exemplarily, a harmonic dense connection network (HardBlock) module can be introduced into the encoder to enhance the feature expression capability. The decoder can gradually restore the feature map resolution by upsampling. In addition, the corresponding layers of the encoder and decoder can fuse information through jump connections. Specifically, the jump connection concatenates the feature map before each downsampling in the encoding path with the symmetrical upsampled feature map in the decoding path in the channel dimension. For example, the output feature map xi of the i-th layer in the encoding path is input to the i+1-th layer after downsampling; the j-th layer in the decoding path is the corresponding symmetric layer of the i-th layer, which first upsamples to obtain the feature map yi, and then concatenates xi and yi in the channel dimension, that is, z = Concatenate (x_i, y_i), and z is used as the input of the j-th layer, which incorporates the detail information of the i-th layer. Through the concatenation operation, the low-level detail features of the encoding path can be directly added to the high-level semantic features, so that the decoding path can retain the semantic information of the encoding path while restoring the details, and finally output a multi-modal feature map (i.e., a multi-channel feature map).
[0041] Combine the following Figure 2 The structure of the above-mentioned fully convolutional segmentation network is described exemplarily. Figure 2 As shown in the figure, the input image 3*640*640 (where 3 is the number of channels and 640*640 is the image size) is input through 6 convolution layers, that is, 6 convolution (conv) downsampling operations are performed, and the step size of each downsampling is 2. Therefore, the input image 3*640*640 is subjected to 6 convolution downsampling operations, and the feature maps of 48*640*640, 64*80*80, 96*40*40, 160*20*20, 224*10*10, and 320*5*5 are obtained in turn. Among them, the HardBlock module is added in the convolution downsampling operation, so that the feature expression ability of the obtained feature map is stronger. After that, through 6 deconvolution upsamplings with an upsampling step of 2, feature maps of 272*10*10, 216*20*20, 156*40*40, 110*80*80, 79*160*160, and 19*160*160 are obtained respectively. In addition, the corresponding layers (layers of the same size) can fuse information through skip connections (Concaten), such as Figure 2As shown. Through the concatenation operation, the low-level detail features of the encoding path can be directly added to the high-level semantic features, so that the decoding path can retain the semantic information of the encoding path while restoring the details. The features after each concatenation are convolved to fuse the channel information, and finally output the prediction results of 19 channels (i.e., the feature maps of 19 channels). The model has a clear structure, does not require detection branches, and the backbone network can be replaced by any commonly used network structure (such as ResNet-FPN, UNet, etc.), which can be compatible with different task requirements. In other words, the model is suitable for any segmentation model, has no preset shape for parking spaces, and supports parking space detection in any form.
[0042] In the embodiment of the present application, the multimodal feature map output by the fully convolutional segmentation network includes four corner point heat maps, and the corner point prediction information in step S130 is obtained based on threshold filtering and maximum pooling of the corner point heat maps. Figure 2 The example shown is used to describe, assuming that channels 0 to 3 in the feature map of 19 channels are four corner point heat maps. The four predicted corner points of the parking space to be detected can be defined as corner point 0, corner point 1, corner point 2 and corner point 3, which correspond to the four corner point heat maps of channels 0 to 3, respectively. For each corner point heat map, all points on the heat map can be first threshold filtered. In an example, the threshold (heatmap_threshold) is set to 0.2, and the points with heat values below the threshold are set to -1. Then, the heat map can be subjected to a maximum pooling operation, and non-maximum points are set to -1. Finally, the coordinates of the points with heat values greater than 0 are output as the candidate point set for the corner point. In this way, the candidate point sets of the four corner points can be obtained.
[0043] In an embodiment of the present application, the multimodal feature map output by the fully convolutional segmentation network may also include a corner occlusion probability map, and the corner occlusion attribute information described in step S130 is obtained based on the corner occlusion probability map. Figure 2 The example shown is used to describe the feature map of channel 4 in the feature map of 19 channels. It is assumed that the feature map of channel 4 is a corner point occlusion probability map. In the corner point occlusion probability map, the pixel value of each pixel marks the probability of each pixel being occluded, and the probability can be compared with a threshold (for example, 0.5) to judge: when the probability of a point being occluded is greater than the threshold, the corner point occlusion attribute information of the point can be determined as 1, indicating that the point is occluded; when the probability of a point being occluded is not greater than the threshold, the corner point occlusion attribute information of the point can be determined as 0, indicating that the point is not occluded. In this way, the corner point occlusion attribute information of each point in the candidate point set of each corner point is obtained.
[0044] In an embodiment of the present application, the multimodal feature map output by the fully convolutional segmentation network may also include a corner point truncation probability map, and the corner point truncation attribute information described in step S130 is obtained based on the corner point truncation probability map. Figure 2 The example shown is used to describe that it is assumed that the feature map of channel 5 in the feature map of 19 channels is a corner point truncation probability map. In the corner point truncation probability map, the pixel value of each pixel marks the probability of each pixel being truncated, and the probability can be judged with a threshold (for example, 0.5): when the probability of a point being truncated is greater than the threshold, the corner point truncation attribute information of the point can be determined as 1, indicating that the point is truncated; when the probability of a point being truncated is not greater than the threshold, the corner point truncation attribute information of the point can be determined as 0, indicating that the point is not truncated. In this way, the corner point truncation attribute information of each point in the candidate point set of each corner point is obtained.
[0045] In an embodiment of the present application, the multimodal feature map output by the fully convolutional segmentation network may also include four parking line segmentation probability maps, so as to obtain parking line position information based on binarization of the parking line segmentation probability map. Specifically, obtaining parking line position information based on binarization of the parking line segmentation probability map may include: for each point in the parking line segmentation probability map, obtaining the pixel value of the point; when the pixel value is greater than a preset threshold, binarizing the pixel value into a first value (e.g., 1); when the pixel value is not greater than the preset threshold, binarizing the pixel value into a second value (e.g., 0); obtaining a set of points whose pixel values are binarized into the first value, and obtaining the parking line position information based on the position of each point in the set of points.
[0046] For example, Figure 2 The example shown is used to describe the above. It is assumed that the feature maps of channels 7-10 in the feature map of 19 channels are parking line segmentation probability maps. In the parking line segmentation probability map, the pixel value of each pixel marks the probability that each pixel in the probability map is located on the parking line. Therefore, the parking line segmentation probability map can be binarized to obtain a set of points located on the parking line, thereby predicting the parking line position information. The binarization calculation formula is as follows:
[0047]
[0048] Among them, P l (x, y) is the pixel value of the probability map at position (x, y), t is the threshold, and in this embodiment, t=0.5. l When (x,y)>t, the corresponding pixel point attribute is judged to be a point on the parking space line, and the binary output is 1; otherwise, it is judged to be a point not on the parking space line, and the output is 0.
[0049] In an embodiment of the present application, the multimodal feature map output by the fully convolutional segmentation network may further include an occupancy probability map, which may be used to obtain parking space entrance line occupancy attribute information based on the parking space line position information and the occupancy probability map, thereby determining the parking space occupancy state. Specifically, obtaining the parking space entrance line occupancy attribute information based on the parking space line position information and the occupancy probability map may include: obtaining a set of points corresponding to the parking space entrance line in the occupancy probability map based on the parking space line position information; calculating the average value of the pixel values of all points in the set of points; when the average value is greater than a preset threshold, binarizing the average value into a first value (e.g., 1); when the average value is not greater than the preset threshold, binarizing the average value into a second value (e.g., -1); wherein the first value and the second value are both parking space entrance line occupancy attribute information, the first value indicates that the parking space is occupied, and the second value indicates that the parking space is not occupied.
[0050] In an embodiment of the present application, the multimodal feature map output by the fully convolutional segmentation network may also include eight parking line direction segmentation maps, so as to obtain parking line direction information based on the parking line direction segmentation maps. The parking line direction true value may be defined as: for each point on the parking line, the direction true value of the point is defined as the angle of the unit direction vector of the two endpoints (corner points) of the parking line where the point is located. The calculation formula is:
[0051]
[0052] Where (u A ,v A )、(u B ,v B ) are the coordinates of point A and point B, respectively, |(u B ,v B )-(u A ,v A )| is the Euclidean distance between point A and point B. Finally, the angle θ is taken as the true value of the direction of the point on the corresponding parking space line in the form of sinθ and cosθ.
[0053] For example, Figure 2 The example shown is used for description, assuming that the feature maps of channels 11-18 in the feature map of 19 channels are parking line direction segmentation maps. Then, extracting parking line direction information based on the parking line direction segmentation map may include: processing the output value of the parking line direction segmentation map (channels 11-18), the pixel value output range of each channel is [-1,1], and each two consecutive channels respectively output the sinθ and cosθ values of the direction of a point on the parking line. The direction of each point on the parking line is represented by an angle value θ, and the angle value θ is calculated as the inverse tangent function value of the ratio of sinθ and cosθ, which is expressed by the following formula:
[0054]
[0055] Finally, for a parking line, assuming that the parking line includes two corner points, one corner point is i and the other corner point is j, then the direction of each pixel point on the parking line from i to j is the direction in which each pixel point points to the corner point j. More specifically, the direction information of the parking line is recorded by two parking line direction segmentation maps. For example, for the corner point i of the parking line, in the two parking line direction segmentation maps, the pixel values of point i are sinθi and cosθi, respectively, so the direction information of point i can be calculated, that is, the direction from point i to point j is θi. For a non-corner point m on the parking line, in the two parking line direction segmentation maps, the pixel values of point m are sinθm and cosθm, respectively, so the direction information of point m can be calculated, that is, the direction from point m to point j is θm. In this way, the parking line direction information of the parking line can be obtained, that is, the direction of each point on the parking line pointing to the corner point j.
[0056] After obtaining the above-mentioned parking space information such as corner point prediction information, corner point occlusion attribute information, corner point truncation attribute information, occupancy attribute information, parking space line position information, parking space line direction information based on the multimodal feature map, corner point matching can be performed based on the parking space information to obtain the parking space detection result. It should be understood that the above-mentioned corner point prediction information, corner point occupancy attribute information, corner point truncation attribute information, occupancy attribute information, parking space line position information, and parking space line direction information are not all necessary. For example, in some scenarios, it is not necessary to have parking space line position information, parking space line direction information, and occupancy attribute information. Based on the corner point prediction information, as well as the corner point occupancy attribute information and / or the corner point truncation attribute information, the corner point matching can also be performed to obtain the four corner points of the parking space. The four corner points can also be connected to obtain the parking space line of the parking space, so the parking space detection result can also be obtained. Of course, the more information contained in the parking space information, the more information contained in the parking space detection result, and the information in the parking space information can also verify each other, thereby improving the accuracy of the parking space detection result.
[0057] The following describes the corner point matching based on the parking space information. Corner point matching refers to matching and associating the four corner points of the parking space. Only when the corner points are matched successfully can the parking space detection result be obtained.
[0058] In an embodiment of the present application, as described above, the corner point prediction information may include a candidate point set for each of the four corner points of the parking space. Based on this, the corner point matching based on the parking space information in step S140 may include: matching points in the candidate point set for each of the four corner points with points in the candidate point set for other corner points based on a preset corner point matching threshold; wherein the points in the candidate point set are divided into non-special corner points and special corner points due to corner point occlusion attribute information and / or corner point truncation attribute information, and the special corner points include occluded corner points and / or truncated corner points; wherein the corner point matching threshold preset for the special corner points is greater than the corner point matching threshold preset for the non-special corner points. In this embodiment, by setting a larger matching threshold for special corner points such as occluded corner points and truncated corner points, the robustness of the corner points in invisible positions such as occlusion / truncation can be increased, so that when the parking space line or corner point is in a state where the detection performance is jittered due to occlusion or other reasons, the corner points can be stably matched to generate the parking space.
[0059] In an embodiment of the present application, when the parking space information also includes parking space line position information and parking space line direction information, as mentioned above, the parking space line direction information includes the direction from one corner point on the parking space line position to another corner point and the direction from a non-corner point to another corner point. Based on this, performing corner point matching based on the parking space information may include: taking a corner point as a starting point, translating to a non-corner point based on the direction from one corner point to another corner point, and performing corner point matching on the non-corner point with another corner point; when the matching fails, translating to the next non-corner point based on the direction from the non-corner point to another corner point, and performing corner point matching on the next non-corner point with another corner point, until the matching succeeds. In this embodiment, based on the dense parking space line direction information, even if the direction regression of a single point is inaccurate, the correct parking space corner point can be matched in a cascade manner.
[0060] Combine the following Figure 3 , Figure 4 To describe the above corner point matching process in more detail. The specific strategy is: match corner point i and corner point j based on the dense direction information of the points on the parking line i->j (here i->j refers to the parking line from corner point i to corner point j) (that is, the direction of each point on the parking line pointing to j), the segmentation position information, and the occlusion and truncation attribute information of corner points i and j: starting from a corner point, keep moving forward according to the direction information in the feature map (such as Figure 4 As shown in the figure, the next corner point is matched. Then, the newly matched corner point is used as the starting point to continue matching the next corner point. By analogy, the four corner points on each parking space can be matched one by one.
[0061] In a specific implementation, the above corner point matching process is as follows: Figure 3 shown.
[0062] In step S301: the starting point k is placed at point i. For a starting point k, its position and direction information must first be determined, and then the next corner point is moved to match. Because the current step is to match two corner points i and j, it is necessary to start from the position of corner point i and find the corresponding corner point j. Therefore, in this example, the point k is initialized through the information of corner point i, specifically:
[0063] x k =x i
[0064] y k =y i
[0065] sinθ k = sinθ i
[0066] cosθ k = cosθ i
[0067] Among them, (x k ,y k ) is the position information of point k; (sinθ k ,cosθ k ) is the direction information of point k; (x i ,y i ) is the position information of corner point i, which is obtained by extracting the thermal map of corner point i in this example; (sinθ i ,cosθ i ) is the direction information of corner point i, which is obtained by extracting the segmentation map of parking space line i->j in this example.
[0068] Step S302: Point k is translated by t pixels according to the parking space line direction result i->j on point k. The schematic diagram of translation of matching points based on dense parking space line direction information is as follows: Figure 4 Specifically, according to the parking space line direction information of the current point k, k is translated by t pixels, and the coordinates of point k after the movement are:
[0069] x k =x k +t·cosθ k
[0070] y k =y k +t·sinθ k
[0071] In one example, t is set to 2 pixels.
[0072] In step S303: calculate the distance between point k and all j points, and set a matching threshold for all j points according to their occlusion and truncation properties (the matching threshold can generally be preset, and after calculating the distance here, matching can be performed directly based on the preset matching threshold). The distance between point k and all points in the set of corner points j is calculated by the Euclidean distance method, and the distance calculation formula is:
[0073]
[0074] In this example, the set of corner points j is obtained through the extraction results of the heat map of corner point i. For key points whose positions cannot be accurately predicted due to occlusion, truncation, etc., by increasing the matching threshold of key points under occlusion and truncation, multiple key points in the same parking space can be correctly matched when the positions of the key points are uncertain. In this example, for the set of corner points j, different matching thresholds are set for each j point according to the extraction results of the segmentation map of its occlusion and truncation attributes. Exemplarily, the distance matching threshold of non-occluded and truncated corner point j is 3 pixels, and the matching threshold of occluded or truncated corner point j is 5 pixels.
[0075] In step S304: determine whether the distance between k and all j points is greater than the matching threshold of point j: if the results are all greater than the matching threshold, continue to determine whether the parking line segmentation score of point k is greater than the threshold (that is, determine whether the probability of point k being on the parking line is greater than the preset threshold through the parking line segmentation probability map. The threshold here is the segmentation score threshold, which is not the same as the distance matching threshold): if the segmentation score is greater than the threshold, it means that point k has moved out of the parking line, then point i is a false alarm and is deleted; if the segmentation score is less than the threshold, continue to return to S302 to try matching; if the distance between a certain k and j point is less than the matching threshold of the point, then point i is successfully matched with point j, and return to S301 to place the starting point at point j.
[0076] More specifically, based on the matching threshold, it is determined whether point k is successfully matched with corner point j. Specifically, it is determined whether the distance between point k and all j points is greater than the matching threshold of point j. If the dist(k,j) value of point k and all points in the set of corner point j is greater than the matching threshold, that is, the dist(k,j) value of point k and all non-occluded and truncated corner points j is greater than 3 pixels away and the dist(k,j) value of point k with all occluded or truncated points j is greater than 5 pixels away, it is determined that there is no matching pair at this time.
[0077] For the current situation where there is no match, it can be further determined whether point k is on the parking line or has been removed from the parking line range. In this example, the segmentation score of the current k position is obtained through the extraction result of the parking line i->j segmentation probability map, and the segmentation score threshold is set to 0.2. If the segmentation score of point k is less than 0.2, it means that point k has moved out of the potential parking line range, and point i is judged to be a false alarm and deleted. If the segmentation score of point k is greater than the threshold, point k is still on the parking line, and point k is continued to be moved according to the direction information of the current point k to try to match. At this time, the direction information of point k is obtained through the extraction result of the parking line i->j direction segmentation map.
[0078] If the distance between point k and a corner point j after a certain movement is less than the matching threshold of the point, then corner point i and corner point j are successfully matched, and corner point i and corner point j and parking space line i->j are added to a parking space instance result, and continue to match the next corner point with j as the starting point.
[0079] In step S305: after matching all corner points 0->1, 1->2, 2->3, 3->0, complete parking space instance information can be obtained. Specifically, the parking space instance information can include four corner points, the segmentation result and direction of the parking space line, and the parking space occupancy attribute. After generating the parking space instance, the correctness of the parking space instance can be verified according to the matching situation, and parking space instances with less than 3 matching points are excluded (for example, i can match j, but j cannot match the next corner point, and i cannot match other corner points, then only two corner points are matched, and the matching result at this time is invalid).
[0080] In general, when the instance-level parking spaces are generated through the above post-processing logic, the dense parking space line direction information can be used to match the correct parking space corner points in a cascaded manner even when the single-point direction regression is inaccurate. At the same time, when there are corner points in a parking space instance whose positions cannot be accurately predicted due to occlusion, truncation, etc., this module improves the robustness of the matching results by setting a larger matching threshold for the occluded and truncated corner points.
[0081] Based on the above description, the parking space detection method according to the embodiment of the present application can obtain pixel-level corner point prediction information through a fully convolutional segmentation network, improve the accuracy of corner point prediction, and can also obtain corner point occlusion attribute information and / or corner point truncation attribute information. For parking space detection scenarios where corner points are occluded or truncated, the problem of failure to successfully achieve matching due to occlusion or truncation can be avoided, thereby improving the robustness of parking space detection. In addition, the parking space detection method according to the embodiment of the present application can use dense parking space line direction information to match the correct parking space corner points in a cascade manner even when the direction regression of a single point is inaccurate, thereby further improving the robustness of parking space detection.
[0082] Through experimental verification, the angle accuracy (Precision) and recall rate (Recall) of the parking space detection method of this application are both above 95%; at the same time, the angle accuracy of the three visible parking lines is above 99%; for the occlusion and truncation attributes, the prediction accuracy rate also reaches 99%; this information enables the corner points to be better associated through the direction of the parking line. For scenes where parking spaces are blocked due to multiple groups of vehicles occupying, difficult scenes such as night, after rain, motion blur, irregular parking spaces, etc., the parking space detection method of this application can also achieve an accuracy rate and recall rate of 90% and 83% in parking space assessment, respectively. It is further explained that the method of this embodiment can stably match corner points to generate parking spaces when the parking space lines or corner points are in a state of jitter in detection performance due to occlusion or other reasons.
[0083] Figure 5 FIG. 2 is an exemplary schematic diagram showing a parking space instance generation result obtained by the parking space detection method according to an embodiment of the present application. Figure 5 As shown, for occlusion and truncation scenarios, the method of the present application can still well predict the complete parking space results; in addition, the existing post-processing methods based on parking space types do not support uncommon parking space types such as curved parking spaces. When encountering undefined parking space types, such methods will have parking space line or direction errors. The model proposed by the method of the present application supports the detection of parking spaces of any form including curved parking spaces by outputting dense parking space line directions and the robustness of corner point matching strategies.
[0084] Finally, the algorithm based on this application has been successfully deployed on a real vehicle platform and tested on a real vehicle. The test results show that the actual processing rate of the algorithm on the vehicle side reaches 20FPS, and the parking space detection accuracy reaches 98%, meeting the performance requirements of actual applications.
[0085] Combine the following Figure 6 A parking space detection device according to another aspect of the present application is described. Figure 6 FIG. 6 shows a schematic structural block diagram of a parking space detection device 600 according to an embodiment of the present application. Figure 6 As shown, the parking space detection device 600 includes a memory 610 and a processor 620, wherein: the memory 610 stores a computer executable program executed by the processor 620, and when the computer executable program is executed by the processor 620, the processor 620 executes the above parking space detection method 100. Those skilled in the art can understand the structure and specific operation of each module in the parking space detection device 600 according to the embodiment of the present application in combination with the above contents, and for the sake of brevity, they are not described here.
[0086] According to yet another aspect of the present application, a vehicle is provided, which includes the parking space detection device 600 according to the embodiment of the present application described above.
[0087] In addition, the present application also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the processor executes the parking space detection method according to the embodiment of the present application described above. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0088] In addition, the present application also provides a computer program product, which, when executed by a processor, enables the processor to execute the above-mentioned parking space detection method 100 .
[0089] Based on the above description, the parking space detection method, device and vehicle according to the embodiment of the present application can obtain pixel-level corner point prediction information through a fully convolutional segmentation network to improve the accuracy of corner point prediction, and can also obtain corner point occlusion attribute information and / or corner point truncation attribute information. For parking space detection scenarios where corner points are occluded or truncated, the problem of failure to successfully achieve matching due to occlusion or truncation of corner points can be avoided, thereby improving the robustness of parking space detection. In addition, according to the parking space detection method according to the embodiment of the present application, dense parking space line direction information can be used to match the correct parking space corner points in a cascade manner even when the direction regression of a single point is inaccurate, thereby further improving the robustness of parking space detection.
[0090] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present application to this. Those of ordinary skill in the art may make various changes and modifications therein without departing from the scope and spirit of the present application. All these changes and modifications are intended to be included within the scope of the present application as required by the appended claims.
[0091] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0092] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0093] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.
[0094] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present application should not be interpreted as reflecting the following intention: the claimed application requires more features than the features clearly stated in each claim. More specifically, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with features less than all the features of a single disclosed embodiment. Therefore, the claims following the specific embodiment are hereby explicitly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.
[0095] It will be understood by those skilled in the art that, except for mutually exclusive features, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this specification may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0096] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present application and form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.
[0097] The various component embodiments of the present application can be implemented in hardware, or implemented in software modules running on one or more processors, or implemented in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all functions of some modules according to the embodiments of the present application. The application can also be implemented as a part or all of a program (e.g., a computer program and a computer program product) for performing the method described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0098] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim listing a number of parking space detection devices, several of these parking space detection devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names.
[0099] The above is only a specific implementation or description of a specific implementation of the present application, and the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. The protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A parking space detection method, It is characterized in that The method comprises: Acquire parking space images; Processing the parking space image using a trained fully convolutional segmentation network to obtain a multimodal feature map; Obtaining parking space information based on the multimodal feature map, the parking space information including corner point prediction information, and the parking space information further including corner point occlusion attribute information and / or corner point truncation attribute information; Corner point matching is performed based on the parking space information to obtain a parking space detection result.
2. The method according to claim 1, It is characterized in that The corner point prediction information includes a set of candidate points for each of the four corner points of the parking space, and performing corner point matching based on the parking space information includes: Based on a preset corner point matching threshold, matching points in the candidate point set of each corner point of the four corner points with points in the candidate point sets of other corner points; The points in the candidate point set are divided into non-special corner points and special corner points due to the corner point occlusion attribute information and / or the corner point truncation attribute information, and the special corner points include occluded corner points and / or truncated corner points; The corner point matching threshold preset for the special corner point is greater than the corner point matching threshold preset for the non-special corner point.
3. The method according to claim 1 or 2, It is characterized in that The parking space information further includes parking space line position information and parking space line direction information, wherein the parking space line direction information includes a direction from a corner point on the parking space line position to another corner point and a direction from a non-corner point to the another corner point, and performing corner point matching based on the parking space information includes: Taking the one corner point as a starting point, translating to a non-corner point based on the direction in which the one corner point points to the other corner point, and performing corner point matching between the non-corner point and the other corner point; When the matching fails, the non-corner point is translated to the next non-corner point based on the direction in which the non-corner point points to the other corner point, and the next non-corner point is matched with the other corner point until the matching succeeds.
4. The method according to claim 3, It is characterized in that When the matching fails, performing corner point matching based on the parking space information further includes: Determining whether the non-corner point is located on the parking space line; When the non-corner point is located on the parking space line, translating to the next non-corner point based on the direction in which the non-corner point points to the other corner point; When the non-corner point is not located on the parking space line, the one corner point serving as the starting point is deleted as a false alarm point.
5. The method according to claim 3, It is characterized in that The multimodal feature map includes eight parking line direction segmentation maps, and the parking line direction information is obtained based on the parking line direction segmentation maps, wherein every two of the parking line direction segmentation maps respectively record the sine value and cosine value of each point on a parking line, and the direction of each point on the parking line pointing to the other corner point is obtained based on the sine value and the cosine value.
6. The method according to claim 5, It is characterized in that The obtaining the direction of each point on the parking space line pointing to the other corner point based on the sine value and the cosine value includes: For each of the points on the parking space line, obtain the sine value and the cosine value corresponding to the point, calculate the ratio of the sine value to the cosine value, and calculate the inverse tangent function value of the ratio to obtain an angle value, wherein the angle value represents the direction of the point pointing to the other corner point.
7. The method according to claim 3, It is characterized in that The multimodal feature map includes four parking space line segmentation probability maps, and the parking space line position information is obtained based on binarization processing of the parking space line segmentation probability maps.
8. The method according to claim 7, It is characterized in that The parking space line position information is obtained based on binarization processing of the parking space line segmentation probability map, including: For each point in the parking space line segmentation probability map, obtaining a pixel value of the point; When the pixel value is greater than a preset threshold, binarizing the pixel value into a first value; When the pixel value is not greater than the preset threshold, binarizing the pixel value into a second value; A set of points whose pixel values are binarized into the first value is obtained, and the parking space line position information is obtained based on the position of each point in the set of points.
9. The method according to claim 2, It is characterized in that The multimodal feature map includes four corner point heat maps, and the corner point prediction information is obtained based on threshold filtering and maximum pooling of the corner point heat maps.
10. The method according to claim 3, It is characterized in that The parking space information also includes parking space entrance line occupancy attribute information, and the method further includes: A parking space occupancy state is determined based on the parking space entrance line occupancy attribute information, and the parking space occupancy state serves as a part of the parking space detection result.
11. The method according to claim 10, It is characterized in that The multimodal feature map includes an occupancy probability map, and the parking space entrance line occupancy attribute information is obtained based on the parking space line position information and the occupancy probability map.
12. The method according to claim 11, It is characterized in that Obtaining the parking space entrance line occupancy attribute information based on the parking space line position information and the occupancy probability map includes: Acquire a set of points on the parking space entrance line corresponding to the parking space entrance line in the occupancy probability map based on the parking space line position information; Calculate the average of the pixel values of all points in the set of points; When the average value is greater than the preset threshold, binarizing the average value into a first value; When the average value is not greater than the preset threshold, binarizing the average value into a second value; The first value and the second value are both the parking space entrance line occupation attribute information, the first value represents that the parking space is occupied, and the second value represents that the parking space is not occupied.
13. The method according to claim 1, It is characterized in that The multimodal feature map also includes a corner occlusion probability map and / or a corner truncation probability map, wherein the corner occlusion attribute information is obtained based on the corner occlusion probability map, and the corner truncation attribute information is obtained based on the corner truncation probability map.
14. The method according to claim 13, It is characterized in that The corner point occlusion attribute information is obtained based on the corner point occlusion probability map, including: The corner point occlusion probability map is the probability of each pixel being occluded, and when the probability of the pixel being occluded is greater than a predetermined threshold, the corner point occlusion attribute information of the pixel is determined to be 1; The corner point truncation attribute information is obtained based on the corner point truncation probability map, including: The corner point truncation probability map is the probability of each pixel being truncated. When the probability of the pixel being truncated is greater than a predetermined threshold, the corner point truncation attribute information of the pixel is determined to be 1.
15. The method according to claim 1, It is characterized in that The step of acquiring the parking space image comprises: A plurality of images captured by a vehicle-mounted camera are acquired, the plurality of images are spliced to obtain a spliced image, and the spliced image is transformed into a bird's-eye view to obtain the parking space image.
16. A parking space detection device, It is characterized in that The device includes a processor and a memory, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the processor executes the parking space detection method according to any one of claims 1 to 15.
17. A vehicle, It is characterized in that The vehicle includes the parking space detection device according to claim 16.