A method, apparatus and equipment for detecting three-dimensional parking spaces
By combining the generation of panoramic top-down views with a three-dimensional parking space detection network, the problem of low detection accuracy of three-dimensional parking spaces is solved, enabling high-precision automatic parking of three-dimensional parking spaces, avoiding scratch accidents, and improving detection robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-03
AI Technical Summary
Automatic parking assist systems cannot accurately detect and recognize parking space frames in multi-level parking spaces, causing vehicles to fail to park automatically and increasing the risk of scratches and collisions.
A panoramic top-down view is generated based on images captured by multiple cameras. A 3D parking space detection network is used to adjust the predicted parking space bounding box and the features of the predicted target entrance point. By combining the parking space depth value and corner features, the target parking space bounding box of the 3D parking space is accurately determined.
It improves the accuracy of automated parking space detection, ensuring that vehicles can safely and automatically park in automated parking spaces, avoiding scratches and reducing reliance on additional sensors. It is also robust for parking space detection under different conditions.
Smart Images

Figure CN121260039B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent transportation, and in particular to a method, apparatus and equipment for detecting three-dimensional parking spaces. Background Technology
[0002] Automated Parking Assist (APA) is an intelligent driving assistance system that integrates perception, decision-making, path planning, and vehicle control. It enables the vehicle to automate the entire process from a stationary position to parking or leaving a parking space without the driver needing to directly control the steering wheel, accelerator, or brakes. APA primarily relies on the perception of the vehicle's environment, the identification and judgment of parking spaces, dynamic path planning, and comprehensive scheduling and coordination of the vehicle's underlying control system to ultimately achieve a precise, safe, and smooth automated parking experience. APA systems are widely used.
[0003] As the number of vehicles continues to rise, urban roads are becoming increasingly congested, and the number of available parking spaces can no longer meet the rapid increase in the number of vehicles. As a result, more and more parking lots are starting to use multi-level parking spaces.
[0004] If an automatic parking assist system is used to achieve automatic parking in a multi-level parking space, considering the narrow width of the multi-level parking space, if the automatic parking assist system cannot accurately detect and identify the parking space frame of the multi-level parking space, that is, if the accuracy of the parking space frame is low, then it cannot meet the automatic parking requirements of the multi-level parking space. In other words, it cannot automatically park the vehicle into the multi-level parking space, and scratches are likely to occur during automatic parking. Summary of the Invention
[0005] This application provides a method for detecting three-dimensional parking spaces, the method comprising:
[0006] A panoramic top-down view is generated based on images captured by multiple cameras around the vehicle;
[0007] The panoramic top view is input into the three-dimensional parking space detection network to obtain the predicted parking space frame of the three-dimensional parking space and the target predicted corner features of the target predicted entrance point of the three-dimensional parking space.
[0008] The parking space depth value of the three-dimensional parking space is determined based on the predicted parking space frame; wherein, the parking space depth value represents the distance between the entrance line of the predicted parking space frame and the inner line of the predicted parking space frame;
[0009] The predicted parking space frame is adjusted based on the parking space depth value and the target predicted corner feature to obtain a candidate parking space frame, and the target parking space frame of the three-dimensional parking space is determined based on the candidate parking space frame.
[0010] This application provides a three-dimensional parking space detection device, the device comprising:
[0011] The generation module is used to generate a panoramic top-down view based on images captured by multiple cameras around the vehicle;
[0012] The processing module is used to input the panoramic top view into the three-dimensional parking space detection network to obtain the predicted parking space frame of the three-dimensional parking space and the target predicted corner feature of the target predicted entrance point of the three-dimensional parking space.
[0013] The determination module is used to determine the parking space depth value of the three-dimensional parking space based on the predicted parking space frame, wherein the parking space depth value represents the distance between the entrance line and the inner line of the predicted parking space frame; adjust the predicted parking space frame based on the parking space depth value and the target predicted corner point features to obtain a candidate parking space frame; and determine the target parking space frame of the three-dimensional parking space based on the candidate parking space frame.
[0014] This application provides an electronic device, including: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the three-dimensional parking space detection method of the above example.
[0015] This application provides a computer program product, which includes a computer program that, when executed by a processor, implements the three-dimensional parking space detection method described above.
[0016] This application provides a machine-readable storage medium storing machine-executable instructions that can be executed by a processor; wherein the processor is used to execute the machine-executable instructions to implement the three-dimensional parking space detection method of the above example.
[0017] As can be seen from the above technical solutions, in this embodiment, after obtaining the predicted parking space frame of the three-dimensional parking space, the predicted parking space frame can be adjusted based on the parking space depth value and the target predicted corner point features to obtain candidate parking space frames. Based on the candidate parking space frames, the target parking space frame of the three-dimensional parking space is determined, accurately detecting and identifying the parking space frame of the three-dimensional parking space. The accuracy of the parking space frame is higher, meeting the automatic parking requirements of three-dimensional parking spaces, enabling vehicles to be automatically parked in three-dimensional parking spaces, and avoiding scratches during automatic parking. By adjusting the predicted parking space frame, the inherent characteristics of three-dimensional parking spaces are addressed, eliminating detection errors caused by the height difference of the three-dimensional parking space plane and image coordinate transformation, resulting in higher detection accuracy of the parking space frame. This improves the detection accuracy of three-dimensional parking spaces and is suitable for automatic parking in various states such as parking, entering, and exiting the parking space. It ensures the continuity and stability of parking space information under different states, improves the robustness of parking space detection in dynamic scenarios (such as vehicle movement), eliminates the dependence on additional sensors (such as ultrasonic radar), and achieves high-precision detection based solely on visual data. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a three-dimensional parking space detection method according to one embodiment of this application;
[0019] Figure 2 This is a flowchart illustrating a three-dimensional parking space detection method according to one embodiment of this application;
[0020] Figure 3A This is a schematic diagram of four fisheye images in one embodiment of this application;
[0021] Figure 3B This is a schematic diagram of a panoramic top view generated based on a fisheye image in one embodiment of this application;
[0022] Figure 3C This is a schematic diagram showing the position and angle of the predicted parking space frame, the entrance line of the predicted parking space frame, and the entrance point of the predicted parking space frame in one embodiment of this application.
[0023] Figure 4A This is a schematic diagram of the structure of a three-dimensional parking space detection network in one embodiment of this application;
[0024] Figure 4B This is a schematic diagram of a predicted parking space frame in one embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the parking space stabilization and following process in one embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the structure of a three-dimensional parking space detection device according to one embodiment of this application;
[0027] Figure 7This is a hardware structure diagram of an electronic device according to one embodiment of this application. Detailed Implementation
[0028] This application proposes a three-dimensional parking space detection method, which can be applied to electronic devices, such as in-vehicle devices. The in-vehicle devices support automatic parking assistance systems, and the three-dimensional parking space detection method is implemented based on these systems. See also... Figure 1 The diagram shown is a flowchart of the method, which may include:
[0029] Step 101: Generate a panoramic top-down view based on images captured by multiple cameras around the vehicle.
[0030] Step 102: Input the panoramic top view into the three-dimensional parking space detection network to obtain the predicted parking space bounding box of the three-dimensional parking space and the target predicted corner features of the target predicted entrance point of the three-dimensional parking space.
[0031] Step 103: Determine the parking space depth value of the three-dimensional parking space based on the predicted parking space frame; wherein, the parking space depth value can represent the distance between the entrance line of the predicted parking space frame and the inner line of the predicted parking space frame.
[0032] Step 104: Adjust the predicted parking space frame based on the parking space depth value and the target predicted corner feature to obtain the candidate parking space frame, and determine the target parking space frame of the three-dimensional parking space based on the candidate parking space frame.
[0033] For example, inputting a panoramic top-down view into a three-dimensional parking space detection network to obtain target predicted corner features of the predicted entrance point of the three-dimensional parking space may include, but is not limited to: inputting the panoramic top-down view into the three-dimensional parking space detection network to obtain predicted corner features of multiple predicted entrance points; wherein, for each predicted entrance point, the predicted corner feature of the predicted entrance point includes the offset distance between the predicted entrance point and a reference entrance point; the three-dimensional parking space may include two entrance points, and the reference entrance point is the entrance point that is closest to the predicted entrance point among the two entrance points. The pixel position of the predicted entrance point is determined based on the pixel position of the reference entrance point and the offset distance. Based on the distance between the pixel position of the entrance point of the predicted parking space bounding box and the pixel position of each predicted entrance point, the predicted entrance point corresponding to the minimum distance is determined as the target predicted entrance point, and the predicted corner features of the target predicted entrance point are determined as the target predicted corner features.
[0034] For example, the entrance point of the predicted parking space frame is located at the intersection of the multi-level parking space and the ground; the target predicted corner point features include the side line angle between the target straight line and the side line of the predicted parking space frame, and the target straight line is the line connecting the target predicted entrance point and the entrance point of the predicted parking space frame.
[0035] The predicted parking space bounding box can include an entrance point and an inner point. The predicted parking space bounding box is adjusted based on the parking space depth value and the target predicted corner point features to obtain candidate parking space bounding boxes. This adjustment may include, but is not limited to: determining the position adjustment amount based on the parking space depth value and the side line angle; determining the pixel position of the inner point of the predicted parking space bounding box based on the pixel position of the entrance point and the position adjustment amount; and generating candidate parking space bounding boxes based on the pixel positions of the entrance point and the inner points of the predicted parking space bounding box.
[0036] For example, determining the target parking space frame of a multi-level parking space based on candidate parking space frames may include, but is not limited to: determining the candidate parking space frame as the target parking space frame of the multi-level parking space. Alternatively, if the distance between the candidate parking space frame and the historical parking space frame is greater than a distance threshold, and / or the difference in orientation angle between the candidate parking space frame and the historical parking space frame is greater than an angle threshold, then the historical parking space frame can be determined as the target parking space frame of the multi-level parking space; wherein, the historical parking space frame is the previously determined target parking space frame; if the distance between the candidate parking space frame and the historical parking space frame is not greater than a distance threshold, and the difference in orientation angle between the candidate parking space frame and the historical parking space frame is not greater than an angle threshold, then the candidate parking space frame and the historical parking space frame can be weighted and fused to obtain the target parking space frame of the multi-level parking space.
[0037] For example, a three-dimensional parking space detection network includes at least an initial feature extractor, a global feature extractor, a local feature extractor, a parking space bounding box prediction network, and a corner feature prediction network. Based on this, a panoramic top-down view is input into the three-dimensional parking space detection network to obtain the predicted parking space bounding box and the predicted corner features of the target predicted entry point of the three-dimensional parking space, which may include, but are not limited to:
[0038] The panoramic top-down view is input into an initial feature extractor, which extracts features from the panoramic top-down view to obtain initial features. These initial features are then input into a global feature extractor, which extracts global features from the panoramic top-down view based on the initial features. The global features are then input into a parking space bounding box prediction network, which processes the global features to obtain predicted parking space bounding boxes for the multi-level parking space, and outputs the predicted parking space bounding boxes. Finally, the initial features are input into a local feature extractor, which extracts local features from the panoramic top-down view based on the initial features. These local features are then input into a corner feature prediction network, which processes the local features to obtain predicted corner features for the predicted entrance points of the multi-level parking space, and outputs the predicted corner features.
[0039] For example, in the training process of a three-dimensional parking space detection network, a panoramic top-down view of the samples can be input into the network to be trained, obtaining sample parking space bounding boxes and sample corner features of the sample entrance points of the three-dimensional parking spaces. A heatmap is generated based on the sample parking space bounding boxes, including heat values of multiple pixels. For each pixel in the heatmap, the closer the pixel is to a sample location point, the higher its heat value. The sample location points include the entrance point and inner points of the sample parking space bounding boxes. Candidate pixels matching the sample location points are selected from all pixels in the heatmap, and a first loss value is determined based on the heat values of the candidate pixels. A second loss value is determined based on the sample corner features and the labeled corner feature tags. A target loss value is determined based on the first and second loss values, and the three-dimensional parking space detection network is adjusted based on the target loss value to obtain the trained three-dimensional parking space detection network.
[0040] For example, determining the first loss value based on the thermal values of candidate pixels may include, but is not limited to: determining the predicted value of a candidate pixel based on its thermal value and pixel position; and determining the first loss value based on the distance between the predicted value of the candidate pixel and the pixel position of the sample location. The predicted value of the candidate pixel is determined using the following formula: The first loss value is determined using the following formula: ;in, This represents the predicted value of the candidate pixel. Indicates the pixel position of the candidate pixel. This represents the thermal value of the candidate pixel. Represents the first loss value; where, This indicates the pixel position of the sample location point. This represents the distance between the predicted value of a candidate pixel and the pixel position of the sample location.
[0041] For example, when a panoramic top view of a sample is input into the 3D parking space detection network to be trained, the 3D parking space detection network outputs multiple sample corner features corresponding to the sample entry point; for each sample corner feature, the sample corner feature includes the offset distance between the sample entry point and the reference entry point.
[0042] Determining a second loss value based on sample corner features and labeled corner feature tags can include: for each sample corner feature corresponding to the sample entry point, determining the confidence level of the sample corner feature based on the offset distance in the sample corner feature; and determining the second loss value based on the confidence level of each sample corner feature, the sample corner feature corresponding to the maximum confidence level, and the corner feature tag. The confidence level of the sample corner feature is determined using the following formula: ; Indicates the confidence level. This indicates the lateral offset distance between the sample entry point and the reference entry point. This indicates the longitudinal offset distance between the sample entry point and the reference entry point. This indicates the configured parameter value.
[0043] As can be seen from the above technical solutions, in this embodiment, after obtaining the predicted parking space frame of the three-dimensional parking space, the predicted parking space frame can be adjusted based on the parking space depth value and the target predicted corner point features to obtain candidate parking space frames. Based on the candidate parking space frames, the target parking space frame of the three-dimensional parking space is determined, accurately detecting and identifying the parking space frame of the three-dimensional parking space. The accuracy of the parking space frame is higher, meeting the automatic parking requirements of three-dimensional parking spaces, enabling vehicles to be automatically parked in three-dimensional parking spaces, and avoiding scratches during automatic parking. By adjusting the predicted parking space frame, the inherent characteristics of three-dimensional parking spaces are addressed, eliminating detection errors caused by the height difference of the three-dimensional parking space plane and image coordinate transformation, resulting in higher detection accuracy of the parking space frame. This improves the detection accuracy of three-dimensional parking spaces and is suitable for automatic parking in various states such as parking, entering, and exiting the parking space. It ensures the continuity and stability of parking space information under different states, improves the robustness of parking space detection in dynamic scenarios (such as vehicle movement), eliminates the dependence on additional sensors (such as ultrasonic radar), and achieves high-precision detection based solely on visual data.
[0044] The technical solutions described above in the embodiments of this application will be explained below in conjunction with specific application scenarios.
[0045] This application proposes a method for detecting three-dimensional parking spaces. In a three-dimensional parking space scenario, it proposes a parking space detection, recognition, and following method based on a panoramic vision-assisted system, achieving stable three-dimensional parking space detection and following. Based on a panoramic top-down view, an end-to-end deep neural network (such as a three-dimensional parking space detection network) is used to infer the parking space bounding box, entrance line (i.e., the boundary line between the parking space and the ground area), and entrance corner points of the three-dimensional parking space. By detecting the parking space bounding box, entrance line, and entrance corner points, the parking space position and orientation are inferred, thus determining the final position and orientation of the three-dimensional parking space. Post-processing logic is then used to infer and stabilize the inner corner points of the three-dimensional parking space, stabilizing and following the three-dimensional parking space result. This achieves both accurate localization of the three-dimensional parking space on the image and stable following of the three-dimensional parking space. In this embodiment, a panoramic top-down view is used, and inference is performed directly in the world coordinate system, effectively avoiding losses caused by coordinate transformation and errors caused by image distortion. Predicting parking spaces and corner points based on a single network effectively reduces computational complexity.
[0046] This application proposes a method for detecting three-dimensional parking spaces; see [link to relevant documentation]. Figure 2The diagram shows a flowchart of a three-dimensional parking space detection method, which may include processes such as generating a panoramic top view, inferring a three-dimensional parking space detection network, inferring the position and orientation of the three-dimensional parking space, and stabilizing and following the parking space.
[0047] For the process of generating a panoramic top view, four fisheye images taken simultaneously from the front, back, left, and right can be used to generate the corresponding panoramic top view after distortion removal and coordinate transformation.
[0048] For the inference process of the 3D parking space detection network, a panoramic top-down view can be input into the network. The network then performs inference to obtain the parking space bounding box and the predicted corner features of the predicted entrance point (also known as the parking space entrance corner point). The 3D parking space detection network can also be called a PSD (ParkSlot Detection) network.
[0049] For the reasoning process of the location and orientation of the multi-level parking space, the parking space frame can be post-processed based on the prediction results of the multi-level parking space detection network (i.e., the parking space frame and the predicted corner features of the predicted entrance point of the multi-level parking space) and combined with the inherent features of the multi-level parking space to obtain the optimized parking space frame and orientation of the multi-level parking space.
[0050] To address the stabilization and following process of parking spaces, considering the instability of single-frame detection results, the optimized parking space frame and orientation of the three-dimensional parking space can be recorded to facilitate the stabilization and following of the parking space frame in the subsequent automatic parking stage.
[0051] The following describes the processes of panoramic top-down view generation, three-dimensional parking space detection network inference, three-dimensional parking space position and orientation inference, parking space stabilization and following, etc., in conjunction with specific embodiments.
[0052] First, the panoramic top-down view generation process. During the panoramic top-down view generation process, a panoramic top-down view can be generated based on images captured by multiple cameras (e.g., four cameras) around the vehicle.
[0053] For example, multiple cameras can be installed around a vehicle, such as a camera at the front of the vehicle, a camera at the rear of the vehicle, a camera on the left side of the vehicle, and a camera on the right side of the vehicle. In this way, images can be captured by multiple cameras. These cameras can be fisheye cameras, and the images captured by fisheye cameras are fisheye images, meaning that multiple fisheye images can be obtained at the same time.
[0054] After obtaining multiple fisheye images, a panoramic top-down view can be generated based on them. For example, based on the installation positions of the four fisheye cameras mounted around the vehicle body and the intrinsic and extrinsic parameters of the fisheye cameras, distortion correction processing can be performed on the fisheye images. Then, according to the relationship between the image coordinate system and the real-world coordinate system, an inverse projection transformation method is used to perform a perspective transformation on the distortion-corrected image, realizing the conversion from a perspective view to a top-down view. Finally, an image overlap region fusion algorithm is used to stitch together the multiple top-down views to obtain a panoramic top-down view of the vehicle body. See also Figure 3A The image shown is a schematic diagram of four fisheye images. (See attached image.) Figure 3B The image shown is a schematic diagram of a panoramic top-down view generated based on four fisheye images.
[0055] For example, an Around View Monitor (AVM) system is a driver assistance system that uses cameras around the vehicle to capture environmental images and stitches them together using image processing technology to create a 360-degree panoramic view. AVM systems can eliminate blind spots for the driver, provide 2D / 3D view switching, and pedestrian and obstacle detection functions, and are widely used in scenarios such as parking and low-speed driving to improve safety. Based on this, an AVM system can acquire multiple fisheye images from multiple cameras and generate a panoramic top-down view based on these images. This embodiment does not limit the method of generating this panoramic top-down view.
[0056] Second, regarding the reasoning process of the three-dimensional parking space detection network.
[0057] During the inference process of the 3D parking space detection network, a panoramic top-down view can be input into the network to obtain the predicted parking space bounding box and the predicted corner features of multiple predicted entry points. For example, using the panoramic top-down view as input, the 3D parking space detection network can output the predicted parking space bounding box and the predicted corner features of multiple predicted entry points. The predicted parking space bounding box can include two entry points and two inner points. The line connecting the two entry points of the predicted parking space bounding box is called the entry line of the predicted parking space bounding box, and the line connecting the two inner points of the predicted parking space bounding box is called the inner line of the predicted parking space bounding box.
[0058] Because of the height difference between the 3D parking space plane and the ground plane (the 3D parking space plane is 10cm to 40cm higher than the ground plane), there may be some distortion in the panoramic top-down view. Therefore, when the 3D parking space detection network outputs the predicted parking space bounding box, both entrance points of the predicted parking space bounding box are located at the boundary between the 3D parking space and the ground. This also ensures that the entrance line of the predicted parking space bounding box is located at the boundary between the 3D parking space and the ground. Based on this, detection errors caused by height information can be avoided. Furthermore, since the entrance line of the predicted parking space bounding box, the two entrance points of the predicted parking space bounding box, and the vehicle are on the same plane, the world coordinates can be accurately calculated based on the conversion ratio between pixel distance and actual distance in the panoramic top-down view. See also... Figure 3C The diagram shows the position and angle of the predicted parking space frame, the entrance line of the predicted parking space frame, and the entrance point of the predicted parking space frame.
[0059] In one possible implementation, see Figure 4A The diagram shown illustrates the structure of a 3D parking space detection network (also known as a 3D parking space detection model). This network can include an initial feature extractor, a global feature extractor, a local feature extractor, a parking space bounding box prediction network, and a corner feature prediction network. Of course, Figure 4A This is merely an example; the structure of this 3D parking space detection network is not limited, as long as it can output predicted parking space bounding boxes and predicted corner features of multiple predicted entry points. The training process for the 3D parking space detection network can be found in subsequent embodiments. Based on the trained 3D parking space detection network, after inputting a panoramic top-down view, the network can output predicted parking space bounding boxes and predicted corner features of multiple predicted entry points.
[0060] For example, a panoramic top-down view can be input into an initial feature extractor, which then extracts features from the panoramic top-down view to obtain its initial features. For instance, the initial feature extractor is a network model used for feature extraction, such as a CNN, RNN, or Transformer. After the panoramic top-down view is input into the initial feature extractor, it can automatically learn to extract the features most beneficial to the task, which are denoted as the initial features of the panoramic top-down view.
[0061] For example, after obtaining the initial features, the initial features are input to a global feature extractor, which then extracts global features of the panoramic top-down view based on the initial features. For instance, the global feature extractor is a network model used to extract global features. After the initial features are input to the global feature extractor, the global feature extractor processes the initial features to obtain the global features of the panoramic top-down view.
[0062] Global features are features that describe the entire image. They are obtained from all (or most) pixels in the image. For example, global features can be color histograms, shape descriptors, etc. Global features can be extracted using fully connected layers, global pooling, self-attention, and other techniques.
[0063] For example, after obtaining the initial features, the initial features are input to a local feature extractor, which then extracts local features of the panoramic top-down view based on the initial features. For instance, the local feature extractor is a network model used to extract local features. After the initial features are input to the local feature extractor, the extractor processes them to obtain the local features of the panoramic top-down view.
[0064] Local features refer to features calculated based on local image patches; they are fine-grained, small-scale features within an image. Local features can be scale-invariant feature transforms (SIFT), local binary patterns (LBP), etc. Convolutional kernels, sliding windows, local attention, and other methods can be used to extract local features.
[0065] For example, after obtaining the global features, these features are input into a parking space bounding box prediction network. The network processes these features to obtain predicted parking space bounding boxes for the three-dimensional parking spaces, and then outputs these predicted bounding boxes. For instance, the parking space bounding box prediction network is a network model used to output parking space bounding boxes. The network structure of this network is not limited; it only needs to be able to output parking space bounding boxes. After inputting the global features into the network, it can predict the parking space bounding boxes in a panoramic top-down view (denoted as predicted parking space bounding boxes) based on the global features and output these predicted bounding boxes.
[0066] For example, 8 scalars can be used. , To represent the predicted parking space frame. See also Figure 4B The image shown is a schematic diagram of the predicted parking space frame. and To predict the two entrance points of the parking space frame, and the entrance points With the entrance point The line connecting the two parking spaces is called the entrance line of the predicted parking space frame. and To predict the two inner points of the parking space frame, and and The line connecting the two sides is called the inner line of the predicted parking space frame, which can also be called the tail line.
[0067] In summary, the parking space bounding box prediction network can output a predicted parking space bounding box. Based on this predicted parking space bounding box, we can know the pixel positions of the two entrance points of the predicted parking space bounding box, the pixel positions of the two inner points of the predicted parking space bounding box, the entrance line of the predicted parking space bounding box, and the inner line of the predicted parking space bounding box.
[0068] For example, after obtaining the local features, these features are input into a corner feature prediction network. The network processes these local features to obtain the predicted corner features of the predicted entrance point of the automated parking space, and then outputs these predicted corner features. For instance, the corner feature prediction network is a network model used to output corner features; its structure is not limited, as long as it can output corner features. After inputting the local features, the network can predict the corner features of the panoramic top-down view (denoted as predicted corner features) based on these local features and output these predicted corner features. The predicted corner features are specific to the predicted entrance point (i.e., the current predicted point) and can represent the corner features of the current predicted point.
[0069] For example, when outputting predicted corner features, the corner feature prediction network can output predicted corner features for multiple predicted entry points, such as predicted corner features for K predicted entry points, where K is a positive integer greater than 1, indicating that all K predicted entry points could potentially be actual entry points. Clearly, a multi-level parking space only has two actual entry points, see [reference needed]. Figure 4B As shown, and These correspond to the two actual entry points. Therefore, two predicted entry points need to be selected from the K predicted entry points as target predicted entry points. These two target predicted entry points correspond to the actual entry points of the automated parking spaces. The selection method is detailed in subsequent embodiments.
[0070] For example, for each predicted entry point, the corner feature prediction network can output multiple predicted corner features for that entry point. Assuming the predicted parking space bounding box is a rectangle within an M*N region (e.g., if the panoramic top-down view is a 320*320 image, and the 3D parking space detection network performs a 4x downsampling during processing, then the predicted parking space bounding box is a rectangle within an 80*80 region; or, if the panoramic top-down view is a 320*320 image, then the predicted parking space bounding box is a rectangle within a 320*320 region), then the corner feature prediction network outputs M*N predicted corner features for that predicted entry point. The first predicted corner feature is the predicted corner feature of the predicted entry point for the first pixel position (i.e., the first row and first column of the M*N region). The second predicted corner feature is the predicted corner feature of the predicted entry point for the second pixel position (i.e., the first row and second column), and so on. The M*Nth predicted corner feature is the predicted corner feature of the predicted entry point for the M*Nth pixel position (i.e., the Mth row and Nth column).
[0071] For example, for each predicted corner feature of the predicted entry point, the predicted corner feature may include the offset distance (offset) between the predicted entry point and the reference entry point, the side line angle between the target line and the side line of the predicted parking space frame, and the confidence level corresponding to the predicted corner feature.
[0072] For example, a multi-level parking space may include two entrance points (i.e., actual entrance points), and a reference entrance point is the entrance point that is closest to the predicted entrance point. The offset distance between the predicted entrance point and the reference entrance point may include... and , This indicates the lateral offset distance between the predicted entry point and the reference entry point. This represents the vertical offset distance between the predicted entry point and the reference entry point. Based on this, for the first predicted corner feature, the first pixel position is used as the reference entry point. This represents the lateral offset distance between the predicted entry point and the first pixel position. This represents the vertical offset distance between the predicted entry point and the first pixel position. For the second predicted corner feature, the second pixel position is used as the reference entry point, and so on.
[0073] For example, regarding the sideline angle between the target straight line and the sideline of the predicted parking space frame, the target straight line is the line connecting the predicted entrance point to the entrance point of the predicted parking space frame, while the sideline is the line connecting the entrance point of the predicted parking space frame to an inner point. The sideline angle can be a first sideline angle or a second sideline angle. The first sideline angle represents the sideline angle between the first target straight line and the first sideline, and the second sideline angle represents the sideline angle between the second target straight line and the second sideline. See also... Figure 4B As shown, the first target line is the line connecting the predicted entrance point and the first entrance point of the predicted parking space frame. The first entrance point is... The first side line is the line connecting the first entry point and the first inner point. The first inner point is... The second target line is the line connecting the predicted entry point and the second entry point of the predicted parking space frame. The second entry point is... The second side line is the line connecting the second entrance point and the second inner point. The second inner point is... For example, if the reference entrance point corresponds to the first entrance point of the predicted parking space frame, then the side line angle can be the first side line angle; or, if the reference entrance point corresponds to the second entrance point of the predicted parking space frame, then the side line angle can be the second side line angle.
[0074] The first sideline angle can include and , This represents the X-direction angle between the straight line of the first target and the first side line. This represents the Y-direction angle between the first target line and the first side line. The angle of the second side line can include... and , This indicates the X-direction angle between the straight line of the second target and the second side line. This indicates the Y-direction angle between the second target line and the second side line.
[0075] For example, regarding the confidence level corresponding to a predicted corner feature, the confidence level represents the degree of credibility of that predicted corner feature. If the predicted entry point corresponds to M*N predicted corner features, then the predicted entry point corresponds to M*N confidence levels, with each predicted corner feature corresponding to one confidence level; that is, there is a one-to-one correspondence between predicted corner features and confidence levels. Obviously, the closer the reference entry point corresponding to a certain predicted corner feature (such as the 100th pixel position) is to the actual entry point of the automated parking space, the higher the confidence level corresponding to that predicted corner feature.
[0076] In summary, for each predicted corner feature of the predicted entry point, this predicted corner feature can be represented by the following five parameters, namely, , and This represents the offset distance (i.e., the amount of offset) between the predicted entry point and the reference entry point. and It represents the side line angle (i.e., the side line angle of the parking space where the entrance point is located; if the reference entrance point corresponds to the first entrance point of the predicted parking space frame, it represents the first side line angle; if the reference entrance point corresponds to the second entrance point of the predicted parking space frame, it represents the second side line angle). This represents the confidence level of the predicted corner feature; the closer to the actual entrance point of the multi-level parking space, the higher the confidence level.
[0077] Third, the reasoning process regarding the location and orientation of multi-level parking spaces.
[0078] For example, since inference is performed on a panoramic top-down view, the 3D parking space detection network may output multiple predicted parking space bounding boxes (multiple predicted parking space bounding boxes correspond to multiple entrance lines, i.e., a one-to-one correspondence between predicted parking space bounding boxes and entrance lines) and multiple predicted entrance points (such as predicted corner features of multiple predicted entrance points). In this way, the predicted parking space bounding box corresponding to the vehicle can be selected from the multiple predicted parking space bounding boxes, indicating that the vehicle needs to drive into the 3D parking space corresponding to this predicted parking space bounding box. Hereafter, we will take a single predicted parking space bounding box as an example.
[0079] Based on this, it is necessary to select two predicted entry points (e.g., two predicted entry points) from all predicted entry points that match the predicted parking space bounding box. These two predicted entry points correspond to the two entry points of the predicted parking space bounding box. For example, the predicted entry points can be matched with the entry lines of the predicted parking space bounding box. The matching relationship between the points and lines is determined based on the distance between the points and lines and a prior threshold. Then, discrete points with no matching relationship are deleted to ensure that a single entry line of a predicted parking space bounding box matches at most two predicted entry points.
[0080] For example, the following steps can be used to select the predicted entry point that matches the predicted parking space frame:
[0081] Step S11: Select candidate predicted corner features from all predicted corner features.
[0082] After inputting the panoramic top-down view into the 3D parking space detection network, the network can output predicted corner features for multiple predicted entry points. For each predicted entry point, the network outputs multiple predicted corner features. Each predicted corner feature corresponds to a confidence level. Based on this, candidate predicted corner features can be selected from all predicted corner features according to their confidence levels. For example, for each predicted corner feature, if the confidence level is greater than a preset confidence threshold, then the predicted corner feature is considered a candidate predicted corner feature; if the confidence level is not greater than the preset confidence threshold, then the predicted corner feature is not considered a candidate predicted corner feature. In this way, candidate predicted corner features can be selected from all predicted corner features of all predicted entry points, and there can be multiple candidate predicted corner features.
[0083] Step S12: For each candidate predicted corner feature (taking one candidate predicted corner feature as an example), the candidate predicted corner feature includes the offset distance between the predicted entry point and the reference entry point. The pixel position of the predicted entry point is determined based on the pixel position of the reference entry point and the offset distance. The three-dimensional parking space includes two entry points, and the reference entry point is the entry point that is closest to the predicted entry point among the two entry points.
[0084] For example, when using a pixel location as a reference entry point—that is, assuming this pixel location is the closest entry point to the predicted entry point in the parking space—then the candidate predicted corner feature includes the offset distance between the predicted entry point and the reference entry point. Clearly, the pixel location of the reference entry point can be determined, such as pixel location (a, b). For instance, when using the pixel location in the 3rd row and 3rd column of an M*N region as the reference entry point, then the pixel location of the reference entry point could be (3, 3).
[0085] Because the offset distance between the predicted entry point and the reference entry point includes and , Indicates the lateral offset distance. This represents the vertical offset distance; therefore, based on the pixel position (a, b) of the reference entry point and this offset distance ( , This can determine the pixel position of the predicted entry point. For example, the pixel position of the predicted entry point can be (a+). b+ ).
[0086] Step S13: Based on the distance between the pixel position of the predicted parking space frame's entrance point and the pixel position of each predicted entrance point, determine the predicted entrance point corresponding to the minimum distance as the target predicted entrance point for matching the predicted parking space frame, and determine the predicted corner feature of the target predicted entrance point as the target predicted corner feature.
[0087] For example, for each candidate predicted corner feature, the pixel position of the predicted entry point can be obtained based on that candidate predicted corner feature. A predicted entry point may correspond to one pixel position (i.e., this predicted entry point corresponds to one candidate predicted corner feature), a predicted entry point may correspond to multiple pixel positions (i.e., this predicted entry point corresponds to multiple candidate predicted corner features), or a predicted entry point may not correspond to any pixel position (i.e., this predicted entry point does not correspond to any candidate predicted corner feature).
[0088] The predicted parking space bounding box can include two entry points, the pixel position of the first entry point is The pixel position of the second entry point is Based on this, for the first entry point, the distance between the pixel position of the first entry point and the pixel position of the predicted entry point can be calculated. After obtaining the distance between the pixel position of the first entry point and the pixel position of each predicted entry point, the predicted entry point corresponding to the minimum distance can be used as the target predicted entry point, and the candidate predicted corner feature corresponding to the minimum distance can be used as the target predicted corner feature. For the second entry point, the distance between the pixel position of the second entry point and the pixel position of the predicted entry point can be calculated. After obtaining the distance between the pixel position of the second entry point and the pixel position of each predicted entry point, the predicted entry point corresponding to the minimum distance can be used as the target predicted entry point, and the candidate predicted corner feature corresponding to the minimum distance can be used as the target predicted corner feature.
[0089] For example, after obtaining the predicted parking space frame, the depth of the multi-level parking space can be preliminarily determined based on the detection results of the predicted parking space frame. That is, the parking space depth value of the three-dimensional parking space is determined based on the predicted parking space frame. This parking space depth value can represent the distance between the entrance line of the predicted parking space frame and the inner line of the predicted parking space frame.
[0090] For example, see Figure 4B As shown, the predicted parking space frame's entrance line is the entrance point. With the entrance point The line connecting the two points predicts the inner line of the parking space frame as the inner point. With the inner point The lines connecting them. Based on this, the entry point is known. , entry point Inner point and inner point The pixel position can be used to determine the center point of the predicted parking space frame's entrance line and the center point of the predicted parking space frame's inner line. The distance between these two center points can be the parking space depth value. .
[0091] For example, for multi-level parking spaces, the predicted parking space bounding box output by the multi-level parking space detection network may not be accurate. Based on this, in this embodiment, the predicted parking space bounding box can be adjusted based on the parking space depth value and the target predicted corner feature to obtain a candidate parking space bounding box, thereby obtaining an accurate candidate parking space bounding box.
[0092] For example, the predicted parking space bounding box can include an entrance point and an inner point. The position adjustment amount can be determined based on the parking space depth value and the side line angle in the predicted corner feature. The pixel position of the inner point of the predicted parking space bounding box is determined based on the pixel position of the entrance point and the position adjustment amount. On this basis, a candidate parking space bounding box is generated based on the pixel positions of the entrance point and the inner point of the predicted parking space bounding box. That is, the candidate parking space bounding box is a parking space bounding box composed of the entrance point and the inner point.
[0093] For example, for the first entry point of the predicted parking space frame, such as the entry point... The sideline angle in the target prediction corner feature corresponding to the first entry point is the first sideline angle, which is the angle between the first target line and the first sideline. The first target line is the line connecting the predicted entry point and the first entry point. The first sideline is the line between the first entry point and the first inner point. The lines connecting them.
[0094] Based on this, the pixel position of the first inner point can be determined using the following formula (1):
[0095] Formula (1)
[0096] In formula (1), This indicates the depth of the parking space. This represents the X-direction angle of the first side edge line. This indicates the position adjustment amount in the X direction. This represents the Y-direction angle of the first side edge. This indicates the position adjustment amount in the Y direction. This indicates the pixel position of the first entry point of the predicted parking space frame. This indicates the pixel position of the first inner point of the predicted parking space bounding box, i.e., using pixel position. Replace the pixel position of the first inner point .
[0097] For example, for the second entrance point of the predicted parking space frame, such as the entrance point The sideline angle in the target prediction corner feature corresponding to the second entry point is the second sideline angle, which is the angle between the second target line and the second sideline. The second target line is the line connecting the predicted entry point and the second entry point. The second sideline is the line between the second entry point and the second inner point. The lines connecting them.
[0098] Based on this, the pixel position of the second inner point can be determined using the following formula (2):
[0099] Formula (2)
[0100] In formula (2), This indicates the depth of the parking space. This represents the X-direction angle of the second sideline. This indicates the position adjustment amount in the X direction. This represents the Y-direction angle of the second side edge. This indicates the position adjustment amount in the Y direction. This indicates the pixel position of the predicted second entry point of the parking space frame. This indicates the pixel position of the second inner point of the predicted parking space frame, i.e., using pixel position. Replace the pixel position of the second inner point .
[0101] Based on the pixel position of the first entry point of the predicted parking space frame Predict the pixel position of the second entry point of the parking space frame. Predict the pixel position of the first inner point of the parking space frame. and the pixel position of the second inner point of the predicted parking space frame The region formed by the aforementioned four pixel locations is called the candidate parking space bounding box. At this point, the location of the 3D parking space has been successfully deduced, and this candidate parking space bounding box is taken as the location of the 3D parking space.
[0102] Furthermore, the orientation angle of the multi-level parking space can be determined based on the sideline angles of the two entrance points of the predicted parking space frame, thus successfully inferring the orientation angle of the multi-level parking space. For example, the orientation angle of the multi-level parking space includes the first sideline angle corresponding to the first entrance point and the second sideline angle corresponding to the second entrance point.
[0103] Fourth, regarding parking space stability and the following process.
[0104] After obtaining the candidate parking space frames of the multi-level parking space, the candidate parking space frames can be determined as the target parking space frames of the multi-level parking space, and the target parking space frames can be output. Parking functions can be implemented based on the target parking space frames.
[0105] Alternatively, considering the instability of inference results from a single-frame panoramic top-down view, especially during the parking phase when a vehicle enters the parking space, causing the parking space features to be obscured by the vehicle, leading to certain errors in the network inference results, we can use previously recorded historical data to stabilize the parking space results, thereby further improving the accuracy and precision of three-dimensional parking space detection. For example, the target parking space frame of a three-dimensional parking space can be determined based on candidate parking space frames and historical parking space frames, where the historical parking space frame is the previously determined target parking space frame. For instance, the target parking space frame for time 2 is determined based on the candidate parking space frame and historical parking space frame at time 2 (the target parking space frame at time 1 is used as the historical parking space frame for time 2); the target parking space frame for time 3 is determined based on the candidate parking space frame and historical parking space frame (the target parking space frame at time 2), and so on.
[0106] For example, see Figure 5 The diagram shown illustrates the parking space stabilization and following process, which may include:
[0107] Step 501: Determine candidate parking space frames based on a single-frame panoramic top-down view. This process is described in the above embodiment.
[0108] Step 502: Determine whether the distance between the candidate parking space frame and the historical parking space frame is greater than the distance threshold (configured according to actual needs, such as 0.5 meters). If yes, proceed to step 503; otherwise, proceed to step 504.
[0109] For example, a first distance can be calculated between the first entry point of the candidate parking space frame and the first entry point of the historical parking space frame, and a second distance can be calculated between the second entry point of the candidate parking space frame and the second entry point of the historical parking space frame. If the first distance and / or the second distance is greater than a distance threshold, then step 503 is executed; if the first distance is not greater than the distance threshold and the second distance is not greater than the distance threshold, then step 504 is executed.
[0110] Step 503: Determine the historical parking space frame as the target parking space frame for the three-dimensional parking space. That is, the candidate parking space frame determined based on the single-frame panoramic top view is invalid and is not used for output.
[0111] Step 504: Determine whether the difference in orientation angle between the candidate parking space frame and the historical parking space frame is greater than the angle threshold (configured according to actual needs, such as 10 degrees). If yes, proceed to step 503; otherwise, proceed to step 505.
[0112] For example, the orientation angle of a candidate parking space frame may include the angle of the first side line corresponding to the first entrance point and the angle of the second side line corresponding to the second entrance point. The orientation angle of a historical parking space frame may include the angle of the first side line corresponding to the first entrance point and the angle of the second side line corresponding to the second entrance point.
[0113] The difference between the first side line angle of the candidate parking space frame and the first side line angle of the historical parking space frame can be calculated (i.e., the difference between the two). The difference between the second side line angle of the candidate parking space frame and the second side line angle of the historical parking space frame can also be calculated. If the difference between the first and / or second orientation angles is greater than an angle threshold, then step 503 is executed. If the difference between the first and second orientation angles is not greater than an angle threshold, then step 505 is executed.
[0114] Step 505: Perform weighted fusion on the candidate parking space frame and the historical parking space frame to obtain the target parking space frame of the three-dimensional parking space. During weighted fusion, the weighting coefficient of the candidate parking space frame can be greater than the weighting coefficient of the historical parking space frame, or it can be less than the weighting coefficient of the historical parking space frame.
[0115] For example, the pixel positions of the first entrance point of the candidate parking space frame and the first entrance point of the historical parking space frame are weighted and fused to obtain the pixel position of the first entrance point of the target parking space frame. The pixel positions of the second entrance point of the candidate parking space frame and the second entrance point of the historical parking space frame are weighted and fused to obtain the pixel position of the second entrance point of the target parking space frame. The pixel positions of the first inner points of the candidate parking space frame and the first inner points of the historical parking space frame are weighted and fused to obtain the pixel position of the first inner point of the target parking space frame. The pixel positions of the second inner points of the candidate parking space frame and the second inner points of the historical parking space frame are weighted and fused to obtain the pixel position of the second inner point of the target parking space frame. Thus, the target parking space frame of the three-dimensional parking space can be obtained.
[0116] In one possible implementation, before performing the above operations, the automated parking space detection network can be pre-trained. The training process for the automated parking space detection network may include:
[0117] Step S21: Obtain the 3D parking space detection network to be trained and a panoramic top view of the samples.
[0118] The structure of the three-dimensional parking space detection network can be found in [reference]. Figure 4A As shown, the 3D parking space detection network can include an initial feature extractor, a global feature extractor, a local feature extractor, a parking space bounding box prediction network, and a corner feature prediction network. The sample panoramic top-down view is the panoramic top-down view used to train the 3D parking space detection network; it is the training data. There can be multiple sample panoramic top-down views; the following example uses only one sample panoramic top-down view.
[0119] Step S22: Input the panoramic top view of the sample into the three-dimensional parking space detection network to be trained to obtain the sample parking space frame of the three-dimensional parking space and the sample corner features of the sample entrance point of the three-dimensional parking space.
[0120] For example, after inputting a sample panoramic top-down view into a 3D parking space detection network, the network can process the data using an initial feature extractor, a global feature extractor, a local feature extractor, a parking space bounding box prediction network, and a corner feature prediction network. The parking space bounding box prediction network can output the parking space bounding boxes of the 3D parking space, recording the bounding boxes obtained during training as sample parking space bounding boxes. The corner feature prediction network can output sample corner features for multiple sample entry points (recording the entry points obtained during training as sample entry points). For each sample entry point, the corner feature prediction network can output multiple sample corner features for that entry point. Assuming the sample parking space bounding box is a rectangle within an M*N region, the corner feature prediction network can output M*N sample corner features for that entry point. Each sample corner feature for that entry point can include information such as offset distance (offset amount), side line angle, and confidence level.
[0121] Step S23: Generate a heatmap based on the sample parking space frame. The heatmap can include the heat values of multiple pixels. For each pixel in the heatmap, the closer the pixel is to the sample location point, the greater the heat value of the pixel. The sample location point includes the entrance point and the inner point of the sample parking space frame.
[0122] For example, assuming the sample parking space bounding box is a rectangle within an M*N area (e.g., if the sample panoramic top view is a 320*320 image, and the 3D parking space detection network performs a 4x downsampling during processing, then the sample parking space bounding box is a rectangle within an 80*80 area; or, if the sample panoramic top view is a 320*320 image, then the sample parking space bounding box is a rectangle within a 320*320 area), then the heatmap can include the heat values of M*N pixels. That is, the size of the heatmap is M*N, and the value of each pixel in the heatmap is a heat value.
[0123] The sample parking space frame includes four pixel locations (two entrance points and two inner points). Therefore, these four pixel locations can be found on the heatmap and recorded as sample location points. That is, the sample location points include the two entrance points and two inner points of the sample parking space frame. Based on this, for each pixel on the heatmap, the closer the pixel is to a sample location point, the higher its heat value. For example, if the pixel is closer to the sample location point... The closer the distance, the higher the thermal value of that pixel. The larger the sample location point It can be the first entry point Second entrance point First inner point Second inner point .
[0124] Furthermore, the sum of the heat values of each pixel in the heatmap can be 1, such as... For example, when k is 1, This represents the heat value of the first pixel (row 1, column 1), when k is 2. This represents the heat value of the second pixel (row 1, column 2), and so on. W represents the total number of pixels in the heatmap. In this way, we can obtain the heat values of a total of W pixels.
[0125] Step S24: Select candidate pixels that match the sample location from all pixels in the heatmap.
[0126] For example, the sample parking space frame includes 4 pixel positions (the pixel positions of the two entrance points and the pixel positions of the two inner points). These 4 pixel positions are recorded as sample position points. Therefore, the pixel points that match these 4 sample position points can be selected from all the pixels in the heat map. The selected 4 pixel points are called candidate pixel points, that is, 4 candidate pixel points are selected from all the pixels in the heat map.
[0127] Step S25: Determine the first loss value based on the thermal values of the candidate pixels.
[0128] For example, the location representation of parking space boxes can be considered as a Dirac distribution. However, in real-world scenarios, parking space boxes often have ambiguity and uncertainty, such as unclear boundaries, making it difficult to model the Dirac distribution. Furthermore, a Gaussian distribution can be used to model parking space boxes, but it cannot capture the location distribution because the actual distribution of parking space boxes may be arbitrary and flexible, not requiring the symmetry of a Gaussian distribution. Therefore, in this embodiment, the location representation of parking space boxes is adjusted, using an arbitrary distribution to model them, and the probability distribution is directly learned in continuous space without introducing prior knowledge similar to a Gaussian distribution. The 3D parking space detection network predicts the probability distribution vector from the center point to the boundary of the parking space box, which can effectively perform parking space box regression and simultaneously perceive the potential distribution of the parking space boxes.
[0129] Based on the above method, in this embodiment, the first loss value is determined based on the thermal value of the candidate pixel, such as the predicted value of the candidate pixel is determined based on the thermal value and the pixel position of the candidate pixel, and the first loss value is determined based on the predicted value of the candidate pixel and the pixel position of the sample position point, such as the first loss value is determined based on the distance between the predicted value of the candidate pixel and the pixel position of the sample position point.
[0130] For example, the predicted value of a candidate pixel can be determined using the following formula: ; This represents the predicted value of the candidate pixel. Indicates the pixel position of the candidate pixel. This represents the heat value of the candidate pixel. For example, if the candidate pixel's position is (3, 3) and its heat value is 0.1, then the predicted value of the candidate pixel is (3.3, 3.3). Similarly, if the candidate pixel's position is (3, 4) and its heat value is 0.1, then the predicted value of the candidate pixel is (3.3, 4.4).
[0131] For example, the first loss value can be determined using the following formula: ; This represents the first loss value. This indicates the pixel position of the sample location point. This represents the distance between the predicted value of a candidate pixel and the pixel position of a sample location, such as Euclidean distance.
[0132] Considering the existence of 4 sample location points and 4 candidate pixels matching the 4 sample location points, the first loss value 1 can be calculated based on the predicted value of candidate pixel 1 and the pixel position of sample location point 1 using the above formula. Similarly, the first loss value 2, first loss value 3, and first loss value 4 can be obtained. Thus, the average (or summation) of these four first loss values can be used as the first loss value.
[0133] Step S26: Determine the second loss value based on the sample corner features and the labeled corner feature labels.
[0134] For example, suppose the automated parking space detection network outputs K sample corner features for each entry point, where K is a positive integer greater than 1. The labeled corner features can then include these K features. Considering that an automated parking space has two entry points (e.g., a first entry point and a second entry point), the K corner features will include a first corner feature corresponding to the first entry point and a second corner feature corresponding to the second entry point. All other corner features besides the first and second corner features will be zero.
[0135] The first corner feature and the second corner feature include calibrated parameters. , and This indicates the offset distance (i.e., the offset amount). and Indicates the angle of the side line. Indicates the confidence level.
[0136] The remaining corner features include the calibrated parameters. , and It can be 0, and It can be 0, It can be 0.
[0137] For example, for sample corner features, the 3D parking space detection network can output sample corner features of K sample entry points. For each sample entry point, the 3D parking space detection network can output multiple sample corner features corresponding to that sample entry point. One sample corner feature can be selected from the multiple sample corner features as the sample corner feature corresponding to that sample entry point (to participate in the calculation of the second loss value).
[0138] For example, for each sample corner feature corresponding to the sample entry point, the confidence level of the sample corner feature can be calculated. Based on the confidence level of each sample corner feature, the sample corner feature corresponding to the highest confidence level is used as the sample corner feature corresponding to the sample entry point (to participate in the calculation of the second loss value).
[0139] For each sample corner feature, the sample corner feature can include the offset distance between the sample entry point and the reference entry point, i.e. and In this way, the confidence level of the corner feature of the sample can be determined based on this offset distance. For example, the confidence level of the corner feature of the sample can be determined using the following formula: ; This indicates the confidence level of the corner feature of the sample. This indicates the lateral offset distance between the sample entry point and the reference entry point. This indicates the longitudinal offset distance between the sample entry point and the reference entry point. This indicates the configured parameter value.
[0140] In summary, we can obtain the sample corner features with the highest confidence for each sample entry point, that is, obtain the K sample corner features for K sample entry points. For example, each sample corner feature includes parameters. , and Indicates the offset distance. and Indicates the angle of the side line. This indicates the confidence level, specifically the maximum confidence level mentioned above.
[0141] In summary, we can obtain K sample corner features and K labeled corner features (i.e., corner feature labels) output by the 3D parking space detection network, with a one-to-one correspondence between the K sample corner features and the K corner feature labels. Thus, the second loss value can be determined based on the K sample corner features and the corner feature labels.
[0142] For example, you can substitute the corner features and corner feature labels of K samples into the MSE (Mean Square Error) loss function to obtain the second loss value. Alternatively, you can substitute the corner features and corner feature labels of K samples into other loss functions to obtain the second loss value. There are no restrictions on this.
[0143] Step S27: Determine the target loss value based on the first loss value and the second loss value. For example, the target loss value can be obtained by performing a weighted calculation based on the first loss value and the second loss value.
[0144] Step S28: Adjust the three-dimensional parking space detection network based on the target loss value (e.g., adjust the network parameters of the three-dimensional parking space detection network) to obtain the trained three-dimensional parking space detection network.
[0145] Based on the target loss value, the network parameters of the three-dimensional parking space detection network can be adjusted using methods such as gradient descent. There are no restrictions on this adjustment process, and the goal is to make the target loss value smaller and smaller.
[0146] After adjusting the 3D parking space detection network based on the target loss value, it can be determined whether the adjusted 3D parking space detection network has converged. If so, the adjusted 3D parking space detection network is used as the trained 3D parking space detection network, and the training process ends. If not, based on the adjusted 3D parking space detection network, return to step S22 and repeat the above steps until the 3D parking space detection network has converged.
[0147] For example, if the target loss value is less than a preset threshold, the adjusted 3D parking space detection network has converged; if the target loss value is not less than the preset threshold, the adjusted 3D parking space detection network has not converged. Alternatively, if the number of iterations reaches a threshold, the adjusted 3D parking space detection network has converged; if the number of iterations does not reach the threshold, the adjusted 3D parking space detection network has not converged. Or, if the iteration duration reaches a threshold, the adjusted 3D parking space detection network has converged; if the iteration duration does not reach the threshold, the adjusted 3D parking space detection network has not converged. Of course, the above are just a few examples of determining whether the adjusted 3D parking space detection network has converged, and there are no restrictions on the determination method.
[0148] As can be seen from the above technical solutions, in this embodiment, both the training and inference processes are end-to-end 3D parking space detection networks (neural networks), directly outputting single-frame parking space prediction results after inputting an image. Using an end-to-end network to perform 3D parking space detection and recognition using a panoramic top-down view reduces system computational complexity while eliminating detection errors caused by height differences in the 3D parking space plane and image coordinate transformations, thus improving the detection accuracy of parking space bounding boxes. Post-processing logic further enhances the accuracy of 3D parking space detection by stabilizing and tracking the detection results across multiple frames, making it suitable for various states of automatic parking, including parking, entry, and exit. Based on the above approach, detection accuracy can be improved. The end-to-end network directly processes the panoramic top-down view, avoiding accumulated errors caused by height differences in the 3D parking space and image coordinate transformations, significantly improving the geometric accuracy of parking space detection. Algorithm optimization targeting the structural features of 3D parking spaces avoids false detections or missed detections caused by ignoring the characteristics of 3D space. Based on the above approach, it can adapt to multiple scenarios. The post-processing logic effectively suppresses the jitter of single-frame detection through multi-frame stabilization and following, improves the robustness of parking space detection in dynamic scenarios (such as when the vehicle is moving), covers the entire process of automatic parking (inspection, entry, and exit), ensures the continuity and stability of parking space information under different states, eliminates the dependence on additional sensors (such as ultrasonic radar), and can achieve high-precision detection by relying solely on visual data.
[0149] Based on the same concept as the above method, this application proposes a three-dimensional parking space detection device, see [link to relevant documentation]. Figure 6 The diagram shown is a structural schematic of the device, which may include:
[0150] Generation module 61 is used to generate a panoramic top view based on images captured by multiple cameras around the vehicle;
[0151] Processing module 62 is used to input the panoramic top view into the three-dimensional parking space detection network to obtain the predicted parking space frame of the three-dimensional parking space and the target predicted corner feature of the target predicted entrance point of the three-dimensional parking space.
[0152] The determining module 63 is used to determine the parking space depth value of the three-dimensional parking space based on the predicted parking space frame, wherein the parking space depth value represents the distance between the entrance line and the inner line of the predicted parking space frame; adjust the predicted parking space frame based on the parking space depth value and the target predicted corner point features to obtain a candidate parking space frame; and determine the target parking space frame of the three-dimensional parking space based on the candidate parking space frame.
[0153] For example, when the processing module 62 inputs the panoramic top view to the three-dimensional parking space detection network to obtain the target predicted corner features of the target predicted entrance point of the three-dimensional parking space, it specifically performs the following: inputting the panoramic top view to the three-dimensional parking space detection network to obtain the predicted corner features of multiple predicted entrance points; wherein, for each predicted entrance point, the predicted corner feature of the predicted entrance point includes the offset distance between the predicted entrance point and the reference entrance point; the three-dimensional parking space includes two entrance points, and the reference entrance point is the entrance point that is closest to the predicted entrance point among the two entrance points;
[0154] The pixel position of the predicted entry point is determined based on the pixel position of the reference entry point and the offset distance; based on the distance between the pixel position of the predicted parking space frame's entry point and the pixel position of each predicted entry point, the predicted entry point corresponding to the minimum distance is determined as the target predicted entry point, and the predicted corner feature of the target predicted entry point is determined as the target predicted corner feature.
[0155] For example, the entrance point of the predicted parking space frame is located at the intersection of the multi-level parking space and the ground; the target predicted corner point feature includes the side line angle between the target straight line and the side line of the predicted parking space frame, and the target straight line is the line connecting the target predicted entrance point and the entrance point of the predicted parking space frame. The predicted parking space frame includes an entrance point and an inner point. When the determining module 63 adjusts the predicted parking space frame based on the parking space depth value and the target predicted corner point feature to obtain a candidate parking space frame, it specifically performs the following: determining the position adjustment amount based on the parking space depth value and the side line angle; determining the pixel position of the inner point of the predicted parking space frame based on the pixel position of the entrance point and the position adjustment amount; and generating the candidate parking space frame based on the pixel position of the entrance point and the pixel position of the inner point of the predicted parking space frame.
[0156] For example, when determining the target parking space frame of the multi-level parking space based on the candidate parking space frame, the determining module 63 is specifically used to: if the distance between the candidate parking space frame and the historical parking space frame is greater than a distance threshold, and / or the difference in orientation angle between the candidate parking space frame and the historical parking space frame is greater than an angle threshold, then the historical parking space frame is determined as the target parking space frame of the multi-level parking space; wherein, the historical parking space frame is the target parking space frame determined last time; if the distance between the candidate parking space frame and the historical parking space frame is not greater than a distance threshold, and the difference in orientation angle between the candidate parking space frame and the historical parking space frame is not greater than an angle threshold, then the candidate parking space frame and the historical parking space frame are weighted and fused to obtain the target parking space frame of the multi-level parking space.
[0157] For example, the three-dimensional parking space detection network includes at least an initial feature extractor, a global feature extractor, a local feature extractor, a parking space bounding box prediction network, and a corner feature prediction network. When the processing module 62 inputs a panoramic top-down view to the three-dimensional parking space detection network to obtain the predicted parking space bounding box and the target predicted corner features of the target predicted entry point of the three-dimensional parking space, it specifically performs the following steps: inputting the panoramic top-down view to the initial feature extractor, and extracting features from the panoramic top-down view using the initial feature extractor to obtain the initial features of the panoramic top-down view; inputting the initial features to the global feature extractor, and extracting features from the global feature extractor based on... The initial features are used to extract global features from the panoramic top-down view. These global features are then input into the parking space bounding box prediction network, which processes the global features to obtain the predicted parking space bounding box of the three-dimensional parking space, and outputs the predicted parking space bounding box. The initial features are also input into the local feature extractor, which extracts local features from the panoramic top-down view based on the initial features. These local features are then input into the corner feature prediction network, which processes the local features to obtain the predicted corner features of the predicted entrance point of the three-dimensional parking space, and outputs the predicted corner features.
[0158] For example, the three-dimensional parking space detection device further includes: a training module, used to input a sample panoramic top view into the three-dimensional parking space detection network to be trained, to obtain sample parking space frames and sample corner features of sample entrance points of the three-dimensional parking spaces; to generate a heat map based on the sample parking space frames, the heat map including heat values of multiple pixels; wherein, for each pixel in the heat map, the closer the pixel is to the sample location point, the larger the heat value of the pixel; wherein, the sample location point includes the entrance point and the inner point of the sample parking space frame; to select candidate pixels that match the sample location point from all pixels in the heat map, and to determine a first loss value based on the heat value of the candidate pixels; to determine a second loss value based on the sample corner features and the labeled corner feature tags; to determine a target loss value based on the first loss value and the second loss value, and to adjust the three-dimensional parking space detection network based on the target loss value, to obtain a trained three-dimensional parking space detection network.
[0159] When the training module determines the first loss value based on the thermal values of the candidate pixels, it specifically performs the following steps: determining the predicted value of the candidate pixel based on its thermal value and pixel position; and determining the first loss value based on the distance between the predicted value of the candidate pixel and the pixel position of the sample location point. The training module uses the following formula to determine the predicted value of the candidate pixel: The training module determines the first loss value using the following formula: ; This represents the predicted value of the candidate pixel. This indicates the pixel position of the candidate pixel. This represents the thermal value of the candidate pixel. This represents the first loss value; This indicates the pixel position of the sample location point. This represents the distance between the predicted value of the candidate pixel and the pixel position of the sample location point.
[0160] For example, when a panoramic top-down view of a sample is input into the 3D parking space detection network to be trained, the 3D parking space detection network outputs multiple sample corner features corresponding to the sample entry point; for each sample corner feature, the sample corner feature includes the offset distance between the sample entry point and the reference entry point; when the training module determines the second loss value based on the sample corner features and the labeled corner feature labels, it specifically performs the following: for each sample corner feature, it determines the confidence level of the sample corner feature based on the offset distance in the sample corner feature; based on the confidence level of each sample corner feature, it determines the second loss value based on the sample corner feature corresponding to the maximum confidence level and the corner feature label; the training module uses the following formula to determine the confidence level of the sample corner feature: ; Indicates the confidence level. This indicates the lateral offset distance between the sample entry point and the reference entry point. This indicates the longitudinal offset distance between the sample entry point and the reference entry point. This indicates the configured parameter value.
[0161] Based on the same concept as the above method, this application proposes an electronic device, see [link to previous application]. Figure 7 As shown, the electronic device includes a processor 71 and a machine-readable storage medium 72, the machine-readable storage medium 72 storing machine-executable instructions that can be executed by the processor 71; the processor 71 is used to execute the machine-executable instructions to implement the three-dimensional parking space detection method disclosed in the above example of this application.
[0162] Based on the same concept as the above method, this application also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the three-dimensional parking space detection method disclosed in the above examples of this application.
[0163] The aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0164] Based on the same concept as the methods described above, this application also provides a computer program product, which includes a computer program. When executed by a processor, the computer program implements the three-dimensional parking space detection method disclosed in the examples above.
[0165] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0166] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for detecting three-dimensional parking spaces, characterized in that, The method includes: A panoramic top-down view is generated based on images captured by multiple cameras around the vehicle; The panoramic top view is input into the three-dimensional parking space detection network to obtain the predicted parking space frame of the three-dimensional parking space and the target predicted corner feature of the target predicted entrance point of the three-dimensional parking space; wherein, the target predicted corner feature includes the side line angle between the target line and the side line of the predicted parking space frame, and the target line is the line connecting the target predicted entrance point and the entrance point of the predicted parking space frame. The parking space depth value of the three-dimensional parking space is determined based on the predicted parking space frame; wherein, the parking space depth value represents the distance between the entrance line of the predicted parking space frame and the inner line of the predicted parking space frame; The predicted parking space frame is adjusted based on the parking space depth value and the target predicted corner feature to obtain a candidate parking space frame, and the target parking space frame of the three-dimensional parking space is determined based on the candidate parking space frame. The panoramic top-down view is input into the 3D parking space detection network to obtain the target predicted corner features of the target predicted entrance point of the 3D parking space, including: The panoramic top-down view is input into the three-dimensional parking space detection network to obtain the predicted corner features of multiple predicted entrance points; wherein, for each predicted entrance point, the predicted corner feature of the predicted entrance point includes the offset distance between the predicted entrance point and the reference entrance point; the three-dimensional parking space includes two entrance points, and the reference entrance point is the entrance point that is closest to the predicted entrance point among the two entrance points; The pixel position of the predicted entry point is determined based on the pixel position of the reference entry point and the offset distance; Based on the distance between the pixel position of the predicted parking space frame's entrance point and the pixel position of each predicted entrance point, the predicted entrance point corresponding to the minimum distance is determined as the target predicted entrance point, and the predicted corner feature of the target predicted entrance point is determined as the target predicted corner feature.
2. The method according to claim 1, characterized in that, The entrance point of the predicted parking space frame is located at the junction of the three-dimensional parking space and the ground. The predicted parking space bounding box includes an entrance point and an inner point. The process of adjusting the predicted parking space bounding box based on the parking space depth value and the target predicted corner feature to obtain a candidate parking space bounding box includes: The position adjustment amount is determined based on the parking space depth value and the side line angle; The pixel position of the inner point of the predicted parking space frame is determined based on the pixel position of the entrance point of the predicted parking space frame and the position adjustment amount, and the candidate parking space frame is generated based on the pixel position of the entrance point of the predicted parking space frame and the pixel position of the inner point of the predicted parking space frame.
3. The method according to claim 1, characterized in that, Determining the target parking space frame of the multi-level parking space based on the candidate parking space frame includes: The candidate parking space frame is determined as the target parking space frame for the multi-level parking space; or... If the distance between the candidate parking space frame and the historical parking space frame is greater than a distance threshold, and / or the difference in orientation angle between the candidate parking space frame and the historical parking space frame is greater than an angle threshold, then the historical parking space frame is determined as the target parking space frame of the three-dimensional parking space; wherein, the historical parking space frame is the previously determined target parking space frame. If the distance between the candidate parking space frame and the historical parking space frame is not greater than a distance threshold, and the difference in orientation angle between the candidate parking space frame and the historical parking space frame is not greater than an angle threshold, then the candidate parking space frame and the historical parking space frame are weighted and fused to obtain the target parking space frame of the three-dimensional parking space.
4. The method according to claim 1, characterized in that, The three-dimensional parking space detection network includes at least an initial feature extractor, a global feature extractor, a local feature extractor, a parking space bounding box prediction network, and a corner feature prediction network. The step of inputting the panoramic top-down view into the 3D parking space detection network to obtain the predicted parking space bounding box and the target predicted corner features of the target predicted entrance point of the 3D parking space includes: The panoramic top view is input to the initial feature extractor, and the initial feature extractor extracts features from the panoramic top view to obtain the initial features of the panoramic top view; The initial features are input to the global feature extractor, which extracts global features of the panoramic top view based on the initial features. The global features are then input to the parking space prediction network, which processes the global features to obtain the predicted parking space frame of the three-dimensional parking space and outputs the predicted parking space frame. The initial features are input to the local feature extractor, which extracts local features of the panoramic top view based on the initial features. The local features are then input to the corner feature prediction network, which processes the local features to obtain the predicted corner features of the predicted entrance point of the three-dimensional parking space, and outputs the predicted corner features.
5. The method according to claim 1, characterized in that, The training process for the aforementioned 3D parking space detection network specifically includes: The sample panoramic top view is input into the three-dimensional parking space detection network to be trained to obtain the sample parking space bounding box of the three-dimensional parking space and the sample corner features of the sample entrance point of the three-dimensional parking space. A heatmap is generated based on the sample parking space frame. The heatmap includes the heat values of multiple pixels. For each pixel in the heatmap, the closer the pixel is to the sample location point, the greater the heat value of the pixel. The sample location point includes the entrance point and the inner point of the sample parking space frame. Candidate pixels that match the sample location points are selected from all pixels in the heatmap, and a first loss value is determined based on the heat values of the candidate pixels. The second loss value is determined based on the sample corner features and the labeled corner feature tags; A target loss value is determined based on the first loss value and the second loss value. The three-dimensional parking space detection network is then adjusted based on the target loss value to obtain a trained three-dimensional parking space detection network.
6. The method according to claim 5, characterized in that, Determining the first loss value based on the thermal values of the candidate pixels includes: determining a predicted value for the candidate pixel based on the thermal values and the pixel position of the candidate pixel; and determining the first loss value based on the distance between the predicted value of the candidate pixel and the pixel position of the sample location point. The predicted value of the candidate pixel is determined using the following formula: ; The first loss value is determined using the following formula: ; in, This represents the predicted value of the candidate pixel. This indicates the pixel position of the candidate pixel. This represents the thermal value of the candidate pixel. This represents the first loss value; in, This indicates the pixel position of the sample location point. This represents the distance between the predicted value of the candidate pixel and the pixel position of the sample location point.
7. The method according to claim 5, characterized in that, When a panoramic top view of a sample is input into the 3D parking space detection network to be trained, the 3D parking space detection network outputs multiple sample corner features corresponding to the sample entry point. For each sample corner feature, the sample corner feature includes the offset distance between the sample entry point and the reference entry point; Determining a second loss value based on the sample corner features and the labeled corner feature tags includes: For each sample corner feature corresponding to the sample entry point, the confidence level of the sample corner feature is determined based on the offset distance in the sample corner feature; based on the confidence level of each sample corner feature, a second loss value is determined based on the sample corner feature corresponding to the maximum confidence level and the corner feature label; The confidence level of the corner feature of the sample is determined using the following formula: ; Indicates the confidence level. This indicates the lateral offset distance between the sample entry point and the reference entry point. This indicates the longitudinal offset distance between the sample entry point and the reference entry point. This indicates the configured parameter value.
8. A three-dimensional parking space detection device, characterized in that, The device includes: The generation module is used to generate a panoramic top-down view based on images captured by multiple cameras around the vehicle; The processing module is used to input the panoramic top view into the three-dimensional parking space detection network to obtain the predicted parking space frame of the three-dimensional parking space and the target predicted corner feature of the target predicted entrance point of the three-dimensional parking space; wherein, the target predicted corner feature includes the side line angle between the target line and the side line of the predicted parking space frame, and the target line is the line connecting the target predicted entrance point and the entrance point of the predicted parking space frame. The determination module is used to determine the parking space depth value of the three-dimensional parking space based on the predicted parking space frame, wherein the parking space depth value represents the distance between the entrance line and the inner line of the predicted parking space frame; adjust the predicted parking space frame based on the parking space depth value and the target predicted corner point features to obtain a candidate parking space frame; and determine the target parking space frame of the three-dimensional parking space based on the candidate parking space frame. Specifically, when the processing module inputs the panoramic top-down view into the 3D parking space detection network to obtain the target predicted corner features of the target predicted entrance point of the 3D parking space, it is used for: The panoramic top-down view is input into the three-dimensional parking space detection network to obtain the predicted corner features of multiple predicted entrance points; wherein, for each predicted entrance point, the predicted corner feature of the predicted entrance point includes the offset distance between the predicted entrance point and the reference entrance point; the three-dimensional parking space includes two entrance points, and the reference entrance point is the entrance point that is closest to the predicted entrance point among the two entrance points; The pixel position of the predicted entry point is determined based on the pixel position of the reference entry point and the offset distance; Based on the distance between the pixel position of the predicted parking space frame's entrance point and the pixel position of each predicted entrance point, the predicted entrance point corresponding to the minimum distance is determined as the target predicted entrance point, and the predicted corner feature of the target predicted entrance point is determined as the target predicted corner feature.
9. An electronic device, characterized in that, include: A processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Three-dimensional parking space detection method and device based on fisheye image
CN113269163A
Parking space angular point identification method and device, storage medium and electronic equipment
CN117351457A
Parking garage location identification method and device, computer equipment and readable storage medium
CN119229419A