Parking space detection method, device, electronic device and readable storage medium

By acquiring and processing panoramic top view, determining parking space information and using re-identification features to predict parking spaces, the problem of low parking space detection accuracy is solved, and high-precision parking space recognition and stability detection are achieved.

CN118314763BActive Publication Date: 2025-08-22CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410407188.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-07
Publication Date
2025-08-22
Estimated Expiration
2044-04-07

AI Technical Summary

Technical Problem

In the prior art, the parking space detection accuracy is low, resulting in the problem of vehicles being unable to accurately identify parking spaces or scratches during parking.

Method used

By obtaining the panoramic top view of the Nth and N+1st frame video frames of the target vehicle, the cameras installed around the vehicle collect environmental information, perform distortion correction and image splicing, determine the parking space information and re-identification features, use the feature re-identification algorithm to predict the next frame of parking space information, and verify the accuracy of the detection results through historical detection information.

Benefits of technology

It improves the accuracy and stability of parking space identification, effectively identifies obstructed parking spaces, reduces the probability of false inspection and missed inspection, and ensures the accuracy of the inspection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118314763B_ABST
    Figure CN118314763B_ABST
Patent Text Reader

Abstract

The present application relates to the field of automotive technology and provides a method, device, electronic device, and readable storage medium for parking space detection. The method comprises: obtaining a first panoramic bird's-eye view corresponding to the Nth video frame and a second panoramic bird's-eye view corresponding to the N+1th video frame of the target vehicle; determining first parking space information based on the first panoramic bird's-eye view, and determining second parking space information based on the second panoramic bird's-eye view; determining a first-level recognition feature of the first parking space based on the first parking space information; predicting third parking space information of the N+1th video frame based on the first-level recognition feature, wherein the third parking space information includes the size and third center point of the predicted third bounding box of the third parking space; and determining whether the detection result of the second parking space information is the final result based on the third parking space information. The present application solves the technical problem of low parking space detection accuracy in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automotive technology, and in particular to a method, device, electronic device, and readable storage medium for parking space detection. Background Art

[0002] Currently, when a vehicle performs automatic parking, it will travel at a low speed and use the surround-view camera on the vehicle body to search for available parking spaces in the surrounding area during driving. After selecting an available parking space, it will perform path planning. However, due to the diversity of current parking pavement materials, changing lighting conditions in the surrounding environment, and irregular shapes of parking spaces, these factors will affect the accuracy of parking space detection, and the vehicle may be unable to park or scratched during parking due to the inability to recognize the parking space or incorrect parking space recognition.

[0003] Therefore, a high-precision parking space detection method is urgently needed. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a method, device, electronic device, and readable storage medium for parking space detection to solve the problem of low parking space detection accuracy in the prior art.

[0005] According to a first aspect of an embodiment of the present application, a method for detecting a parking space is provided, comprising:

[0006] Obtain a first panoramic bird's-eye view corresponding to the Nth video frame and a second panoramic bird's-eye view corresponding to the N+1th video frame of the target vehicle, where N is greater than or equal to 1;

[0007] Determining first parking space information based on the first panoramic bird's-eye view, and determining second parking space information based on the second panoramic bird's-eye view, wherein the first parking space information includes a size and a first center point of a first bounding box of at least one first parking space in the first panoramic bird's-eye view, and the second parking space information includes a size and a second center point of a second bounding box of at least one second parking space in the second panoramic bird's-eye view;

[0008] Determining a first-level identification feature of the first parking space based on the first parking space information;

[0009] Predicting, based on the first-recognition features, third parking space information of the (N+1)th video frame, wherein the third parking space information includes a size and a third center point of a predicted third bounding box of the third parking space;

[0010] According to the third parking space information, it is determined whether the detection result of the second parking space information is a final result.

[0011] According to a second aspect of the embodiments of the present application, a parking space detection device is provided, comprising:

[0012] A first acquisition module is configured to acquire a first panoramic bird's-eye view corresponding to an N-th video frame and a second panoramic bird's-eye view corresponding to an N+1-th video frame of the target vehicle, where N is greater than or equal to 1;

[0013] a second acquisition module, configured to determine first parking space information based on the first panoramic bird's-eye view, and to determine second parking space information based on the second panoramic bird's-eye view, wherein the first parking space information includes a size and a first center point of a first bounding box of at least one first parking space in the first panoramic bird's-eye view, and the second parking space information includes a size and a second center point of a second bounding box of at least one second parking space in the second panoramic bird's-eye view;

[0014] a third acquisition module, configured to determine a first-level identification feature of the first parking space based on the first parking space information;

[0015] a fourth acquisition module, configured to predict third parking space information of the (N+1)th video frame based on the first-recognition feature, wherein the third parking space information includes a size and a third center point of a third bounding box of the predicted third parking space;

[0016] The result detection module is used to determine whether the detection result of the second parking space information is a final result based on the third parking space information.

[0017] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0018] According to a fourth aspect of an embodiment of the present application, a readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0019] The beneficial effects of the embodiments of the present application include at least:

[0020] By obtaining the first panoramic bird's-eye view corresponding to the Nth frame of video and the second panoramic bird's-eye view corresponding to the N+1th frame of video of the target vehicle, the surrounding environment of the vehicle is collected using cameras installed around the vehicle, and a panoramic bird's-eye view is obtained after processing, so that the system can determine the current surrounding drivable area and parking area of ​​the vehicle, and can provide an image basis for obtaining data such as re-identification features of the parking space in the future; by determining the first parking space information according to the first panoramic bird's-eye view, and determining the second parking space information according to the second panoramic bird's-eye view, using the panoramic bird's-eye view to determine the coordinates of the corner points of the collected parking spaces, and according to the coordinates of the corner points to obtain the area of ​​these parking spaces, that is, the size of the bounding box and the center point of the parking space, providing a data basis for predicting the next frame of parking space information in the future; by determining the first parking space information according to the first parking space information The first-recognition feature is used, and the third parking space information of the N+1th frame of video frame is predicted based on the first-recognition feature. The feature re-recognition algorithm is used to extract the features of the parking space, which can effectively identify the feature information of the parking space, and then accurately identify each parking space, thereby improving the accuracy of parking space recognition; by determining whether the detection result of the second parking space information is the final result based on the third parking space information, the predicted third parking space information is compared with the second parking space information obtained by actual measurement, thereby determining the accuracy of the second parking space information, and can effectively identify parking space targets that are blocked due to image distortion or field of view transformation, while also ensuring the accuracy of the detection results, effectively utilizing historical detection information to improve the accuracy of the current detection results, and solving the problem of low parking space detection accuracy in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 This is a flow chart of a method for parking space detection provided by an embodiment of the present application;

[0023] Figure 2 is a schematic diagram of a parking space detection provided by an embodiment of the present application;

[0024] Figure 3 1 is a schematic structural diagram of a parking space detection device provided in an embodiment of the present application;

[0025] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0027] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein. Furthermore, the objects distinguished by "first," "second," and the like are generally of a class, and do not limit the number of objects. For example, the first object may be one or more.

[0028] In addition, it should be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, the elements defined by the phrase "comprises..." do not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the elements.

[0029] A method and device for parking space detection according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0030] Figure 1 This is a flow chart of a parking space detection method provided by an embodiment of the present application. This method can be executed by the vehicle side. Figure 1 As shown, the parking space detection method includes:

[0031] Step 101: Acquire a first panoramic bird's-eye view corresponding to an Nth video frame and a second panoramic bird's-eye view corresponding to an N+1th video frame of a target vehicle, where N is greater than or equal to 1.

[0032] The image information in the panoramic bird's-eye view may include images of the vehicle's current drivable area, obstacles, pedestrians, and parking spaces. The video or images required to obtain the panoramic bird's-eye view can be obtained through a standard video decoding library. The number of video frames can be 500, 700, 900, etc., which can be determined based on the duration and frame rate of the currently used video.

[0033] Specifically, when a vehicle obtains a panoramic bird's-eye view, it can first use a standard video decoding library to obtain image data captured by the vehicle's front, rear, left, and right cameras. These image data are then corrected for distortion using pre-acquired intrinsic parameters. Camera distortion can be divided into radial distortion and tangential distortion. For the image data captured by any camera, the following formula can be used to obtain the radial distortion correction coordinates:

[0034]

[0035] Among them, x dr and y dr The image coordinates after radial distortion correction are represented by x and y, the image coordinates before correction, and k1, k2, and k3 are the radial distortion parameters. The tangential distortion correction coordinates can also be obtained using the following formula:

[0036]

[0037] Among them, x dt and y dt The coordinates of the image after tangential distortion correction are expressed as x and y, the coordinates of the image before correction are expressed as p1 and p2, and the tangential distortion parameters are expressed as p2. After obtaining the coordinates after the two distortion corrections, the following formula can be used to obtain the final corrected image coordinates:

[0038]

[0039] Among them, x d and y d Expressed as the final rectified image coordinates, x dr and y dr Expressed as the image coordinate after radial distortion correction, x dt and y dt Represented as image coordinates after tangential distortion correction. After obtaining the final corrected image coordinates, the image coordinates of key corner points in the corrected image and their expected corresponding real-world 3D coordinates can be selected. Using the solvePNP algorithm, the corresponding camera extrinsics (roll, pitch, and yaw) are obtained. The panoramic images are then stitched together using the corresponding corner points. Finally, geometric transformations such as perspective transformation are used to convert the panoramic image into a panoramic bird's-eye view. Key corner points can be the mid-boundary corners of the drivable area, image boundary points, or randomly selected points, without specific limitations.

[0040] By using cameras installed around the vehicle to collect information about the vehicle's surroundings and obtaining a panoramic bird's-eye view after processing, the system can determine the vehicle's current surrounding drivable area and parking area, and provide an image basis for subsequently obtaining data such as re-identification features of parking spaces.

[0041] Step 102: determining first parking space information based on the first panoramic bird's-eye view, and determining second parking space information based on the second panoramic bird's-eye view.

[0042] The first parking space information includes the size and first center point of a first bounding box of at least one first parking space in the first panoramic bird's-eye view, and the second parking space information includes the size and second center point of a second bounding box of at least one second parking space in the second panoramic bird's-eye view. It should be noted that the second parking space information is determined based on the second panoramic bird's-eye view and is the parking space information corresponding to the N+1th frame detected based on the actual second panoramic bird's-eye view, and is the parking space detection result corresponding to the N+1th frame.

[0043] The corner coordinates of the parking spaces in the panoramic bird's-eye view can be determined from the image. The bounding boxes of these parking spaces can be determined using the corner coordinates. The bounding box of any parking space can be expressed as bbox = (x1, y1, x2, y2), where x1 and y1 represent the coordinates of the corner of the upper left corner of the parking space in the image, and x2 and y2 represent the coordinates of the corner of the lower right corner of the parking space in the image. The size of the bounding box of the parking space can then be obtained using the following formula: s = (x2-x1, y2-y1). The center point of the parking space can be calculated using the following formula:

[0044]

[0045] Among them, c x ,c y It represents the center point coordinates of the parking space in the panoramic top view, x1, y1 represents the upper left corner coordinates of the parking space in the panoramic top view, and x2, y2 represents the lower right corner coordinates of the parking space in the panoramic top view.

[0046] By determining the bounding box and center point of the parking space in the image based on the panoramic bird's-eye view, a data basis is provided for predicting the parking space information of the next frame.

[0047] Step 103: Determine a first-level identification feature of the first parking space based on the first parking space information.

[0048] Specifically, re-identification features can effectively represent the unique appearance characteristics of each parking space, such as its shape, location, and surroundings. Re-identification features can be used to assign a unique identifier to each parking space, allowing the identification of the parking space even while the vehicle is in motion. Subsequent detections of the same parking space can also be assigned the same identifier, linking parking space detections with historical detection results, improving the accuracy and stability of parking space recognition.

[0049] Step 104 : predicting the third parking space information of the (N+1)th video frame based on the first recognition feature.

[0050] The third parking space information includes the predicted size of a third bounding box and a third center point of the third parking space.

[0051] By using the re-identification features, we can effectively predict the re-identification features of the parking space in the next frame, and then predict the position coordinates of the center point of the parking space with the same identity as the parking space in the next frame. Since the size of the parking space does not change, the size of the bounding box of the parking space is the same as the size of the bounding box of the current parking space. Finally, the third parking space information is obtained based on the position coordinates of this center point and the size of the bounding box.

[0052] By using the re-identification features to predict the parking space information of the next frame and adopting the feature re-identification algorithm to extract the features of the parking space, the feature information of the parking space can be effectively identified, and then each parking space can be accurately identified, thereby improving the accuracy of parking space recognition.

[0053] Step 105: Determine whether the detection result of the second parking space information is a final result based on the third parking space information.

[0054] Specifically, the system detects the panoramic overhead view of the next frame to obtain the parking space information in the next frame and obtains the re-identification features of the center point of the parking space in the next frame. Then, the re-identification features of the center point in the third parking space information obtained by prediction can be compared with the re-identification features of the center point of the parking space in the next frame to determine whether the appearance features of the two parking spaces are the same and whether their positions are correct, thereby determining whether the detection result of the parking space information in the next frame (i.e., the second parking space information) is correct. If correct, the detection result is determined to be the final detection result.

[0055] By determining whether the parking space information detection result of the next frame (i.e., the N+1th frame) of the video frame is correct based on the predicted parking space information, the continuity of the video content and the regularity of vehicle movement are utilized to conduct a secondary review of the parking space detection results, effectively avoiding missed detection and false detection of vacant parking spaces, and improving the accuracy and stability of parking space detection.

[0056] According to the technical solution provided by the embodiments of the present application, by acquiring a first panoramic bird's-eye view corresponding to the Nth video frame of the target vehicle and a second panoramic bird's-eye view corresponding to the N+1th video frame, the vehicle's surrounding environment is captured using cameras installed around the vehicle. After processing, a panoramic bird's-eye view is obtained, enabling the system to determine the vehicle's current drivable area and parking area, and providing an image basis for subsequently obtaining data such as parking space re-identification features. By determining first parking space information based on the first panoramic bird's-eye view and second parking space information based on the second panoramic bird's-eye view, the panoramic bird's-eye view is used to determine the coordinates of the parking space corners around the vehicle, and the area, i.e., the size of the bounding box, and the center point of the parking space are obtained based on the corner coordinates, providing a data basis for subsequently predicting the parking space information for the next frame. By determining the first re-identification feature of the first parking space based on the first parking space information, and predicting the third parking space information for the N+1th video frame based on the first re-identification feature, a feature re-identification algorithm is used to extract parking space features, effectively identifying the feature information of the parking space, and then accurately identifying each parking space, thereby improving the accuracy of parking space recognition. By determining whether the detection result of the second parking space information is the final result based on the third parking space information, and comparing the predicted third parking space information with the second parking space information obtained by actual measurement, the accuracy of the second parking space information is determined. This can effectively identify parking space targets that are obscured due to image distortion or field of view transformation, and at the same time ensure the accuracy of the identity assignment of the detection results. It effectively utilizes historical detection information to improve the accuracy of the current detection results, and solves the problem of low parking space detection accuracy in the existing technology.

[0057] In some embodiments, determining a first-factor identification feature of the first parking space based on the first parking space information includes:

[0058] According to the first panoramic bird's-eye view, a first semantic feature map corresponding to the first panoramic bird's-eye view is obtained using a preset semantic segmentation model;

[0059] Determining a center point position of a first center point of the first parking space in the first semantic feature map according to an image size ratio between the first semantic feature map and the first panoramic bird's-eye view; the center point position is a floor-rounded value or a ceiling-rounded value of a ratio of a coordinate of the first center point to the image size ratio;

[0060] Extracting, from the first semantic feature map, a re-identification feature of each pixel in the first semantic feature map using a feature re-identification algorithm to obtain a first re-identification feature map corresponding to the first semantic feature map;

[0061] Based on all center point positions, the weight of each pixel in the first semantic feature map is calculated, and a center point heat map is obtained based on the weight; wherein the weight is the probability that the pixel is the center point position;

[0062] According to the first re-recognition feature map and the center point heat map, the re-recognition feature of the center point position is extracted, and the re-recognition feature is determined as the first re-recognition feature.

[0063] Specifically, using the semantic segmentation model, each pixel in the panoramic bird's-eye view can be assigned to a specific semantic category, thereby distinguishing the vehicle's drivable area, parking area, obstacle area, etc. contained in the bird's-eye view; the extraction process can be to use a deep convolutional network based on deformable convolution as the backbone network to extract the semantic feature map of the panoramic bird's-eye view.

[0064] The image size ratio between the semantic feature map and the panoramic bird's-eye view can be 1:10, 1:15, 1:20, etc., which can be determined according to actual conditions. Since the minimum value of the obtained feature map is one pixel, and the pixel cannot be divided, the coordinates of the center point converted to the semantic feature map must be integers. After the image size ratio conversion, the value obtained can be an integer or a non-integer value. Therefore, for non-integer values, it is necessary to round them up or down. For example, the rounding method can be set to round down. In this case, the center point position can be obtained using the following formula:

[0065]

[0066] in, It can be expressed as the coordinates of the center point of the parking space in the panoramic bird's-eye view corresponding to the center point position in the semantic feature map. It can be expressed as the ratio of the coordinates of the center point of the parking space in the panoramic bird's-eye view to the image size ratio, and the ratio is rounded down.

[0067] After obtaining the semantic feature map corresponding to the panoramic bird's-eye view, a convolutional network with 128 cores can be used to extract the re-identification features of each pixel in the semantic feature map based on the semantic feature map, thereby obtaining a re-identification feature map corresponding to the semantic feature map. The re-identification feature map can be expressed as Among them, E represents the re-identification feature map, H represents the height, and W represents the width; for the re-identification feature map, each point in it is a vector with a length of 128 dimensions.

[0068] Then, based on these center point positions, the probability of each pixel in the semantic feature map being the center point position can be calculated, and the center point heat map of the semantic feature map can be obtained based on these probabilities. Through these probabilities, it is possible to accurately determine which pixel is the center point of the parking space, which can effectively improve the accuracy of detection. Specifically, the weight of a single pixel can be calculated according to the following formula:

[0069]

[0070] Among them, M xy It represents the weight of the pixel as the center point, N represents the number of parking spaces in the semantic feature map, x, y represent the coordinates of the pixel, Represented as the coordinates of the center point position in the semantic feature map, σ c Expressed as standard deviation; This formula uses a two-dimensional Gaussian formula to calculate the pixel weights within a circular area centered on the center point. The farther the pixel coordinates are from the center point, the smaller the probability that the point is the center point. The standard deviation in this formula can be used to describe the degree of probability dispersion centered on the center point. The smaller the standard deviation, the more concentrated the weights are, and the smaller the final center point heat map is, and the larger the peak in the center point heat map is. The larger the standard deviation, the more dispersed the weights are, and the larger the final center point heat map is, and the smaller the peak in the center point heat map is. This standard deviation can be set using prior knowledge, and the optimal value can be adopted through continuous testing and correction. In addition, after obtaining the center point heat map, the loss function of the center point heat map can be calculated using the following formula:

[0071]

[0072] Among them, M xy It is expressed as the weight of the pixel as the center point. It is represented as the predicted center point heat map, and N is the number of parking spaces.

[0073] The re-identification feature map is used to extract the re-identification features corresponding to the center point position. After that, the feature vector corresponding to the heat map corresponding to the center point position can be converted into a one-hot distribution vector through a fully connected network through the center point heat map. The one-hot distribution vector can be expressed as P = {p(k), k∈[1,K]}, where p(k) represents the probability that the point is the kth parking space, and K represents the number of existing parking space trajectories. Finally, the re-identification features of the center point position are obtained.

[0074] This embodiment obtains a semantic feature map through a panoramic bird's-eye view, and uses the semantic feature map to obtain a center point heat map and a re-identification feature map. Finally, the center point heat map and the re-identification feature map are used to determine the re-identification feature corresponding to the center point position, making the final re-identification feature more accurate, thereby improving the accuracy of subsequent parking space detection.

[0075] In some embodiments, predicting the third parking space information of the N+1th video frame based on the first-level recognition feature includes:

[0076] Obtain the motion trajectory of the target vehicle before the Nth video frame, and use the Kalman filter algorithm to determine the motion characteristics of the first parking space relative to the target vehicle based on the motion trajectory;

[0077] Based on the first-level recognition features, the appearance features of the third parking space are predicted;

[0078] determining a target re-identification feature from a second-re-identification feature map based on the motion feature and the appearance feature, wherein a similarity between the target re-identification feature and the motion feature is greater than a first preset value and a similarity between the target re-identification feature and the appearance feature is greater than a second preset value, and the second-re-identification feature map corresponds to the second panoramic bird's-eye view;

[0079] The third parking space information is determined based on the target re-identification features.

[0080] Specifically, the Kalman filter algorithm can predict the coordinates and velocity of an object's position from a finite, noisy sequence of observations of the object's position. Since the target vehicle is in motion, the parking space can be considered to be moving relative to the vehicle, with the vehicle as the stationary point. Therefore, the Kalman filter algorithm can be used to determine the motion characteristics of the first parking space relative to the target vehicle, thereby predicting the movement of the parking space. The first-level recognition features can be used to determine the current appearance of the parking space. Since the appearance characteristics of the same parking space are almost unchanged, the first-level recognition features can be used to predict the appearance characteristics of the third parking space.

[0081] The first preset value and the second preset value can be 60%, 70%, 80%, etc. The two values ​​can be the same or different. However, it should be noted that when determining the target re-identification feature, the similarity between the re-identification feature and the motion feature and the appearance feature must be greater than the preset value in order to confirm it as the target re-identification feature. If only the appearance feature similarity is high or only the motion feature similarity is high, the re-identification feature cannot be considered as the target re-identification feature. After confirming the target re-identification feature, the target re-identification feature can be used for reverse deduction to finally obtain the size of the bounding box of the third parking space and the center point of the parking space.

[0082] Specifically, the historical trajectory before the Nth frame of video can be used as a reference to predict the parking space information of the next frame. By re-identifying features, a certain number of parking spaces with similar appearance features can be determined. These parking spaces can be used as prediction options. Then, the motion features generated by the Kalman filter algorithm can be used to perform a second screening of the parking spaces in the options. Finally, the most similar parking space is found from these alternative parking spaces. This parking space is the predicted third parking space.

[0083] This embodiment obtains the motion and appearance features of the parking space, finds the most similar re-identification feature from the second-re-identification feature map, and determines this feature as the target re-identification feature. The target re-identification feature is then used to obtain predicted parking space information, i.e., the third parking space information. The embodiment also utilizes the continuity of video content and the regularity of vehicle motion to predict parking spaces between adjacent video frames, ensuring prediction accuracy and reducing the probability of missed or false detections. This provides a comparison basis for subsequent determination of parking space detection results in real images, thereby improving the accuracy and reliability of parking space detection.

[0084] In some embodiments, determining whether the detection result of the second parking space information is a final result based on the third parking space information includes:

[0085] Calculating the Mahalanobis distance between the third parking space and the second parking space based on the third parking space information and the second parking space information;

[0086] Determine a second-level recognition feature of the second parking space based on the second parking space information, and calculate a cosine distance between the target re-recognition feature and the second-level recognition feature;

[0087] According to Mahalanobis distance and cosine distance, the target distance between the third parking space and the second parking space is obtained;

[0088] When the target distance is greater than or equal to the preset distance threshold, the third parking space and the second parking space are matched as the same parking space, and the detection result of the second parking space information is determined as the final result.

[0089] Specifically, the Mahalanobis distance is an indicator used to assess the similarity between data. It can address the problem of non-independent and identically distributed dimensions in high-dimensional linearly distributed data. The cosine distance is used to determine the similarity between the re-identified feature vectors corresponding to the third and second parking spaces. It can be calculated by taking the cosine of the angle between the target re-identified feature vector and the second re-identified feature vector. This cosine value is the cosine distance. The Mahalanobis distance can be calculated using the following formula:

[0090]

[0091] Among them, D m It is expressed as Mahalanobis distance, x represents the predicted value, that is, the coordinate of the third parking space, and y represents the true value, that is, the coordinate of the second parking space.

[0092] It should be noted that the target distance is obtained based on the Mahalanobis distance and the cosine distance. The target distance actually refers to the similarity between the third parking space and the second parking space, rather than the length distance in the actual sense; the preset distance threshold can also be a pre-set similarity threshold between the third parking space and the second parking space. The value of the threshold can be 0.4, 0.5, 0.6, etc., which can be determined according to actual conditions, or the system default value can be directly used. No specific limitation is made here.

[0093] In some embodiments, obtaining a target distance between the third parking space and the second parking space based on the Mahalanobis distance and the cosine distance includes:

[0094] According to the Mahalanobis distance and cosine distance, the target distance is obtained using the following formula:

[0095] D=λD r +(1-λ)D m

[0096] Where D represents the target distance, D r represents the cosine distance, D m represents the Mahalanobis distance, and λ represents the weighting parameter.

[0097] The weighting parameter may be 0.8, 0.9, etc., and needs to be determined according to actual needs or through a large number of experiments.

[0098] Specifically, it can be assumed that the value of the weighted parameter is 0.9, and the value of the preset distance threshold can be 0.4. At this time, the Mahalanobis distance between the second parking space and the third parking space obtained is 0.5, and the cosine distance is 0.8. Therefore, the target distance can be calculated as 0.9*0.5+0.1*0.8=0.53, and 0.53 is greater than 0.4, that is, the target distance is greater than the preset distance threshold. At this time, the third parking space and the second parking space can be considered to be the same parking space, and the same parking space identity can be assigned to determine that the detection result of the second parking space information is correct, which can be used as the final detection result.

[0099] This embodiment uses Mahalanobis distance and cosine distance to judge the similarity between the third parking space and the second parking space. When the similarity is greater than or equal to a preset distance threshold, the parking space detection result can be considered correct. When the vehicle performs parking space detection, the parking space can be predicted based on historical information, and the prediction information can be used to verify the detection result to determine whether the detection result is incorrect, thereby effectively improving the accuracy and stability of parking space detection.

[0100] In some embodiments, after obtaining the target distance between the third parking space and the second parking space based on the Mahalanobis distance and the cosine distance, the method further includes:

[0101] When the target distance is less than a preset distance threshold, obtaining an overlapping area between the bounding box of the second parking space and the bounding box of the third parking space and a total area after the overlap, and calculating a ratio of the overlapping area to the total area;

[0102] When the ratio is greater than the preset ratio, the third parking space and the second parking space are matched as the same parking space, and the detection result of the second parking space information is determined as the final result;

[0103] When the ratio is less than or equal to the preset ratio, determining the second parking space as an unmatched parking space;

[0104] If the unmatched parking space is still not matched to the corresponding third parking space after the preset number of frames, the unmatched parking space is determined to be a misdetected parking space, and the information of the second parking space is deleted from the second parking space information.

[0105] Specifically, the preset ratio may be 0.7, 0.6, 0.5, etc., and the preset number of frames may be 20 frames, 30 frames, 40 frames, etc., which may be determined based on actual conditions such as vehicle model, video frame rate, surrounding environment, etc., and no specific limitation is given here.

[0106] Specifically, we can assume that the current target distance is less than the preset distance threshold, the preset ratio is 0.5, and the preset number of frames is 30. At this time, we need to determine the bounding box of the second parking space and the bounding box of the third parking space, and calculate based on their overlapping area and the total area after overlapping. Finally, we get the overlapping area of ​​1.5m. 2 The total area after overlapping is 2.7m 2 , at this time, the ratio is 1.5 / 2.7=0.556, and 0.556 is greater than 0.5, that is, the ratio is greater than the preset ratio. Therefore, it can be determined that the third parking space and the second parking space are the same parking space and have the same identity. The detection result of the second parking space information can be determined as the final detection result. For another example, the calculated overlapping area is 1.2m 2 , and the total area after overlapping is 2.7m 2, the ratio at this time is 1.2 / 2.7=0.444, and 0.444 is less than 0.5, that is, the ratio is less than the preset ratio, because it can be determined that the detected second parking space is an unmatched parking space. At this time, it is necessary to maintain the existence of the second parking space and continue to predict and detect the parking space. After 30 frames, the parking space still has not been matched with the corresponding third parking space, indicating that the parking space may be a parking space obtained by false detection, so the information of the parking space can be deleted from the second parking space information; if the parking space is matched with the third parking space corresponding to the parking space in the 20th frame, it can be indicated that the parking space may be a new parking space that appeared during the vehicle's driving process, so the detection and prediction of the parking space can be continued. The reason why the parking space was not matched with the third parking space in the detection process of the first 20 frames may be that it was blocked by image distortion or field of view transformation, and it cannot be predicted.

[0107] By performing a secondary detection on the second parking space that is not matched to the third parking space, it is determined whether the parking space is an erroneous parking space resulting from a detection error or a newly appeared parking space. This allows for effective identification of parking spaces that are blocked due to image distortion or field of view change. The detection results can be maintained for a period of time to further determine the reason for non-detection, effectively improving the accuracy and stability of the current detection results and preventing detection errors due to field of view obstruction.

[0108] In some embodiments, before extracting the re-identification feature at the center point location based on the first re-identification feature map and the center point heat map and determining the re-identification feature as the first re-identification feature, the method further includes:

[0109] Determining, based on the image size ratio, the converted coordinates of the first center point of the first parking space in the first semantic feature map; wherein the converted coordinates are a ratio of the coordinates of the first center point to the image size ratio;

[0110] The difference between the coordinates corresponding to the center point position and the converted coordinates is calculated, and the error between the first re-recognition feature map and the first panoramic bird's-eye view is corrected based on the difference.

[0111] Specifically, since the center point position in the calculated semantic feature map is obtained by rounding up or down, there will be a certain error compared to the actual center point. Moreover, since the re-identification feature map is downsampled compared to the real image, that is, the panoramic bird's-eye view, the center point of the parking space frame has a quantization error compared to the real image. By calculating the offset of the center point, this error caused by continuous offset can be effectively corrected, reducing the impact of downsampling. The offset of the center point can be calculated using the following formula:

[0112]

[0113] Among them, o can be expressed as the offset of the center point, that is, the difference between the coordinates corresponding to the center point position and the converted coordinates. can be expressed as transformed coordinates, It can be expressed as the coordinates corresponding to the center point position, and r can be expressed as the image size ratio.

[0114] This embodiment calculates the offset of the center point before identifying and obtaining the re-identification features, and uses the offset to correct the error, thereby avoiding the distance error from the real image caused by multiple conversions and feature extractions, reducing the impact of downsampling on the subsequent determination and prediction of re-identification features, and improving the accuracy of parking space detection.

[0115] The present application is described below with reference to a specific embodiment:

[0116] Specifically, Figure 2 As shown, when parking is needed, the automatic parking function can be turned on. The vehicle will then detect the surrounding parking spaces to find a suitable parking space.

[0117] First, the vehicle can use the cameras around the vehicle to capture the surrounding environment, and process the captured images through methods such as distortion correction, corner point correspondence, perspective transformation, etc. to obtain a panoramic bird's-eye view of the current frame. The parking space information in the panoramic bird's-eye view, such as the size of the parking space frame, the position of the center point, etc., will be obtained based on the panoramic bird's-eye view. Then, the re-identification algorithm is used to calculate the re-identification features of the center points of these parking spaces based on these parking space information. For example, there are four parking spaces numbered 1, 2, 3, and 4 in the current frame. At this time, their re-identification features can be obtained based on these four parking spaces and saved. Then, the Kalman filter algorithm is used to obtain the motion features of these parking spaces based on their movement trajectories. At the same time, the appearance features of the corresponding parking spaces in the next frame are predicted based on the re-identification features of the parking spaces in the current frame. Then, the position of the parking spaces in the next frame is predicted based on the motion features and appearance features of these parking spaces.

[0118] Then, in the next frame of video, the vehicle will compare the prediction result obtained based on the data of the previous frame with the current detection result, and use the cosine distance and Mahalanobis distance between the re-identified features to calculate the similarity between the prediction result and the detection result. If the similarity is higher than the set threshold, the detection result is judged to be correct and assigned the same identity as the corresponding detection result of the previous frame, and the two are determined to be the same parking space. If the similarity is lower than the set threshold, a second match can be performed to determine whether the unmatched parking space is a misdetected parking space or a newly appeared parking space; the video processing results of subsequent frames are similar.

[0119] By using this method, parking space targets that are blocked due to image distortion or field of view transformation can still be correctly identified when there are a certain number of video frames between each frame. This effectively utilizes historical information for prediction, thereby improving the accuracy of the current detection results.

[0120] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0121] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0122] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0123] Figure 3 Schematic diagram of a parking space detection device provided in an embodiment of the present application. Figure 3 As shown, the device includes:

[0124] The first acquisition module 301 is configured to acquire a first panoramic top view corresponding to an N-th video frame and a second panoramic top view corresponding to an N+1-th video frame of the target vehicle, where N is greater than or equal to 1;

[0125] a second acquisition module 302 configured to determine first parking space information based on the first panoramic bird's-eye view, and to determine second parking space information based on the second panoramic bird's-eye view, wherein the first parking space information includes a size and a first center point of a first bounding box of at least one first parking space in the first panoramic bird's-eye view, and the second parking space information includes a size and a second center point of a second bounding box of at least one second parking space in the second panoramic bird's-eye view;

[0126] A third acquisition module 303 is configured to determine a first-level identification feature of the first parking space based on the first parking space information;

[0127] a fourth acquisition module 304 configured to predict third parking space information of the (N+1)th video frame based on the first-recognition feature, wherein the third parking space information includes a size and a third center point of a third bounding box of the predicted third parking space;

[0128] The result detection module 305 is configured to determine whether the detection result of the second parking space information is a final result based on the third parking space information.

[0129] According to the technical solution provided in the embodiment of the present application, a first panoramic bird's-eye view corresponding to the Nth video frame and a second panoramic bird's-eye view corresponding to the N+1th video frame of the target vehicle are obtained through a first acquisition module 301. The vehicle's surrounding environment is captured using cameras installed around the vehicle, and a panoramic bird's-eye view is obtained after processing, enabling the system to determine the vehicle's current surrounding drivable area and parking area, and providing an image basis for subsequently obtaining data such as re-identification features of the parking space. A second acquisition module 302 determines the first parking space information based on the first panoramic bird's-eye view, and determines the second parking space information based on the second panoramic bird's-eye view. The panoramic bird's-eye view is used to determine the coordinates of the corner points of the parking spaces around the vehicle, and the area of ​​these parking spaces, i.e., the size of the bounding box, and the center point of the parking space are obtained based on the corner point coordinates, providing a data basis for subsequently predicting the parking space information of the next frame. The third acquisition module 303 and the fourth acquisition module 304 determine the first-level identification features of the first parking space based on the first parking space information, and predict the third parking space information of the N+1th frame of the video based on the first-level identification features. The feature re-identification algorithm is used to extract the features of the parking space, which can effectively identify the characteristic information of the parking space, and then accurately identify each parking space, thereby improving the accuracy of parking space identification. The result detection module 305 determines whether the detection result of the second parking space information is the final result based on the third parking space information. The predicted third parking space information is compared with the second parking space information actually measured to determine the accuracy of the second parking space information. It can effectively identify parking space targets that are obscured due to image distortion or field of view transformation, and at the same time, it can ensure the accuracy of the identity assignment of the detection result, and effectively utilize historical detection information to improve the accuracy of the current detection result. This solves the problem of low parking space detection accuracy in the existing technology.

[0130] In some embodiments, the third acquisition module 303 is specifically used to: obtain a first semantic feature map corresponding to the first panoramic bird's-eye view using a preset semantic segmentation model based on the first panoramic bird's-eye view; determine the center point position of the first center point of the first parking space in the first semantic feature map based on the image size ratio between the first semantic feature map and the first panoramic bird's-eye view; the center point position is the floor-rounded value or ceiling-rounded value of the ratio of the coordinates of the first center point to the image size ratio; extract the re-identification features of each pixel point in the first semantic feature map based on the feature re-identification algorithm based on the first semantic feature map, and obtain a first re-identification feature map corresponding to the first semantic feature map; calculate the weight of each pixel point in the first semantic feature map based on all center point positions, and obtain a center point heat map based on the weight; wherein the weight is the probability that the pixel point is the center point position; extract the re-identification features of the center point position based on the first re-identification feature map and the center point heat map, and determine the re-identification features as the first re-identification features.

[0131] In some embodiments, the fourth acquisition module 304 is specifically used to: obtain the motion trajectory of the target vehicle before the Nth frame of the video, and determine the motion characteristics of the first parking space relative to the target vehicle based on the motion trajectory using the Kalman filter algorithm; predict the appearance characteristics of the third parking space based on the first-recognition characteristics; determine the target re-identification feature from the second-recognition feature map based on the motion characteristics and the appearance characteristics, wherein the similarity between the target re-identification feature and the motion characteristics is greater than a first preset value and the similarity between the target re-identification feature and the appearance characteristics is greater than a second preset value, and the second-recognition feature map corresponds to the second panoramic bird's-eye view; determine the third parking space information based on the target re-identification feature.

[0132] In some embodiments, the result detection module 305 is specifically used to: calculate the Mahalanobis distance between the third parking space and the second parking space based on the third parking space information and the second parking space information; determine the second-recognition feature of the second parking space based on the second parking space information, and calculate the cosine distance between the target re-recognition feature and the second-recognition feature; obtain the target distance between the third parking space and the second parking space based on the Mahalanobis distance and the cosine distance; when the target distance is greater than or equal to a preset distance threshold, match the third parking space and the second parking space as the same parking space, and determine the detection result of the second parking space information as the final result.

[0133] In some embodiments, the result detection module 305 is specifically configured to obtain the target distance using the following formula based on the Mahalanobis distance and the cosine distance:

[0134] D=λD r +(1-λ)D m

[0135] Where D represents the target distance, D r represents the cosine distance, D m represents the Mahalanobis distance, and λ represents the weighting parameter.

[0136] In some embodiments, the result detection module 305 is further used to: when the target distance is less than a preset distance threshold, obtain the overlapping area between the bounding box of the second parking space and the bounding box of the third parking space and the total area after overlapping, and calculate the ratio of the overlapping area to the total area; when the ratio is greater than a preset ratio, match the third parking space and the second parking space as the same parking space, and determine the detection result of the second parking space information as the final result; when the ratio is less than or equal to the preset ratio, determine the second parking space as an unmatched parking space; when the unmatched parking space still has not been matched to the corresponding third parking space after a preset number of frames, determine the unmatched parking space as a falsely detected parking space, and delete the information of the second parking space from the second parking space information.

[0137] In some embodiments, the third acquisition module 303 is also used to: determine the converted coordinates of the first center point of the first parking space in the first semantic feature map based on the image size ratio; wherein the converted coordinates are the ratio of the coordinates of the first center point to the image size ratio; calculate the difference between the coordinates corresponding to the center point position and the converted coordinates, and correct the error between the first re-identification feature map and the first panoramic overhead view based on the difference.

[0138] It should be noted that the device provided in this application can implement all the method steps executed by the above method and can achieve the same technical effect, which will not be repeated here.

[0139] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0140] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0141] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0142] Figure 4 Schematic diagram of the electronic device 4 provided in the embodiment of the present application. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable by the processor 401. When the processor 401 executes the computer program 403, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the above-described device embodiments are implemented.

[0143] The electronic device 4 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 4 may include but is not limited to a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 This is merely an example of the electronic device 4 and does not limit the electronic device 4 . The electronic device 4 may include more or fewer components than shown in the figure, or different components.

[0144] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0145] Memory 402 can be an internal storage unit of electronic device 4, such as a hard disk or memory of electronic device 4. Memory 402 can also be an external storage device of electronic device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on electronic device 4. Memory 402 can also include both an internal storage unit of electronic device 4 and an external storage device. Memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0146] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0147] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The readable storage medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0148] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for parking space detection, characterized in that: include: Obtain a first panoramic bird's-eye view corresponding to the Nth video frame and a second panoramic bird's-eye view corresponding to the N+1th video frame of the target vehicle, where N is greater than or equal to 1; determining first parking space information based on the first panoramic bird's-eye view, and determining second parking space information based on the second panoramic bird's-eye view, wherein the first parking space information includes a size and a first center point of a first bounding box of at least one first parking space in the first panoramic bird's-eye view, and the second parking space information includes a size and a second center point of a second bounding box of at least one second parking space in the second panoramic bird's-eye view; determining a first-recognition feature of the first parking space according to the first parking space information; Predicting third parking space information of the (N+1)th video frame based on the first-recognition feature, wherein the third parking space information includes a size and a third center point of a third bounding box of the predicted third parking space; determining, based on the third parking space information, whether a detection result of the second parking space information is a final result; Determining a first-recognition feature of the first parking space according to the first parking space information includes: Based on the first panoramic bird's-eye view, a first semantic feature map corresponding to the first panoramic bird's-eye view is obtained; based on the image size ratio between the first semantic feature map and the first panoramic bird's-eye view, the center point position of the first center point of the first parking space in the first semantic feature map is determined; based on the first semantic feature map, the re-identification feature of each pixel point in the first semantic feature map is extracted to obtain a first re-identification feature map corresponding to the first semantic feature map; based on all the center point positions, a center point heat map is obtained; based on the first re-identification feature map and the center point heat map, the re-identification feature of the center point position is extracted, and the re-identification feature of the center point position is determined as the first re-identification feature.

2. The method according to claim 1, characterized in that The determining, based on the first parking space information, a first-recognition feature of the first parking space specifically includes: According to the first panoramic bird's-eye view, using a preset semantic segmentation model, obtaining a first semantic feature map corresponding to the first panoramic bird's-eye view; Determining a center point position of a first center point of the first parking space in the first semantic feature map according to an image size ratio between the first semantic feature map and the first panoramic bird's-eye view; the center point position is a floor-rounded value or a ceiling-rounded value of a ratio of a coordinate of the first center point to the image size ratio; Extracting, from the first semantic feature map, a re-identification feature of each pixel in the first semantic feature map using a feature re-identification algorithm to obtain a first re-identification feature map corresponding to the first semantic feature map; Calculating the weight of each pixel in the first semantic feature map based on all the center point positions, and obtaining the center point heat map based on the weight; wherein the weight is the probability that the pixel is the center point position; According to the first re-identification feature map and the center point heat map, the re-identification feature of the center point position is extracted, and the re-identification feature of the center point position is determined as the first re-identification feature.

3. The method according to claim 1, characterized in that The predicting, based on the first recognition feature, the third parking space information of the N+1th video frame includes: Obtaining a motion trajectory of the target vehicle before the Nth video frame, and determining a motion characteristic of the first parking space relative to the target vehicle using a Kalman filter algorithm based on the motion trajectory; Predicting appearance features of the third parking space based on the first-recognition features; determining a target re-identification feature from a second re-identification feature map based on the motion feature and the appearance feature, wherein a similarity between the target re-identification feature and the motion feature is greater than a first preset value and a similarity between the target re-identification feature and the appearance feature is greater than a second preset value, and the second re-identification feature map corresponds to the second panoramic bird's-eye view; The third parking space information is determined according to the target re-identification feature.

4. The method according to claim 3, characterized in that The determining, based on the third parking space information, whether the detection result of the second parking space information is a final result includes: Calculating a Mahalanobis distance between the third parking space and the second parking space according to the third parking space information and the second parking space information; determining a second-recognition feature of the second parking space according to the second parking space information, and calculating a cosine distance between the target re-recognition feature and the second-recognition feature; Obtaining a target distance between the third parking space and the second parking space according to the Mahalanobis distance and the cosine distance; When the target distance is greater than or equal to a preset distance threshold, the third parking space and the second parking space are matched as the same parking space, and the detection result of the second parking space information is determined as the final result.

5. The method according to claim 4, characterized in that Obtaining a target distance between the third parking space and the second parking space according to the Mahalanobis distance and the cosine distance includes: According to the Mahalanobis distance and the cosine distance, the target distance is obtained using the following formula: in, represents the target distance, represents the cosine distance, represents the Mahalanobis distance, represents the weighting parameter.

6. The method according to claim 4, characterized in that After obtaining the target distance between the third parking space and the second parking space according to the Mahalanobis distance and the cosine distance, the method further includes: When the target distance is less than the preset distance threshold, obtaining an overlapping area between the bounding box of the second parking space and the bounding box of the third parking space and a total area after the overlap, and calculating a ratio of the overlapping area to the total area; When the ratio is greater than a preset ratio, matching the third parking space and the second parking space as the same parking space, and determining the detection result of the second parking space information as the final result; If the ratio is less than or equal to a preset ratio, determining that the second parking space is an unmatched parking space; If the unmatched parking space is still not matched to the corresponding third parking space after a preset number of frames, the unmatched parking space is determined to be a misdetected parking space, and the information of the second parking space is deleted from the second parking space information.

7. The method according to claim 2, characterized in that Before extracting the re-identification feature of the center point position according to the first re-identification feature map and the center point heat map, and determining the re-identification feature of the center point position as the first re-identification feature, the method further includes: Determining, according to the image size ratio, converted coordinates of a first center point of the first parking space in the first semantic feature map; wherein the converted coordinates are a ratio of the coordinates of the first center point to the image size ratio; The difference between the coordinates corresponding to the center point position and the converted coordinates is calculated, and the error between the first re-identified feature map and the first panoramic bird's-eye view is corrected according to the difference.

8. A parking space detection device, characterized in that: include: A first acquisition module is configured to acquire a first panoramic bird's-eye view corresponding to an N-th video frame and a second panoramic bird's-eye view corresponding to an N+1-th video frame of the target vehicle, where N is greater than or equal to 1; a second acquisition module, configured to determine first parking space information based on the first panoramic bird's-eye view, and to determine second parking space information based on the second panoramic bird's-eye view, wherein the first parking space information includes a size and a first center point of a first bounding box of at least one first parking space in the first panoramic bird's-eye view, and the second parking space information includes a size and a second center point of a second bounding box of at least one second parking space in the second panoramic bird's-eye view; a third acquisition module, configured to determine a first-recognition feature of the first parking space based on the first parking space information; a fourth acquisition module, configured to predict third parking space information of the (N+1)th video frame based on the first-recognition feature, wherein the third parking space information includes a size and a third center point of a third bounding box of the predicted third parking space; a result detection module, configured to determine whether a detection result of the second parking space information is a final result based on the third parking space information; Determine a first re-identification feature of the first parking space based on the first parking space information, including: obtaining a first semantic feature map corresponding to the first panoramic bird's-eye view based on the first panoramic bird's-eye view; determining a center point position of a first center point of the first parking space in the first semantic feature map based on an image size ratio between the first semantic feature map and the first panoramic bird's-eye view; extracting a re-identification feature of each pixel in the first semantic feature map based on the first semantic feature map to obtain a first re-identification feature map corresponding to the first semantic feature map; obtaining a center point heat map based on all the center point positions; extracting a re-identification feature of the center point position based on the first re-identification feature map and the center point heat map, and determining the re-identification feature of the center point position as the first re-identification feature.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Driving speed determination method and device

    CN110399664A

  • Parking space tracking method and device, vehicle and storage medium

    CN115223135A