Parking position detection method and system
The method and system use a camera and machine learning to generate bird's-eye view images for parking space detection, addressing inefficiencies in existing systems by enabling accurate and autonomous parking space identification.
Patent Information
- Application Number
- JP2025539966
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-06
- Filing Date
- 2023-12-14
- Publication Date
- 2026-01-23
AI Technical Summary
Existing parking space detection systems struggle to accurately and efficiently identify available parking spaces under various conditions, leading to a general distrust of driver assistance systems due to their slow computational methods and limited detection capabilities.
A method and system utilizing a first camera to acquire images, generate a bird's-eye view image, and process it with a machine learning model to predict parking space data, including center coordinates, corner displacements, and reliability, enabling the vehicle to autonomously park in suitable spaces.
The system provides robust and efficient detection of parking spaces regardless of orientation or environmental conditions, enhancing the reliability of driver assistance systems by accurately identifying and parking in available spaces.
Smart Images

Figure 2026502481000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION Embodiments of the present invention relate to a parking location detection method and system. [Background technology]
[0002] Parking a vehicle is known to be a stressful task for drivers, and as such, various automakers have been working to automate the parking task through in-vehicle driver assistance systems. To accurately park a vehicle, it is necessary to identify available parking slots.
[0003] Typically, parking space detection is accomplished by combining one or more sensor systems installed on a vehicle with computational techniques that process sensor data acquired from the sensor systems. A sensor system may include a transmitter (e.g., laser light) and one or more sensors. Exemplary sensor systems include, but are not limited to, cameras, LiDAR, radar, and ultrasonic sensors. However, even with available sensor data, there are many challenges to accurately and efficiently detecting available parking spaces.
[0004] Computational methods for processing vehicle sensor data to identify available parking spaces are often slow and unsuitable for real-time use. Furthermore, many parking space detection systems may only be able to detect parking spaces under certain conditions, such as parking spaces with a specific orientation, clearly delineated boundaries, or adjacent to parking spaces occupied by other vehicles. The inability to quickly and consistently identify available parking spaces under a variety of conditions contributes to a general distrust of driver assistance systems. Summary of the Invention [Means for solving the problem]
[0005] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in limiting the scope of the claimed subject matter.
[0006] An embodiment of the present disclosure relates to a method for acquiring a first image from a first camera mounted on a vehicle and generating a bird's-eye view (BEV) image using the first image. The method further includes processing the bird's-eye view image using a machine learning model to generate parking space prediction data. The parking space prediction data includes a first center coordinate of a first available parking space, a first parking space reliability, and first corner displacement data. The first corner displacement data includes a first relative coordinate pair indicating a position of a first corner relative to the first center coordinate, and a second relative coordinate pair indicating a position of a second corner relative to the first center coordinate. The method further includes determining a first position of the first available parking space using the parking space prediction data, and parking the vehicle in the first available parking space if the first parking space reliability satisfies a threshold.
[0007] An embodiment of the present disclosure further relates to a computer-implemented method for training a machine learning model. The method includes acquiring a plurality of bird's-eye view (BEV) images and identifying available parking spaces in the bird's-eye view images. For each parking space identified in the bird's-eye view images, the method further includes determining a ground truth representation of the identified parking space, the ground truth representation including a center coordinate and a corner displacement; calculating an envelope width and an envelope height using the center coordinate and the corner displacement; and matching the identified parking space with one or more anchor boxes using the envelope height and the envelope width. Once the ground truth representation of each parking space identified in a bird's-eye view image has been determined, the method further includes generating a target data structure including the center coordinate and the corner displacement of each identified parking space in the bird's-eye view image. The method further includes generating a training dataset including a plurality of bird's-eye view images and their associated target data structures, and training a machine learning model using the training dataset. The machine learning model is configured to directly receive one or more of the bird's-eye view images.
[0008] The embodiments disclosed herein also relate to a system including a vehicle, a first camera, a bird's-eye view (BEV) image, a machine learning model, and a computer. The computer includes one or more processors and is configured to acquire a first image from the first camera, construct a bird's-eye view image from the first image, process the bird's-eye view image using the machine learning model, and generate available parking space prediction data. The available parking space prediction data includes a first center coordinate of the first available parking space, a first parking space reliability, and first corner displacement data. The first corner displacement data includes a first relative coordinate pair that positions the first corner relative to the first center coordinate and a second relative coordinate pair that positions the second corner relative to the first center coordinate. The computer is further configured to determine a first position of the first available parking space using the available parking space prediction data, and park the vehicle in the first available parking space without driver assistance if the first parking space reliability meets a threshold.
[0009] Other aspects and advantages of the claimed subject matter will become apparent from the following description and appended claims. [Brief explanation of the drawings]
[0010] Certain embodiments disclosed herein will now be described in detail with reference to the accompanying drawings, in which like elements in the various figures are designated with the same reference numerals for consistency.
[0011] [Figure 1] 1 illustrates a parking lot according to one or more embodiments. [Figure 2A] 1 illustrates a right angle parking stall orientation according to one or more embodiments. [Figure 2B] 1 illustrates a parallel parking stall orientation arrangement according to one or more embodiments. [Figure 2C] 1 illustrates a fishbone parking stall orientation arrangement according to one or more embodiments. [Figure 3] 1 illustrates the fields of view of multiple cameras mounted on a vehicle according to one or more embodiments. [Figure 4] 1 illustrates images captured from multiple cameras mounted on a vehicle according to one or more embodiments. [Figure 5] 1 illustrates an example of a bird's eye view (BEV) image according to one or more embodiments. [Figure 6] 1 illustrates a system according to one or more embodiments. [Figure 7A] 1 illustrates a parking space defined in absolute coordinates according to one or more embodiments. [Figure 7B] 1 illustrates a parking space defined in center-relative coordinates according to one or more embodiments. [Figure 7C] 1 illustrates a fishbone parking space defined in center-relative coordinates according to one or more embodiments. [Figure 8] 1 illustrates an output data structure of a machine learning model according to one or more embodiments. [Figure 9] 1 illustrates an example of a bird's-eye view image with predicted parking space data according to one or more embodiments. [Figure 10]1 illustrates a neural network according to one or more embodiments. [Figure 11] 1 illustrates a ground truth parking representation and two anchor boxes according to one or more embodiments. [Figure 12] 1 illustrates ground truth and predicted parking representations according to one or more embodiments. [Figure 13] 1 shows a flowchart according to one or more embodiments. [Figure 14] 1 shows a flowchart according to one or more embodiments. [Figure 15] 1 shows a flowchart according to one or more embodiments. [Figure 16] 1 illustrates a system according to one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0012] In the detailed description of the embodiments of the present disclosure, many specific details are set forth to enable a fuller understanding of the present disclosure. However, it may be apparent to those skilled in the art that the present disclosure can be practiced without these specific details. Also, in some cases, detailed descriptions of well-known functions are omitted to avoid unnecessarily complicating the description.
[0013] Throughout this specification, ordinal numbers (e.g., "first," "second," "third," etc.) may be used as adjectives to refer to elements (i.e., any nouns herein). The use of ordinal numbers is not intended to imply or establish a particular order for the elements or to limit an element to a single element unless expressly disclosed (e.g., using terms such as "before," "after," "single," etc.). Rather, the use of ordinal numbers is intended to distinguish one element from another. For example, a first element is different from a second element, which may include multiple elements or may be located after (or before) the second element.
[0014] As used herein, the singular forms "a," "an," and "the" are understood to include the plural forms as well, unless the context clearly dictates otherwise. Thus, for example, a reference to an "acoustic signal" includes one or more such acoustic signals.
[0015] Terms such as "approximately" and "substantially" mean that the stated characteristic, parameter, or value need not be exactly achieved, but rather that variations and deviations, including, for example, tolerances, measurement errors, limits of measurement accuracy, and other factors known to those skilled in the art, are allowed to the extent that the characteristic does not interfere with the intended effect.
[0016] It should be understood that one or more steps shown in the flowcharts may be omitted, repeated, and / or performed in a different order than that shown, and therefore the scope of the disclosure herein should not be limited to the particular order of steps shown in the flowcharts.
[0017] Although multiple dependent claims are not introduced herein, it will be clear to those skilled in the art that the subject matter of a dependent claim in one or more embodiments can be combined with other dependent claims.
[0018] In the description of Figures 1 through 16 herein, any component described with respect to a figure may be equivalent to one or more components of the same name described with respect to other figures in other embodiments disclosed herein. For the sake of brevity, the description of these components will not be repeated for each figure. Accordingly, embodiments of the component in each figure are incorporated by reference as if they were optionally present in all other figures that include the same-named component. Furthermore, according to various embodiments disclosed herein, the description of a component in a figure should be interpreted as any embodiment that can be implemented in addition to, in combination with, or instead of other embodiments described with respect to the corresponding like-named component.
[0019] The embodiments disclosed herein describe a method and system for detecting parking spaces around a vehicle using one or more images acquired from one or more cameras mounted on the vehicle. The parking space detection method and system disclosed herein identify available or partially available parking spaces using a parking space representation consisting of at least two corners. Each corner of the parking space representation represents a specific corner of the parking space. For example, if the parking space representation is a polygon with four corners, each of the four corners is explicitly designated as either entrance-left, entrance-right, back-left, or back-right. Furthermore, the corner designations of the parking space are mutually exclusive. Therefore, the parking space representation can directly adapt to any parking space orientation, regardless of the vehicle's relative viewpoint, without rotation or affine transformation, while simultaneously encoding the relationship between the corners.
[0020] FIG. 1 illustrates an example parking lot 100 including parking spaces 102. To reduce clutter, each parking space 102 is not labeled. The parking spaces 102 may be occupied by a parked vehicle 204 or other object (e.g., a trash can 104), or may be unoccupied and available. In some cases, the parking lot 100 may be outdoors (as shown) or indoors, such as in a parking garage (not shown). In other cases, the parking spaces 102 may be located near a road. The surface material of the parking spaces 102 varies depending on their location and construction. Non-limiting examples of surface materials for the parking spaces 102 include asphalt, concrete, brick, paving stone, etc.
[0021] Generally, parking spaces (102) have an associated frame orientation. Figures 2A-2C show parking spaces (102) with various frame orientations. Again, for clarity, not all parking spaces (102) are labeled. Figures 2A-2C each show a vehicle (202) searching for an available parking space or intending to park, and one or more parked vehicles (204) already in the parking space. Additionally, Figures 2A-2C show at least one vacant and available parking space (205). Figure 2B also shows another vehicle (203) that is neither parked nor searching for an available parking space. Specifically, Figure 2A shows the parking space (102) with a rectangular frame orientation (206) when occupied by a parked vehicle (204) and when vacant. Figure 2B shows the parking space (102) adjacent to a road (209) and with a parallel frame orientation (208). FIG. 2C shows a parking stall with a fishbone or diagonal stall orientation (210).
[0022] 2A-2C further illustrate that the parking stall (102) may be marked in a variety of ways. For example, as seen in FIG. 2A, the parking stall (102) may be marked or surrounded by a solid line. Alternatively, as seen in FIGS. 2B and 2C, the parking stall (102) may be marked by a dashed line. In other cases, some or all of the markings may be implicit (e.g., a curb, a parking stall entrance, etc.), explicitly indicated, or indistinct (e.g., blurred). Furthermore, the markings defining the parking stall (102) may partially surround the parking stall (102), as shown in FIGS. 2A and 2C, or completely surround the parking stall (102), as shown in FIG. 2B. Additionally, as shown in FIG. 2A, one or more corners of the parking stall (102) may be indicated by an "L-corner" marking (212) or a "T-corner" marking (214). In other cases, corner markings may not be used (e.g., FIG. 2C).
[0023] In general, a parking space (102) may have any combination of the above-described characteristics. That is, a parking space (102) may have a combination of surface material (e.g., asphalt, concrete), installation environment (e.g., indoor or outdoor, parking lot or adjacent to a road), space orientation (perpendicular (204), parallel (208), fishbone (210)), line marking style (dashed, solid, mixed), corner type (e.g., T-corner, L-corner), and marking enclosure (partially enclosed, fully enclosed). Those skilled in the art will appreciate that alternative and / or additional characteristics, such as vehicle approach angle and space width, may be used to define a parking space. Therefore, due to the wide variety of configurations and characteristics of a parking space (102), a comprehensive list is not necessary and does not limit the scope of the present disclosure.
[0024] As will be described below, the parking space detection method and system disclosed herein accommodates any configuration of parking spaces 102. Additionally, the method and system are robust to environmental conditions such as nighttime and rainy weather, and when the parking space 102 is partially obstructed.
[0025] FIG. 3 illustrates a surrounding camera view (300) of a vehicle (301). According to one or more embodiments, one or more cameras may be installed on the vehicle (301) to detect parking spaces (102) around the vehicle (301) (e.g., around the search vehicle (202)). According to one or more embodiments, one or more cameras may be equipped with fisheye lenses. In other embodiments, the cameras installed on the vehicle may be pinhole (i.e., regular) cameras. The cameras themselves are not explicitly depicted in FIG. 3. However, the locations of four cameras located on the vehicle (301) are indicated by circles. Typically, the cameras installed on the vehicle (301) are small compared to the vehicle itself. As such, their small size allows them to be installed in a discreet location that does not interfere with the vehicle's functionality. Furthermore, each of the one or more cameras may be concealed, such as under the vehicle's trim, and may be constructed with a robust structure that can withstand vibration, operate over a wide temperature range, and be waterproof.
[0026] 3 shows a front camera (304), two side cameras (306), and a rear camera (308). The number and locations of the cameras installed are not limited to those shown in FIG. 3. For example, depending on the embodiment, a fewer or more number of cameras may be used. For example, according to one or more embodiments, a camera or camera system with all-around visibility may be installed on top of the vehicle (301).
[0027] A camera has a field of view (FOV). Here, field of view is a general term meaning the extent of the observable world captured by the camera. More technically, it refers to the solid angle over which the camera's sensor can sense electromagnetic waves. The field of view of each of the four cameras shown in Figure 3 is indicated by the curved dashed lines. The front camera (304) has a forward field of view (310) and can capture the environment in front of the vehicle (301). Similarly, the side cameras (306) are associated with either a left field of view (312) or a right field of view (314), where left and right are relative directions to the vehicle (301) as viewed from above in Figure 3. Finally, the rear camera (308) has a rear field of view (316) and can capture the environment behind the vehicle (301). As shown in Figure 3, the camera fields of view overlap, allowing the entire surroundings of the vehicle (301) to be captured. As mentioned above, it is also possible to capture the entire surroundings all at once using a camera system equipped with one or more cameras installed on top of the vehicle (301).
[0028] Continuing with the example of Figure 3, it should be noted that the installation location of one or more cameras mounted on the vehicle (301) (e.g., the height of the one or more cameras relative to the ground) can directly affect the field of view of the one or more cameras and the visibility of parking stalls (102) around the vehicle (301). Those skilled in the art will appreciate that the height of one or more cameras mounted on the vehicle (301) can be adjusted and / or selected to optimize the camera's field of view without departing from the scope of the present disclosure. Additionally, the one or more cameras can be positioned such that the field of view of the one or more cameras is not obstructed by structures on the vehicle (301).
[0029] Figure 4 shows four example images captured using one of four cameras installed on the vehicle (301) shown in Figure 3. Because each image corresponds to a specific camera installed on the vehicle (301), the images can be described as a front image (402) (from the front camera (304)), a left image (404) (from the side camera (306) with a left-facing field of view (312)), a right image (406) (from the side camera (306) with a right-facing field of view (314)), and a rear image (408) (from the rear camera (308)). These example images have overlapping fields of view and form the ambient camera view (300). As an example of overlapping fields of view, a parked vehicle (204) is shown appearing in both the left image (404) and the rear image (408).
[0030] The example image of FIG. 4 was captured at a moment when a vehicle (301) was traveling through a parking lot (100). As shown in FIG. 4, the parking lot (100) is outdoors and parking spaces (102) are present, although not all of the parking spaces are labeled in FIG. 4. The parking spaces (102) in FIG. 4 are vertically oriented (206) and are demarcated by partially enclosed solid lines. Specifically, the boundaries of each parking space (102) are indicated by solid lines, with the rear side implicitly marked by a curb and the entrance side unmarked. Additionally, an unavailable area (420) of the parking lot (100) is visible in the example image of FIG. 4.
[0031] In one or more embodiments, images acquired from one or more cameras mounted on the vehicle (301) are stitched together to create a bird's-eye view (BEV) image or representation of the vehicle's surroundings. In one or more embodiments, the bird's-eye view image is often created from images acquired from one or more cameras positioned on the vehicle (301) using a technique known as inverse perspective projection (IPM). IPM typically assumes that the vehicle's (301) surroundings are flat, and after calibration, maps the image pixels onto a plane using homography projection. Figure 5 shows an example bird's-eye view image (500) created from the example image of Figure 4. The example bird's-eye view image (500) shows the vehicle's (301) surroundings (parking lot (100)) including parking stalls (102), unusable areas (420), and parked vehicles (204). Note that for simplicity, not all parking stalls are labeled. As such, IPM may distort non-planar objects. For example, in Figure 5, the parked vehicle (204) appears to stretch to infinity. However, the bird's-eye view image generally has little effect on the representation of parking spaces (102) and other planar objects, and can therefore be used to identify and locate parking spaces.
[0032] In one or more embodiments, the vehicle (301) may be equipped with additional sensor systems. These sensor systems may include ultrasonic systems and light detection and ranging (LiDAR) systems, which are comprised of one or more sources (e.g., ultrasonic transmitters, lasers) and receivers or sensors. In one or more embodiments, sensor data from the additional sensor systems is used to complement the bird's-eye view image and to assist in locating the vehicle (301) and parking stalls (102) in the bird's-eye view image.
[0033] FIG. 6 illustrates an overview of a parking space detection method according to one or more embodiments. As shown, one or more cameras mounted on a vehicle (301) are used to capture one or more ambient images (602). The term "ambient image" refers to an image capturing at least a portion of the vehicle's (301) surroundings. In one or more embodiments, the camera mounted on the vehicle (301) includes a fisheye lens and has an image resolution of 1280 x 800 pixels. For simplicity, the following description assumes there are at least two ambient images (602), and plural references are not intended to create ambiguity. However, embodiments of the present disclosure can operate with a single ambient image (602). Therefore, the ambient image (602) need not consist of multiple images. The ambient image (602), as the name suggests, is an image of the vehicle's (301) local environment at a given point in time. The surroundings image 602 is processed by an inverse perspective projection (IPM) module 604, which applies IPM techniques to generate a bird's-eye view image 606. Note that in one or more embodiments, the bird's-eye view image 606 may be generated without the IPM module 604, for example, when a surroundings view camera system mounted on top of the vehicle 301 is used. In one or more embodiments, the generated bird's-eye view image 606 has a resolution of 640 x 640 pixels and covers a physical area of 25 meters x 25 meters. In this embodiment, the length of one pixel corresponds to 3.9 centimeters. Furthermore, in one or more embodiments, the vehicle 301 is positioned at the center of the bird's-eye view image 606. Therefore, the bird's-eye view image 606 can depict parking stalls 102 located up to 12.5 meters away from the vehicle 301. That is, in one or more embodiments, the detection area for the vehicle 301 is 25 meters by 25 meters, and the detection area is centered on the vehicle 301. Note that the detection area can be changed by adjusting one or more cameras (e.g., camera type, camera installation position, etc.), and is not necessarily limited to 25 meters by 25 meters, and may be smaller or larger in other embodiments.As shown in FIG. 6 , the bird's-eye view image (606) is processed by a machine learning model (608) to generate parking space prediction data (610). The machine learning model (608) is described in more detail below. However, for now, the machine learning model (608) is trained and configured to receive the bird's-eye view image (606) as input and identify the locations of parking spaces (102) depicted in the bird's-eye view image (606), regardless of the placement of the parking spaces (102) or other environmental factors (e.g., time of day, weather conditions). The output of the machine learning model (608) is parking space prediction data (610), which includes, for all identified parking spaces (102) depicted in the target bird's-eye view image (606), the locations of the corners (two or more) of the parking representation, the road surface material of each parking space (102), and a confidence level for each parking space (102). Additionally, the parking stall prediction data (610) can be used to determine the approach line data and stall direction for each parking stall (102).
[0034] To better understand the output included in the parking space prediction data (610), it is useful to illustrate how a parking space (102) is represented by a parking representation. According to one or more embodiments, a four-sided parking representation is used to represent the parking space (102) in an enclosed form. The four-sided parking representation has four corners. In one or more embodiments, each corner of the parking representation represents a specific corner of the parking space (102). That is, the four corners are explicitly designated as either entrance-left, entrance-right, end-left, or end-right. Furthermore, the corner designations of the parking space are mutually exclusive. Figures 7A through 7C each illustrate a parking space (102) represented by a four-sided parking representation and a search vehicle (202) attempting to enter and park in the detected parking space (102). Note that not all parking stalls (102) are labeled in Figures 7A-7C to avoid cluttering the illustrations.
[0035] As shown in FIG. 7A , the parking representation has four uniquely labeled corners: entrance-left corner (702), entrance-right corner (704), back-left corner (706), and back-right corner (708). The designations of "left" and "right" may be relative to the vehicle's orientation, relative to the bird's-eye view image (606), or relative to some other fixed reference, but the relative designations must be applied consistently. Generally, the parking representation is fully defined by describing the corner designations, if more than one corner can be used. Because the pixels of the bird's-eye view image (606) are spatially distributed on a plane, the coordinate system is a two-dimensional coordinate system with axes parallel to the edges of the bird's-eye view image (606). For example, in one or more embodiments, the x-axis is defined along the width direction of the bird's-eye view image (606), and the y-axis is defined along the height direction of the bird's-eye view image (606). The origin of the coordinate system can be arbitrarily defined but must be fixed across all provided bird's-eye view images (606). In one or more embodiments, the origin of the coordinate system is located at the center of each bird's-eye view image (606). It should be noted that the center of a bird's-eye view image (606) is typically aligned with the center of the rear axle of the vehicle (301) as a common standard. The units of the coordinate system are easily convertible between pixel space and physical space, and the parking space detection method and system herein are not limited to the unit selection. Furthermore, in one or more embodiments, the coordinate system may use normalized units. For example, the x-axis and y-axis coordinates are given relative to the width and height of the bird's-eye view image (606), respectively, and may be in either pixel space or physical space. In one or more embodiments, as will be described in more detail below, the coordinate system used for the bird's-eye view image (606) may be piecewise defined using a grid of cells overlaid on the bird's-eye view image (606). In this case, the coordinate system may be normalized according to the width and height of each grid cell.
[0036] The location of corners in a parking representation can be described using absolute coordinates or center-relative coordinates. The parking representation shown in FIG. 7A, which represents a parking stall (102), uses absolute coordinates (701). In this case, each corner is directly specified on the x- and y-axes using a coordinate pair (x, y). These coordinates may be normalized and may refer to an origin defined in the bird's-eye view image (606) or an origin defined by a grid cell in the bird's-eye view image (606). To distinguish corners from other corners in the parking representation, each coordinate pair is assigned an index corresponding to the corresponding corner. Four corners are used in FIGS. 7A-7C: the entrance left corner (702) is assigned index 1, the entrance right corner (704) is assigned index 2, the back left corner (706) is assigned index 3, and the back right corner (708) is assigned index 4. Thus, for example, using the absolute coordinates (701) shown in FIG. 7A, the location of the back right corner (708) is indicated by the coordinate pair (x4, y4). Assigning indices to corners is arbitrary, and those skilled in the art will understand that any mutually exclusive index selection for the corners of a parking representation can be used, and the indices shown herein are not intended to be limiting, so long as the selection is consistently applied. In absolute coordinates (701), an indexed coordinate pair (e.g., (x1, y1)) indicates both the referenced corner and its location, so that the parking representation encompassing the area of the parking space (102) is fully defined using the indexed coordinate pairs corresponding to each corner of the parking representation. In Figure 7A, a rectangular parking representation is depicted using coordinate pairs (x1, y1), (x2, y2), (x3, y3), and (x4, y4).
[0037] Figure 7B shows the same rectangular parking representation as Figure 7A, but with center-relative coordinates (703). As before, the parking representation is fully defined by indicating the locations of designated (e.g., indexed) corners. As in Figure 7A, the entrance left corner (702) is index 1, the entrance right corner (704) is index 2, the back left corner (706) is index 3, and the back right corner (708) is index 4. Additionally, Figure 7B shows a coordinate pair (x c , y c ) is given. Here, this coordinate pair is written as center coordinate (712). The center coordinate (712) has the following relationship with the corner coordinate pair in absolute coordinate (701):
number
[0038] As shown in FIG. 7B, using the center-relative coordinates (703), the corners of the parking representation are positioned relative to the center coordinates (712). That is, the coordinate pair of each corner is expressed as a relative coordinate pair indicating the displacement in the x-axis and y-axis directions from the center coordinates (712). For example, using the center-relative coordinates (703), the position of the entrance left corner (702) is given by the relative coordinate pair (Δx1, Δy1). Similarly, the remaining corners can be expressed by relative coordinate pairs. Mathematically, the relative coordinate pair using the center-relative coordinates (703) of the nth corner has the following relationship with the coordinate pair of the absolute coordinates (701):
number
number
[0039] According to one or more embodiments, the machine learning model (608) is configured using center-relative coordinates (703). FIG. 7C shows another example of a parking representation of a parking space (102) using center-relative coordinates (703). For simplicity, not all of the parking spaces (102) or corners of the parking representation are labeled in FIG. 7C. The parking representation in FIG. 7C is not different in definition from the rectangular parking representation in FIG. 7B, but its shape is not rectangular. FIG. 7C is shown to highlight an advantage of the present disclosure: the parking representation can be directly adapted to represent parking spaces (102) of any orientation and shape (including non-rectangular shapes), regardless of the relative viewpoint of the vehicle (301), without requiring additional rotations or affine transformations. This provides a significant advantage over existing parking detection systems that are limited to representing parking spaces (102) as rectangles or that can only detect parking spaces (102) in specific orientations relative to the vehicle (301).
[0040] Additionally, one advantage of providing a specific designation for each corner (e.g., the left entrance corner (702) rather than simply a corner) is that it allows for easy identification of the approach line (710) that marks the boundary of the parking space (102) into which the vehicle (301) (search vehicle (202)) should enter. According to one or more embodiments, the approach line (710) is defined as the straight line connecting the left entrance corner (702) and the right entrance corner (704). Mathematically, the approach line (710) of the parking representation is represented by the following set of points:
number
[0041] In one or more embodiments, multiple approach lines may be determined for any detected parking space (102). For example, in one or more embodiments, the parking representation may be used in conjunction with one or more additional object detectors capable of locating other objects, such as vehicles or roads. In this case, the spatial relationship between the parking space (102) and surrounding objects is used to determine one or more approach lines for that parking space (102). For example, in FIG. 7A , the search vehicle (202) has detected an available parking space (102) bounded on the top and bottom by open areas (e.g., roads) and on the left and right by other available parking spaces. Therefore, using this information, any line segment directly connecting adjacent corners of the parking representation may be treated as a valid approach line.
[0042] 8 illustrates in more detail the information contained in the parking space prediction data (610). As shown in the exemplary bird's-eye view image (500) of FIG. 5, the bird's-eye view image (606) may include multiple parking spaces (102). In general, the bird's-eye view image (606) may include zero or more parking spaces (102). The parking space prediction data (610) includes individual information about each parking space (801) detected in the bird's-eye view image (606) that is input to the machine learning model (608). Specifically, for each detected parking space (801), the parking space prediction data (610) includes a predicted center coordinate (712), a center relative coordinate pair corresponding to each corner of the parking representation (collectively referred to as corner displacement (804)), a confidence level ranging from 0 to 1 indicating the visibility of each corner (collectively referred to as corner visibility (806)), a parking space confidence level (810) ranging from 0 to 1 indicating the confidence level that the parking representation formed using the center coordinates (712) and corner displacement (804) accurately surrounds the available parking space (102), and a category prediction (814) regarding the surface material of the parking space. Additionally, the parking space prediction data (610) may be used to calculate a category prediction regarding the orientation of each parking space (102) (i.e., perpendicular, parallel, fishbone). That is, in one or more embodiments, the orientation of each parking space (102) is determined by a post-processing technique using the parking space prediction data (610).
[0043] Upon receiving the bird's-eye view image (606), the machine learning model (608) outputs parking space prediction data (610). The parking space prediction data (610) contains all the information necessary to unambiguously identify and locate available parking spaces (102) within the bird's-eye view image (606) based on their confidence (parking space confidence (810)). The bird's-eye view image (606) represents the vehicle's (301) surrounding environment. Furthermore, the bird's-eye view image (606) can be spatially mapped to the physical space surrounding the vehicle. Therefore, the parking space prediction data (610) can be used to physically detect (identify and locate) available parking spaces (102) in the vicinity of the vehicle (301). Furthermore, the parking space prediction data (610) can be combined with an on-board driver assistance system to automatically park the vehicle (301) (i.e., without driver input).
[0044] FIG. 9 shows an example of marking the bird's-eye view image (500) of FIG. 5 using the parking space prediction data (610) output by the machine learning model (608). As shown in FIG. 9, a quadrilateral parking space representation surrounding each detected available parking space (102) can be drawn using the center coordinates (712) and corner displacements (804) (four corners are used in this example). Furthermore, the approach line (710) of each parking space (102) is determined, and the parking space reliability (810) of each parking space (102) is shown in FIG. 9. FIG. 9 also shows the detection of an occluded parking space (902). This occluded parking space is partially occluded by a parked vehicle (204) from the viewpoint of the vehicle (301). Therefore, in the parking space representation representing the occluded parking space (902), the corner visibility of the far-left corner (706) and the far-right corner (708) is set to a value close to zero or thresholded to zero, indicating that these corners are not visible. Finally, as can be seen from Figure 9, the parking space representation can be used to easily calculate the surface area and approach line width of the detected parking space (102). Note that, to avoid cluttering Figure 9, not all corners, center coordinates (712), approach lines (710), and parking space confidence values (810) of all parking spaces (102) are assigned numerical labels or lines.
[0045] As previously mentioned, a machine learning model (608) is used to generate parking space prediction data (610) from the bird's-eye view image (606). Machine learning, broadly defined, refers to the extraction of patterns and insights from data. Terms such as "artificial intelligence," "machine learning," "deep learning," and "pattern recognition" are often confused, used interchangeably, and treated synonymously in the literature. This ambiguity stems from the fact that the field of "extracting patterns and insights from data" has developed simultaneously and independently in multiple classical fields, such as mathematics, statistics, and computer science. For consistency, the terms "machine learning" and "machine-learned" are used herein. However, those skilled in the art will understand that this choice of terminology does not limit the concepts and techniques described below.
[0046] Those skilled in the art will appreciate that machine learning encompasses a field and concepts so broad and profound that it would be impossible to fully describe them in this specification. However, to provide necessary context for the machine learning models used in one or more embodiments of the present invention, the following paragraphs provide a minimal description of neural networks and convolutional neural networks. Note that the following description is intended to provide a superficial understanding of some machine learning techniques and models and should not be construed as limiting the present disclosure.
[0047] Neural networks are a type of machine learning model. Neural networks are often used as subcomponents of larger machine learning models. A diagram of a neural network is shown in Figure 10. Roughly speaking, a neural network (1000) consists of nodes (1002) and edges (1004). In Figure 10, the nodes (1002) are represented by circles, and the edges (1004) are represented by directed lines. The nodes (1002) are grouped into layers (1005). In Figure 10, four layers (1008, 1010, 1012, 1014) represent the nodes (1002), and the nodes (1002) are grouped into columns, but this grouping is not limited to what is shown in Figure 10. Edges (1004) connect the nodes (1002). An edge (1004) may or may not connect to any node (1002), regardless of which layer (1005) the connected node (1002) belongs to. That is, nodes (1002) may be loosely or residually connected. However, if all nodes (1002) in a layer are connected to all nodes (1002) in an adjacent layer, the layer is called densely or fully connected. If all layers of a neural network (1000) are densely connected, the neural network (1000) is called a densely (fully connected) neural network. A neural network (1000) has at least two layers (1005), with the first layer (1008) called the "input layer" and the last layer (1014) called the "output layer." The intermediate layers (1010, 1012) are usually called "hidden layers." A neural network (1000) with one or more hidden layers (1010, 1012) is referred to as a "deep neural network" or "deep learning method." Thus, in some embodiments, the machine learning model is a deep neural network. In general, the output layer (1014) of the neural network (1000) can have one or more nodes (1002). In this case, the neural network (1000) is referred to as a "multi-target" or "multi-output" network.
[0048] There is an additional association between the nodes (1002) and the edges (1004): every edge has a numerical value associated with it. This edge value, or the edges (1004) themselves, are often called "weights" or "parameters." When training the neural network (1000), a numerical value is assigned to each edge (1004). Furthermore, every node (1002) has a numerical variable and an activation function associated with it. The activation function is not limited to a particular class of functions, but traditionally it often takes the following form:
number
number
number
number
[0049] When the neural network (1000) receives an input, the input propagates through the network according to the activation function, the values of the input nodes (1002) and the values of the edges (1004), and the value of each node (1002) is calculated. In other words, the value of each node (1002) may change with each input it receives. In some cases, the node (1002) is assigned a fixed value, such as 1, that is not affected by the input and does not change depending on the values of the edges (1004) or the activation function. The fixed node (1002) is often called the "bias" or "bias node" (1006) and is indicated by a dashed circle in Figure 10.
[0050] In some implementations, the neural network (1000) may include special layers (1005), such as normalization layers (when batch or layer-wise normalization is performed) or dropout layers, and may also include additional connection processes such as concatenation. Those skilled in the art will understand that these modifications do not depart from the scope of the present disclosure.
[0051] As previously mentioned, the neural network (1000) training procedure involves assigning values to edges (1004). At the start of training, edges (1004) are assigned initial values. These initial values may be assigned randomly, according to a predetermined distribution, manually, or by other assignment methods. Once the values of edges (1004) are initialized, the neural network (1000) can function, receive inputs, and generate outputs. To this end, at least one input is propagated through the neural network (1000), which generates an output. During training, a data set, called a training set or training data, is provided to the neural network (1000). This training set consists of inputs and their associated targets, which represent the "ground truth" or desired output. Inputs are processed by the neural network (1000), and the outputs are compared to the corresponding targets. This comparison of the neural network (1000) output with the target is typically performed using a function known as a "loss function," although other terms such as "error function," "objective function," "value function," and "cost function" are also commonly used. While there are various loss functions, such as the root mean square error function, a common feature of a loss function is that it numerically evaluates the similarity between the neural network (1000) output and the target. In some implementations, a loss function may be constructed by applying multiple loss functions to different parts of the output-target comparison. Loss functions may also be designed to impose additional constraints on the values assumed by the edges (1004), for example, by adding penalty terms or regularization terms based on physical laws. Generally, the goal of training is to adjust the values of the edges (1004) to increase the similarity between the neural network (1000) output and the target across a dataset. Therefore, loss functions are used to guide changes in the values of the edges (1004), typically performed using a technique known as "backpropagation."
[0052] A detailed description of the backpropagation process is beyond the scope of this disclosure, but briefly summarized: Backpropagation consists of computing the gradient of the loss function with respect to the edge (1004) values. The gradient indicates the direction of change of the edge (1004) values that will most significantly change the loss function. Because the gradient depends locally on the current edge (1004) values, the edge (1004) values are typically updated by a "step" in the direction indicated by the gradient. The size of this step is often referred to as the "learning rate" and need not be fixed during the training process. Furthermore, the size and direction of the step may also be determined by previously observed edge (1004) values or previously computed gradients. Such step direction determination methods are commonly referred to as "momentum"-based techniques.
[0053] When the edge (1004) values are updated or changed from their initial values through the backpropagation step, the neural network (1000) is likely to generate a different output. Therefore, at least one input is forward propagated through the neural network (1000), the neural network's (1000) output is compared with the corresponding target using a loss function (the comparison may be performed partially using multiple loss functions), the gradient of the loss function with respect to the edge (1004) values is calculated, and the edge (1004) values are updated in steps guided by the gradient. This process is repeated until a termination condition is met. Typical termination conditions include a certain number of edge (1004) updates (iterations), a decreasing learning rate, little change in the loss function between iterations, or the performance indicators specified by the training data or a separate validation data set are met. Once the termination condition is met and the edge (1004) value updates have ceased, the neural network (1000) is considered "trained." Depending on the configuration of the loss function, the loss function is often minimized to increase the similarity between the target and the neural network (1000) output. In other forms, the goal may be to maximize the loss function (in which case the loss function is often called an objective function or value function). Those skilled in the art will appreciate that the tasks of maximization and minimization can be treated equivalently using techniques such as negation (sign reversal). In other words, updating the edge (1004) values in gradient-guided steps may proceed in the direction of the gradient or in the opposite direction, depending on the configuration of the loss function.
[0054] The architecture of a machine learning model defines the overall structure of the machine learning model. For example, in the case of a neural network (1000), the number of hidden layers, the type of activation function used, and the number of outputs must be specified. Furthermore, the use and placement of special layers such as batch normalization are also defined. These choices, such as the number of hidden layers in a neural network (1000), are called the hyperparameters of the machine learning model. In other words, the architecture of a machine learning model specifies the hyperparameters surrounding the machine learning model. However, the architecture of a machine learning model does not describe the edge values (weights, parameters) of the model. These values are learned during training or specified separately when using a pre-trained model.
[0055] Another type of machine learning model is the convolutional neural network (CNN). A CNN can technically be represented as a diagram with edges (1004) and nodes (1002) grouped by layers, similar to a neural network (1000). However, it is more useful to think of a CNN as a structural grouping of weights. The term "structural" here refers to the interrelationship of weights within the same group. CNNs are widely used when the input data also has structural relationships, such as spatial relationships where a given element of the input data can always be considered to be "to the left" of another element. For example, an image has structural relationships because each pixel (element) has a directional relationship to its neighbors.
[0056] A structural grouping of weights, or a group of weights, is referred to herein as a "filter." The number of weights in a filter is typically much smaller than the number of elements in the input (e.g., the number of pixels in an image). In a CNN, filters "slide," or convolve, over the input data, forming intermediate outputs or representations that preserve the same structural relationships as the input data. As with neural networks (1000), the intermediate outputs are often further processed using activation functions. Multiple filters are applied to the input data, generating multiple intermediate representations. Additional filters may be formed that operate on the intermediate representations, creating even more intermediate representations. This process can be repeated as specified by the user. Filters may have a stride, skipping some elements (e.g., pixels) of the input before convolution. Grouping of intermediate output representations can involve pooling, for example, by considering only the largest ratio of groups in the computation of the coefficients. Stride and pooling can be used to downsample the intermediate representations. Similar to neural networks (1000), additional operations such as normalization, concatenation, dropout, and residual connections may be applied to the intermediate representations. Furthermore, the intermediate representations may be upsampled, for example, by techniques such as transposition convolution. CNNs have a "final" group of intermediate representations, to which no further filters are applied. In some cases, the structural relationships in the final layer are destroyed, a process known as "flattening." The flattened representations are typically passed to the neural network (1000) and at least one fully connected layer to generate the final output. Note that in this context, the neural network (1000) is still considered part of the CNN. In other cases, the structural relationships in the final layer (i.e., the output of the CNN) are preserved or reorganized or reshaped for interpretation. Similar to neural networks (1000), CNNs are trained using backpropagation according to a loss function after initializing the filter weights and, if present, the edges (1004) of the internal neural network (1000).
[0057] According to one or more embodiments, the machine learning model (608) used in the parking space detection methods and systems disclosed herein is a CNN. In particular, in one or more embodiments, the architecture of this CNN is similar to the well-known You Only Look Once (YOLO) object detection model. There are various versions of YOLO, which differ in the type of layers used and the resolution of the training data. However, a common feature of all YOLO versions is their ability to detect multiple objects at different scales in a single pass. Furthermore, recent YOLO architectures divide the input image into grid cells, each with one or more associated anchor boxes. In one or more embodiments, the machine learning model (608) follows a layer structure similar to YOLOv4.
[0058] The machine learning model (608) of the present disclosure may be modeled after YOLO, but with a number of key differences. The first difference is how the target parking representations in the training data are matched with anchor boxes. To match the target parking representations with anchor boxes, a training dataset must be provided. According to one or more embodiments, the private dataset is curated and manually annotated using one or more vehicles (301) equipped with one or more cameras capable of generating images of the vehicle's (301) surroundings (see, e.g., FIG. 3 ). The private dataset may be updated and / or expanded with additional data (e.g., bird's-eye view images) as they become available. In one or more embodiments, the private dataset includes annotated bird's-eye view images (606), each covering a 25-meter by 25-meter area, providing significantly greater coverage than publicly available datasets. Furthermore, the private dataset encompasses a wide variety of parking space configurations. The private dataset is used to train, configure, and evaluate the machine learning model (608). As a general procedure, the private data set is split into a training data set and a testing data set. In one or more embodiments, the private data set is split into a training data set, a validation data set, and a testing data set.
[0059] The training and testing datasets consist of a number of input-target pairs, where the input is a bird's-eye view image (606) (e.g., the exemplary bird's-eye view image (500)), and the corresponding target is a data structure having a similar format to the parking space prediction data (610), but containing ground truth values (determined manually or semi-automatically) that indicate the exact location of the parking space (102) represented by each parking representation in the bird's-eye view image (606). To properly format the target data structure for each bird's-eye view image (606), each ground truth parking representation in the bird's-eye view image (606) must be matched with at least one anchor box.
[0060] FIG. 11 shows a ground truth parking representation (1102) from the annotated bird's-eye view image (606). In this example, the ground truth parking representation (1102) is defined using four corners. The center coordinates (712) and corner displacements (804) of the ground truth parking representation (1102) are known. Therefore, the exact locations of the entrance left corner (702), entrance right corner (704), back left corner (706), and back right corner (708) are known (or can be calculated). Given the center coordinates (712), the ground truth parking representation (1102) is associated with a grid cell (not shown). Each grid cell is associated with K anchor boxes, where K is one or more integers. FIG. 11 shows two anchor boxes: a first anchor box (1104) and a second anchor box (1106). Note that the ground truth parking representation (1102) is a non-rectangular polygon, while the anchor box is rectangular according to one or more embodiments. To match the ground truth parking representation (1102) to the anchor box, first calculate the envelope width (1110) and envelope height (1112) of the ground truth parking representation (1102). The envelope width (1110) and envelope height (1112) are the width and height, respectively, of the smallest rectangular box (1108) that completely encloses the ground truth parking representation (1102). The envelope width (1110) w is calculated as follows:
number
number
[0061] Each anchor box is rectangular, and therefore has a width and a height. To determine which anchor box to associate with the ground truth parking representation (1102), the width and height of each anchor box are compared with the ground truth parking representation's envelope width (1110) and envelope height (1112). Specifically, for each of the K anchor boxes, the anchor box relevance score is calculated as follows:
number
[0062] A second difference in the parking space detection system of the present disclosure is that during training, the parameters of the machine learning model (608) are updated using a custom loss function L, which takes the form:
number
[0063] Loss term L CIoU and L CD is used to evaluate the accuracy of the predicted parking space representation in the parking space prediction data (610) by comparing it with the ground truth parking space representation (1102) represented in the target data structure. CIoU is referred to herein as the aggregate corner intersection over union loss. The aggregate corner intersection over union loss is defined as follows:
number
[0064] To better understand the aggregate corner crossing rate loss and explain how the IoU-based function is applied to each corner of the parking space representation, we use the corner crossing rate value (C i Here we show an example using the generalized intersection over union (GIoU) function as an IoU-based function for calculating the intersection over union (IoU). To understand this example, knowledge of the generalized intersection over union (GIoU) function is useful. GIoU is an index that evaluates the similarity of the shape and position of two convex shapes. In general, given two arbitrary convex shapes A and B, the intersection over union (IoU) of these two shapes is defined as follows:
number
number
[0065] When the GIoU function is used for the aggregate corner intersection rate loss, in the comparison between a single ground truth parking space representation (1102) and its predicted parking space representation, the aggregate corner intersection rate loss is calculated as follows:
number
[0066] For a given pair of ground truth parking space representation (1102) and its corresponding predicted parking space representation (1201), the aggregate corner intersection rate loss L CIoU is the intersection rate value C of each corner that defines the parking space. i It is emphasized that while FIG. 12 only shows an example of calculating the generalized intersection rate (GIoU) for the fourth corner, a similar process can be applied to other corners to calculate the aggregate corner intersection rate loss for the parking space shown in FIG. 12. It is also emphasized that while the example above uses a generalized intersection rate (GIoU) function, any intersection rate-based function (e.g., IoU, DIoU, GIoU, etc.) can be used for the aggregate corner intersection rate loss without departing from the scope of this disclosure.
[0067] L CD is referred to herein as the corner distance loss. Again using FIG. 12 as an example, the corner distance d4 (1220) for the fourth corner is shown. In general, the i-th corner distance d i denotes the Euclidean norm between the i-th corner of the correct parking space representation (1102) and the i-th corner of the predicted parking space representation (1201). Mathematically, the i-th corner distance d i is calculated as follows:
number
number
[0068] In one or more embodiments, the corner distance loss further includes clamping, thresholding, and scaling operations. Corner Distance Threshold Thr CD Given , in one or more embodiments, the corner distance loss is defined as:
number
[0069] L SC is the frame confidence loss. The frame confidence loss operates on the frame confidence (810) output by the machine learning model (608). The frame confidence (810) is a value between 0 and 1, representing the confidence that the predicted parking representation (1201) indicates the correct parking representation (1102). The frame confidence (810) corresponds to the objectness score in a standard object detection network. In one or more embodiments, the frame confidence loss applies a standard binary cross-entropy loss to the frame confidence (810) corresponding to the predicted parking representation (1201) using information on whether the corresponding correct parking representation (1102) is associated with it. In one or more embodiments, the frame confidence loss is a binary cross-entropy loss function using a logit loss function.
[0070] L CVis the corner visibility loss. Corners of a parking representation, or parking space (102), can be defined as either visible or occluded. The corner visibility (806) in the parking space prediction data (610) indicates the predicted visibility or occlusion state of each corner of the detected parking space (102). In practice, corner visibility is expressed as a continuous value from 0 to 1, indicating the likelihood that the corner is visible. Therefore, for each corner of the predicted parking representation (1201), if a corresponding ground truth parking representation (1102) exists, a binary cross-entropy loss can be applied to each of them. Corner visibility loss L CV is computed as the average of the binary cross-entropy losses for each corner.
[0071] Finally, L SM is the road surface material loss. The road surface material loss operates on the road surface material of the predicted parking space (814) and the road surface material of the known parking space (102). In one or more embodiments, the road surface material loss uses a categorical cross-entropy loss function.
[0072] According to one or more embodiments, the parking space prediction data (610) includes only the center coordinates (712), corner displacements (804), and space confidence (810) for each detected parking space (102). In this case, the custom loss function is defined as follows:
number
[0073] The parking space detection data (610) is likely to predict a large number of parking spaces (102) with low space confidence (810) values. According to one or more embodiments, parking spaces (102) with low space confidence (810) values may be removed from consideration based on a user-specified confidence threshold. For example, in one or more embodiments, only parking spaces (102) with space confidence (810) values greater than 0.6 are retained and considered as detected parking spaces (102). In other embodiments, post-processing techniques such as non-maximum suppression (NMS) may be applied to remove or filter predicted parking representations with low confidence and to merge overlapping predicted parking representations.
[0074] FIG. 13 illustrates the training step of the machine learning model (608) according to one or more embodiments. Training assumes that a training dataset including one or more bird's-eye view images (606) and corresponding target, or ground truth data (annotations) is provided. To perform the training step, at least one bird's-eye view image input (1302) is required. During training, the bird's-eye view image input (1302) is associated with ground truth parking representation data (1304) and bird's-eye view image input metadata (1306). The ground truth parking representation data (1304) describes the ground truth parking representation of the available parking spaces (102) in the bird's-eye view image input (1302). The bird's-eye view image input metadata (1306) includes any additional information related to the parking spaces in the bird's-eye view image input (1302). For example, the bird's-eye view image input metadata (1306) may indicate the road surface material and corner visibility of the parking spaces (102).
[0075] A target data structure (1307) must be created for the bird's-eye view image input (1302). To create the target data structure (1307), information about each ground truth parking representation (1308) in the bird's-eye view image input (1302) is encoded into the target data structure (1307). This is done by referencing the ground truth parking representation data (1304) for each ground truth parking representation (1308) in the bird's-eye view image input (1302) according to the following procedure: First, as shown in block 1310, a grid cell corresponding to the ground truth parking representation is identified. Next, the center coordinate of the ground truth parking representation is scaled relative to the grid cell. Subsequently, in block 1312, the envelope width and envelope height of the ground truth parking representation are calculated. Next, in block 1314, the ground truth parking representation is matched to at least one rectangular anchor box according to the method described above. Finally, for each ground truth parking representation, the corner displacements are transformed to match the scale of the corresponding anchor box, as shown in block 1316. Once these steps have been applied to each ground truth parking representation in the bird's eye view image input (1302), the scaled center coordinates and corner displacements, along with the associated bird's eye view image input metadata (1306), are inserted into the appropriate locations in the target data structure (1307) according to the identified grid cell and corresponding anchor box.
[0076] The bird's-eye view image input (1302) is then processed by the machine learning model (608) to generate parking space prediction data (610). In block 1318, corner displacements in the parking space prediction data (610) are represented, scaled, or transformed based on the dimensions of their predicted corresponding anchor boxes. In block 1320, the parking space prediction data (610) is compared to the target data structure (1307) using a custom loss function L. As shown in block 1322, the gradient of the custom loss function is calculated with respect to the parameters of the machine learning model (608). Then, in block 1324, the parameters of the machine learning model are updated according to the gradient.
[0077] While FIG. 13 illustrates updating the parameters of the machine learning model upon evaluation of a single bird's-eye view image input (1302), it should be noted that updates may generally occur after processing any number of bird's-eye view image inputs (1302). That is, those skilled in the art will appreciate that the training process may be applied to batches of inputs, and this is not intended to limit the scope of the present disclosure. Furthermore, the process of generating the target data structure (1307) and the process of generating the parking space prediction data (610) need not occur simultaneously or in parallel. In one or more embodiments, the target data structure (1307) corresponding to each bird's-eye view image input (1302) in the training dataset may be determined and stored in advance, prior to training the machine learning model (608). Additionally, in one or more embodiments, the training dataset may be augmented using any data augmentation technique known in the art. For example, each bird's-eye view image (606) in the training dataset may be augmented by randomly applying one or more of the following: vertical flip, horizontal flip, random rotation, and hue-saturation-value (HSV) color space adjustment.
[0078] Although the preceding examples refer to a CNN (e.g., a YOLO architecture), the methods and techniques described herein (e.g., custom loss functions, anchor box determination, etc.) are not limited to this choice of machine learning model. In one or more embodiments, the machine learning model (608) is a Vision Transformer (ViT) trained using the custom loss function described above.
[0079] In one or more embodiments, to improve the speed and efficiency of the parking space detection method and system of the present disclosure and enable real-time parking space (102) detection on an onboard computing system of a vehicle (301), the machine learning model (608) is implemented in a compiled computer language. In one or more embodiments, the machine learning model is implemented in C++.
[0080] In one or more embodiments, the predictive parking representation (1201) determined using the parking space prediction data (610) is post-processed using a Canny filtering method. The Canny filtering method slightly shifts each corner of the predictive parking representation (1201) to the nearest point output by the Canny filter. As mentioned above, the vehicle (301) may be equipped with additional sensor systems, such as an ultrasonic system or a LiDAR system. In one or more embodiments, sensor data obtained from the additional sensor systems is used to improve the accuracy of the predictive parking representation (1201).
[0081] FIG. 14 is a flowchart illustrating a general process for training a machine learning model (608) according to one or more embodiments. In block 1402, a plurality of bird's-eye view images are collected. Each bird's-eye view image is composed of one or more images captured by one or more cameras mounted on a vehicle (301). The plurality of bird's-eye view images include a variety of parking space (102) configurations (e.g., parking space orientation, parking space surface material, marking line type, etc.) under various environmental conditions (e.g., night, day, rain, etc.) and settings (indoor, outdoor). In block 1404, available parking spaces (102) are identified in each of the plurality of bird's-eye view images. Generally, a bird's-eye view image may include zero or more available parking spaces (102). In one or more embodiments, the available parking spaces (102) are identified through a manual or semi-automated iterative process. In block 1406, each identified available parking space is represented using a ground truth parking representation. Each ground truth parking representation is represented by the center coordinate of the identified parking space (102) and two or more pairs of center-relative coordinates (corresponding to corners of the ground truth parking representation). In block 1408, a grid cell in the bird's-eye view image corresponding to each identified available parking space (102) in the given bird's-eye view image is identified. In block 1410, an envelope width and an envelope height are calculated for each identified parking space (102). The envelope width and envelope height correspond to the width and height of the smallest rectangle that encloses the ground truth parking representation. In block 1412, for each identified parking space (102) in the given bird's-eye view image, the ground truth parking representation is associated with one or more anchor boxes. That is, each identified vacant parking space (102) in the given bird's-eye view image is associated with at least one anchor box. Given a particular single ground truth parking representation, the association is performed by comparing the envelope width and envelope height of the ground truth parking representation with the widths and heights of the candidate anchor boxes. Note that in one or more embodiments, the candidate anchor boxes are anchor boxes associated with the grid cells of the corresponding ground truth parking representation.However, in other embodiments, any anchor box may be considered a candidate anchor box, regardless of the association of any grid cell with a given ground truth representation or anchor box. At block 1416, a target data structure (1307) is generated for each of the plurality of bird's-eye view images. Each target data structure (1307) includes at least a center coordinate and a center-relative coordinate pair (i.e., corner displacement) for each identified vacant parking space (102) in the corresponding bird's-eye view image. In other words, each target data structure (1307) includes a ground truth representation for each identified vacant parking space (102) in the corresponding bird's-eye view image. The target data structure (1307) is configured with respect to grid cells and anchor boxes of the bird's-eye view image. At block 1418, a training dataset is generated. The training dataset includes a plurality of bird's-eye view images and their corresponding target data structures (1307). Finally, at block 1420, a machine learning model (608) is trained using at least the training dataset. In one or more embodiments, a portion of the training dataset may be reserved as a validation dataset and / or a test dataset (not used for training). The machine learning model (608) is trained using a backpropagation process, during which the parameters of the machine learning model (608) are updated based on a custom loss function. The resulting trained machine learning model (608) can directly receive bird's-eye view images and output parking space prediction data (610). The parking space prediction data (610) includes at least predicted center coordinates and corner displacements of the predicted parking space (102). In one or more embodiments, the parking space prediction data (610) is post-processed using scaling and Canny filtering to further refine and localize the predicted parking space (102). Because the multiple bird's-eye view images contain parking spaces (102) with diverse configurations, the trained machine learning model is robust and can accurately detect parking spaces (102) regardless of their orientation, road surface material, and other factors.
[0082] FIG. 15 is a flowchart illustrating a procedure for using the trained machine learning model (608) according to one or more embodiments. In block 1502, a first image is received from a first camera mounted on the vehicle (301). In block 1504, a bird's-eye view image is generated from the first image. In one or more embodiments, the bird's-eye view image is generated from the first image using an inverse perspective projection (IPM) method. In block 1506, the bird's-eye view image is processed by the machine learning model (608). For purposes of FIG. 15, it is assumed that the machine learning model (608) has already been trained, for example, according to the flowchart of FIG. 14. Thus, the machine learning model (608) is configured to process the bird's-eye view image to generate or output parking space prediction data (610). The parking space prediction data includes a first center coordinate where a first available parking space is visible in the bird's-eye view image and first corner displacement data. The first corner displacement data includes relative coordinate pairs corresponding to each corner of a parking representation of the first available parking space. That is, the first corner displacement data includes at least a first relative coordinate pair indicating a first corner relative to the first center coordinate and a second relative coordinate pair indicating a second corner relative to the first center coordinate. In one or more embodiments, the parking representation may be defined using four corners, in which case the first corner displacement data further includes a third relative coordinate pair indicating a third corner relative to the first center coordinate and a fourth relative coordinate pair indicating a fourth corner relative to the first center coordinate. Each of these coordinate pairs has an exclusive corner designation. For example, in one or more embodiments, the first corner is the entrance left corner (702). In one or more embodiments, the corner designation is provided by an index. For example, the corners can be designated as corner 1, corner 2, ... (the number of corners may be two or more) up to the final corner. The first center coordinate and first corner displacement data are sufficient to completely define a first parking representation (i.e., a predictive parking representation) that is a prediction indicating the location of the first available parking stall. The parking space prediction data (610) further includes a first parking space confidence indicating the machine learning model's (608) confidence that the first parking space representation accurately matches an available parking space.Next, in block 1506, a first location of the first available parking space is determined using the parking space prediction data (610). Specifically, the first location is determined using a first parking representation, where the first parking representation indicates a predicted location (i.e., the first location) of the first available parking space. In one or more embodiments, the first location indicates a location of the first available parking space in the physical space of the vehicle (301). Determining the first location may require scaling of the first parking representation, but those skilled in the art will recognize that coordinate transformation between model space and physical space is within the scope of this disclosure. Once the predicted first location is obtained, in block 1510, the vehicle (301) is parked in the first available parking space. In one or more embodiments, the vehicle (301) is parked in the first available parking space when the first parking space reliability is equal to or greater than a predetermined threshold. For example, in one or more embodiments, the threshold is set to 0.5. In one or more embodiments, the vehicle (301) is parked in the first available parking space without driver assistance. That is, the vehicle (301) is parked automatically. In one or more embodiments, the first available parking space (102) is proposed to a user (e.g., the driver) of the vehicle, who can accept or reject the parking space. Finally, it should be noted that as the vehicle (301) is parked, bird's-eye view images may be continuously generated using one or more images acquired from one or more cameras mounted on the vehicle (301) upon detection of a first available parking space that meets a threshold. The bird's-eye view images acquired during parking may be processed by a machine learning model (608) to continuously update first center coordinates and first corner displacement data (i.e., a predicted parking representation and a first position) for the vehicle (301) in real time to assist in the parking process (e.g., determining and monitoring the vehicle's proposed trajectory).
[0083] Embodiments of the parking detection method and system of the present disclosure have at least the following advantages: Embodiments of the present disclosure identify available parking spaces (102) using a parking representation that encloses the area of the parking space (102). The parking representation is defined using two or more corners. Therefore, the representation of the parking space (102) is not limited to a rectangular shape. Furthermore, by representing available parking spaces (102) using a parking representation, embodiments of the present disclosure can detect available parking spaces (102) regardless of the relative viewpoint of the vehicle without performing rotation operations or affine transformations. Furthermore, embodiments of the present disclosure assign corner designations to each corner of the parking representation of the parking space (102). This facilitates determining one or more approach lines for available parking spaces (102) directly from the parking representation without requiring a specialized and separate approach line detector. Additionally, embodiments of the present disclosure output parking space prediction data (610) directly from the machine-learned model (608) without requiring multiple neural network heads. That is, there is no need to divide the machine-learned model (608) into separate prediction tasks. Additionally, in one or more embodiments, the parking space prediction data (610) classifies the visibility of each corner of the parking representation, allowing the parking space (102) to be represented using the parking representation even if one or more corners of the parking space (102) are occluded. Another significant advantage is that, because bird's-eye view images are acquired and processed by the machine-learned model (608) throughout the parking procedure, the relative position of the parking space (102) to the vehicle (301) is continuously monitored to ensure successful parking of the vehicle. That is, in one or more embodiments, the parking space prediction data (610) is used by the vehicle's (301) in-vehicle driver assistance system to proactively park or "park guide" the vehicle into a detected available parking space (102). Finally, the parking representation allows the area enclosed by the parking space (102) to be easily determined, regardless of whether the parking space (102) is occluded.
[0084] Embodiments of the present invention can be implemented on virtually any type of computer system, regardless of the platform used. For example, the parking space detection methods and systems described herein may be implemented as or include one or more computer systems, such as the one shown in FIG. 16 . Such a computer system may be one or more mobile devices (e.g., laptops, smartphones, personal digital assistants, tablet computers, or other mobile devices), desktop computers, servers, blades in a server chassis, one or more ECUs in a vehicle, or any other type of computing device or apparatus with the minimum processing power, memory, and input / output devices necessary to execute one or more embodiments of the present invention. For example, as shown in FIG. 16 , a computer system (1600) may include one or more computer processors (1602), associated memory (1604) (e.g., random access memory (RAM), cache memory, flash memory, etc.), one or more storage devices (1606) (e.g., hard disks, optical drives (e.g., compact disc (CD) drives, digital versatile disc (DVD) drives), flash memory sticks, etc.), and many other elements and functions. The computer processor (1602) may be an integrated circuit for processing instructions. For example, a computer processor may have one or more cores, or micro-cores, of a processor. The computer system (1600) may include one or more input devices (1610), including a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or other type of input device. Additionally, the computer system (1600) may include one or more output devices (1608), including a screen (e.g., a liquid crystal display (LCD), plasma display, touchscreen, cathode ray tube (CRT) monitor, projector, or other display device), printer, external storage device, or other output device. The output device(s) may be the same as or different from the input device(s).The computer system 1600 may be connected to a network 1612 (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, a mobile network, or other type of network) via a network interface connection (not shown). Input and output devices may be connected locally or remotely (e.g., via the network 1612) to the computer processor 1602, memory 1604, and storage device 1606. Many different types of computer systems exist, and the input and output devices described above may take other forms.
[0085] The software instructions for carrying out embodiments of the present invention are in the form of computer-readable program code and may be stored, in whole or in part, temporarily or permanently on a non-transitory computer-readable medium, including a CD, DVD, storage device, diskette, tape, flash memory, physical memory, or other computer-readable storage medium. Specifically, these software instructions correspond to computer-readable program code that, when executed by a processor, is configured to carry out embodiments of the present invention.
[0086] Additionally, one or more elements of the computing system (1600) may be located remotely and connected to other elements via a network (1612). Furthermore, one or more embodiments of the present invention may be implemented on a distributed system having multiple nodes, with portions of the present invention located on different nodes within the distributed system. In one embodiment of the present invention, a node may correspond to a separate computing device. Alternatively, a node may correspond to a computer processor with associated physical memory. Alternatively, a node may correspond to a computer processor, or micro-cores thereof, with shared memory and / or resources.
[0087] Although the embodiments described herein use a convolutional neural network-based machine learning model, those skilled in the art will understand that the parking space detection method and system disclosed herein can be readily used in conjunction with other types of machine learning models. Thus, the embodiments disclosed herein are not limited to the use of a convolutional neural network-based machine learning model. Machine learning models such as the Vision Transformer (ViT) can be easily incorporated into this framework and do not depart from the scope of the present disclosure.
[0088] Moreover, while the present invention has been described with respect to a limited number of embodiments, those skilled in the art, having the benefit of this disclosure, will appreciate that they may devise other embodiments without departing from the scope of the invention as disclosed herein. Accordingly, the scope of the present invention should be limited only by the appended claims.
Claims
1. A first image is acquired from a first camera mounted on the vehicle; generating a bird's-eye view image using the first image; Processing the bird's-eye view image using a machine-learned model to generate parking space prediction data; The parking space prediction data is First center coordinates of the first available parking space; first corner displacement data including a first relative coordinate pair indicating a position of a first corner relative to the first center coordinate, and a second relative coordinate pair indicating a position of a second corner relative to the first center coordinate; The first parking space reliability; Including, Using the parking space prediction data, determine a first location of the first available parking space; When the first parking space reliability satisfies a threshold value, the vehicle is parked in the first available parking space. method.
2. The parking space prediction data is Second center coordinates of the second available parking space; second corner displacement data including a third relative coordinate pair indicating a position of a third corner relative to the second center coordinates, and a fourth relative coordinate pair indicating a position of a fourth corner relative to the second center coordinates; Equipped with Further, a second position of the second available parking space is determined using the parking space prediction data. The method of claim 1.
3. moreover, acquiring a second image from a second camera mounted on the vehicle; generating the bird's-eye view image using the first image and the second image; The method of claim 1.
4. the bird's-eye view image is generated by inverse perspective projection; The method of claim 1.
5. While the vehicle is being parked, the first center coordinate and the first corner displacement data are continuously updated by processing a newly acquired bird's-eye view image using the machine-learned model. The method of claim 1.
6. Expressing the first available parking space with a first parking expression; the first parking representation is fully specified by the first center coordinate and the first corner displacement data; determining one or more entry lines for the first available parking stall; determining an area enclosed by said first parking representation; The method of claim 1.
7. The parking space prediction data further includes: first corner visibility data indicating whether each corner of the first corner displacement data is visible or occluded; A first frame road surface material that identifies the material class of the first parking space; Equipped with The method of claim 1.
8. Determine the frame direction of the first available parking space; The frame direction is either perpendicular, parallel, or fishbone-shaped. The method of claim 1.
9. 1. A computer-implemented method for training a machine learning model, comprising: Acquire multiple bird's-eye view images, For each of the plurality of bird's-eye view images, Identify available parking spaces, For each identified parking slot, determining a ground truth parking representation of the identified available parking space, the ground truth representation including a center coordinate and a corner displacement; Calculating an envelope width and an envelope height using the center coordinates and the corner displacements; Using the envelope width and the envelope height, the identified available parking space is associated with one or more anchor boxes; generating a target data structure including the center coordinates and the corner displacements corresponding to each of the identified available parking spaces; generating a training data set comprising the plurality of bird's-eye view images and the corresponding target data structures; training the machine learning model using the training dataset; Prepare for this. the machine learning model directly receives one or more of the plurality of bird's-eye view images; method.
10. The target data structure for each of the bird's-eye view images of the plurality of bird's-eye view images further includes frame road surface material data and corner visibility data for each of the identified available parking spaces. The method of claim 9.
11. the machine learning model is trained using a loss function; The loss function is The aggregate corner crossing rate loss, Corner distance loss and Frame reliability loss and Including, The method of claim 9.
12. increasing the training data by applying at least one of vertical flip, horizontal flip, rotation, and color adjustment to one or more of the plurality of bird's-eye view images; The method of claim 9.
13. For each of the identified available parking spaces, scaling the center coordinates relative to a grid cell; transforming the corner displacements to the scale of one or more corresponding anchor boxes; The method of claim 9.
14. For a given predicted parking representation and a corresponding ground truth representation, the aggregate corner intersection rate loss is an average of a function based on intersection rates calculated for each corner of the predicted parking representation and the corresponding ground truth representation; For a given predicted parking representation and the corresponding ground truth representation, the corner distance loss is the average of Euclidean norms calculated for each corner of the predicted parking representation and the corresponding ground truth representation. The method of claim 11.
15. Vehicles and a first camera mounted on the vehicle; Bird's-eye view images, Machine learning models and a computer comprising one or more computer processors; Equipped with The computer acquiring a first image from the first camera; constructing the bird's-eye view image from the first image; Processing the bird's-eye view image using the machine-learned model to generate parking space prediction data; The parking space prediction data is First center coordinates of the first available parking space; first corner displacement data including a first relative coordinate pair indicating a position of a first corner relative to the first center coordinate, and a second relative coordinate pair indicating a position of a second corner relative to the first center coordinate; The first parking space reliability; Including, Using the parking space prediction data, determine a first location of the first available parking space; When the first parking space reliability satisfies a threshold, the vehicle is parked in the first available parking space without driver assistance. system.
16. The parking space prediction data is Second center coordinates of the second available parking space; second corner displacement data including a third relative coordinate pair indicating a position of a third corner relative to the second center coordinates, and a fourth relative coordinate pair indicating a position of a fourth corner relative to the second center coordinates; The second parking space reliability; further comprising The computer further determines the location of the second available parking space using the parking space prediction data.
16. The system of claim 15.
17. a second camera mounted on the vehicle; Furthermore, The computer acquiring a second image from the second camera; generating the bird's-eye view image using the first image and the second image; 16. The system of claim 15.
18. While the vehicle is being parked, the first center coordinate and the first corner displacement data are continuously updated by processing a newly acquired bird's-eye view image using the machine-learned model.
16. The system of claim 15.
19. The parking space prediction data further includes: first corner visibility data indicating whether each corner included in the first corner displacement data is visible or occluded; A first parking stall road surface material that identifies the material class of the first parking stall; Equipped with 16. The system of claim 15.
20. The computer further determines a frame direction of the first available parking frame; The frame direction is either perpendicular, parallel, or fishbone-shaped.
16. The system of claim 15.
Citation Information
Patent Citations
Road maintenance management system, pavement type determination device, pavement deterioration determination device, repair priority determination device, road maintenance management method, pavement type determination method, pavement deterioration determination method, and repair priority determination method
JP2020147961A
Parking spot detection method and parking spot detection system
US20220245952A1
Parking space detection method and apparatus, and device and storage medium
WO2021184616A1