Parking assistance system
The system uses external cameras to generate and display top-view images for vehicles without on-board cameras or user terminals, addressing accessibility limitations and enhancing parking assistance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- 株式会社シーディアイ
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing parking assistance systems require a vehicle-mounted camera and a user terminal to provide top-view images, which not all vehicles have, limiting their accessibility.
A parking assistance system using multiple external cameras installed at different locations in a parking lot to generate a top-view image visible to drivers from a display unit outside the vehicle, without the need for on-board cameras or user terminals.
Enables drivers of vehicles without on-board cameras or user terminals to access top-view images, facilitating easy parking by providing real-time, intuitive guidance and warning notifications.
Smart Images

Figure 2026074541000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a parking assistance system.
Background Art
[0002] Conventionally, a technique is known in which a top view image of a vehicle as seen from above is generated based on an image captured by a camera mounted on the vehicle and provided to a driver (for example, Patent Document 1). Further, Patent Document 2 discloses that a composite video synthesized based on an image captured by a surveillance camera may be displayed on a user terminal such as a smartphone.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] When parking a vehicle in a parking lot, it is necessary to move the vehicle to a predetermined position while paying attention to the surroundings. According to the top view image, since the vehicle and its surroundings can be confirmed in a state where the vehicle is looked down upon, it is very useful for the driver to be able to provide the top view image in the parking lot or around the parking lot. In order to provide a top view image to the driver in the vehicle by the technique disclosed in Patent Document 1, it is necessary to mount a camera on the vehicle, and it is also necessary to convert the image of the camera into a top view image by an in-vehicle device. In order to display a composite video on a user terminal by the technique disclosed in Patent Document 2, it is necessary to prepare in advance a user terminal capable of displaying the composite video. The present invention has been made in view of the above problems, and aims to provide top-view images to drivers of vehicles that do not have an in-vehicle camera for displaying top-view images, and to drivers who do not have a user terminal for displaying top-view images. [Means for solving the problem]
[0005] A parking assistance system according to one embodiment of the invention comprises: a plurality of cameras installed at different locations and having a field of view of a predetermined location in a parking lot; a conversion unit that converts images from the plurality of cameras into a single top-view image of the predetermined location viewed from a predetermined direction; and a display control unit that displays the top-view image on a first display unit that is visible from inside a vehicle moving toward the predetermined location and is located outside the vehicle.
[0006] In other words, the parking assistance system generates a top-view image based on images from multiple cameras installed outside the vehicle to capture images of a predetermined location in the parking lot, and displays the top-view image on a first display unit visible to the driver attempting to park, located outside the vehicle. Therefore, the system can provide a top-view image to drivers of vehicles that do not have an on-board camera for displaying top-view images, or to drivers who do not have a user terminal for displaying top-view images. [Brief explanation of the drawing]
[0007] [Figure 1] This is a diagram showing the configuration of the parking assistance system in this embodiment. [Figure 2] This is a block diagram of a parking assistance system. [Figure 3] This diagram shows the camera image, the corrected image, and the world coordinate system. [Figure 4] This is a flowchart for the parking assistance process. [Figure 5] Figures 5A to 5D are camera images showing the state after the vehicle image has been identified by the first machine learning model. [Figure 6] This figure shows an example of a top-view image segment. [Figure 7] Figures 7A and 7B show top-view images, Figure 7C shows a camera image taken of a rendered 3D model, and Figure 7D shows the outline of the vehicle and the outline of the 3D model. [Figure 8] This is a flowchart for the parking assistance process. [Figure 9] This diagram shows an example of displaying multiple camera images alongside a top-down view image. [Modes for carrying out the invention]
[0008] Here, embodiments of the present invention will be described in the following order. (1) Configuration of the parking assistance system: (2) Parking assistance processing: (3) Other embodiments:
[0009] (1) Configuration of the parking assistance system: Figure 1 shows the configuration of a parking assistance system according to one embodiment. The parking assistance system in this embodiment is installed in a multi-story parking garage. In this embodiment, the parking garage is one in which vehicles are stopped on pallets, and after the occupants get out of the vehicles, the pallets transport the vehicles to a storage location. The pallets are associated with a position where the vehicle should be stopped, and in this embodiment, this position is called a predetermined position Pf. The predetermined position Pf is installed in a waiting space Sp where vehicles are temporarily kept, and a gate may be provided in front of the waiting space Sp. When parking a vehicle in the parking garage, the driver of the vehicle moves the vehicle to the predetermined position Pf, parks the vehicle in the predetermined position Pf, gets out of the vehicle, and then exits from the waiting space Sp. The parking assistance system is a system that assists the driver in moving the vehicle to the predetermined position Pf.
[0010] To assist the driver, the parking assistance system includes a first display unit 11. The parking assistance system assists the driver of the vehicle in moving the vehicle to a predetermined position Pf by displaying a top-view image of the vehicle on the first display unit 11. In this embodiment, the parking assistance system includes four cameras 21, 22, 23, and 24. The cameras 21, 22, 23, and 24 are installed at four different locations above the vehicle's movement plane where the predetermined position Pf is located, and are installed so as to include the predetermined position Pf in their field of view.
[0011] Cameras 21, 22, 23, and 24 only need to be installed in positions where there are no (or almost no) blind spots when the vehicle is moving within the predetermined position Pf. In this embodiment, when the vehicle's movement plane is divided into four quadrants with the center of the predetermined position Pf as the origin, cameras 21, 22, 23, and 24 are installed so that they exist in each of the four quadrants.
[0012] The parking assist system generates a top-view image based on images from cameras 21, 22, 23, and 24 and displays it on the first display unit 11. To perform this processing, the parking assist system is equipped with a control system that performs various processing tasks. Figure 2 is a block diagram of the parking assist system including the control system.
[0013] The parking assistance system includes the first display unit 11, a second display unit 12 different from the first display unit 11, cameras 21, 22, 23, and 24, a storage medium 30, and a control unit 40. The first display unit 11 is a display installed on one of the walls forming the waiting space Sp, and is visible from vehicles located around a predetermined position Pf. That is, the first display unit 11 is visible from inside a vehicle moving towards the predetermined position Pf, or from inside a vehicle located within the predetermined position Pf.
[0014] The second display unit 12 is a display installed at a position different from the first display unit 11, and in the present embodiment, it is installed in a management room where the parking lot administrator is present. That is, the second display unit 12 is a display visually recognized by the parking lot administrator. In the present embodiment, the same image as the image displayed on the first display unit 11 is also displayed on the second display unit 12.
[0015] As described above, the cameras 21, 22, 23, 24 are cameras that include a predetermined position Pf in their fields of view and capture the predetermined position Pf from different positions. In the present embodiment, the cameras 21, 22, 23, 24 are visible light cameras. The cameras 21, 22, 23, 24 are connected to the control unit 40 via an interface not shown.
[0016] The storage medium 30 can store various types of information. In the present embodiment, camera parameters 30a, a first machine learning model 30b, a second machine learning model 30c, a third machine learning model 30d, and a 3D model 30e are stored in the storage medium 30. The camera parameters 30a are parameters for defining the correspondence between the coordinate system of the camera images captured by the cameras 21, 22, 23, 24 and the world coordinate system. In the present embodiment, the camera parameters 30a are parameters for converting the coordinate systems of the respective cameras 21, 22, 23, 24 into the world coordinate system set for the parking lot. That is, according to the camera parameters 30a, any coordinate of the camera image can be converted into a coordinate in the world coordinate system, and any coordinate in the world coordinate system can be converted into any coordinate of the camera image.
[0017] Specifically, the camera parameters 30a are composed of internal parameters and external parameters. The internal parameters are parameters corresponding to the characteristics of the cameras 21, 22, 23, 24, and for example, are parameters for correcting the influence of lens distortion. In FIG. 3, an example of a camera image captured by the cameras 21, 22, 23, 24 is shown in the upper row, an example of the image after correction by the internal parameters is shown in the middle row, and the world coordinate system is shown in the lower row. Here, the coordinate system of the camera image is set as a two-axis uv coordinate system.
[0018] When cameras 21, 22, 23, and 24 capture a predetermined position Pf, due to the influence of lens distortion and the like, the frame of the rectangular predetermined position Pf is captured as a curve. When correction is performed using the internal parameters, the camera image is converted into an image having the shape shown in the middle row, and in the converted image, the frame of the predetermined position Pf becomes a straight line. Such internal parameters can be specified by known calibration. For example, a calibration board in which rectangles colored in two colors, black and white, are arranged alternately up, down, left, and right in a grid pattern, for example, a chessboard, is captured by cameras 21, 22, 23, and 24, and the internal parameters can be specified by calibrating using a known method such as OPENCV. In this embodiment, the internal parameters are defined for each of cameras 21, 22, 23, and 24. Note that the calibration board is not limited to a chessboard, and may be a board in which circular markers are regularly arranged vertically and horizontally, or may be a board showing a two-dimensional code including encoded information.
[0019] The external parameters are parameters corresponding to the positions and orientations of cameras 21, 22, 23, and 24 in the world coordinate system. When conversion is performed using the external parameters, it is specified at which coordinate values in the world coordinate system the subject of each pixel in the images of cameras 21, 22, 23, and 24 exists. The external parameters can be specified by known calibration. For example, a plurality of reference points with known coordinate values in the world coordinate system, such as the four corners of the predetermined position Pf, are specified, and the external parameters can be specified by calibrating according to a known procedure such as the DLT method (Direct linear transformation method). Note that the external parameters are defined for each of cameras 21, 22, 23, and 24. In this embodiment, with the camera parameter 30a, it is possible to convert between camera coordinates (u, v) and world coordinates (X, Y, Z) for any coordinate within the field of view of the camera.
[0020] The first machine learning model 30b is a machine learning model that takes each of the images captured by cameras 21, 22, 23, and 24 as input and separates the vehicle image from the background contained in each image. In this embodiment, the first machine learning model 30b is a model trained using semantic segmentation. Therefore, when each image is input to the first machine learning model 30b, it is possible to identify whether the subject of each pixel in the image is a vehicle or a background other than a vehicle.
[0021] The first machine learning model 30b described above can be generated by running a known machine learning algorithm using training data that associates images of the vehicle with the pixels of the vehicle. The first machine learning model 30b may be a model corresponding to each of the cameras 21, 22, 23, and 24, or it may be a common model.
[0022] The second machine learning model 30c is a machine learning model that takes each of the images captured by cameras 21, 22, 23, and 24 as input and outputs the objects contained in each image. In this embodiment, the second machine learning model 30c is a model trained using semantic segmentation. Therefore, when each image is input to the second machine learning model 30c, the type of object that is the subject can be identified for each pixel of each image. In this embodiment, the types of objects output include at least vehicle parts (door mirrors, hood, windows, doors, etc.) and objects around the vehicle (people, luggage, etc.).
[0023] The second machine learning model 30c described above can be generated by running a known machine learning algorithm using training data that associates images of the object to be recognized with the label of that object. The second machine learning model 30c may be a model corresponding to each of the cameras 21, 22, 23, and 24, or it may be a common model.
[0024] The third machine learning model 30d is a machine learning model for identifying rectangles that circumsect the area where the vehicle image is masked in a top-down view image (details described later). Specifically, when a top-down view image with the vehicle image masked is input to the third machine learning model 30d, it outputs a rectangle (called a bounding box) that circumsects the mask of the vehicle image. The third machine learning model 30d described above can be generated by running a known machine learning algorithm using training data that associates an image containing the mask of the object to be recognized with the bounding box circumsecting the said mask.
[0025] The 3D model 30e is a model for reproducing the shape of a vehicle in a virtual 3D space. In this embodiment, it is a model that reproduces the surfaces of the vehicle's components using multiple polygons. That is, the 3D model 30e is a model that represents the surface of a vehicle, which is composed of typical vehicle components such as bumpers, hoods, windows, doors, and side mirrors, using polygons, and is defined by information indicating the coordinates of each polygon. A 3D model 30e is prepared for each vehicle type. That is, a 3D model 30e is defined for each vehicle type that can use a parking lot (sedan, RV, wagon, etc.) and stored in the storage medium 300. The form of the 3D model 30e is not limited to the form in this embodiment and may be defined in various forms.
[0026] The control unit 40 includes a CPU, RAM, and ROM (not shown), and can execute various programs stored in the storage medium 30, ROM, etc. The programs executed by the control unit 40 include various programs. In this embodiment, it includes a parking assistance program 41 that generates a top-view image based on camera images captured by cameras 21, 22, 23, and 24, and assists the vehicle driver based on the top-view image. When the parking assistance program 41 is executed, the control unit 40 functions as a conversion unit 41a, a specification unit 41b, an acquisition unit 41c, a recognition unit 41d, and a display control unit 41e.
[0027] The conversion unit 41a has the function of converting images from multiple cameras 21, 22, 23, and 24 into a single top-view image of a predetermined position Pf viewed from a predetermined direction. In this embodiment, the top-view image is an image showing the state of the predetermined position Pf viewed from above in a direction along the vertical. In this embodiment, the control unit 40, using the function of the conversion unit 41a, refers to the camera parameter 30a and converts each of the camera images captured by cameras 21, 22, 23, and 24 into an image viewed from a predetermined direction. Here, each of the four images obtained by converting the camera images of each camera 21, 22, 23, and 24 is called a top-view image segment.
[0028] This transformation is achieved, for example, by positioning a pixel at an arbitrary camera coordinate (u,v) in the camera image at the (X,Y,0) position in the world coordinate system based on the camera parameter 30a (where Z=0 in the world coordinate system is the Z coordinate of the plane where the predetermined position Pf exists). Once four top-view image segments are obtained, the control unit 40 performs predetermined processing on each top-view image segment and combines them to generate a single top-view image. In the generated top-view image, the image of the vehicle is masked. Details of the processing of the transformation unit 41a will be described later.
[0029] The identification unit 41b has the function of determining the position of the vehicle in the world coordinate system based on camera images captured by cameras 21, 22, 23, and 24 and the top view image. In this embodiment, the control unit 40 identifies the masked position in the top view image based on the top view image and determines the position in the world coordinate system corresponding to the masked position in the top view image based on the camera parameter 30a.
[0030] Furthermore, the control unit 40 virtually places a 3D model 30e of the vehicle at the given location and virtually identifies the camera image obtained when the virtually placed 3D model 30e is captured by cameras 21, 22, 23, and 24, based on the camera parameters 30a. The control unit 40 then adjusts the position, shape, and size of the 3D model 30e so that the virtual model position, which is the position of the 3D model 30e within the camera image, best matches the image of the vehicle in the actual camera image. Finally, the control unit 40 considers the position of the 3D model 30e where the virtual model position best matches the position of the vehicle in the actual camera image to be the vehicle's position in the world coordinate system.
[0031] The acquisition unit 41c has the function of acquiring information about the vehicle's color from images from multiple cameras. The control unit 40 identifies the part of the vehicle image in each of the camera images captured by cameras 21, 22, 23, and 24, and identifies the color of the vehicle body.
[0032] The recognition unit 41d has the function of recognizing objects contained in images from multiple cameras. The control unit 40 recognizes objects present around a predetermined position Pf by inputting camera images captured by cameras 21, 22, 23, and 24 to the second machine learning model 30c. The objects to be recognized here are objects that may be subject to warnings, and include vehicle parts and objects around the vehicle.
[0033] The display control unit 41e has the function of displaying a top-view image on a first display unit located outside the vehicle, and which is visible from inside the vehicle as it moves toward a predetermined position. In other words, the control unit 40 controls the first display unit 11 to display the top-view image using the function of the display control unit 41e. The control unit 40 generates a top-view image from camera images captured by cameras 21, 22, 23, and 24 and repeats the process of displaying it on the first display unit 11 at predetermined intervals (for example, 100 ms). As a result, the top-view image is displayed live on the first display unit 11 in virtually real time.
[0034] Since the vehicle driver can see the first display unit 11, they can easily compare the positional relationship between the vehicle's position and the predetermined position Pf, and easily move the vehicle into the predetermined position Pf. Furthermore, since the cameras 21, 22, 23, 24 and the first display unit 11 are installed in the parking lot, there is no need for the vehicle to be equipped with a camera or display for displaying top-view images, nor is it necessary for the driver to own a device for displaying top-view images. For this reason, according to this embodiment, top-view images can be provided to drivers of vehicles that do not have an on-board camera for displaying top-view images, or to drivers who do not have a user terminal for displaying top-view images.
[0035] Furthermore, the control unit 40, using the functions of the display control unit 41e, colors the masked parts of the vehicle in the top-view image with the vehicle's color. That is, the control unit 40 uses the vehicle's color acquired by the acquisition unit 41c to color the masked parts, i.e., parts of the vehicle, in the top-view image. As a result, the driver can intuitively identify the relationship between the vehicle they are driving and a predetermined position Pf in the top-view image.
[0036] Furthermore, in this embodiment, the control unit 40 causes the second display unit 12 to display a top-view image using the functions of the display control unit 41e. In this embodiment, the control unit 40 also displays the same image in the second display unit 12 as the image displayed in the first display unit 11. Since the image displayed in the first display unit 11 includes a top-view image, the administrator can view the top-view image on the second display unit 12. Therefore, the administrator can determine whether the driver's driving is normal and can issue warnings, etc. Since the top-view image displayed in the second display unit 12 is the same as the top-view image displayed in the first display unit 11, the vehicle in the top-view image on the second display unit 12 is also colored in the same way as the actual vehicle.
[0037] Furthermore, in this embodiment, the display control unit 41e functions to display a warning if an object is subject to a warning. That is, the control unit 40 determines whether the object recognized by the recognition unit 41d is in a state subject to a warning. If it is subject to a warning, the control unit 40 displays a warning on the first display unit 11 using the function of the display control unit 41e. State subject to a warning includes, for example, a state where the side mirror is not folded, a state where the door is open, a state where luggage is placed around the vehicle, or a state where there is a person around the vehicle. With this configuration, the driver can easily understand whether or not a warning-worthy event has occurred. The control unit 40 also displays the same warning on the second display unit 12 as on the first display unit 11. Therefore, the administrator can easily understand whether or not a warning-worthy event has occurred.
[0038] (2) Parking assistance processing: Figures 4 and 8 are flowcharts illustrating an example of the parking assistance process performed by the control unit 40. In this embodiment, once the parking assistance process is completed up to the end of Figure 8, it returns to the beginning of Figure 4 and is repeatedly executed at predetermined fixed intervals (e.g., 100 ms). When the parking assistance process starts, the control unit 40 communicates with the multiple cameras 21, 22, 23, and 24 using the function of the conversion unit 41a and acquires camera images (step S100).
[0039] Next, the control unit 40 separates the vehicle from the background in each camera image using the function of the conversion unit 41a (step S105). That is, the control unit 40 inputs each camera image into the first machine learning model 30b and identifies the pixels that represent the vehicle image. Then, the control unit 40 masks the vehicle image by converting the grayscale value of the pixels that represent the vehicle image to a constant value (for example, a grayscale value that represents black). As a result, the vehicle image is masked, and everything except the masked part is separated as the background.
[0040] Figures 5A to 5D are camera images showing the state in which the image of a vehicle has been identified by the first machine learning model 30b. Figure 5A shows camera image I21 from camera 21, Figure 5B shows camera image I22 from camera 22, Figure 5C shows camera image I23 from camera 23, and Figure 5D shows camera image I24 from camera 24. In these camera images, pixels recognized as images of a vehicle are colored with a transparent light gray.
[0041] Next, the control unit 40 converts each camera image into a top-view image segment using the function of the conversion unit 41a (step S110). That is, the control unit 40 converts each masked camera image into a top-view image segment by placing a pixel at an arbitrary camera coordinate (u,v) in the masked camera image at the position of coordinate (X,Y,0) in the world coordinate system, based on the camera parameter 30a.
[0042] Figures 6A to 6D show examples of top-view image segments. Figure 6A shows top-view image segment Its21 generated based on camera image I21 from camera 21, and Figure 6B shows top-view image segment Its22 generated based on camera image I22 from camera 22. Figure 6C shows top-view image segment Its23 generated based on camera image I23 from camera 23, and Figure 6D shows top-view image segment Its24 generated based on camera image I24 from camera 24. In these figures, the uniform black areas are either masked vehicles or parts outside the field of view of cameras 21, 22, 23, and 24.
[0043] Next, the control unit 40 synthesizes the top-view image segments using the function of the conversion unit 41a to generate a single top-view image (step S115). The pixel positions of the top-view image segments correspond to coordinate positions in the world coordinate system. Therefore, the control unit 40 synthesizes the top-view image segments so that pixels with the same coordinates in the world coordinate system overlap. Figure 7A shows the top-view image It generated by synthesizing the top-view image segments Its21 to Its24 from Figures 6A to 6D.
[0044] Next, the control unit 40 uses the function of the identification unit 41b to identify the position of the bounding box that circumscribes the vehicle in the top-view image (step S120). Specifically, the control unit 40 inputs the top-view image to the third machine learning model 30d and obtains the bounding box that circumscribes the mask of the vehicle's image. The bounding box is defined within the top-view image. The position of each pixel in the top-view image corresponds to the X and Y coordinates of the world coordinate system. Therefore, once the position of the bounding box is identified, the position of the vehicle in the world coordinate system can be determined based on the position in the world coordinate system corresponding to the position of the bounding box. Figure 7B shows the top-view image It with the bounding box BB identified.
[0045] Next, the control unit 40 places the 3D model 30e in the world coordinate system using the function of the identification unit 41b (step S125). Specifically, the control unit 40 identifies the vehicle type based on the camera images captured by cameras 21, 22, 23, and 24, and obtains the 3D model 30e corresponding to that vehicle type by referring to the storage medium 30. The identification of the vehicle type can be performed using, for example, a machine learning model or pattern matching. Once the 3D model 30e is obtained, the control unit 40 places the obtained 3D model 30e at the position of the bounding box in the virtually set world coordinate system. That is, since the position of the bounding box in the world coordinate system corresponds to the horizontal position of the vehicle, the control unit 40 places the 3D model 30e in the world coordinate system to match the horizontal position of the vehicle.
[0046] Steps S125 to S155 are a loop process, and the control unit 40 adjusts the vehicle's position by repeating this loop process. Therefore, in the initial loop process or when the number of repetitions is small, the shape and size of the 3D model 30e do not necessarily match the actual vehicle.
[0047] Next, the control unit 40 renders the 3D model 30e using the functions of the specific unit 41b (step S130). Specifically, the 3D model 30e, which is composed of polygons, is a model that virtually reproduces the vehicle in a wireframe form, and the control unit 40 colors each polygon. The color is not limited, but for example, rendering with a uniform gray color is conceivable.
[0048] Next, the control unit 40 acquires images of the 3D model 30e when it is photographed by multiple cameras using the function of the specific unit 41b (step S135). Specifically, the control unit 40 refers to the camera parameters 30a and converts the coordinates of each surface of the rendered 3D model 30e placed in the world coordinate system, and the coordinates of parts other than the vehicle in the world coordinate system, into coordinates in the camera image. This conversion acquires images of the 3D model 30e when it is photographed by a camera. The control unit 40 performs this process for each of the cameras 21, 22, 23, and 24 to simulate the camera images that would be taken when the 3D model 30e is photographed by each of the cameras 21, 22, 23, and 24. Figure 7C shows an example of a camera image taken when the rendered 3D model 30e is photographed by camera 21.
[0049] Next, the control unit 40 extracts the contour of the 3D model 30e using the function of the identification unit 41b (step S140). Specifically, the control unit 40 extracts the boundary between the rendered portion and the unrendered portion as a contour from each image acquired in step S135. Various processes can be used for extracting the contour; for example, a process that identifies the rendered portion based on the fact that the color of the rendered portion is almost the same, or a process using a machine learning model can be used. The process of extracting the contour corresponds to the process of extracting the virtual model position, which is the position in the camera image when the 3D model 30e is captured by multiple cameras.
[0050] Next, the control unit 40 obtains the difference between the contour of the 3D model 30e and the contour of the vehicle part using the function of the identification unit 41b (step S145). Figure 7D is a diagram in which lines representing the contour of the 3D model 30e and lines representing the contour of the vehicle part are superimposed on the camera image. Here, the contour of the vehicle part can be defined, for example, by the boundary between the masked part and the unmasked part by the processing in step S105. The difference between the contour of the 3D model 30e and the contour of the vehicle part can be defined by various methods, for example, by the number or area of pixels existing between the two contours. The control unit 40 obtains the difference between the contour of the 3D model 30e and the contour of the vehicle part by identifying the difference for each of the cameras 21, 22, 23, and 24 and obtaining statistical values (for example, the mean or cumulative value).
[0051] Next, the control unit 40 determines whether the difference is below a threshold using the function of the identification unit 41b (step S150). Here, the threshold is an index used to determine whether the difference is small enough that the contour of the 3D model 30e and the contour of the vehicle can be considered to match, and it is predetermined. In other words, the process of determining whether the difference in contours is below a threshold can be said to be the process of determining whether the virtual model position matches the position in the images of the vehicle taken by multiple cameras.
[0052] In step S150, if the difference is not determined to be below a threshold, the control unit 40 adjusts the size and position of the 3D model 30e (step S155). Here, the adjustment of the size of the 3D model 30e includes adjustment of the whole and parts of the 3D model 30e. Specifically, the control unit 40 adjusts the front-to-back and left-to-right size of the 3D model 30e. The control unit 40 also adjusts the positions of parts such as tires and windows.
[0053] The adjustment of the position of the 3D model 30e involves adjusting its front-to-back and left-to-right position. That is, the control unit 40 adjusts it in the front-to-back and left-to-right directions. These adjustments can be performed using various methods. For example, a configuration can be adopted in which the amount of adjustment for size and position for each adjustment is predetermined, and this adjustment amount is adjusted in each loop process of steps S125 to S155. Furthermore, if the difference increases after one adjustment, the adjustment direction may be reversed. In addition, the amount of adjustment may decrease as the difference decreases. Furthermore, in a single adjustment, either the size or the position may be adjusted, or the position may be adjusted first, followed by the size.
[0054] When step S155 is executed, the control unit 40 repeats the processing from step S125 onward using the size and position of the adjusted 3D model 30e. That is, in the first step S125, the 3D model 30e is placed at the bounding box position, but from the second time onward, the 3D model 30e is placed at the position adjusted in step S155.
[0055] On the other hand, if it is determined in step S150 that the difference is less than or equal to a threshold, the control unit 40 uses the function of the identification unit 41b to determine the position of the vehicle based on the position of the 3D model 30e in the world coordinate system (step S200). That is, the control unit 40 considers the position of the 3D model 30e in the world coordinate system to be the position of the vehicle in the world coordinate system. Through the above process, the horizontal position of the vehicle is determined in the world coordinate system, as well as the positions of the tires and windows. Therefore, it becomes possible to perform various processes based on the positions of each part of the vehicle.
[0056] Next, the control unit 40 identifies the color of the vehicle based on the camera images (step S205). Specifically, the control unit 40 extracts the image of the part that was masked in step S105 from the camera images of each camera 21, 22, 23, and 24, and identifies the color of the vehicle body based on that image. Furthermore, the control unit 40 obtains statistical values (average value, etc.) of the color identified from the camera images of each camera 21, 22, 23, and 24 and considers them to be the color of the vehicle body.
[0057] Next, the control unit 40 colors the vehicle in the top-view image using the functions of the display control unit 41e (step S210). That is, the control unit 40 identifies the masked portion from the top-view image generated in step S115 and superimposes an image representing the vehicle onto that portion. This image can be generated, for example, by adjusting the size of an image created in advance based on the 3D model 30e. Then, the control unit 40 colors the vehicle body portion of the image with the color identified in step S205. The control unit 40 also colors the parts other than the vehicle body, such as the windows and lights, with predetermined colors. In a configuration where the masked portion is composited without masking, there is a high possibility that the vehicle image will be distorted, but according to this embodiment, the vehicle image can be superimposed on the top-view image without distortion.
[0058] Next, the control unit 40 displays the camera image and the top-view image side by side using the functions of the display control unit 41e (step S215). In this embodiment, the control unit 40 displays the camera and the top-view image on the first display unit 11 and the second display unit 12, but images other than the top-view image may also be displayed for parking assistance. In this embodiment, camera images taken by multiple cameras are displayed side by side with the top-view image.
[0059] Figure 9 shows an example of a display where multiple camera images I21-I24 and a top-view image It are shown side by side. In the example shown in Figure 9, the top-view image It is placed in the center, camera images I21 and I24 are placed to the left of the top-view image It, and camera images I22 and I23 are placed to the right of the top-view image It. In this example, camera images I21-I24 are enlarged portions of images taken by cameras 21, 22, 23, and 24.
[0060] With the above-described display, the vehicle driver can very easily grasp the positional relationship between the vehicle C and the predetermined position Pf by viewing the first display unit 11. Furthermore, according to this embodiment, the vehicle driver can grasp in detail the situation to the left and right in front of and behind the vehicle by viewing the camera images I21 to I24. The administrator can also obtain the same information as the vehicle driver by viewing the second display unit 12.
[0061] Next, the control unit 40 displays a message corresponding to the vehicle's position using the functions of the display control unit 41e (step S220). In this embodiment, a message box MB is formed within the screen where the top-view image It is displayed, and a message for guiding the vehicle to the correct position within a predetermined position Pf is displayed within the message box MB. To display this message, the control unit 40 identifies the relationship between the vehicle's position and the correct position. In this embodiment, the correct position is defined as the vehicle's position when the vehicle's front wheels are within a specific area.
[0062] Specifically, as shown in Figure 9, a region Zt in which the vehicle's front wheels should be located is predefined. The control unit 40 identifies the position of the vehicle's front wheels in the world coordinate system based on the position of the vehicle's 3D model 30e in the world coordinate system and compares it with the position in region Zt. If the front wheels are located behind region Zt, the control unit 40 displays a message in message box MB prompting the vehicle to move forward. Figure 9 shows an example where a message prompting the vehicle to move forward is displayed in message box MB.
[0063] If the front wheels are in front of area Zt, the control unit 40 displays a message in the message box MB prompting the vehicle to move in reverse. Alternatively, the control unit 40 may also display a message in the message box MB prompting the vehicle to move to the right or left based on the relationship between a predetermined position Pf and the vehicle's position. The message is not limited to words or sentences, but may also include images explaining steering operations. Arrows indicating the vehicle's path may also be displayed.
[0064] Next, the control unit 40 determines whether the vehicle's position is correct (step S225). Specifically, the control unit 40 considers the vehicle's position to be correct if the vehicle's front wheels are within region Z in the world coordinate system. If the vehicle's front wheels are not within region Z in the world coordinate system, the control unit 40 does not consider the vehicle's position to be correct. If the control unit 40 does not determine that the vehicle's position is correct in step S225, the control unit 40 terminates the parking assistance process. However, since the parking assistance process is executed repeatedly at regular intervals, the camera image and top-view image are effectively displayed live, and the message display also continues.
[0065] On the other hand, if it is determined in step S225 that the vehicle is in the correct position, the control unit 40 displays a stop message using the function of the display control unit 41e (step S230). That is, the control unit 40 controls the first display unit 11 and the second display unit 12 to display a message in the message box MB prompting the driver to stop the vehicle.
[0066] Next, the control unit 40 performs object recognition using the functions of the recognition unit 41d (step S235). Specifically, the control unit 40 inputs the camera images I21, I22, I23, and I24 captured by cameras 21, 22, 23, and 24 into the second machine learning model 30c and identifies the objects contained in each camera image. As a result, for example, door mirrors, luggage, people, etc., contained in each camera image are identified.
[0067] Next, the control unit 40 determines whether or not the object is subject to a warning based on the functions of the display control unit 41e (step S240). That is, the control unit 40 determines whether or not the object recognized in step S235 is in a state that is subject to a warning. If it is in a state that is subject to a warning, the control unit 40 controls the first display unit 11 and the second display unit 12 based on the functions of the display control unit 41e to display a warning (step S245). As a result, the vehicle driver and manager can easily understand whether or not a warning-worthy event has occurred. If the object is not determined to be subject to a warning in step S240, the control unit 40 skips step S245.
[0068] (3) Other embodiments: The above embodiments are merely examples for carrying out the present invention, and various other embodiments can be adopted. For example, at least a portion of each part constituting the parking assistance system 10 may be divided into multiple devices or systems. At least a portion of the multiple devices may be connected to other devices via a network, and these devices may cooperate to constitute the parking assistance system. Furthermore, at least a portion of the functions of the conversion unit 41a, the identification unit 41b, the acquisition unit 41c, the recognition unit 41d, and the display control unit 41e may be divided into multiple devices. In addition, some of the configurations of the above embodiments may be omitted, and the order of processing may be changed or omitted.
[0069] Furthermore, various alternative technologies may be used for the configuration adopted in the above-described embodiment. For example, the first machine learning model 30b and the second machine learning model 30c are not limited to semantic segmentation, but may also be instance segmentation or panoptic segmentation, and various other machine learning models may be employed. The third machine learning model 30d is also not limited to the above-described model, and the position of the bounding box may be determined by pattern recognition technology other than machine learning models.
[0070] Multiple cameras should be positioned at different locations within the parking lot, and their fields of view should include the designated location. In other words, the cameras should be positioned so that a top-down view image of the designated location can be generated based on the images from multiple cameras. The number of cameras is not limited to four as in the embodiment described above; it may be fewer or more. Furthermore, cameras that look down on the designated location from above may also be used.
[0071] The designated location can be any location that is the target for generating the top-view image, and may be a parking lot pallet or parking space, or a frame indicating a waiting position for vehicles located in front of the pallet or parking space. Multiple cameras will include the designated location, but of course, the area around the designated location will also be included in the field of view, and it is preferable that a top-view image including the area around the designated location is generated.
[0072] The conversion unit only needs to be able to convert images from multiple cameras into a single top-view image taken from a predetermined direction at a predetermined position. In other words, the conversion unit only needs to be able to generate a top-view image by converting images from multiple cameras into top-view image segments and combining them. The image of the vehicle included in the top-view image may be replaced with a top-view image of a 3D model colored with a color equivalent to the vehicle's color, as in the embodiment described above, or it may be replaced with another image. Furthermore, the image of the vehicle included in the top-view image may be displayed without being masked.
[0073] The display control unit only needs to be able to display a top-view image on a first display unit located outside the vehicle, and visible from inside the vehicle as it moves toward a predetermined position. In other words, it is sufficient that a top-view image is displayed on a first display unit installed outside the vehicle and visible from inside the vehicle, so that the driver of the vehicle moving toward the predetermined position can see the top-view image. The installation location of the first display unit is not limited to the front wall surface of the vehicle moving toward the predetermined position Pf, as in the embodiment described above, but may also be installed on the side wall surface or at any other location.
[0074] The second display unit can be any display unit different from the first display unit that is visible to the parking lot manager. In other words, if the parking lot manager can see the top-view image, the manager can provide warnings to drivers and guide vehicles. The image displayed on the second display unit may be different from that on the first display unit.
[0075] The identification unit only needs to be able to identify the position of the bounding box that circumscribes the vehicle in the top-view image, and then determine the vehicle's position based on the correspondence between the coordinate systems of the images from multiple cameras and the world coordinate system, and the position in the world coordinate system corresponding to the position of the bounding box. In other words, the identification unit only needs to be able to use the bounding box to determine the vehicle's position in the world coordinate system. Therefore, in addition to the above configuration in which a 3D model of the vehicle is virtually placed in the world coordinate system based on the position of the bounding box, various other configurations may be adopted. For example, the identification unit may identify the position in the world coordinate system corresponding to the position of the bounding box as the horizontal position of the vehicle in the world coordinate system.
[0076] The acquisition unit only needs to be able to acquire information about the vehicle's color from images from multiple cameras. That is, the acquisition unit only needs to be able to acquire the vehicle's color in the top-view image based on images from multiple cameras. The information about the vehicle's color can be defined in various ways, for example, by the gradation values of each of the multiple color components of each pixel. Of course, the information about the vehicle's color may also be statistically determined based on the gradation values that make up the image portion of the vehicle. Furthermore, color management technology may be applied so that the colors in the camera images are accurately reproduced.
[0077] The recognition unit only needs to be able to recognize objects contained in images from multiple cameras. The method for recognizing objects is not limited to semantic segmentation methods as in the embodiments described above. For example, various other machine learning models may be applied, or recognition may be performed by pattern matching, image differences between a state where an object does not exist and a state where an object may exist, etc. The specific state that triggers a warning may be a variety of states. For example, the presence of an object may trigger a warning, or the object may be in a specific posture (mirrors not folded, doors open, etc.). Furthermore, a warning may be issued if the vehicle is in a specific state. For example, in the embodiments described above, the dimensions of the vehicle may be determined based on the 3D model 30e, and a warning may be issued if the dimensions exceed the size that can be parked in the parking lot.
[0078] Furthermore, the present invention is also applicable as a program or method. Moreover, such systems, programs, and methods may be implemented as standalone devices or using shared components, encompassing various embodiments. For example, it is possible to provide methods and programs implemented using such systems. Furthermore, the components can be modified as appropriate, such as being partly software and partly hardware. The invention also functions as a recording medium for a program that controls a device. Of course, the recording medium for the software may be a magnetic recording medium, a semiconductor memory, or any recording medium developed in the future; the same principle applies. [Explanation of Symbols]
[0079] 10...Parking assistance system, 11...First display unit, 12...Second display unit, 21...Camera, 22...Camera, 23...Camera, 24...Camera, 30...Storage medium, 30a...Camera parameters, 30b...First machine learning model, 30c...Second machine learning model, 30d...Third machine learning model, 30e...3D model, 40...Control unit, 41...Parking assistance program, 41a...Conversion unit, 41b...Identification unit, 41c...Acquisition unit, 41d...Recognition unit, 41e...Display control unit, 300...Storage medium
Claims
1. Multiple cameras are installed at different locations and include the designated area of the parking lot in their field of view. A conversion unit that converts images from multiple cameras into a single top-view image of the predetermined position viewed from a predetermined direction, A display control unit that displays the top-view image on a first display unit located outside the vehicle, which is visible from inside the vehicle as it moves toward the predetermined position, A parking assistance system equipped with the following features.
2. The display control unit, The top-view image is displayed on a second display unit, different from the first display unit, which is visible to the manager of the aforementioned parking lot. The parking assistance system according to claim 1.
3. The system further includes a unit that identifies the position of a bounding box circumscribing the vehicle in the top-view image, and identifies the position of the vehicle based on the correspondence between the coordinate systems of the images from the multiple cameras and the world coordinate system, and the position in the world coordinate system corresponding to the position of the bounding box. A parking assistance system according to claim 1 or claim 2.
4. The specified part is, A pre-prepared 3D model of the vehicle is virtually placed in the world coordinate system, and based on the correspondence, the virtual model position, which is the position of the 3D model virtually placed in the world coordinate system when it is photographed by multiple cameras, is identified, and the position of the vehicle is identified based on the position of the 3D model in the world coordinate system where the virtual model position and the position of the vehicle in the images taken by the multiple cameras best match. The parking assistance system according to claim 3.
5. The system further includes an acquisition unit that acquires information regarding the color of the vehicle from images of multiple cameras, The display control unit colors the vehicle in the top-view image based on the information regarding the vehicle's color acquired by the acquisition unit. The parking assistance system according to claim 1.
6. The parking assistance system according to claim 1, wherein the first display unit displays images from a plurality of cameras and the top-view image side by side.
7. The parking assistance system according to claim 2, wherein the second display unit displays images from multiple cameras and the top-view image side by side.
8. The system further includes a recognition unit that recognizes objects included in images from multiple cameras, The display control unit, A warning is displayed when the aforementioned object is in a specific state. The parking assistance system according to claim 1.
Citation Information
Patent Citations
Around-view providing device
JP2020516100A
Driving assistance device and driving assistance method
JP2023066220A