Object position estimation device and parking system

The object position estimation device improves vehicle position detection accuracy by using a trained model to estimate reference points and orientation vectors, addressing inaccuracies in existing methods and enhancing precision in world coordinate system projections.

JP7729154B2Active Publication Date: 2025-08-26AISIN CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021162185
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-08-26
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Existing vehicle position estimation methods using machine learning often inaccurately represent the vehicle's location due to the midpoint of the applied frame not aligning with the vehicle's actual position, leading to errors when projected onto a world coordinate system.

Method used

An object position estimation device that utilizes a trained model to estimate the position and orientation of a vehicle by learning the relationship between reference points on the vehicle's bottom surface and orientation vectors, converting image coordinates into a world coordinate system, and averaging multiple estimates for improved accuracy.

Benefits of technology

Enhances the accuracy of vehicle position detection on a world coordinate system by accurately estimating the vehicle's position and orientation using reference points and orientation vectors, even when the vehicle is positioned diagonally or facing differently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007729154000001
    Figure 0007729154000001
  • Figure 0007729154000002
    Figure 0007729154000002
  • Figure 0007729154000003
    Figure 0007729154000003
Patent Text Reader

Abstract

To provide an object position estimation device and a parking system that can more easily and accurately estimate a position and an orientation of a vehicle body (vehicle) from image information using machine learning and accurately detect the position and the orientation of the vehicle on a world coordinate system.SOLUTION: An object position estimation device as an example of the present disclosure includes an acquisition unit, an estimation unit, and a processing unit. The acquisition unit acquires image information from an imaging unit that images an image of a detection target area. The estimation unit estimates, on the basis of a comparison of a learning object included in the image information, a learned model as a result of learning a relationship between at least one reference point indicating a reference position included in a bottom surface of a learning object and a direction vector of the learning object, and an estimation object included in the image information acquired by the acquisition unit, a corresponding position and an orientation of a bottom surface of the estimation object corresponding to the position of the reference point. The processing unit transforms the corresponding position into a world coordinate system and detects the position of the estimation object on the world coordinate system.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an object position estimation device and a parking system. [Background technology]

[0002] Conventionally, a technique has been known in which a vehicle (car) captured in an image captured by a fixed camera or the like is detected using a technique such as machine learning, and the position of the detected vehicle is projected onto a world coordinate system. In this technique, first, an area in the image that includes the characteristics (e.g., shape, etc.) of the vehicle is identified using a machine learning technique, and a rectangular "frame" of a size corresponding to the area is drawn. Then, the coordinates of the midpoint of, for example, the bottom side of the frame are considered to be the location of the vehicle, and the coordinates are projected onto the world coordinate system to detect the location of the vehicle. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-26281 Summary of the Invention [Problem to be solved by the invention]

[0004] However, because a "frame" is applied to surround the vehicle in the image, the coordinates assumed to represent the vehicle's location often do not accurately indicate the actual location of the vehicle. For example, depending on the vehicle's position (e.g., facing diagonally), the midpoint of the base may be located on the road surface, not on the vehicle itself. Therefore, when projected onto the world coordinate system, there is a problem in that an error occurs between the actual vehicle position and the displayed (recognized) position.

[0005] Therefore, one of the objectives of the present disclosure is to provide an object position estimation device and a parking system that can more easily and accurately estimate the position and attitude of a vehicle body (vehicle) from image information using machine learning, and accurately detect the position and attitude of the vehicle on a world coordinate system. [Means for solving the problem]

[0006] An example of an object position estimation device according to the present disclosure includes an acquisition unit that acquires image information from an imaging unit that captures an image of a detection target area; a trained model resulting from learning a relationship between a training object included in the image information, at least one reference point indicating a reference position included on the bottom surface of the training object, and the orientation vector of the training object; an estimation unit that estimates a corresponding position and orientation of the bottom surface of the training object corresponding to the position of the reference point based on a comparison with the training object included in the image information acquired by the acquisition unit; and a processing unit that transforms the corresponding position into a world coordinate system and detects the position and orientation of the training object in the world coordinate system. With this configuration, for example, the trained model is created including at least one reference point indicating a reference position included on the bottom surface of the training object, and the position and orientation of the training object are estimated based on the reference point. As a result, the accuracy of estimating the position and orientation of the training object from image information can be improved, enabling more accurate position detection in the world coordinate system.

[0007] Furthermore, the trained model of the object position estimation device described above may be the result of learning the relationship between a direction vector, which is the shape of a training object included in image information and indicates a predetermined direction of the training object, and at least one reference point indicating a reference position included on the bottom surface of the training object, and the processing unit may match the corresponding position with a corresponding location of a template prepared in advance depending on the type of the estimation object, convert the image coordinates of any two points on the direction vector into a world coordinate system to obtain the direction vector in the world coordinate system, rotate the template around the corresponding position as a rotation center so that the predetermined direction coincides with the direction vector in the world coordinate system, and convert the image coordinates of any two points on the direction vector into the world coordinate system to obtain the direction vector in the world coordinate system, thereby identifying the position and orientation of the estimation object. This configuration improves the accuracy of estimating the position and orientation of the estimation object from image information, enabling more accurate detection of its position in the world coordinate system.

[0008] Furthermore, when the estimation unit estimates multiple corresponding positions, the processing unit of the object position estimation device may individually match the multiple corresponding positions with templates and identify the position and orientation of the object for estimation by averaging the matching results. This configuration improves the accuracy of estimating the position and orientation of the object for estimation from image information, enabling more accurate detection of the position on the world coordinate system.

[0009] Furthermore, when a plurality of corresponding positions are estimated by the estimation unit, the processing unit of the object position estimation device may specify the position and orientation of the object for estimation using one of the corresponding positions that is most likely to be estimated and a direction vector. This configuration can further improve the accuracy of estimating the position and orientation of the object for estimation from image information, enabling more accurate detection of the position on the world coordinate system.

[0010] In addition, when converting the image coordinates of any two points on the orientation vector into the world coordinate system, the processing unit of the above-mentioned object position estimation device may move the orientation vector so that the starting point of the orientation vector matches the corresponding position of the object being estimated.

[0011] In addition, when converting the image coordinates of any two points on the orientation vector into the world coordinate system, the processing unit of the above-mentioned object position estimation device may move the orientation vector so that the starting point of the orientation vector coincides with any point on the object to be estimated.

[0012] Furthermore, the reference point of the object position estimation device described above may be, for example, at least one of the ground contact point of the training object, a first projection point when a portion of the front end of the training object is projected onto the road surface, and a second projection point when a portion of the rear end of the training object is projected onto the road surface. With this configuration, for example, regardless of the direction in which the orientation of the object captured in the image information is facing, the corresponding position and orientation of the bottom of the estimation object can be estimated using the reference point. For example, even if the object is facing directly to the side, directly ahead, directly behind, etc. in the image information, the corresponding position and orientation of the bottom of the estimation object can be estimated. As a result, more accurate position detection on the world coordinate system becomes possible.

[0013] Furthermore, in the above-described object position estimation device, the first projection point may be a projection point at approximately the center of the front end of the training object, and the second projection point may be a projection point at approximately the center of the rear end of the training object. This configuration, for example, makes it easier to identify the reference point. As a result, when creating a trained model in machine learning or when performing estimation using the trained model, it is possible to improve the accuracy of setting the reference point and the accuracy of recognizing the corresponding position of the reference point.

[0014] Furthermore, the estimation unit of the object position estimation device may perform estimation when it detects an object to be estimated entering the detection target area. With this configuration, it is possible to perform accurate object position detection processing at an appropriate time.

[0015] Furthermore, a parking system as an example of the present disclosure includes the object position estimation device according to any one of claims 1 to 9, and a parking control device that controls the movement of a vehicle within a parking lot based on the position of the vehicle estimated by the object position estimation device. This configuration improves the accuracy of estimating the position and attitude of the vehicle based on image information, enabling more accurate detection of the position on a world coordinate system. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 is an exemplary schematic diagram showing an automatic valet parking system that can use a vehicle position detection device according to an embodiment. [Figure 2] FIG. 2 is an exemplary schematic block diagram illustrating the configuration of a vehicle position detection device according to an embodiment. [Figure 3] FIG. 3 is an exemplary schematic top view illustrating the positions and orientation vectors of reference points in machine learning used in the vehicle position detection device according to the embodiment. [Figure 4] FIG. 4 is an exemplary and schematic perspective view showing the position and orientation vector of a reference point set during learning on a vehicle included in image information, or a specific point detected during estimation, in machine learning used in a vehicle position detection device according to an embodiment. [Figure 5] FIG. 5 is an exemplary and schematic perspective view showing other positions of reference points set during learning for a vehicle included in image information, or specific points detected during estimation, in machine learning used in a vehicle position detection device according to an embodiment. [Figure 6] FIG. 6 is an exemplary schematic perspective view illustrating the position and orientation vector of a specific point detected when an estimation mode using machine learning is executed in the vehicle position detection device according to the embodiment. [Figure 7] FIG. 7 is an exemplary schematic diagram showing a case where the position and attitude of the vehicle are estimated based on the specific points detected in FIG. 6 and projected onto the world coordinate system. [Figure 8] FIG. 8 is a diagram showing a method for projecting the position of the estimation vehicle onto a map in the world coordinate system. [Figure 9] FIG. 9 is an exemplary flowchart illustrating a process for executing a learning mode in machine learning used in the vehicle position detection device according to the embodiment. [Figure 10] FIG. 10 is an exemplary flowchart illustrating processing when an estimation mode is executed in machine learning used in the vehicle position detection device according to the embodiment. [Figure 11] FIG. 11 is an exemplary flowchart illustrating the flow of a vehicle body position identification process when executing the estimation mode in machine learning used in the vehicle position detection device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0017] Hereinafter, embodiments and modifications of the present disclosure will be described with reference to the drawings. The configurations of the embodiments and modifications described below, as well as the actions and effects brought about by the configurations, are merely examples and are not limited to the following description.

[0018] In this embodiment, a vehicle position detection device that estimates the position and attitude of a vehicle (four-wheeled automobile) that is an object is applied as the object position estimation device.

[0019] 1 is an exemplary schematic diagram showing an automatic valet parking system 100 that can use a vehicle position detection device 10 according to an embodiment. The vehicle position detection device 10 can be incorporated into, for example, a parking control device 101 that manages and monitors the automatic valet parking system 100. In another embodiment, the vehicle position detection device 10 and the parking control device 101 may be provided independently, and may perform control while transmitting and receiving information to and from each other.

[0020] First, the automated valet parking system 100 will be described.

[0021] The automatic valet parking system 100 is a system for realizing automatic valet parking, including automatic parking and automatic departure, in a parking lot P having one or more parking areas R (parking spaces) demarcated by predetermined demarcation lines L, such as white lines on the road surface. The parking control device 101 generates parking guidance information based on, for example, detection results from sensors located within the area of ​​the parking lot P. Then, using the generated parking guidance information, the system smoothly moves the vehicle to an appropriate parking area R or controls the vehicle's departure from the parking area R.

[0022] As shown in FIG. 1, in automatic valet parking, after occupant X gets off vehicle V in a predetermined drop-off area P1 in parking lot P, automatic parking is executed in which vehicle V automatically moves (drives) from drop-off area P1 to an empty parking area R (e.g., parking area R1) and parks in response to a predetermined instruction (see guided route C1 indicated by a dashed-dotted arrow in FIG. 1). After automatic parking is completed, automatic exit is executed in which vehicle V automatically leaves parking area R (parking area R1) and moves to a predetermined boarding area P2 and stops there in response to a predetermined call (see guided route C2 indicated by a dashed-dotted arrow in FIG. 1). The predetermined instruction and the predetermined call are realized, for example, by occupant X operating terminal device T.

[0023] Furthermore, in the automatic valet parking system 100, the automatic driving of the vehicle V is realized through cooperation between a parking control device 101 provided in the parking lot P and a vehicle control device 102 mounted on the vehicle V. The parking control device 101 and the vehicle control device 102 are configured to be able to communicate with each other via wireless communication.

[0024] Here, the parking control device 101 monitors the situation within the parking lot P by receiving image information obtained from one or more surveillance cameras 103 (imaging units) that capture images of the situation within the parking lot P, and data output from various sensors (not shown) installed within the parking lot P.

[0025] The surveillance camera 103 is a digital camera incorporating an imaging element such as a charge coupled device (CCD) or a CMOS image sensor (CIS). The surveillance camera 103 can output video data (image information) at a predetermined frame rate. The parking control device 101 is configured to manage multiple parking areas R based on the monitoring results.

[0026] 1 shows an example in which two surveillance cameras 103 are installed on the walls or the like of parking lot P. In another example, multiple surveillance cameras 103 may be placed on pillars or ceiling surfaces near parking area R so as to accurately monitor the usage status of one or more parking areas R and the status of the driving paths, thereby improving monitoring accuracy. In addition to surveillance cameras 103, infrared sensors, ultrasonic sensors, etc. may be installed near parking area R, and the usage status of parking area R may be managed in a similar manner.

[0027] The vehicle V is equipped with, for example, an on-board camera 308 (cameras 308R, 308L, etc. attached to the door mirrors), and the vehicle control device 102 detects the dividing lines L on the left and right sides of the vehicle V and the parking area R, while checking the surrounding conditions and realizing automatic driving.

[0028] On the other hand, the vehicle position detection device 10 detects the position of the vehicle V present in the parking lot P on the world coordinate system based on image data captured by the surveillance camera 103 or a dedicated imaging unit, and provides the detected position information of the vehicle V to the parking control device 101. At this time, the vehicle position detection device 10 can detect (estimate) the positions and attitudes of all vehicles V present in the parking lot P, including vehicles V under automatic parking control, vehicles V under automatic exit control, vehicles V parked in the parking area R, and vehicles V stopped in positions other than the parking area R for some reason.

[0029] The vehicle position detection device 10 uses machine learning when detecting (estimating) the position and attitude of the vehicle V from the imaging data. The vehicle position detection device 10 executes a learning mode in machine learning and an estimation mode in which the position and attitude of the vehicle V are detected (estimated) using a trained model constructed in the learning mode. The learning mode and the estimation mode will be described in detail later.

[0030] Based on the position information of each vehicle V provided by the vehicle position detection device 10, the parking control device 101 controls, for example, a vehicle V in automatic driving so that the automatic driving is not hindered by other vehicles V (other vehicles V in automatic driving or parked vehicles V, etc.) or the vehicle does not come into contact with other vehicles V.

[0031] In the embodiment, the number and arrangement of the drop-off area P1, the pick-up area P2, and the parking area R in the parking lot P are not limited to the example shown in Fig. 1. The technology of the embodiment is applicable to parking lots with various configurations different from the parking lot P shown in Fig. 1.

[0032] 2 is an exemplary schematic block diagram showing the configuration of the vehicle position detection device 10. The vehicle position detection device 10 can be realized by a general personal computer configured with hardware such as a processor and memory. The processor reads and executes programs stored in a memory such as a storage unit, thereby realizing each functional module in the vehicle position detection device 10.

[0033] The vehicle position detection device 10 includes, as functional modules, for example, a mode switching unit 12, an acquisition unit 14, a learning mode execution unit 16, and an estimation mode execution unit 18. The learning mode execution unit 16 includes detailed modules such as a reference point and direction vector setting unit 16a and a model creation unit 16b. The estimation mode execution unit 18 includes a reference point and direction vector estimation unit 18a and a processing unit 18b, and the processing unit 18b further includes detailed modules such as a conversion unit 18b1 and a position identification unit 18b2. These modules may be separated or integrated by function, or may be configured as hardware.

[0034] In addition, the vehicle position detection device 10 is connected to a learned model storage unit 20 that stores an updatable learned model constructed by machine learning, and a display device 22 that displays a vehicle V detected in a parking lot P on a map of a world coordinate system, etc.

[0035] The trained model storage unit 20 is a non-volatile rewritable storage device such as a hard disk drive (HDD) or a solid state drive (SSD).

[0036] The display device 22 includes a display unit 22a configured, for example, by an LCD (liquid crystal display) or an OLED (organic electroluminescent display). The display device 22 may be configured only by the display unit 22a, or may be integrated with an operation unit 22b that performs input operations when creating training data for a trained model or when performing mode switching operations via the mode switching unit 12. When the display unit 22a and the operation unit 22b are integrated, the display unit 22a of the display device 22 is covered with a transparent operation unit 22b such as a touch panel. An operator of the vehicle position detection device 10 can view an image displayed on the display unit 22a through the operation unit 22b. In addition, the operator can perform operation input by touching, pressing, or moving the operation unit 22b with a finger, a touch pen, or the like at a position corresponding to an image displayed on the display screen of the display unit 22a.

[0037] In addition, as mentioned above, the vehicle position detection device 10 is connected to a surveillance camera 103 that acquires image information of the parking lot P (detection target area), an entry detection unit 23 that detects the entry of a moving object into the parking lot P, and the like.

[0038] As mentioned above, the surveillance camera 103 is capable of capturing image information for detecting the dividing lines L and parking areas R within the parking lot P and monitoring the presence or absence of parked vehicles, and is also capable of capturing image information showing vehicles V (moving or stopped vehicles) present within the parking lot P.

[0039] The entry detection unit 23 is a sensor capable of detecting an object (e.g., a moving object) entering the parking lot P. The type of sensor for the entry detection unit 23 can be selected appropriately, for example, an optical sensor, an electromagnetic sensor, a capacitance sensor, or the like, as long as it can detect an object entering the parking lot P. The detection result of the entry detection unit 23 can be used as a signal for starting the position detection process of the vehicle V by the vehicle position detection device 10. The entry detection unit 23 may be arranged outside the drop-off area P1 and the boarding area P2 (outside the parking lot P) so that the vehicle position detection device 10 can detect the location of the vehicle V in the drop-off area P1 and the boarding area P2 as well.

[0040] The mode switching unit 12 switches between the learning mode execution unit 16 and the estimation mode execution unit 18 of the vehicle position detection device 10. For example, an operator can select the learning mode or the estimation mode in machine learning by operating the operation unit 22b.

[0041] The acquisition unit 14 acquires image information (video information or still image information) of an area (detection target area) within the parking lot P acquired by the surveillance camera 103. In this case, the image information acquired by the surveillance camera 103 is assumed to include the entire area of ​​the parking lot P. When two surveillance cameras 103 are installed as shown in FIG. 1, it is sufficient that the entire area of ​​the parking lot P is covered by two types of image information acquired by the two cameras. The same applies when three or more surveillance cameras 103 are installed. The acquisition unit 14 may constantly acquire image information acquired by the surveillance cameras 103, or may start acquiring image information from the surveillance cameras 103 only when the entry detection unit 23 detects that an object (i.e., a moving object) has entered the parking lot P. In this case, the acquisition of image information related to the vehicle V, whose position has been estimated by the estimation mode execution unit 18, is confirmed to have completed parking. The same applies when an imaging unit dedicated to the vehicle position detection device 10 is provided separately from the surveillance cameras 103. In another embodiment, the parking control device 101 may control the acquisition of image information by the acquisition unit 14 .

[0042] The learning mode execution unit 16 executes learning mode processing in machine learning to estimate accurate position information of a vehicle V present in a parking lot P. The learning mode execution unit 16 is executed when the learning mode is selected by the mode switching unit 12.

[0043] As described above, the learning mode execution unit 16 includes a reference point and orientation vector setting unit 16a and a model creation unit 16b. The reference point and orientation vector setting unit 16a sets a correct answer value for training data to construct a learned model used in machine learning to estimate the position and orientation of the vehicle V in the world coordinate system. That is, an operator performs annotation (correct answer value assignment work) on the vehicle V captured in the image information of the parking lot P acquired by the acquisition unit 14. Note that the image information used at this time may be referred to as a "learning image," and the vehicle V captured in the image information may be referred to as a "learning vehicle." If the image information captured by the surveillance camera 103 is a video, a frame (i.e., a still image) in which a portion that can be an effective reference point for estimating the vehicle V is clearly captured is selected as the learning image from among multiple frames. If the image information captured by the surveillance camera 103 is a still image, the learning image is selected from the multiple still images. These selections may be made automatically by the learning mode execution unit 16, or may be made by an annotating worker (a worker who performs annotations) using the operation unit 22b.

[0044] The reference point and orientation vector setting unit 16a sequentially sets positions on the learning image specified by the annotator using the operation unit 22b as reference points. Furthermore, in order to accurately detect the position and orientation of the vehicle V in the world coordinate system from the camera image, the reference point and orientation vector setting unit 16a associates the reference points with a template TP (see FIG. 8 ), which is the shape of the learning vehicle. The templates TP are classified into, for example, four-door vehicles, two-door vehicles, standard cars, light cars, minivans, two-box cars, sports cars, SUVs, recreational vehicles, minivans, etc., and each template TP is tagged with image information. By associating the templates TP with reference points, the reliability of the training data can be improved.

[0045] The reference point and direction vector setting unit 16a associates the direction vector, which indicates the forward direction of the learning vehicle and is specified by the annotator using the operation unit 22b, with the reference point and sets it. The annotator operates the operation unit 22b in advance to specify the correct value of the direction vector, which indicates the forward direction.

[0046] The reason for associating the orientation vector with the reference point in this manner is as follows: When considering identifying the position of a vehicle using only the reference point, in a situation where only one reference point (bottom point) of the vehicle can be detected due to occlusion or the like, the vehicle's posture cannot be uniquely determined, making it impossible to identify the vehicle's position. Therefore, in this embodiment, at least one reference point (bottom point) and orientation vector of the vehicle are targeted for estimation by deep learning.

[0047] 3 is a top view of an exemplary training vehicle V0 (vehicle V) showing the positions and orientation vector Z0 of reference points 26, 28 in machine learning used by the vehicle position detection device 10. As shown in FIG. 3, in the estimation mode of this embodiment, the reference points 26, 28 used to estimate the position of the vehicle V are set at, for example, six locations on the training vehicle V0 (specifically, six locations on the bottom of the vehicle). As shown in FIG. 3, the reference points 26, 28 can be the ground contact points (four points: ground contact points 26a to 26d) of each wheel 24 of the training vehicle V0, a ​​first projection point 28a when a portion of the front end portion VF of the vehicle body is projected onto the road surface, and a second projection point 28b when a portion of the rear end portion VR of the vehicle body is projected onto the road surface.

[0048] Here, reference points 26, 28 are positions at which the position and posture of training vehicle V0 relative to the road surface can be easily recognized. In FIG. 3, reference point 26 of right front wheel 24FR is ground contact point 26a, and reference point 26 of right rear wheel 24RR is ground contact point 26b. Similarly, reference point 26 of left front wheel 24FL is ground contact point 26c, and reference point 26 of left rear wheel 24RL is ground contact point 26d. As described above, the installation position and angle of view of surveillance camera 103 are set so that it can monitor vehicle V in addition to demarcation lines L and parking area R in parking lot P. Therefore, if training vehicle V0 on the road surface is captured in image information captured by surveillance camera 103, it is highly likely that one of wheels 24 is captured, and it is highly likely that one of ground contact points 26a to 26d can be identified as reference point 26.

[0049] Furthermore, if the learning vehicle V0 is captured in the image information captured by the monitoring camera 103, it is highly likely that either the front end VF or the rear end VR of the learning vehicle V0 is captured. In other words, it is highly likely that either the first projection point 28a, which is obtained when a portion of the front end VF of the vehicle is projected onto the road surface, or the second projection point 28b, which is obtained when a portion of the rear end VR of the vehicle is projected onto the road surface, can be confirmed. It is easy to determine whether the front end or the rear end of the learning vehicle V0 is captured in the captured image. Therefore, the front end VF of the learning vehicle V0 is easy to recognize in either position. Annotators who perform annotation while viewing image information may deviate from the position they consider to be the correct value of the training data depending on their own experience and intuition. However, for example, in the front end portion VF of the vehicle body, the right corner portion FRc of the front bumper, the left corner portion FLc of the front bumper, or the approximate center portion (approximately central position) FC, etc., have high geometrical characteristics, so there is little discrepancy in recognition by the annotator. In particular, since the vehicle V is generally symmetrical, the discrepancy in recognition of the approximate center portion FC of the front end portion VF of the vehicle body can be considered very small. The same is true for the rear end portion VR of the vehicle body, where the right corner portion RRc of the rear bumper, the left corner portion RLc of the rear bumper, or the approximate center portion (approximately central position) RC are easy to recognize, and the discrepancy in recognition of the approximate center portion RC of the rear end portion RF of the vehicle body can be considered very small. In the example of Figure 3, the first projection point 28a is set to the approximate center portion FC of the front bumper, and the second projection point 28b is set to the approximate center portion RC of the rear bumper. Therefore, when annotating the training vehicle V0, by using the ground contact points 26a to 26d of each wheel 24 and the first projection point 28a and the second projection point 28b as the correct values ​​for identifying the position of the vehicle V, more accurate annotation can be easily achieved without being influenced by the level of skill of the annotator, etc.

[0050] For example, Fig. 4 is an exemplary schematic perspective view showing the positions and orientation vector Z0 of reference points 26, 28 that are set for a training vehicle V0 included in an actual training image IM0 (image information) during training in machine learning used by the vehicle position detection device 10. Fig. 4 shows, for example, a training image IM0a captured by the surveillance camera 103 when the training vehicle V0 passes through the gate of a parking lot P and travels on an access road Pin.

[0051] The reference point + direction vector setting unit 16a sequentially sets (registers) reference points 26, 28 at positions on the training vehicle V0 that the annotator specifies by operating the operation unit 22b while checking the training vehicle V0 included in the training image IM0a displayed on the display unit 22a. In the example of FIG. 4, as the reference point 26 for the wheels 24 of the training vehicle V0, the ground contact point 26a of the right front wheel 24FR is set as the correct value for the ground contact point of the right front wheel 24FR of the training vehicle V0, and the ground contact point 26b of the right rear wheel 24RR is set as the correct value for the ground contact point of the right rear wheel 24RR. Also, in FIG. 4, the annotator projects the position of the approximate center FC of the vehicle front end VF of the training vehicle V0 displayed on the display unit 22a onto the road surface G, and sets a first projection point 28a as the correct value for the road surface projection position of the approximate center FC of the vehicle body front end VF of the vehicle V. In this way, the reference point and direction vector setting unit 16a can accurately and easily set the reference points 26 and 28 as correct values ​​at positions on the learning vehicle V0 that are easy to confirm.

[0052] FIG. 5 is an exemplary schematic perspective view showing the position of a reference point 28 that is set on a training vehicle V0 contained in an actual training image IM0 (image information) during training in machine learning used by the vehicle position detection device 10. FIG. 5 shows, for example, a training image IM0b captured by a monitoring camera 103 while the training vehicle V0 is traveling in a parking lot P. In the example shown in FIG. 5, the training vehicle V0 is facing directly ahead with respect to the monitoring camera 103. In this case, most of the wheels 24 (front wheel FR, left front wheel 24FL) are hidden by the vehicle body, and the contact points are not visible.

[0053] The reference point + direction vector setting unit 16a sets (registers) a reference point 28 at a position on the learning vehicle V0 that the annotator specifies by operating the operation unit 22b while checking the learning vehicle V0 displayed on the display unit 22a. In the example of FIG. 5, the annotator projects the position of the approximate center FC of the vehicle front end VF of the learning vehicle V0 displayed on the display unit 22a onto the road surface G, and sets the first projection point 28a as the correct value for the road surface projection position of the approximate center FC of the vehicle body front end VF of the vehicle V. When the area directly behind the learning vehicle V0 is shown in the learning image IM0, only the second projection point 28b can be set by the annotator.

[0054] In this way, regardless of the posture of the learning vehicle V0, the reference point + direction vector setting unit 16a can easily and accurately register at least one of the ground contact points 26a-26d of the wheels 24, the first projection point 28a, and the second projection point 28b as a correct value. The reference point + direction vector setting unit 16a may set reference points for hidden portions of the learning vehicle V0 through estimation (prediction) by an annotator. For example, in the case of FIG. 4, the ground contact point 26c of the left front wheel FL, the ground contact point 26d of the left rear wheel RL, and the second projection point 28b, which is the road surface projection position of the approximate center RC of the rear end VR of the vehicle body, may be set through estimation by an annotator. In this case, even if a learning image IM0 in which the left side or rear of the vehicle body of the learning vehicle V0 can be confirmed is not obtained when the learning mode is executed, the estimation mode can be executed using the reference points 26 and 28 estimated during learning. 5, even when the wheel 24 is not visible, the annotator may set the reference points 26a to 26d, etc., through estimation. In this case, the positional accuracy may be lower than that of the reference points 26, 28 that are visible on the learning image IM0, so the reference points 26, 28 set through estimation may be distinguished as reference values.

[0055] The reference point + direction vector setting unit 16a sets correct values ​​for a plurality of learning images IM0. The more learning images IM0 processed by the reference point + direction vector setting unit 16a, the more accurately a learned model can be constructed in the model creation unit 16b to realize position estimation of the vehicle V to be estimated when the estimation mode is executed.

[0056] It should be noted that the learning image IM0 annotated by the reference point + direction vector setting unit 16a may or may not have been captured in a parking lot P where the position of an actual vehicle V is estimated, as shown in FIGS. 4 and 5. In other words, image information captured in another location may be used as the learning image IM0 as long as it is a vehicle type that may enter the parking lot P. It may also be acquired from an existing database that stores vehicle images. Note that, when using learning images captured in a parking lot P where the position of an actual vehicle V is estimated, a trained model that corresponds to the tendency of vehicles using the parking lot P can be constructed at an early stage, enabling efficient construction of the trained model.

[0057] The model creation unit 16b constructs a trained model and sequentially updates the contents of the trained model storage unit 20. Well-known techniques can be used to construct the trained model, and detailed description thereof will be omitted here. For example, the trained model can be created by deep learning, a type of machine learning technique. The model creation unit 16b performs training using the training data created by the reference point + orientation vector setting unit 16a. In this case, for example, a function is constructed using parameters, a loss is defined for the correct data, and training is performed by minimizing this loss.

[0058] The estimation mode execution unit 18 executes estimation mode processing using a trained model in order to estimate accurate position information of a vehicle V present in a parking lot P. The estimation mode execution unit 18 is executed when the estimation mode is selected by the mode switching unit 12.

[0059] For example, when acquisition unit 14 acquires a signal indicating that an object (moving body) has entered parking lot P from entry detection unit 23, estimation mode execution unit 18 acquires image information from surveillance camera 103 via acquisition unit 14. In the following explanation, FIGS. 4 and 5 will be referred to as estimation images IM1 and IM2.

[0060] The reference point and direction vector estimation unit 18a estimates a corresponding position (specific point) and direction vector of the underside of the estimation vehicle corresponding to the position of the reference point based on a comparison between the learned model stored in the learned model storage unit 20 and the estimation vehicle V1 included in the estimation image acquired by the acquisition unit 14. That is, the estimation image IM1 captured by the monitoring camera 103 and acquired by the acquisition unit 14 is input to the learned model, an area that can be considered to be the estimation vehicle V1 is extracted, and specific points 26T and 28T corresponding to the reference points 26 and 28 are detected. Note that the reference point and direction vector estimation unit 18a does not need to detect specific points 26T and 28T corresponding to all of the six reference points 26 and 28 described above; it is sufficient to detect at least one of them. As described above, specific point 26Ta corresponding to ground contact point 26a as reference point 26 is the ground contact point of right front wheel 24FR, specific point 26Tb corresponding to ground contact point 26b is the ground contact point of right rear wheel 24RR, specific point 26Tc corresponding to ground contact point 26c is the ground contact point of left front wheel 24FL, and specific point 26Td corresponding to ground contact point 26d is the ground contact point of left rear wheel 24RL. Furthermore, of the reference points 28, specific point 28Ta corresponding to first projection point 28a is a road surface projection point of approximately the center FC of the vehicle front end VF, and specific point 28Tb corresponding to second projection point 28b is a road surface projection point of approximately the center RC of the vehicle rear end VR. Therefore, if at least one of the six specific points 26T and 28T can be detected, the position of the bottom of vehicle V1 can be determined, and the position of vehicle V1 can be identified.

[0061] If two or more of the specific points 26T and 28T can be detected, it becomes possible to further identify the orientation vector Z0 and size of the estimation vehicle V1 on the world coordinate system. For example, by inputting the estimation image IM1 shown in FIG. 4 into a learned model and comparing it, the estimation vehicle V1 is extracted, and if specific point 26Ta, contact point 26Tb, and specific point 28Ta are detected, the position (pixel position) of the estimation vehicle V1 on the estimation image IM1 (image) is identified. Furthermore, by detecting specific point 28Ta, it is possible to estimate the orientation vector Z0 indicating that the front end VF of the estimation vehicle V1 is facing the monitoring camera 103. Furthermore, the length of the wheelbase can be estimated from the distance (number of pixels) between specific point 26Ta and specific point 26Tb.

[0062] 5 is input to the trained model for comparison, the estimation vehicle V1 is extracted, and when only the specific point 28Ta is detected, the position (pixel position) of the estimation vehicle V1 on the estimation image IM2 (image) is identified. Furthermore, by detecting the specific point 28Ta, a direction vector Z0 indicating that the front end VF of the estimation vehicle V1 is facing the monitoring camera 103 can be detected.

[0063] Next, processing unit 18b converts the corresponding positions (specific points 26T, 28T) and orientation vectors corresponding to reference points 26, 28 estimated by reference point+orientation vector estimation unit 18a into a world coordinate system. Well-known techniques can be used for the conversion to the world coordinate system, and details will be omitted. For example, conversion unit 18b1 projects the estimated coordinates of estimation vehicle V1 onto the world coordinate system using an affine projection technique or the like. Position identification unit 18b2 identifies the position (coordinates) of estimation vehicle V1 on the world coordinate system and outputs the position (coordinates) to display device 22 or parking control device 101, for example.

[0064] For example, Fig. 6 is an exemplary schematic perspective view showing the positions and orientation vector Z0 of specific points 26T and 28T corresponding to reference points 26 and 28, detected when an estimation mode using machine learning is executed in the vehicle position detection device 10. In the example shown in Fig. 6, an estimation vehicle V1 is extracted on the road surface G of a parking lot P by comparing an estimation image IM2 captured by a surveillance camera 103 with a trained model. For the estimation vehicle V1, specific points 26Ta of the right front wheel 24FR, 26Tc of the left front wheel 24FL, and 28Ta and orientation vector Z0 corresponding to a first projection point 28a, which is a road surface projection point at approximately the center FC of the front end VF of the vehicle body, are detected.

[0065] 6, when the bottom point is estimated, processing unit 18b outputs the estimated orientation vector Z0 of estimation vehicle V1 to the center of the rectangle. In addition, processing unit 18b moves the starting point of orientation vector Z0 to the bottom point, and converts the image coordinates of the bottom point and any one point on orientation vector Z0 into the world coordinate system, thereby acquiring orientation vector Z1 in the world coordinate system.

[0066] FIG. 7 is an exemplary schematic diagram illustrating a case where the position of the estimation vehicle V1 is projected onto a map M in the world coordinate system based on the specific points and orientation vector detected in FIG. 6. In this case, since the road surface G of the parking lot P is at the same height, the height direction of the map M in the world coordinate system is omitted. FIG. 8 is a diagram illustrating a method for projecting the position of the estimation vehicle V1 onto the map M in the world coordinate system. As shown in FIG. 8, the processing unit 18b first aligns the position of the template TP with a reliable point (the most probable detected point) projected onto the world coordinate system. Next, the processing unit 18b rotates the template TP so that it coincides with the direction of the orientation vector in the world coordinate system. As a result, as shown in FIG. 7, the estimation vehicle V1 is indicated on the map M by corresponding positions (specific points 26T, 28T) corresponding to reference points 26, 28 indicating the position of the bottom of the vehicle and the orientation vector Z1. In the case of FIG. 7, the location of the estimation vehicle V1 can be easily and accurately detected from a bird's-eye view.

[0067] As described above, according to this embodiment, when one or more reference points (bottom points) are estimated, the detected detection points (bottom points) are matched with corresponding locations on the template TP, and then the template TP is rotated around the detection points as the rotation center so that the forward direction matches the orientation vector of the world coordinate system, thereby correcting the position and determining the position and attitude of the vehicle.

[0068] The estimation mode execution unit 18 may feed back the estimation image and estimation results used in the estimation mode processing to the learning mode execution unit 16 and use them to update the trained model.

[0069] In this way, the vehicle position detection device 10 of this embodiment estimates (identifies) specific points 26T, 28T corresponding to the reference points 26, 28 set on the opposite side of the vehicle to be used for estimating the position of the vehicle V. As a result, the position (coordinates) of the estimation vehicle V1 on the world coordinate system can be detected more accurately and easily using only image information captured by the monitoring camera 103, without using a special sensor (for example, a sensor that identifies the position with high accuracy, such as a radar).

[0070] The operation of the vehicle position detection device 10 configured as above will be described with reference to the flowcharts of FIGS.

[0071] FIG. 9 is an exemplary flowchart showing a process performed when the learning mode in the machine learning used in the vehicle position detection device 10 is executed.

[0072] The vehicle position detection device 10 first switches between a learning mode and an estimation mode via the mode switching unit 12 in response to an operation of the operation unit 22b by the annotator (S100). If the learning mode is not selected (No in S100), this flow is temporarily terminated. On the other hand, if the learning mode is selected (Yes in S100), learning images (information) captured by the monitoring camera 103 are acquired via the acquisition unit 14 (S102). Then, the reference point and orientation vector setting unit 16a of the learning mode execution unit 16 creates training data in response to an operation of the operation unit 22b by the annotator (S104). That is, for multiple learning images IM0, reference points 26 and 28 and orientation vectors for the learning vehicle V0 are set and associated with the learning vehicle V0. Then, the model creation unit 16b executes a learning process to create a trained model using a well-known machine learning (e.g., deep learning) technique (S106). Then, the model creation unit 16b provides the created trained model to the trained model storage unit 20, and constructs (updates) the stored trained model (S108).

[0073] The vehicle position detection device 10 checks whether the learning termination condition is met based on the state of the mode switching unit 12 and the operation state of the operation unit 22b (S110). If the learning termination condition is not met (No in S110), the process proceeds to S100 and checks the current mode state. If the learning mode continues, the next multiple learning images IM0 are acquired, and the subsequent processes are repeatedly executed to continue building the trained model. If the learning termination condition is met in S110 (Yes in S110), for example, if a command to end learning is input from the operation unit 22b or if imaging by the monitoring camera 103 stops, the flow is temporarily terminated.

[0074] FIG. 10 is an exemplary flowchart showing processing when the estimation mode in machine learning used in the vehicle position detection device 10 is executed.

[0075] The vehicle position detection device 10 first switches between the learning mode and the estimation mode via the mode switching unit 12 in response to an operation of the operation unit 22b by the annotator (S200). If the estimation mode is not selected (No in S200), this flow is temporarily terminated. On the other hand, if the estimation mode is selected (Yes in S200), the acquisition unit 14 causes the entry detection unit 23 to check whether an object (moving object) has entered the parking lot P (S202). If the entry of an object cannot be confirmed (No in S202), that is, if there is no entry of the estimation vehicle V1 or the like, the vehicle position detection device 10 temporarily terminates this flow. As a result, if there is no entry of the estimation vehicle V1 or the like, the processing is paused, which makes it possible to reduce the processing load or to use the pause period to execute the learning mode processing, thereby contributing to efficient operation of the vehicle position detection device 10.

[0076] In S202, if the intrusion of an object is confirmed (Yes in S202), that is, if the intrusion of an estimation vehicle V1 or the like is detected, the reference point + direction vector estimation unit 18a acquires the estimation image (information) captured by the surveillance camera 103 via the acquisition unit 14 (S204).

[0077] Furthermore, the reference point+orientation vector estimation unit 18a acquires a learned model from the learned model storage unit 20 (S206). Then, the reference point+orientation vector estimation unit 18a inputs the estimation image to the learning model, extracts the estimation vehicle V1 from the estimation image IM2 etc., and executes a process of identifying specific points 26T, 28T, which are positions corresponding to the reference points 26, 28, and the orientation vector (S208).

[0078] Next, reference point+orientation vector estimation unit 18a checks whether specific points 26T, 28T, which are positions corresponding to reference points 26, 28, have been detected (S210). If no specific points exist (No in S210), the vehicle position cannot be detected, and reference point+orientation vector estimation unit 18a proceeds to S218 to perform a determination process of whether an estimation process termination condition is met. On the other hand, if reference point+orientation vector estimation unit 18a determines that specific point 26T or specific point 28T exists in estimation image IM2 or the like (Yes in S210), conversion unit 18b1 converts the coordinates of specific point 26T or specific point 28T into the world coordinate system and projects them onto map M of the world coordinate system (S212).

[0079] Next, the position specifying unit 18b2 specifies the position of the estimation vehicle on the map M of the world coordinate system (S214). The vehicle body position specifying process for specifying the position of the estimation vehicle will be described in detail below.

[0080] FIG. 11 is an exemplary flowchart illustrating the flow of the vehicle body position identification process. As shown in FIG. 11, first, the position identification unit 18b2 adjusts the position of the template TP (S2141). Specifically, the position identification unit 18b2 fixes the locations of the template TP corresponding to the detection points created according to the shape of the vehicle to the coordinates of the detection points. Here, when multiple detection points are estimated, the coordinates of the most likely detection point are used as the coordinates of the detection point for fixing the template TP. A reliability may be set in advance for each detection point. For example, the position identification unit 18b2 sets the first projection point 28a as the most reliable detection point.

[0081] For example, in an automatic valet parking system 100 that identifies the position of a vehicle, when the contact point of the tire on the front right side of the vehicle is detected, the position identification unit 18b2 fixes the point corresponding to the contact point of the tire on the front right side of the template TP created to fit the vehicle to the coordinates of the detected point.

[0082] Next, the position specifying unit 18b2 moves the direction vector (S2142). Specifically, the position specifying unit 18b2 moves the direction vector in the image coordinate system so that the starting point of the direction vector coincides with the detection point.

[0083] Next, the position specifying unit 18b2 acquires the image coordinates of two arbitrary points on the direction vector (S2143). Specifically, the position specifying unit 18b2 selects the detection point as one of the two points and an arbitrary point on the direction vector as the other point.

[0084] Next, the position specifying unit 18b2 acquires a direction vector in the world coordinate system (S2144). Specifically, the position specifying unit 18b2 converts the coordinates of any two points representing the direction vector acquired in S2143 into the world coordinate system. Then, the position specifying unit 18b2 calculates the direction vector in the world coordinate system using the coordinates of the two points in the world coordinate system.

[0085] Next, the position specifying unit 18b2 rotates the orientation of the template TP (S2145). Specifically, the position specifying unit 18b2 acquires an orientation vector in the world coordinate system. Then, the position specifying unit 18b2 uses the detection point as the center of rotation in the world coordinate system, and matches the orientation of the template TP with the orientation vector of the object, thereby uniquely determining the posture.

[0086] The above processing completes the vehicle body position identification processing for identifying the position of the estimation vehicle.

[0087] Next, returning to FIG. 10, the processing unit 18b executes a result notification process (S216) in which the position of the identified estimation vehicle on the map M of the world coordinate system is output to the display device 22 or the like, and the coordinates are provided to the parking control device 101.

[0088] Then, the vehicle position detection device 10 checks whether the estimation process termination condition is met based on the state of the mode switching unit 12 and the operation state of the operation unit 22b (S218), and if the estimation process termination condition is not met (No in S218), the process proceeds to S204, where the next estimation image is acquired, and the subsequent processes are repeated to continue the estimation process.If the estimation process termination condition is met in S218 (Yes in S218), for example, if an estimation termination operation is input from the operation unit 22b or imaging by the monitoring camera 103 stops, this flow is temporarily ended.

[0089] Thus, according to the vehicle position detection device 10 of the embodiment, the trained model is created in a state including at least one reference point that indicates a reference position that is included on the underside of the training vehicle and is easy to identify as a correct value, and the vehicle position and orientation are estimated based on the reference point and the orientation vector. As a result, the accuracy of estimating the vehicle position and orientation from image information can be improved, and more accurate position and orientation can be detected on the world coordinate system.

[0090] In particular, according to the vehicle position detection device 10 of the embodiment, in order to accurately detect the position and orientation of an object in a world coordinate system from a camera image, a template TP corresponding to the vehicle shape is created in advance, and at least one bottom point (detection point) and orientation vector of the vehicle are used as estimation targets for deep learning. Then, after matching the detected bottom point (detection point) with a corresponding location on the template TP, the template TP is rotated around the detection point as the rotation center so that the forward direction matches the orientation vector, thereby identifying the position and orientation of the vehicle.

[0091] Furthermore, according to the vehicle position detection device 10 of the embodiment, by estimating the bottom points (detection points) of the vehicle for each corresponding position without confusing them, it is possible to determine which position on the vehicle each bottom point (detection point) corresponds to. According to the vehicle position detection device 10 of the embodiment, even when multiple bottom points (detection points) are defined for the vehicle, and only one point is estimated due to occlusion or the like, the position and orientation of the object can be uniquely determined using the orientation vector.

[0092] Furthermore, according to the vehicle position detection device 10 of the embodiment, the type of vehicle is not limited to one, and by using templates TP corresponding to multiple types of vehicles and a means for estimating the vehicle type, it is possible to detect the positions and attitudes of multiple vehicles in the world coordinate system.

[0093] In this embodiment, after the detected bottom point (detection point) is matched with the corresponding portion of the template TP, the template TP is rotated around the bottom point (detection point) as the center of rotation so that the forward direction matches the orientation vector, but this is not limited to this. When multiple bottom points (detection points) are estimated, the image coordinates of the multiple bottom points (detection points) may be converted into a world coordinate system to uniquely identify the position and orientation of the vehicle. Furthermore, to make the system robust, matching with the template TP may be performed individually at multiple bottom points (detection points), and the position and orientation of the vehicle may be estimated by averaging the matching results.

[0094] In this embodiment, the orientation vector is moved so that the starting point of the orientation vector coincides with the bottom point (detection point), but this is not limiting, and the orientation vector may be moved so that it coincides with any other appropriate point on the vehicle. For example, among the contact points with the bottom point (detection point), a point that is not a detection target but is located closest to the vehicle may be applied.

[0095] In the above-described embodiment, an example in which the vehicle position detection device 10 is applied to the automatic valet parking system 100 has been described. However, the vehicle position detection device 10 can be applied to other technologies and provide similar benefits as long as the vehicle position contained in the image information can be effectively utilized by identifying the vehicle position on a world coordinate system. For example, the vehicle position detection device 10 can also be used to estimate the vehicle position in a regular parking lot, monitor the parking status of vehicles on roads, traffic conditions, and whether or not a vehicle is entering a specific area. This can contribute to vehicle monitoring and control based on the accurate vehicle position.

[0096] In the above-described embodiment, an example in which the object is a vehicle (four-wheeled automobile) has been described, but this is not limited to this, and it goes without saying that the object can also be a moving body such as a bicycle or a motorcycle, or a human being.

[0097] In addition, the programs for the learning mode processing and estimation mode processing executed by the processor that realizes the vehicle position detection device 10 of this embodiment may be configured to be provided by being recorded in an installable or executable format on a computer-readable recording medium such as a CD-ROM, a flexible disk (FD), a CD-R, or a DVD (Digital Versatile Disk).

[0098] Furthermore, the program for executing the processing of this embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the program executed in this embodiment may be provided or distributed via a network such as the Internet.

[0099] Although the embodiments and modifications of the present invention have been described, these embodiments and modifications are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and modifications are included within the scope and spirit of the invention, and are also included in the inventions and their equivalents as defined in the claims. [Explanation of symbols]

[0100] 10...vehicle position detection device, 12...mode switching unit, 14...acquisition unit, 16...learning mode execution unit, 16a...reference point + direction vector setting unit, 16b...model creation unit, 18...estimation mode execution unit, 18a...reference point estimation unit, 18b...processing unit, 18b1...conversion unit, 18b2...position identification unit, 20...learned model memory unit, 22...display device, 22a...display unit, 22b...operation unit, 23...entry detection unit, 24...wheel, 26, 28...reference point, 26a, 26b, 26c, 26d...ground contact point, 28a...first projection point, 28b...second projection point, 26T, 28T...specific point, TP...template.

Claims

1. an acquisition unit that acquires image information from an imaging unit that images the detection target area; an estimation unit that estimates a corresponding position and orientation of the object bottom surface of the estimation object corresponding to the position of the reference point based on a comparison between a training object included in the image information, at least one reference point indicating a reference position included in the object bottom surface of the training object, and a training object model obtained by learning the relationship between the orientation vector of the training object and the training object included in the image information acquired by the acquisition unit; a processing unit that transforms the corresponding positions into a world coordinate system and detects the position and orientation of the estimation object on the world coordinate system; Equipped with the trained model is a shape of a training object included in the image information, and is a result of learning a relationship between a direction vector indicating a predetermined direction of the training object and at least one reference point indicating a reference position included in a bottom surface of the training object; the processing unit matches the corresponding position with a corresponding portion of a template prepared in advance according to the type of the object to be estimated, then converts image coordinates of any two points on the orientation vector into a world coordinate system to obtain an orientation vector in the world coordinate system, and rotates the template around the corresponding position as a rotation center so that the predetermined direction coincides with the orientation vector in the world coordinate system, thereby specifying the position and orientation of the object to be estimated. Object position estimation device.

2. when a plurality of corresponding positions are estimated by the estimation unit, the processing unit performs matching with the template individually at the plurality of corresponding positions and identifies the position and orientation of the object for estimation by averaging matching results. The object position estimation device according to claim 1 .

3. when a plurality of the corresponding positions are estimated by the estimation unit, the processing unit specifies the position and orientation of the object for estimation using one of the corresponding positions that is most likely among the plurality of the corresponding positions and the orientation vector. The object position estimation device according to claim 1 .

4. when converting the image coordinates of any two points on the orientation vector into a world coordinate system, the processing unit moves the orientation vector so that a starting point of the orientation vector coincides with the corresponding position of the object for estimation. The object position estimation device according to claim 1 .

5. the processing unit, when converting the image coordinates of any two points on the orientation vector into a world coordinate system, moves the orientation vector so that a start point of the orientation vector coincides with any point on the object for estimation. The object position estimation device according to claim 1 .

6. the reference point is at least one of a ground contact point of the training object, a first projection point when a part of a front end of the training object is projected onto a road surface, and a second projection point when a part of a rear end of the training object is projected onto a road surface. The object position estimation device according to any one of claims 1 to 5.

7. the first projection point is a projection point at a substantially central position of a front end portion of the training object, and the second projection point is a projection point at a substantially central position of a rear end portion of the training object. The object position estimation device according to claim 6 .

8. the estimation unit performs the estimation when detecting entry of the object to be estimated into the detection target area. The object position estimation device according to any one of claims 1 to 7.

9. An acquisition unit that acquires image information from an imaging unit that images a detection target area; an estimation unit that estimates a corresponding position and orientation of the object bottom surface of the estimation object corresponding to the position of the reference point based on a comparison between a training object included in the image information, at least one reference point indicating a reference position included in the object bottom surface of the training object, and a training object model obtained by learning the relationship between the orientation vector of the training object and the training object included in the image information acquired by the acquisition unit; a processing unit that transforms the corresponding positions into a world coordinate system and detects the position and orientation of the estimation object on the world coordinate system; Equipped with the reference point is at least one of a ground contact point of the training object, a first projection point when a part of a front end of the training object is projected onto a road surface, and a second projection point when a part of a rear end of the training object is projected onto a road surface. Object position estimation device.

10. The first projection point is a projection point at approximately the center of the front end of the learning object, and the second projection point is a projection point at approximately the center of the rear end of the learning object. The object position estimation device according to claim 9 .

11. The estimation unit performs the estimation when it detects that the object for estimation has entered the detection target area. The object position estimation device according to claim 9 or 10.

12. An object position estimation device according to any one of claims 1 to 11; a parking control device that controls the movement of the vehicle within the parking lot based on the position of the vehicle estimated by the object position estimation device; A parking system comprising:

Citation Information

Patent Citations

  • Non-inductive charging pile transaction visual management system based on computer vision and CIM

    CN111260852A

  • Vehicle progress status estimating system, vehicle progress status estimating method, and vehicle progress status estimating program

    JP2020190413A

  • Object location estimation device and program of the same, and object location estimation method

    JP2021026281A