Position estimation method, position estimation apparatus, and position estimation system

The method adapts position estimation for moving objects by using bounding boxes or feature points based on camera distance, addressing accuracy issues and reducing processing loads, thus maintaining precision and efficiency.

JP2025108916APending Publication Date: 2025-07-24TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024002453
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The accuracy of position estimation for moving objects, such as vehicles, decreases when the distance between the camera and the object is small due to increased differences between the bounding box and the actual contour, especially when the camera is positioned low, leading to inaccuracies in estimating the object's position.

Method used

A method that switches between using a bounding box and line segments connecting feature points for position estimation based on the distance between the camera and the object, employing machine learning models to accurately determine the object's position without relying solely on bounding boxes when the distance is small.

Benefits of technology

This approach maintains position estimation accuracy by using appropriate methods based on distance, reducing errors and processing requirements, and minimizing annotation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108916000001_ABST
    Figure 2025108916000001_ABST
Patent Text Reader

Abstract

To provide a position estimation method which accurately estimates a position of a moving body.SOLUTION: A position estimation method for estimating a position of a moving body comprises: an acquisition step of acquiring a captured image output from a camera; and a position estimation step of estimating the position of the moving body included in the captured image through image recognition. In the position estimation step, when a distance in the gravity direction between the moving body and the camera is greater than or equal to a predetermined threshold, the position of the moving body is estimated by using at least one of a bounding box generated by estimating a region including the moving body from the captured image, and a contour of the moving body extracted from the captured image, and when the distance is less than the threshold, the position of the moving body is estimated by using a line segment connecting two of a plurality of feature points of the moving body extracted from the captured image.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a position estimation method, a position estimation device, and a position estimation system.

Background Art

[0002] Conventionally, a technique for estimating the position of an object using a bounding box generated by estimating a region including the object from a captured image of a camera installed at a location different from the object is known (Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The smaller the distance in the gravitational direction between a moving object as the object and the camera, the greater the difference between the bounding box and the contour of the actual moving object, and thus the accuracy of the estimated position of the moving object may decrease. For example, when estimating the position of a moving object such as a vehicle moving on a road, if the vertical height of a camera installed on the road surface is low, the camera captures the moving object from an obliquely upper direction. As a result, the difference between the bounding box and the contour of the actual moving object may increase, and the accuracy of the estimated position of the moving object may decrease. Such a problem is not limited to the case of estimating the position of a moving object using a bounding box, but is also common in the case of estimating the position of a moving object using the contour of the moving object extracted from a captured image.

Means for Solving the Problems

[0005] The present disclosure can be realized in the following forms.

[0006] (1) According to the first aspect of the present disclosure, a position estimation method is provided. The position estimation method for estimating the position of a moving body that can move by autonomous driving includes an acquisition step of acquiring a captured image output from a camera installed at a location different from the moving body, and a position estimation step of estimating the position of the moving body included in the captured image by image recognition. In the position estimation step, when the distance in the gravitational direction between the moving body and the camera is equal to or greater than a predetermined threshold, the position of the moving body is estimated using at least one of a bounding box generated by estimating a region including the moving body from the captured image, and the contour of the moving body extracted from the captured image. When the distance is less than the threshold, the position of the moving body is estimated using a line segment connecting two of the plurality of feature points of the moving body extracted from the captured image. According to this aspect, when the distance in the gravitational direction between the moving body and the camera is less than the threshold, the position of the moving body can be estimated using a line segment connecting two of the plurality of feature points of the moving body extracted from the captured image. In this way, the position of the moving body can be estimated without using the bounding box generated by estimating the region including the moving body from the captured image. Thereby, it is possible to suppress a decrease in the accuracy of the estimated position of the moving body due to an increase in the difference between the bounding box and the actual contour of the moving body. Further, according to this aspect, the position of the moving body can be estimated without using the contour of the moving body extracted from the captured image. Thereby, it is possible to suppress a decrease in the accuracy of the estimated position of the moving body due to an increase in the difference between the contour of the moving body extracted from the captured image and the actual contour of the moving body. (2) In the above aspect, the plurality of feature points are obtained by inputting the captured image to a learned machine learning model that outputs the plurality of feature points when the captured image is input. The machine learning model may be pre-learned using one or more learning datasets generated by labeling the plurality of feature points as correct labels in training images. According to this aspect, the plurality of feature points of the moving body can be obtained by inputting the captured image to a machine learning model that outputs the plurality of feature points of the moving body when the captured image is input. (3) According to a second aspect of the present disclosure, a position estimation device is provided. A position estimation device that estimates the position of a moving body that can move by autonomous driving includes an acquisition unit that acquires a captured image output from a camera installed at a location different from the moving body, and a position estimation unit that estimates the position of the moving body included in the captured image by image recognition. The position estimation unit estimates the position of the moving body using at least one of a bounding box generated by estimating a region including the moving body from the captured image and a contour of the moving body extracted from the captured image when a distance in the gravitational direction between the moving body and the camera is equal to or greater than a predetermined threshold value, and estimates the position of the moving body using a line segment connecting two of the plurality of feature points of the moving body extracted from the captured image when the distance is less than the threshold value. According to this aspect, when the distance in the gravitational direction between the moving body and the camera is less than the threshold value, the position of the moving body can be estimated using a line segment connecting two of the plurality of feature points of the moving body extracted from the captured image. By doing so, the position of the moving body can be estimated without using a bounding box generated by estimating a region including the moving body from the captured image. As a result, it is possible to suppress a decrease in the accuracy of the estimated position of the moving body due to an increase in the difference between the bounding box and the actual contour of the moving body. Further, according to this aspect, the position of the moving body can be estimated without using the contour of the moving body extracted from the captured image. As a result, it is possible to suppress a decrease in the accuracy of the estimated position of the moving body due to an increase in the difference between the contour of the moving body extracted from the captured image and the actual contour of the moving body. (4) In the above-described form, the position estimation unit acquires the plurality of feature points by inputting the captured image to a trained machine learning model that outputs the plurality of feature points when the captured image is input. The machine learning model may be pre-trained using one or more training datasets generated by labeling the plurality of feature points as correct labels in training images. According to this form, the plurality of feature points of the moving object can be acquired by inputting the captured image to a machine learning model that outputs the plurality of feature points of the moving object when the captured image is input. (5) According to a third aspect of the present disclosure, a position estimation system is provided. The position estimation system for estimating the position of a moving object includes a moving object movable by autonomous driving, a camera installed at a location different from the moving object, an acquisition unit that acquires a captured image output from the camera, and a position estimation unit that estimates the position of the moving object included in the captured image by image recognition. The position estimation unit estimates the position of the moving object using at least one of a bounding box generated by estimating a region including the moving object from the captured image and a contour of the moving object extracted from the captured image when a distance in the direction of gravity between the moving object and the camera is equal to or greater than a predetermined threshold value. When the distance is less than the threshold value, the position of the moving object is estimated using a line segment connecting two of the plurality of feature points of the moving object extracted from the captured image. According to this aspect, when the distance in the direction of gravity between the moving object and the camera is less than the threshold value, the position of the moving object can be estimated using a line segment connecting two of the plurality of feature points of the moving object extracted from the captured image. In this way, the position of the moving object can be estimated without using the bounding box generated by estimating the region including the moving object from the captured image. Thereby, it is possible to suppress a decrease in the accuracy of the estimated position of the moving object due to an increase in the difference between the bounding box and the actual contour of the moving object. Further, according to this aspect, the position of the moving object can be estimated without using the contour of the moving object extracted from the captured image. Thereby, it is possible to suppress a decrease in the accuracy of the estimated position of the moving object due to an increase in the difference between the contour of the moving object extracted from the captured image and the actual contour of the moving object. The present disclosure can be realized in various forms other than the above-described position estimation method, position estimation device, and position estimation system. For example, it can be realized in forms such as a control method of a position estimation device and a position estimation system, a computer program that realizes the control method, and a non-transitory recording medium on which the computer program is recorded. BRIEF DESCRIPTION OF THE DRAWINGS

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0008] A. First Embodiment: FIG. 1 is a conceptual diagram showing the configuration of the traveling system 50. The traveling system 50 includes a position estimation system 7 and a remote control device 80. The position estimation system 7 is a system for estimating the position of a moving body. The position estimation system 7 includes an external sensor 300, one or more vehicles 100 as moving bodies, and a position estimation device 70 for estimating the position of the vehicle 100. The remote control device 80 remotely controls the operation of the vehicle 100 using the position etc. of the vehicle 100. In the present embodiment, the functions of the position estimation device 70 and the remote control device 80 are realized by a server 200 installed at a location different from the vehicle 100.

[0009] In the present disclosure, a "mobile body" means an object that can move, for example, a vehicle or an electric vertical take-off and landing aircraft (so-called flying car). The vehicle may be a vehicle that runs on wheels or a vehicle that runs on an endless track, and examples thereof include a passenger car, a truck, a bus, a two-wheeled vehicle, a four-wheeled vehicle, a tank, a construction vehicle, and the like. The vehicle includes a battery electric vehicle (BEV), a gasoline vehicle, a hybrid vehicle, and a fuel cell vehicle. When the mobile body is other than a vehicle, the expressions "vehicle" and "car" in the present disclosure can be appropriately replaced with "mobile body", and the expression "running" can be appropriately replaced with "moving".

[0010] The vehicle 100 is configured to be capable of traveling by autonomous driving. "Autonomous driving" means driving without depending on the driving operation of a passenger. The driving operation means an operation related to at least any one of "running", "turning", and "stopping" of the vehicle 100. Autonomous driving is realized by automatic or manual remote control using a device located outside the vehicle 100, or by autonomous control of the vehicle 100. A passenger who does not perform a driving operation may board the vehicle 100 traveling by autonomous driving. Passengers who do not perform a driving operation include, for example, a person simply sitting on the seat of the vehicle 100 and a person performing work different from the driving operation, such as assembly, inspection, and operation of switches, while boarding the vehicle 100. Note that driving by the driving operation of a passenger may be called "manned driving".

[0011] In this specification, "remote control" includes "complete remote control" in which all the operations of the vehicle 100 are completely determined from outside the vehicle 100 and "partial remote control" in which a part of the operations of the vehicle 100 is determined from outside the vehicle 100. Further, "autonomous control" includes "complete autonomous control" in which the vehicle 100 autonomously controls its own operations without receiving any information from a device outside the vehicle 100 and "partial autonomous control" in which the vehicle 100 autonomously controls its own operations using the information received from a device outside the vehicle 100.

[0012] In this embodiment, the driving system 50 is used in a factory FC that manufactures the vehicle 100. The reference coordinate system of the factory FC is the global coordinate system GC, and any position within the factory FC can be represented by the coordinates X, Y, and Z in the global coordinate system GC. A plurality of external sensors 300 are installed along the runway TR in the factory FC. The external sensor 300 is a sensor located outside the vehicle 100. In this embodiment, the external sensor 300 is a camera that images the vehicle 100 and outputs a captured image as a detection result. Hereinafter, the camera as the external sensor 300 is also referred to as the "external camera 310".

[0013] FIG. 2 is a block diagram showing the configuration of the driving system 50 in the first embodiment. The vehicle 100 includes a vehicle control device 110, an actuator group 120 including one or more actuators, and a communication device 130 for communicating with an external device such as a server 200 by wireless communication. The actuator group 120 includes an actuator of a driving device for accelerating the vehicle 100, an actuator of a steering device for changing the traveling direction of the vehicle 100, and an actuator of a braking device for decelerating the vehicle 100.

[0014] The vehicle control device 110 is constituted by a computer including a processor 111, a memory 112, an input / output interface 113, and an internal bus 114. The processor 111, the memory 112, and the input / output interface 113 are connected to be communicable bidirectionally via the internal bus 114. The actuator group 120 and the communication device 130 are connected to the input / output interface 113. The processor 111 functions as a vehicle control unit 115 by executing a program PG1 stored in the memory 112.

[0015] The vehicle control unit 115 controls the actuator group 120 to run the vehicle 100. The vehicle control unit 115 can run the vehicle 100 by controlling the actuator group 120 using the driving control signal received from the server 200. The driving control signal is a control signal for running the vehicle 100. In the present embodiment, the driving control signal includes the acceleration and steering angle of the vehicle 100 as parameters. In other embodiments, the driving control signal may include the speed of the vehicle 100 as a parameter instead of or in addition to the acceleration of the vehicle 100.

[0016] The server 200 is composed of a computer including a processor 201, a memory 202, an input / output interface 203, and an internal bus 204. The processor 201, the memory 202, and the input / output interface 203 are connected to be communicable bidirectionally via the internal bus 204. A communication device 205 for communicating with various external devices outside the server 200 is connected to the input / output interface 203. The communication device 205 can communicate with the vehicle 100 by wireless communication and can communicate with each external sensor 300 by wired communication or wireless communication. The processor 201 functions as an acquisition unit 211, a position estimation unit 212, a direction estimation unit 213, and a remote control unit 214 by executing a program PG2 stored in the memory 202.

[0017] The acquisition unit 211 acquires a captured image output from an external camera 310 that is planned to include the vehicle 100 to be the target of position estimation in the imaging range.

[0018] The position estimation unit 212 estimates the position of the vehicle 100 included in the captured image by image recognition. The position of the vehicle 100 included in the captured image can be estimated, for example, using a bounding box. The bounding box is generated by estimating the region including the vehicle 100 from the captured image. The bounding box is an outer circumscribed rectangle that surrounds the region estimated to be the region indicating the vehicle 100 among the regions constituting the captured image. The bounding box can be obtained, for example, by inputting the captured image into a general-purpose detection model DM that utilizes artificial intelligence. The general-purpose detection model DM is a general-purpose pre-trained machine learning model that detects objects in the captured image. The general-purpose detection model DM realizes, for example, pattern matching. When the captured image is input, the general-purpose detection model DM detects the vehicle 100 and outputs a bounding box. As the general-purpose detection model DM, for example, a convolutional neural network (hereinafter, CNN) can be used. The general-purpose detection model DM is stored in advance in the memory 202 of the server 200, for example.

[0019] FIG. 3 is a diagram schematically showing a bounding box BO output from the general-purpose detection model DM by inputting a first captured image IM1 output from the first external camera 311 into the general-purpose detection model DM. As shown in FIG. 1, the first external camera 311 is an external camera 310 installed at a position where the distance H1 from the vehicle 100 in the gravitational direction is equal to or greater than a predetermined threshold value HS. FIG. 4 is a diagram schematically showing a bounding box BO output from the general-purpose detection model DM by inputting a second captured image IM2 output from the second external camera 312 into the general-purpose detection model DM. As shown in FIG. 1, the second external camera 312 is an external camera 310 installed at a position where the distance H2 from the vehicle 100 in the gravitational direction is less than the threshold value HS.

[0020] As with the first external camera 311 shown in FIG. 1, the higher the vertical height of the external camera 310 with respect to the road surface RS, the greater the distance H in the gravitational direction between the vehicle 100 and the external camera 310, and the external camera 310 can image the vehicle 100 as if looking down from directly above. The vertical height of the external camera 310 with respect to the road surface RS is, for example, the vertical height from the road surface RS where the vehicle 100 is located to the center of the lens of the external camera 310. The greater the distance H in the gravitational direction between the vehicle 100 and the external camera 310, as shown in FIG. 3, the smaller the difference between the area indicated by the bounding box BO and the area actually occupied by the vehicle 100 when the vehicle 100 is projected onto the road surface RS from the upward gravitational direction. That is, the greater the distance H in the gravitational direction between the vehicle 100 and the external camera 310, the smaller the difference between the bounding box BO and the contour of the actual vehicle 100. However, as with the second external camera 312 shown in FIG. 1, the lower the vertical height of the external camera 310 with respect to the road surface RS, the smaller the distance H in the gravitational direction between the vehicle 100 and the external camera 310, and the external camera 310 images the vehicle 100 from an obliquely upward direction. Therefore, the smaller the distance H in the gravitational direction between the vehicle 100 and the external camera 310, as shown in FIG. 4, the greater the difference between the area indicated by the bounding box BO and the area actually occupied by the vehicle 100 when the vehicle 100 is projected onto the road surface RS from the upward gravitational direction. That is, the smaller the distance H in the gravitational direction between the vehicle 100 and the external camera 310, the greater the difference between the bounding box BO and the contour of the actual vehicle 100. Therefore, the smaller the distance H in the gravitational direction between the vehicle 100 and the external camera 310, the more likely the accuracy of the position of the vehicle 100 estimated using the bounding box BO will decrease.

[0021] Therefore, when the distance H between the vehicle 100 and the external camera 310 is greater than or equal to the threshold value HS, the position estimation unit 212 estimates the position of the vehicle 100 using the bounding box BO. Specifically, as shown in FIG. 3, the position estimation unit 212 first obtains the bounding box BO by inputting the first captured image IM1 into the general detection model DM. Next, the position estimation unit 212 calculates the coordinates of the intersection point IS1 of the diagonals DG1 and DG2 of the bounding box BO in the camera coordinate system using the coordinates representing the positions of the vertices VT1 to VT4 of the bounding box BO in the camera coordinate system. Next, the position estimation unit 212 converts the coordinates of the intersection point IS1 represented in the camera coordinate system into coordinates in the image coordinate system by means of perspective transformation or the like. The camera coordinate system is a coordinate system with the focal point of the external camera 310 that outputs the captured images IM1 and IM2 used for position estimation as the origin. The image coordinate system is a coordinate system with a point on the image plane as the origin. Next, the position estimation unit 212 converts the coordinates of the intersection point IS1 represented in the image coordinate system into coordinates in the global coordinate system GC using camera parameters or the like. The camera parameters are information regarding the external camera 310 that outputs the captured images IM1 and IM2 used for position estimation. The camera parameters are, for example, the installation position, installation orientation, and focal length of the external camera 310 that outputs the captured images IM1 and IM2 used for position estimation. Thereby, the position estimation unit 212 estimates the coordinates of the intersection point IS1 represented in the global coordinate system GC as the position of the vehicle 100.

[0022] On the other hand, when the distance H between the vehicle 100 and the external camera 310 is less than the threshold value HS, the position estimation unit 212 estimates the position of the vehicle 100 by a method different from the case where the distance H between the vehicle 100 and the external camera 310 is greater than or equal to the threshold value HS.

[0023] FIG. 5 schematically shows a plurality of feature points FP1 to FP4 output from the dedicated detection model DN by inputting the second captured image IM2 into the dedicated detection model DN, and line segments LS1 to LS4 connecting two of the plurality of feature points FP1 to FP4. When the distance H between the vehicle 100 and the external camera 310 is less than the threshold value HS, the position estimation unit 212 estimates the position of the vehicle 100 using the line segments LS1 to LS4 connecting two of the plurality of feature points FP1 to FP4 of the vehicle 100 extracted from the second captured image IM2.

[0024] Specifically, the position estimation unit 212 first obtains a plurality of feature points FP1 to FP4 of the vehicle 100 by inputting the second captured image IM2 into the dedicated detection model DN. The dedicated detection model DN is a trained machine learning model specialized for estimating the position of the vehicle 100, which is trained to accurately estimate the position of the vehicle 100 when the distance H between the vehicle 100 and the external camera 310 is less than the threshold value HS. The dedicated detection model DN realizes, for example, either semantic segmentation or instance segmentation. When the second captured image IM2 is input, the dedicated detection model DN outputs a plurality of feature points FP1 to FP4 of the vehicle 100. As the dedicated detection model DN, for example, a CNN can be used. As the dedicated detection model DN, for example, a CNN trained by supervised learning using one or more dedicated learning datasets can be used. The dedicated learning dataset is generated, for example, by labeling a plurality of feature points FP1 to FP4 of the vehicle 100 as correct labels in training images. During the learning of the CNN, it is preferable that the parameters of the CNN are updated by backpropagation (error backpropagation method) so as to reduce the error between the output result by the dedicated detection model DN and the label. The dedicated detection model DN is prepared, for example, inside or outside the driving system 50 and is stored in advance in the memory 202 of the server 200.

[0025] In this embodiment, the dedicated detection model DN extracts the four wheels WH1 to WH4 as the feature points FP1 to FP4 of the vehicle 100, and outputs the coordinates representing the positions of the respective feature points FP1 to FP4 in the camera coordinate system. Therefore, the position estimation unit 212 inputs the second captured image IM2 to the dedicated detection model DN to obtain the coordinates representing the positions of the four feature points FP1 to FP4 in the camera coordinate system. Next, the position estimation unit 212 generates four line segments LS1 to LS4 that connect two adjacent feature points FP1 to FP4 so as to form a polygon by the feature points FP1 to FP4. Thereby, the position estimation unit 212 obtains the rectangular data BR formed by the four line segments LS1 to LS4. Next, the position estimation unit 212 calculates the coordinates of the intersection point IS2 of the diagonals DG3 and DG4 of the rectangular data BR in the camera coordinate system using the coordinates representing the positions of the respective vertices of the rectangular data BR, that is, the respective feature points FP1 to FP4 in the camera coordinate system. Next, the position estimation unit 212 converts the coordinates of the intersection point IS2 represented in the camera coordinate system into coordinates in the image coordinate system by means of perspective transformation or the like. Next, the position estimation unit 212 converts the coordinates of the intersection point IS2 represented in the image coordinate system into coordinates in the global coordinate system GC using camera parameters or the like. Thereby, the position estimation unit 212 estimates the coordinates of the intersection point IS2 represented in the global coordinate system GC as the position of the vehicle 100.

[0026] Note that the dedicated detection model DN may be a machine learning model that detects a plurality of feature points FP1 to FP4 of the vehicle 100 and outputs one or more line segments LS1 to LS4 that connect two of the plurality of feature points FP1 to FP4 when the second captured image IM2 is input. In this case, the dedicated learning dataset used for learning the dedicated detection model DN is generated by labeling the training image with line segments LS1 to LS4 that connect two of the plurality of feature points FP1 to FP4 of the vehicle 100 as the correct labels.

[0027] The direction estimation unit 213 estimates the orientation of the vehicle 100 included in the captured images IM1 and IM2. In the present embodiment, the orientation of the vehicle 100 is represented by the direction of a vector along the longitudinal axis of the vehicle 100 from the rear side to the front side of the vehicle 100. When the distance H between the vehicle 100 and the external camera 310 is greater than or equal to the threshold value HS, the direction estimation unit 213 estimates the orientation of the vehicle 100 using, for example, the bounding box BO. When the distance H between the vehicle 100 and the external camera 310 is greater than or equal to the threshold value HS, the direction estimation unit 213 estimates the orientation of the vehicle 100 using, for example, the rectangular data BR. Note that the direction estimation unit 213 may estimate the orientation of the vehicle 100 by estimating, for example, based on the direction of the movement vector of the vehicles 100 and 100v calculated from the position change of the feature points of the vehicle 100 between the frames of the captured images IM1 and IM2 using the optical flow method.

[0028] The remote control unit 214 generates a driving control signal for controlling the actuator group 120 of the vehicle 100, and transmits the driving control signal to the vehicle 100, thereby driving the vehicle 100 by remote control.

[0029] FIG. 6 is a flowchart showing the processing procedure of the driving control of the vehicle 100 in the first embodiment.

[0030] In step S1, the processor 201 of the server 200 acquires vehicle position information using the detection result output from the external sensor 300. The vehicle position information is the position information that serves as the basis for generating the driving control signal. In the present embodiment, the vehicle position information includes the position and orientation of the vehicle 100 in the global coordinate system GC of the factory FC. Specifically, in step S1, the processor 201 acquires the vehicle position information using the captured image obtained from the camera which is the external sensor 300.

[0031] In step S2, the processor 201 of the server 200 determines the target position to which the vehicle 100 should next head. In the present embodiment, the target position is represented by the coordinates of X, Y, and Z in the global coordinate system GC. In the memory 202 of the server 200, a reference route RR, which is the route along which the vehicle 100 should travel, is stored in advance. The route is represented by a node indicating the starting point, a node indicating a passing point, a node indicating the destination, and links connecting each node. The processor 201 determines the target position to which the vehicle 100 should next head using the vehicle position information and the reference route RR. The processor 201 determines the target position on the reference route RR ahead of the current position of the vehicle 100.

[0032] In step S3, the processor 201 of the server 200 generates a driving control signal for driving the vehicle 100 toward the determined target position. The processor 201 calculates the driving speed of the vehicle 100 from the change in the position of the vehicle 100 and compares the calculated driving speed with the target speed. Overall, when the driving speed is lower than the target speed, the processor 201 determines the acceleration so that the vehicle 100 accelerates, and when the driving speed is higher than the target speed, the processor 201 determines the acceleration so that the vehicle 100 decelerates. Also, when the vehicle 100 is located on the reference route RR, the processor 201 determines the steering angle and acceleration so that the vehicle 100 does not deviate from the reference route RR, and when the vehicle 100 is not located on the reference route RR, in other words, when the vehicle 100 has deviated from the reference route RR, the processor 201 determines the steering angle and acceleration so that the vehicle 100 returns to the reference route RR.

[0033] In step S4, the processor 201 of the server 200 transmits the generated driving control signal to the vehicle 100. The processor 201 repeats the acquisition of vehicle position information, determination of the target position, generation of the driving control signal, and transmission of the driving control signal at a predetermined cycle.

[0034] In step S5, the processor 111 of the vehicle 100 receives a driving control signal transmitted from the server 200. In step S6, the processor 111 of the vehicle 100 controls the actuator group 120 using the received driving control signal, thereby driving the vehicle 100 at the acceleration and steering angle represented by the driving control signal. The processor 111 repeats the reception of the driving control signal and the control of the actuator group 120 at a predetermined cycle. According to the driving system 50 in the present embodiment, the vehicle 100 can be driven by remote control, and the vehicle 100 can be moved without using conveying facilities such as a crane or a conveyor.

[0035] FIG. 7 is a flowchart showing a position estimation method in the first embodiment. The flow shown in FIG. 7 is repeatedly executed at predetermined time intervals, for example, during a period in which control by autonomous driving is being executed.

[0036] In the position estimation method, first, an acquisition step is executed. The acquisition step is a step of acquiring the captured images IM1 and IM2. In the acquisition step, when an instruction indicating that it is the timing for estimating the position of the vehicle 100 is received (step S101: Yes), the acquisition unit 211 of the server 200 executes step S102. In step S102, the acquisition unit 211 transmits an image request signal for acquiring the captured images IM1 and IM2 to an external camera 310 capable of capturing an area where the vehicle 100 is expected to be present for position estimation. The external camera 310 that has received the image request signal transmits, in step S103, camera identification information for identifying a plurality of external cameras 310, associated with the captured images IM1 and IM2, to the server 200. When an instruction indicating that it is the timing for estimating the position of the vehicle 100 has not been received (step S101: No), as shown in step S104, the server 200 waits.

[0037] Next to the acquisition process, a position estimation process is executed. The position estimation process is a process of estimating the position of the vehicle 100 included in the captured images IM1 and IM2 by image recognition. In the position estimation process, the position estimation unit 212 of the server 200 identifies the external camera 310 that acquired the captured images IM1 and IM2 using the camera identification information associated with the captured images IM1 and IM2 in step S105. In step S106, the position estimation unit 212 acquires the vertical height of the identified external camera 310 from the road surface RS. Thereby, the position estimation unit 212 acquires the distance H in the gravitational direction between the vehicle 100 and the external camera 310. When the distance H in the gravitational direction between the vehicle 100 and the external camera 310 is equal to or greater than the threshold value HS (step S107: Yes), in step S108, the position estimation unit 212 inputs the first captured image IM1 into the general-purpose detection model DM to acquire the bounding box BO. In step S109, the position estimation unit 212 estimates the position of the vehicle 100 using the bounding box BO. When the distance H in the gravitational direction between the vehicle 100 and the external camera 310 is less than the threshold value HS (step S107: No), in step S110, the position estimation unit 212 inputs the second captured image IM2 into the dedicated detection model DN to acquire a plurality of feature points FP1 to FP4 of the vehicle 100. In step S111, the position estimation unit 212 generates line segments LS1 to LS4 connecting two of the plurality of feature points FP1 to FP4 of the vehicle 100. In step S112, the position estimation unit 212 estimates the position of the vehicle 100 using the line segments LS1 to LS4. In step S113, the position estimation unit 212 outputs the coordinates representing the estimated position of the vehicle 100.

[0038] According to the above embodiment, when the distance H in the gravitational direction between the vehicle 100 and the external camera 310 is less than the threshold value HS, the position of the vehicle 100 can be estimated using the line segments LS1 to LS4 connecting two of the plurality of feature points FP1 to FP4 of the vehicle 100 extracted from the second captured image IM2. By doing so, the position of the vehicle 100 can be estimated without using the bounding box BO. Thereby, it is possible to suppress a decrease in the accuracy of the estimated position of the vehicle 100 due to an increase in the difference between the bounding box BO and the contour of the actual vehicle 100.

[0039] Further, according to the above embodiment, by inputting the second captured image IM2 to the dedicated detection model DN, which is a machine learning model that outputs a plurality of feature points FP1 to FP4 when the second captured image IM2 is input, the plurality of feature points FP1 to FP4 can be obtained. That is, when the distance H in the gravitational direction between the vehicle 100 and the external camera 310 is less than the threshold value HS, the position of the vehicle 100 can be estimated using a machine learning model specialized for estimating the position of the vehicle 100. Thereby, it is possible to suppress a decrease in the accuracy of the estimated position of the vehicle 100 when the distance H in the gravitational direction between the vehicle 100 and the external camera 310 is less than the threshold value HS.

[0040] Further, according to the above embodiment, the function of estimating the region including the vehicle 100 from the second captured image IM2 and the function of extracting the plurality of feature points FP1 to FP4 of the vehicle 100 from the estimated region can be realized by the dedicated detection model DN, which is a single machine learning model. Thereby, the position of the vehicle 100 can be estimated without using a computer with higher processing power. In addition, it is possible to reduce the possibility that the processing time required for estimating the position of the vehicle 100 increases or the annotation cost increases.

[0041] Further, according to the above embodiment, when the distance H in the gravitational direction between the vehicle 100 and the external camera 310 is equal to or greater than the threshold value HS, the position of the vehicle 100 can be estimated using the bounding box BO. That is, when the distance H in the gravitational direction between the vehicle 100 and the external camera 310 is equal to or greater than the threshold value HS, the position of the vehicle 100 can be estimated using the general detection model DM, which is a general machine learning model for detecting an object in the first captured image IM1. By doing so, it is possible to estimate the position of the vehicle 100 without preparing a plurality of training images including vehicles 100 of various vehicle types, without preparing a plurality of training images generated by imaging the vehicle 100 from all angles, and without labeling each training image. As a result, it is possible to reduce the load required for the preprocessing performed to estimate the position of the vehicle 100, such as preparing training images and labeling the training images, and to reduce the possibility of an increase in annotation costs.

[0042] B. Second Embodiment: FIG. 8 is a block diagram showing the configuration of the driving system 50v in the second embodiment. In this embodiment, the driving system 50v is different from the first embodiment in that it does not include the server 200. Further, the vehicle 100v in this embodiment can travel by autonomous control of the vehicle 100v. Other configurations are the same as those in the first embodiment unless otherwise specified.

[0043] In this embodiment, the functions of the position estimation device 70 are realized by the vehicle control device 110v. The processor 111v of the vehicle control device 110v functions as a vehicle control unit 115v, an acquisition unit 116, a position estimation unit 117, and a direction estimation unit 118 by executing a program PG1v stored in the memory 112v. The vehicle control unit 115v generates a driving control signal, outputs the generated driving control signal, and operates the actuator group 120 to drive the vehicle 100v by autonomous control. The functions of the acquisition unit 116, the position estimation unit 117, and the direction estimation unit 118 are the same as those of the acquisition unit 211, the position estimation unit 212, and the direction estimation unit 213 in the first embodiment.

[0044] FIG. 9 is a flowchart showing a processing procedure for travel control of the vehicle 100v in the second embodiment.

[0045] In step S901, the processor 111v of the vehicle control device 110v acquires vehicle position information using the detection result output from a camera which is an external sensor 300. In step S902, the processor 111v determines a target position to which the vehicle 100v should next head. In step S903, the processor 111v generates a travel control signal for causing the vehicle 100v to travel toward the determined target position. In step S904, the processor 111v controls the actuator group 120 using the generated travel control signal, thereby causing the vehicle 100v to travel according to the parameters represented by the travel control signal. The processor 111v repeats the acquisition of vehicle position information, determination of the target position, generation of the travel control signal, and control of the actuator at a predetermined cycle. According to the travel system 50v in the present embodiment, the vehicle 100v can be caused to travel by autonomous control of the vehicle 100v without remotely controlling the vehicle 100v by the server 200.

[0046] C. Other Embodiments: When the distance H in the gravitational direction between the vehicles 100, 100v and the external camera 310 is equal to or greater than the threshold value HS, the position estimation units 117, 212 may estimate the position of the vehicles 100, 100v using the contour of the vehicles 100, 100v determined by the outer shape of the vehicles 100, 100v extracted from the first captured image IM1. For example, the position estimation units 117, 212 detect the outer shape of the vehicles 100, 100v from the first captured image IM1, calculate the coordinates of the measurement points of the vehicles 100, 100v in the coordinate system of the captured image IM1, that is, the local coordinate system, and convert the calculated coordinates into coordinates in the global coordinate system GC, thereby obtaining the positions of the vehicles 100, 100v. The outer shape of the vehicles 100, 100v included in the first captured image IM1 can be detected, for example, by inputting the first captured image IM1 into a general-purpose detection model DM. Examples of the general-purpose detection model DM include a trained machine learning model trained to realize either semantic segmentation or instance segmentation. As this machine learning model, for example, a convolutional neural network (hereinafter, CNN) trained by supervised learning using a general-purpose learning dataset can be used. The general-purpose learning dataset has, for example, a plurality of training images including the vehicles 100, 100v and labels indicating whether each region in the training image is a region indicating the vehicles 100, 100v or a region indicating other than the vehicles 100, 100v. During the learning of the CNN, it is preferable that the parameters of the CNN are updated by backpropagation (error backpropagation method) so as to reduce the error between the output result by the general-purpose detection model DM and the label. In such a form, when the distance H in the gravitational direction between the vehicles 100, 100v and the external camera 310 is equal to or greater than the threshold value HS, the position of the vehicles 100, 100v can be estimated using the contour of the vehicles 100, 100v extracted from the first captured image IM1.

[0047] When the distance H in the gravitational direction between the vehicles 100 and 100v and the external camera 310 is less than the threshold value HS, the position estimation units 117 and 212 may estimate the positions of the vehicles 100 and 100v using one line segment LS4 connecting two feature points FP1 and FP2. In this case, for example, the left rear wheel WH1 and the left front wheel WH2 shown in FIG. 5 visible on the second captured image IM2 are extracted as two feature points FP1 and FP2 of the vehicles 100 and 100v. The position estimation units 117 and 212 estimate the positions of the vehicles 100 and 100v using, for example, the distance between two points, i.e., the first feature point FP1 corresponding to the left rear wheel WH1 and the second feature point FP2 corresponding to the left front wheel WH2. In such a form, the position of the vehicles 100 and 100v can be estimated using one line segment LS3 without using graphic data of a polygon formed by three or more line segments LS1 to LS4, such as the rectangular data BR.

[0048] (C3) The external camera 310 may image the moving body from below the moving body. Even in such a form, when the distance H in the gravitational direction between the vehicles 100 and 100v and the external camera 310 is less than the threshold value HS, the positions of the vehicles 100 and 100v can be estimated using the line segments LS1 to LS4 connecting two feature points. (C4) In the first embodiment described above, the processes from the acquisition of vehicle position information to the generation of a driving control signal are executed by the server 200. In contrast, at least a part of the processes from the acquisition of vehicle position information to the generation of a driving control signal may be executed by the vehicle 100. For example, the following forms (1) to (3) may be adopted.

[0049] (1) The server 200 may acquire vehicle position information, determine a target position to which the vehicle 100 should next head, and generate a route from the current position of the vehicle 100 represented by the acquired vehicle position information to the target position. The server 200 may generate a route to a target position between the current position and the destination, or may generate a route to the destination. The server 200 may transmit the generated route to the vehicle 100. The vehicle 100 may generate a driving control signal so that the vehicle 100 travels on the route received from the server 200, and control the actuator group 120 using the generated driving control signal.

[0050] (2) The server 200 may acquire vehicle position information and transmit the acquired vehicle position information to the vehicle 100. The vehicle 100 may determine a target position to which the vehicle 100 should next head, generate a route from the current position of the vehicle 100 represented by the received vehicle position information to the target position, generate a driving control signal so that the vehicle 100 travels on the generated route, and control the actuator group 120 using the generated driving control signal.

[0051] (3) In the forms (1) and (2) above, an internal sensor is mounted on the vehicle 100, and a detection result output from the internal sensor may be used for at least one of the generation of the route and the generation of the driving control signal. Specifically, the internal sensor may include, for example, a camera, LiDAR, millimeter wave radar, ultrasonic sensor, GPS sensor, acceleration sensor, gyro sensor, etc. For example, in the form (1) above, the server 200 may acquire the detection result of the internal sensor and reflect the detection result of the internal sensor in the route when generating the route. In the form (1) above, the vehicle 100 may acquire the detection result of the internal sensor and reflect the detection result of the internal sensor in the driving control signal when generating the driving control signal. In the form (2) above, the vehicle 100 may acquire the detection result of the internal sensor and reflect the detection result of the internal sensor in the route when generating the route. In the form (2) above, the vehicle 100 may acquire the detection result of the internal sensor and reflect the detection result of the internal sensor in the driving control signal when generating the driving control signal.

[0052] (C5) In the second embodiment described above, the vehicle 100v is equipped with an internal sensor, and the detection result output from the internal sensor may be used for at least one of route generation and generation of a driving control signal. For example, the vehicle 100v may acquire the detection result of the internal sensor and reflect the detection result of the internal sensor in the route when generating the route. The vehicle 100v may acquire the detection result of the internal sensor and reflect the detection result of the internal sensor in the driving control signal when generating the driving control signal.

[0053] (C6) In the first embodiment described above, the server 200 automatically generates the driving control signal to be transmitted to the vehicle 100. In contrast, the server 200 may generate the driving control signal to be transmitted to the vehicle 100 according to the operation of an external operator located outside the vehicle 100. For example, an external operator operates a control device including a display for displaying the captured images IM1 and IM2 output from the external sensor 300, a steering wheel for remotely operating the vehicle 100, an accelerator pedal, a brake pedal, and a communication device for communicating with the server 200 by wired or wireless communication, and the server 200 may generate a driving control signal corresponding to the operation applied to the control device.

[0054] In each of the above embodiments, the vehicles 100 and 100v only need to be configured to be movable by autonomous driving. For example, they may be in the form of a platform having the configuration described below. Specifically, the vehicles 100 and 100v only need to include at least a vehicle control device 110 or 110v and an actuator group 120 in order to perform three functions of "running", "turning", and "stopping" by autonomous driving. When the vehicles 100 and 100v acquire information from the outside for autonomous driving, the vehicles 100 and 100v may further include a communication device 130. That is, for the vehicles 100 and 100v that can be moved by autonomous driving, at least a part of the interior parts such as the driver's seat and the dashboard may not be installed, at least a part of the exterior parts such as the bumper and the fender may not be installed, and the body shell may not be installed. In this case, until the vehicles 100 and 100v are shipped from the factory FC, the remaining parts such as the body shell may be installed on the vehicles 100 and 100v, or after the vehicles 100 and 100v are shipped from the factory FC in a state where the remaining parts such as the body shell are not installed on the vehicles 100 and 100v, the remaining parts such as the body shell may be installed on the vehicles 100 and 100v. Each part may be installed from any direction such as the upper side, the lower side, the front side, the rear side, the right side, or the left side of the vehicles 100 and 100v, and they may be installed from the same direction or from different directions respectively. Note that the positioning can be performed in the same manner as the vehicles 100 and 100v in the first embodiment for the form of the platform.

[0055] (C8) Vehicles 100 and 100v may be manufactured by combining a plurality of modules. A module means a unit composed of one or more parts grouped according to the configuration and functions of the vehicle 100 or 100v. For example, the platform of the vehicle 100 or 100v may be manufactured by combining a front module that constitutes the front part of the platform, a center module that constitutes the central part of the platform, and a rear module that constitutes the rear part of the platform. Note that the number of modules constituting the platform is not limited to three, and may be two or less or four or more. Also, in addition to or instead of the platform, parts of the vehicle 100 or 100v that are different from the platform may be modularized. Further, each type of module may include any exterior parts such as bumpers and grills, and any interior parts such as seats and consoles. Also, not limited to the vehicle 100 or 100v, any type of moving body may be manufactured by combining a plurality of modules. Such modules may be manufactured, for example, by joining a plurality of parts by welding or fixtures, etc., or by integrally molding at least a part of the module by casting as one part. The molding method of integrally molding at least a part of the module as one part is also called gigacasting or megacasting. By using gigacasting, each part of a moving body that was conventionally formed by joining a plurality of parts can be formed as one part. For example, the above-mentioned front module, center module, and rear module may be manufactured using gigacasting.

[0056] (C9) Using the driving of the vehicles 100 and 100v by autonomous driving to transport the vehicles 100 and 100v is also called "self-driving transportation". Also, the configuration for realizing self-driving transportation is also called "vehicle remote control autonomous driving transportation system". Also, the production method of producing the vehicles 100 and 100v using self-driving transportation is also called "self-driving production". In self-driving production, for example, in a factory FC that manufactures the vehicles 100 and 100v, at least a part of the transportation of the vehicles 100 and 100v is realized by self-driving transportation.

[0057] The present disclosure is not limited to the above-described embodiments, and can be implemented in various configurations without departing from the gist thereof. For example, the technical features of the embodiments corresponding to the technical features in each form described in the summary of the invention can be appropriately replaced or combined in order to solve some or all of the above-described problems, or to achieve some or all of the above-described effects. Further, if the technical feature is not described as essential in this specification, it can be appropriately deleted.

Explanation of Reference Numerals

[0058] 7... Position estimation system, 50, 50v... Travel system, 70... Position estimation device, 80... Remote control device, 100, 100v... Vehicle, 110, 110v... Vehicle control device, 111, 111v... Processor, 112, 112v... Memory, 113... Input / output interface, 114... Internal bus, 115, 115v... Vehicle control unit, 116, 211... Acquisition unit, 117, 212... Position estimation unit, 118, 213... Direction estimation unit, 120... Actuator group, 130... Communication device, 200... Server, 201... Processor, 202... Memory, 203... Input / output interface, 204... Internal bus, 205... Communication device, 214... Remote control unit, 300... External sensor, 310 to 312... External camera, BO... Bounding box, BR... Rectangular data, DG1 to DG4... Diagonal line, DM... General-purpose detection model, DN... Dedicated detection model, FC... Factory, FP1 to FP4... Feature point, GC... Global coordinate system, H, H1, H2... Distance, HS... Threshold value, IM1... First captured image, IM2... Second captured image, IS1, IS2... Intersection point, LB1 to LB4, LS1 to LS4... Line segment, PG1, PG1v, PG2... Program, RR... Reference path, RS... Road surface, TR... Track, VT1 to VT4... Vertex, WH1 to WH4... Wheel

Claims

1. A position estimation method for estimating the position of a moving body that can move by autonomous driving, comprising: an acquisition step of acquiring a captured image output from a camera installed at a location different from the moving body; a position estimation step of estimating the position of the moving body included in the captured image by image recognition, wherein in the position estimation step, when the distance in the gravitational direction between the moving body and the camera is equal to or greater than a predetermined threshold, the position of the moving body is estimated using at least one of a bounding box generated by estimating a region including the moving body from the captured image and a contour of the moving body extracted from the captured image; when the distance is less than the threshold, the position of the moving body is estimated using a line segment connecting two of the plurality of feature points of the moving body extracted from the captured image.

2. The position estimation method according to claim 1, wherein the plurality of feature points are obtained by inputting the captured image into a trained machine learning model that outputs the plurality of feature points when the captured image is input, and the machine learning model is pre-trained using one or more learning datasets generated by labeling the plurality of feature points as correct labels in training images.

3. A position estimation device for estimating the position of a moving body that can move by autonomous driving, comprising: an acquisition unit that acquires a captured image output from a camera installed at a location different from the moving body; a position estimation unit that estimates the position of the moving body included in the captured image by image recognition, wherein the position estimation unit when the distance in the gravitational direction between the moving body and the camera is equal to or greater than a predetermined threshold, estimates the position of the moving body using at least one of a bounding box generated by estimating a region including the moving body from the captured image and a contour of the moving body extracted from the captured image; when the distance is less than the threshold, estimates the position of the moving body using a line segment connecting two of the plurality of feature points of the moving body extracted from the captured image.

4. The position estimation device according to claim 3, When the captured image is input, the position estimation unit obtains the plurality of feature points by inputting the captured image into a trained machine learning model that outputs the plurality of feature points, The machine learning model is a position estimation device that has been pre-trained using one or more training datasets generated by labeling the plurality of feature points as correct labels in training images.

5. A position estimation system for estimating the position of a moving body, A moving body capable of moving by autonomous driving, A camera installed at a location different from the moving body, An acquisition unit that acquires a captured image output from the camera, A position estimation unit that estimates the position of the moving body included in the captured image by image recognition, and The position estimation unit, When the distance in the gravitational direction between the moving body and the camera is equal to or greater than a predetermined threshold, the position of the moving body is estimated using at least one of a bounding box generated by estimating a region including the moving body from the captured image and the contour of the moving body extracted from the captured image, A position estimation system that estimates the position of the moving body using a line segment connecting two of the plurality of feature points of the moving body extracted from the captured image when the distance is less than the threshold.

Citation Information

Patent Citations

  • Method and system for automatic driving of vehicle

    JP2021149968A

  • Model production method, model production device, model production program, moving body posture estimation method, and moving body posture estimation device

    JP2023017341A

  • External environment recognition device and external environment recognition method

    WO2022269980A1

  • Position measurement system

    JP2022175768A