Location determination system and vehicle equipped with same
The system allows operators to specify vehicle destinations using image processing, enabling accurate navigation by setting destinations based on key points or target objects, addressing the limitation of existing technologies.
Patent Information
- Application Number
- JP2022047322
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-03-23
AI Technical Summary
Existing technologies do not allow an operator outside the vehicle to specify a destination for autonomous vehicles effectively.
A system that uses an imaging means to capture an image, identify a person, estimate three-dimensional positions of key points based on the image, and set a destination position, either as an intersection with the ground if within a predetermined range or as a target object near the line, allowing remote control of the vehicle.
Enables an operator outside the vehicle to specify a destination with simple operations, ensuring accurate navigation to both near and distant locations.
Smart Images

Figure 0007724178000001 
Figure 0007724178000002 
Figure 0007724178000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a location specifying system for remotely controlling a vehicle, for example, and a vehicle equipped with the same. [Background technology]
[0002] In recent years, autonomous vehicles that travel automatically while detecting obstacles and the like are being developed. In autonomous driving, a route is usually determined from a map to a selected destination, and the vehicle moves along the determined route. Some autonomous driving functions allow an operator outside the vehicle to remotely control the vehicle to move (see, for example, Patent Document 1). In the technology of Patent Document 1, when moving the vehicle toward a parking position, the operator instructs the terminal device whether to continue the movement or stop the vehicle.
[0003] The destination parking location is identified based on obstacles and white lines detected by external sensors equipped on the vehicle. Technology has also been proposed that uses an autonomous driving function to move a vehicle to a destination that has not been selected from a map. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2021-109530 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the technology of Patent Document 1 does not allow an operator outside the vehicle to specify a destination and move the vehicle toward that destination.
[0006] The present invention has been made in consideration of the above-described embodiment, and aims to provide an image processing method that allows an operator outside the vehicle to specify a destination with simple operations, and a vehicle control device and vehicle that use the same. [Means for solving the problem]
[0007] In order to achieve the above object, the present invention has the following configuration. An imaging means for capturing an image; and a specifying means for specifying a destination position, which is a three-dimensional position of a destination, based on the image, the three-dimensional position is a position based on the position and imaging direction of the imaging means in three-dimensional space, The identification means Identifying a person from the image, and if the person is identified, estimating the three-dimensional positions of two key points of the person; If an intersection of a line connecting the two key points with the ground is within a predetermined range from the person, the intersection is identified as a destination position; If the intersection of the line connecting the two key points with the ground is not within the predetermined range from the person, the position of a target object identified from the image that exists within a predetermined distance from the line is identified as the destination position. A location system is provided, comprising: [Effects of the Invention]
[0008] According to the present invention, an operator outside the vehicle can specify a destination with a simple operation. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing a configuration for controlling an autonomously driven vehicle. [Figure 2] 10A and 10B are diagrams illustrating examples of gestures used by an operator to set a destination. [Figure 3] 10A and 10B are diagrams illustrating an example of a gesture performed by an operator to set a long-distance destination. [Figure 4] FIG. 10 is a schematic diagram showing an example of a method for setting a long-distance destination. [Figure 5] 10 is a flowchart of a process for setting a destination. [Figure 6]10 is a flowchart of a process for setting a destination. [Figure 7] FIG. 10 is a schematic diagram showing an example of identifying three-dimensional coordinates from an image. DETAILED DESCRIPTION OF THE INVENTION
[0010] [First embodiment] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all combinations of features described in the embodiments are necessarily essential to the invention. Two or more of the features described in the embodiments may be combined in any desired manner. Furthermore, the same reference numerals are used to designate identical or similar components, and redundant descriptions will be omitted.
[0011] ●System configuration First, we will explain a vehicle control system that includes an autonomous vehicle and an operating terminal device (also called an operating terminal or remote operating terminal). As shown in Fig. 1, the vehicle control system 1 has a vehicle system 2 mounted on the vehicle and an operating terminal 3. The vehicle system 2 has a propulsion device 4, a braking device 5, a steering device 6, a transmission 61, an external sensor 7, a vehicle sensor 8, a communication device 9, a navigation device 10, a driving operation device 11, a driver detection sensor 12, an interface device (HMI device) 13, a smart key 14, and a control device 15. Each component of the vehicle system 2 is connected to each other via an in-vehicle communication network such as a Controller Area Network (CAN) so that signals can be transmitted.
[0012] The propulsion device 4 is a device that applies driving force to the vehicle and includes, for example, a power source. The transmission 61 is, for example, a stepless or stepped transmission, and changes the rotation speed of the driven shaft relative to the rotation speed of the drive shaft. The power source includes at least one of an internal combustion engine such as a gasoline engine or a diesel engine and an electric motor. The brake device 5 is a device that applies braking force to the vehicle and includes, for example, a brake caliper that presses a pad against a brake rotor and an electric cylinder that supplies hydraulic pressure to the brake caliper. The brake device 5 includes a parking brake device that restricts the rotation of the wheels using a wire cable. The steering device 6 is a device that changes the steering angle of the wheels and includes, for example, a rack-and-pinion mechanism that steers the wheels and an electric motor that drives the rack-and-pinion mechanism. The propulsion device 4, the brake device 5, and the steering device 6 are controlled by the control device 15.
[0013] The external sensor 7 is a sensor that detects objects around the vehicle, etc. The external sensor 7 includes a radar 16, a lidar (Light Detection and Ranging: LIDAR) 17, and a camera 18, and outputs the detection results to the control device 15.
[0014] The radar 16 is, for example, a millimeter wave radar, and is capable of detecting objects around the vehicle and measuring the distance to the objects using radio waves. A plurality of radars 16 are provided around the vehicle, for example, one radar 16 is provided in the front center of the vehicle, one at each front corner, and one at each rear corner.
[0015] The LIDAR 17 is capable of detecting objects around the vehicle 1 by light and measuring the distance to the objects. A plurality of LIDARs 17 are provided around the vehicle, for example, one LIDAR 17 is provided at each corner of the front of the vehicle, one LIDAR 17 is provided in the center of the rear, and one LIDAR 17 is provided on each side of the rear.
[0016] Camera 18 is a device that captures images of the surroundings of the vehicle, and is, for example, a digital camera that uses a solid-state image sensor such as a CCD or CMOS. Camera 18 includes a front camera that captures images in front of the vehicle and a rear camera that captures images behind the vehicle. Camera 18 is installed near the door mirror installation locations of the vehicle and includes a pair of door mirror cameras on the left and right sides that capture images of the rear of the left and right sides.
[0017] The vehicle sensor 8 includes a vehicle speed sensor that detects the speed of the vehicle, an acceleration sensor that detects the acceleration, a yaw rate sensor that detects the angular velocity around a vertical axis, a direction sensor that detects the direction of the vehicle, etc. The yaw rate sensor is, for example, a gyro sensor.
[0018] The communication device 9 mediates wireless communication between the control device 15 and the communication unit 35 of the operation terminal 3. That is, the control device 15 can communicate with the operation terminal 3 carried by the user via the communication device 9 using a communication method such as infrared communication or Bluetooth (registered trademark).
[0019] The navigation device 10 is a device that acquires the current position of the vehicle and provides route guidance to the destination, and includes a GPS receiver 20 and a map storage unit 21. The GPS receiver 20 identifies the vehicle's position (latitude and longitude) based on signals received from artificial satellites (positioning satellites). The map storage unit 21 is configured with a storage device such as a flash memory or a hard disk, and stores map information.
[0020] The driving operation device 11 is installed inside the vehicle cabin and accepts input operations performed by the user to control the vehicle. The driving operation device 11 includes, as driving operation units, for example, a steering wheel, an accelerator pedal, a brake pedal, a parking brake device, a shift lever, and a push start switch (engine start button). The push start switch accepts input operations to start the vehicle by driving operations from the user. The driving operation device 11 includes a sensor that detects the amount of operation and outputs a signal indicating the amount of operation to the control device 15.
[0021] The driver detection sensor 12 is a sensor for detecting whether or not a person is seated in the driver's seat. The driver detection sensor 12 is, for example, a seating sensor provided on the seating surface of the driver's seat. The seating sensor may be of a capacitance type, or may be a membrane switch that turns on when a person is seated in the driver's seat. Alternatively, the driver detection sensor 12 may be an interior camera that captures an image of a user seated in the driver's seat. The driver detection sensor 12 may also be a sensor that detects whether or not the tongue buckle of the driver's seat belt is inserted, and thereby detects that a person is seated in the driver's seat and has fastened the seat belt. The driver detection sensor 12 outputs the detection result to the control device 15.
[0022] The interface device 13 (HMI device) provides an interface (HMI: Human Machine Interface) between the control device 15 and the user, notifies the user of various information by display and sound, and accepts input operations by the user. The interface device 13 has a display unit 23 made of liquid crystal, organic EL, or the like, and functions as a touch panel that can accept input operations from the user, and an audio generation unit 24 such as a buzzer or speaker.
[0023] The control device 15 is an electronic control unit (ECU) including a CPU, a non-volatile memory (ROM), a volatile memory (RAM), etc. The control device 15 can perform various vehicle controls by executing arithmetic processing based on a program in the CPU. At least some of the functional units of the control device 15 may be realized by hardware such as an LSI, an ASIC, or an FPGA, or may be realized by a combination of software and hardware.
[0024] The smart key 14 (FOB) is a wireless terminal that can be carried by the user, and is configured to be able to communicate with the control device 15 from outside the vehicle via the communication device 9. The smart key 14 has buttons that the user can use to input data, and by operating the buttons on the smart key 14, the user can lock or unlock the doors, start the vehicle, and so on.
[0025] The operation terminal 3 is a wireless terminal that can be carried by the user, and can communicate with the control device 15 from outside the vehicle via the communication device 9. In this embodiment, the operation terminal 3 is, for example, a portable information processing device such as a smartphone. A predetermined application is installed in the operation terminal 3 in advance, thereby enabling the operation terminal 3 to communicate with the control device 15. Information that can identify the operation terminal 3 (for example, a terminal ID including a predetermined number, character string, or the like for identifying each operation terminal) is set in the operation terminal 3, and the control device 15 can authenticate the operation terminal 3 based on the terminal ID.
[0026] As shown in FIG. 1, the operation terminal 3 has, as functional components, an input / output unit 30, an imaging unit 31, a position detection unit 32, a processing unit 33, and a communication unit .
[0027] The input / output unit 30 presents information to a user operating the operation terminal 3 and accepts input from the user operating the operation terminal 3. The input / output unit 30 functions as, for example, a touch panel, and upon accepting input from the user, outputs a signal corresponding to the input to the processing unit 33. The input / output control unit 30 further includes an audio input / output device and a vibration generating device, both of which are not shown. The audio input / output device can output, for example, a digital signal as audio and convert audio input into a digital signal. The vibration generating device generates vibrations in addition to or instead of audio output, causing the housing of the operation terminal 3 to vibrate.
[0028] The imaging unit 31 is capable of capturing images (still images and moving images) in an imaging mode set by the input / output unit 30, and is, for example, a digital camera configured with a CMOS or the like. The processing unit 33 acquires image features by performing predetermined image processing on an image captured of a user operating the operation terminal 3, and can authenticate the user by comparing the acquired image features with facial image features of pre-registered users.
[0029] The position detection unit 32 includes a sensor capable of acquiring position information of the operation terminal 3. The position detection unit 32 can acquire the position of the operation terminal 3 by receiving a signal from, for example, a geodetic satellite (GPS satellite). The position detection unit 32 can also acquire position information including the relative position of the operation terminal 3 with respect to the vehicle by communicating with the control device 15 via the communication device 9. The position detection unit 32 outputs the acquired position information to the processing unit 33.
[0030] The processing unit 33 transmits to the control device 15 the terminal ID set in the operation terminal 3, a signal from the input / output unit 30, and location information acquired by the location detection unit 32. Furthermore, upon receiving a signal from the control device 15, the processing unit 33 processes the signal and causes the input / output unit 30 to present information to the user operating the operation terminal 3. The information is presented, for example, by display on the input / output unit 30. The communication unit 35 communicates with the communication device 9 wirelessly or via a wire. In this example, the description will be given assuming wireless communication.
[0031] The control device 15 can drive the vehicle based on a signal from the operation terminal 3. The control device 15 can also move the vehicle to a predetermined position to perform remote parking. To control the vehicle, the control device 15 has at least a starting unit 40, an external environment recognition unit 41, a position identification unit 42, a trajectory planning unit 43, a driving control unit 44, and a memory unit 45.
[0032] The activation unit 40 authenticates the smart key 14 based on a signal from the push start switch and determines whether the smart key 14 is inside the vehicle. When the smart key 14 is authenticated and the smart key 14 is inside the vehicle, the activation unit 40 starts driving the propulsion device 4. Furthermore, when the activation unit 40 receives a signal instructing activation from the operation terminal 3, it authenticates the operation terminal 3 and starts driving the vehicle when authenticated. When starting to drive the vehicle, the activation unit 40 turns on the ignition if the propulsion device 4 includes an internal combustion engine.
[0033] The external environment recognition unit 41 recognizes obstacles such as parked vehicles and walls that exist around the vehicle based on the detection results of the external environment sensor 7, and acquires information on the position, size, etc. of the obstacles. The external environment recognition unit 41 can also analyze images acquired by the camera 18 based on an image analysis method such as pattern matching, and acquire the presence or absence of an obstacle and its size. Furthermore, the external environment recognition unit 41 can calculate the distance to the obstacle using signals from the radar 16 and the lidar 17, and acquire the position of the obstacle.
[0034] The position identifying unit 42 is capable of detecting the position of the vehicle based on a signal from the GPS receiving unit 20 of the navigation device 10. In addition to the signal from the GPS receiving unit 20, the position identifying unit 42 is also capable of acquiring the vehicle speed and yaw rate from the vehicle sensor 8 and identifying the position and attitude of the vehicle using so-called inertial navigation.
[0035] The external environment recognition unit 41 analyzes the detection results of the external environment sensor 7, more specifically, the images captured by the camera 18, based on an image analysis method such as pattern matching, and can obtain, for example, the position of white lines painted on the road surface of a parking lot, etc.
[0036] The traveling control unit 44 controls the propulsion device 4, the braking device 5, and the steering device 6 based on a traveling control instruction from the trajectory planning unit 43, and causes the vehicle to travel.
[0037] The storage unit 37 is configured with a RAM or the like, and stores information required for processing by the trajectory planning unit 43 and the traveling control unit 44.
[0038] When the user inputs to the HMI device 13 or the operation terminal 3, the trajectory planning unit 43 calculates a trajectory that will be the vehicle's driving route, as necessary, and outputs driving control instructions to the driving control unit 44.
[0039] After the vehicle is stopped, the trajectory planning unit 43 performs parking assist processing when there is an input from the user indicating a desire for parking assistance by remote operation (remote parking assist).
[0040] ●Setting a destination In the present embodiment, the position identification unit 42 further has a function of setting a destination designated by the operator based on an image captured by the camera 18. The camera capturing the forward view is a monocular camera that is fixed to the vehicle body and has a fixed focal length. The position identification unit 42 can identify (or estimate) a location designated by the operator based on the operator's gestures included in the image captured by the camera 18, and set the location as the destination. The control device 15 then controls the driving, braking, and steering to drive the vehicle toward the set destination. The operator may set the destination and trigger the vehicle to drive via a predetermined application on the mobile terminal 3, for example. Because the system identifies the destination, it is sometimes called a position identification system.
[0041] In the following description, camera 18 is a monocular camera. Although camera 18 is assumed to face forward (i.e., in the forward direction), this is for convenience and camera 18 may face in either direction. For example, once the location of the destination (referred to as the target location) is identified based on an image and using the camera position and the direction of the optical axis as references, it can be converted into a predetermined coordinate system by projective transformation or the like.
[0042] FIG. 2 shows an example of an image 200 captured of an operator indicating a destination. In FIG. 2(A), the operator 210 is indicating a nearby destination. The ground contact area 211 is located at the feet of the operator 210, and the eye 212 is located almost directly above it. The operator 210 extends his arm to indicate the destination, and the destination is indicated by two key points. In this example, the destination is an intersection point 215 where an indication line 214 extending from the eye 212 to the wrist 213 and toward the wrist intersects with the ground surface. While any key point may be selected, it is desirable to select one of the points as the eye, since the operator can accurately specify the position through the line of sight. The other point is desirably a part that is easy to identify from the image, such as the wrist, the fingertip, the tip or center of a clenched fist, or the like. Furthermore, when the operator indicates the destination, the face may be facing the destination, preventing the camera 18 from capturing the eyes. In such cases, the eye position may be estimated. If the face direction can be identified, the eye position can be estimated. The estimation of the eye positions may also be performed using a machine learning model.
[0043] In Figure 2(B), operator 210 is pointing to a distant destination. The arm is raised higher than in Figure 2(A), and indicator line 224 points farther away from operator 210. In this case, even a slight movement of the arm results in a large movement of intersection point 215, reducing the accuracy of the indicated destination. Furthermore, depending on the height of the wrist position, indicator line 224 may not intersect with the ground surface, making it impossible to identify intersection point 215.
[0044] Therefore, if the destination is far away, as shown in FIG. 3, facilities or targets such as vending machines, mailboxes, or buildings near indicator line 314 are identified from the image. Then, from among the targets identified from the image, the target closest to indicator line 314 is identified and set as the destination. Whether the destination is far away may be determined, for example, if the distance from the operator's ground contact part 211 to intersection 215 exceeds a predetermined threshold, or if intersection 215 cannot be identified. The predetermined threshold may be set to a specific value, for example, about 20 to 30 meters, but this is of course just one example.
[0045] FIG. 4 shows an example of identifying a target object as a destination when setting a destination at a long distance. In this case, the indicator line 224 is not identified in three dimensions, but may be treated as a line 400 projected onto the ground surface. The target object closest to the projected indicator line 400 is identified from among the targets 411, 412, and 413. In this case, the distance may be the distance measured from the target position to the indicator line 400 in a direction perpendicular to the indicator line 400 (i.e., the closest distance to the indicator line 400). The target position may be its ground contact position, or, if there is an extension, the center of the extension. Alternatively, if there is an extension, the distance from the indicator line 400 to the closest end point of the extension may be the distance between the indicator line 400 and the target.
[0046] In the example of Fig. 4, the target 413 is identified as the target closest to the indicator line 400, and the position of the target 413 is the destination. If the destination overlaps with the position of the target, the vehicle may be driven to avoid the target using automatic driving control.
[0047] ●Destination setting process 5 and 6 show the destination setting process steps performed by the control unit 15, particularly the location identification unit 42. As described above, the functions of the control unit 15 are realized by the CPU executing a program stored in memory, and therefore the steps in FIGS. 5 and 6 may also be executed by the CPU (or processor).
[0048] 5 is initiated when the operator issues an instruction to set a destination to the vehicle's control unit 15 via the communication device 9 from the mobile terminal 3, for example. The operator may issue the instruction from an operation panel or the like provided on the vehicle body. Note that the vehicle is turned on, and power is being supplied to the control unit 15.
[0049] First, the operator is recognized from an image captured by the camera 18 (S501). In addition to the person who is the operator, landmarks such as vending machines, mailboxes, utility poles, and buildings within the captured range may also be recognized. This recognition of the operator may be performed by determining the similarity between the features of the captured object and the features corresponding to a person, or may be performed using a trained machine learning model.
[0050] If an attempt is made to recognize the operator, it is determined whether the recognition was successful (S503). If a person is recognized, the recognition may be determined to be successful. If a person cannot be recognized, the recognition may be determined to be unsuccessful even if other objects are recognized. Furthermore, at this time, the person's face may be recognized and a match may be determined with a specific person stored in advance who has been granted authority to operate the vehicle, and if there is no match, the recognition may be determined to be unsuccessful. If the recognition fails, step S501 is repeated using a new image taken after the target image.
[0051] If the recognition of the operator is successful, it is determined whether there is a saved image that was taken a predetermined time before the current processing target and designated image (S505). Since the camera 18 takes video and the images to be processed are frames that make up that video, the saved image may be a predetermined number of frames before. If it is determined that there is a saved image, it is determined whether the operator's posture is stable (S507). The destination cannot be correctly identified if the operator is in the middle of indicating the destination. Therefore, if it is determined that the posture is not stable, the target image is saved and the process waits for a predetermined time (S521), and the process is repeated from step S501 using a newly acquired image. If it is determined in step S505 that there is no saved image, there is no material to determine posture stability, so the process branches to step S521.
[0052] In step S507, the image of the operator contained in the saved image is compared with the image of the operator contained in the current image to be processed to determine whether the posture is stable. For example, the amount of deviation between the people contained in the two images may be identified, and if the amount of deviation does not exceed a predetermined threshold, the posture may be determined to be stable. For example, if the ratio of the area of the person contained in the image to be processed to the area of the person obtained by combining the people contained in the two images is within a predetermined value, the amount of deviation may be within the predetermined threshold, and the posture may be determined to be stable.
[0053] If it is determined that the operator's posture is stable, the two-dimensional positions of the operator and the key points are identified (S509). The operator's position may be the operator's ground contact point, i.e., the position of the feet. The two-dimensional position is the position on the image.
[0054] Next, the two-dimensional positions of the identified key points are converted into three-dimensional positions (S511). The three-dimensional positions are positions in a three-dimensional space where the operator and key points exist, expressed in a predetermined coordinate system. This conversion process will be described with reference to FIG. 7, but any commonly used method may be used. Once the positions of the two key points in three-dimensional space have been identified, an instruction line passing through the positions of the key points is identified (S513). Furthermore, the intersection of the instruction line and the ground surface is identified (S515).
[0055] Next, it is determined whether the intersection identified in step S513 is within a predetermined range from the vehicle (S517). It may also be determined whether the intersection is within a predetermined range from the operator, rather than from the vehicle. If the intersection cannot be identified, it may be determined that the intersection is not within the predetermined range. If it is determined that the intersection is within the predetermined range, the position of the identified intersection is set as the destination (S519). On the other hand, if it is determined that the intersection is not within the predetermined range, the process branches to the long-distance destination setting procedure shown in FIG. 6(A).
[0056] ● Procedure for setting a long-distance destination In FIG. 6(A), the position identification unit 42 first recognizes targets other than people from the target image (S601). This recognition may also be performed using pattern matching or a machine learning model. Next, it is determined whether the recognition was successful (S603). If at least one target is recognized, it may be determined that the recognition was successful. If the recognition failed, it is determined that the setting of the destination by gesture has failed, and this is notified to the operator (S611). The notification may be, for example, a message sent to the mobile terminal 3, or may be performed by flashing the vehicle's lights or emitting a warning sound.
[0057] If it is determined that the recognition is successful, the position of the recognized target is identified (S605). The position can be identified by converting the two-dimensional position identified on the screen into a three-dimensional position as in steps S509-S511 of FIG. 5, but here, these are performed together. Although the three-dimensional position is used in this embodiment, since it is assumed that all targets are on the ground surface, the height value may be a constant corresponding to the height above the ground surface. Note that in step S605, only targets within a predetermined distance from the vehicle may be identified. This can prevent a target that is far away and unlikely to be a target from being mistakenly set as the destination.
[0058] Once the three-dimensional positions of the targets have been identified, the target positions and the indicator lines identified in step S513 are projected onto the ground surface, and targets within a predetermined distance from the projected indicator lines, for example, the closest target, are identified (S607). If there are multiple targets closest to the projected indicator lines, one of them, for example, the target closest to the operator (or vehicle), may be selected.
[0059] Finally, the position of the identified target object is set as the destination (S609). The above procedure makes it easy to specify not only destinations close to the operator but also distant destinations by gesture. Furthermore, if the operator is aware that distant destinations can be specified in this manner, even distant destinations can be specified with high accuracy by specifying a distant target object.
[0060] It is not desirable for the vehicle to wait for the destination to be set even if the operator leaves the destination setting operation suspended. Therefore, a time limit may be set, and if the destination has not been set even after the time limit has expired, a notification of failure may be sent as shown in step S621 of FIG. 6(B). This notification may be the same as in step S611. The time limit may start, for example, when the operator notifies the operation unit 15 that a destination will be set, and for example, a timer with a time limit set may be started at the start of FIG.
[0061] In order to move the vehicle to the destination set as described above, automatic driving control toward the set destination is performed as shown in step S631 of Fig. 6(C). Fig. 6(C) may be executed immediately after the destination setting is completed in Fig. 5 or Fig. 6(A), or may be started in response to a signal from the operator.
[0062] ● Identifying 2D positions and converting them to 3D positions A specific example of the processing in steps S509 and S511 is shown in Figure 7. The camera 18 is fixed to the vehicle at a height H. For simplicity of explanation, it is assumed that the camera is attached so that its optical axis is parallel to the ground surface. Figure 7(A) shows an orthogonal coordinate system with the camera position as the origin, the height direction as the Y axis, the direction of the optical axis A as the Z axis, and the direction perpendicular to these as the X axis. The X axis direction is sometimes called width, the Y axis direction as height, and the Z axis direction as depth. Next, consider a virtual frame Fv. The virtual frame Fv is an image enlarged so that the camera's optical axis intersects the virtual frame Fv at its center O' and the bottom edge of the virtual frame Fv is located at a position on the ground surface that corresponds to the bottom edge of the camera's angle of view in the height direction. Let Lb be the Z-direction distance from the origin O to the virtual frame Fv. This distance Lb is determined by the optical axis direction and the angle of view. However, it is also possible to identify the position where the bottom edge of the image is located and actually measure the distance from that position to the camera position. In other words, the distance Lb is a known value. Also, in Figure 7(A), the optical axis is parallel to the ground surface, so the height of the virtual frame Fv is 2H. The length in the virtual frame Fv is proportional to the length in the captured image. The proportionality constant of the virtual frame Fv relative to the actual image frame is Cf. The proportionality constant, i.e., the magnification factor Cf, may be a constant that indicates the distance of the virtual frame Fv corresponding to one pixel in the actual image frame. In this case, if the pixel densities in the vertical and horizontal directions of the image frame are different, the magnification factor may be set separately for the vertical and horizontal directions.
[0063] In Figure 7(A), the operator is standing at the ground contact point Pf and giving instructions for the destination. In other words, the ground contact point Pf is the position of the operator's feet. The wrist is at wrist position Pw. Here, we will use the identification of the wrist position as an example, but the same applies to the eye position, and it is also the same when other parts of the body are used as key points. In this case, consider the vector Vpf from the origin O, which is the camera position, to the ground contact point Pf. Since the ground contact point Pf is on the ground surface, its height yf is -H, and its coordinates can be expressed as (xf, -H, zf). This value is also the vector Vpf itself.
[0064] The ground contact point Pf is projected onto point Pf' on the virtual frame Fv. The height position of a point on the image can be associated with a position in the Z-axis direction in the actual three-dimensional space, assuming that the point is on the ground surface (ground) in the actual three-dimensional space. In other words, the image height in the image can be converted into a position in the depth direction. The position of point Pf' on the virtual frame Fv is expressed by coordinates (xf', yf') with the origin O'. xf', yf' can be determined from the position in the actual image and the proportionality coefficient Cf. Assuming that the optical axis A is parallel to the ground surface, Lb:yf'=zf:H. Therefore, zf=Lb·H / yf'. For xf, Lb:xf'=zf:xf. Therefore, xf=xf'·zf / Lb=H·xf' / yf'.
[0065] In this way, the position of the contact point Pf can be determined. Next, the wrist position Pw is identified. The coordinates of the position Pw are (xw, yw, zw). The position of the point Pw' when the position Pw is projected onto the virtual frame Fv is (xw', yw'). Considering the vector Vpw' from the point Pf' to the point Pw', this vector Vpw' is obtained by projecting the vector Vpw from the point Pf to the point Pw onto the virtual frame Fv. Because the depth component cannot be identified from the vector Vpw' that can be observed on the virtual frame Fv, it is not possible to directly identify the vector Vpw. However, it is possible to identify the vector Vpwp by projecting the vector Vpw onto a plane that is parallel to the virtual frame fV and includes the point Pf.
[0066] To do this, we can use the same method as for determining the x-component of point Pf. That is, The x and y components of the end point (xwp, ywp, zf) of the vector Vpwp are respectively xwp=xw'·zf / Lb=H·xw' / yf', ywp=yw'·zf / Lb=H·yw' / yf' This becomes:
[0067] If the position of the operator's eyes in three-dimensional space is the same as the ground contact point Pf, i.e., the standing position, in the Z-axis direction, it can be identified from the eye position in the image in the same manner as (xwp, ywp, zf). However, the position of the wrist must be considered due to the shift in the depth direction. In Figure 7(A), this shift in the depth direction is shown by the vector Vd. The vector Vd is a vector along the line that projects the wrist position Pw onto the virtual frame Fv, and does not appear in the virtual frame Fv. Therefore, as shown in FIG. 7B, the vector Vd is estimated using the estimated arm length La. In addition to being able to identify the operator's ground contact point Pf in the above-described manner, if the position of the operator's head is identified in the image, its position in three-dimensional space can be identified in the same manner as point PwP. If two points, the ground contact point Pf and the head contact point, can be identified, the apparent height in the virtual frame Fv can be determined. Based on the assumption that the optical axis A is parallel to the ground surface, the actual height of the operator, who is located at a distance zf (= Lb · H / yf') from the origin O, can be estimated by multiplying the apparent height by zf / Lb. Furthermore, the arm length La can be estimated by previously storing the ratio of the arm length (for example, from the base to the wrist) to the height in the control unit 15. The vector Vd can be estimated using this value La. The method for this is as follows.
[0068] As shown in Figure 7(A), the point Pwp and the vector Vpwp to it are identified. In the same way, the position Ps of the base of the arm (see Figure 7(B)) can be identified, and the vector Vps to it can be identified. Vpwp-Vps+Vd is the vector from the shoulder position Ps to the wrist position Pw, |Vpwp-Vps+Vd|=La Here, Vpwp and Vd are both vectors in the line of sight, Vd=k Vpwp (k is a scalar constant) There is. Therefore, |(k+1)·Vpwp-Vps|=La In the above equation, all the variables except for the constant k are known, so the constant k can be determined, and therefore the vector Vd can be determined. In this way, the 3D position of the wrist in the image can be estimated by shifting the wrist position in the image to a position corresponding to the arm length based on the estimated height of the operator. However, the procedure for determining the constant k includes square rooting, so two values can be obtained instead of just one.
[0069] Therefore, in this embodiment, the value to be used is determined based on the recognition result of the operator's eyes, i.e., the direction of the face. For example, if the eyes cannot be recognized from the face image of the operator, i.e., if the operator's face is not facing the camera, the larger value is adopted as the constant k. Conversely, if the eyes can be recognized from the face image, i.e., if the operator's face is facing the camera, the smaller value is adopted as the constant k. In this way, an appropriate location can be set as the destination in accordance with the operator's action of specifying the destination. Note that a trained machine learning model may also be used for eye recognition.
[0070] By identifying vector Vd in the above manner, vector Vpw, i.e., wrist position Pw, can be identified by Vpw = Vpwp + Vd. Since the eye positions have already been identified, the instruction line along which the operator indicated the destination can be identified based on the positions of the two determined key points.
[0071] Another example of how to determine the vector Vd The vector Vd can also be determined in a simpler way. Since the magnitude of the vector Vd is considered to be small, it can be approximately determined from the actual length La and the apparent length La' of the arm by assuming that it is perpendicular to the virtual frame Fv. In this case, the magnitude |Vd| of the vector Vd and La and La' are expressed as La 2 =La' 2 +|Vd| 2 That is, |Vd|=√(La 2 -La' 2) If both the x and y components are set to 0, the z-component zvd of Vd may be zvd = |Vd| or -|Vd|. From the vector Vd determined in this way, the vector Vpw, i.e., the wrist position Pw, can also be identified by Vpw = Vpwp + Vd. Even with this method, Vd cannot be uniquely determined because the sign of zvd may be either positive or negative. Therefore, as described above, for example, if the operator's eyes can be recognized, the sign of the z value may be negative, and if the eyes cannot be recognized, the sign of the z value may be positive.
[0072] Variations In Figure 7(A), an example has been explained in which the coordinates of the optical system and the coordinates on the ground are in a parallel translation relationship. However, if the camera has a depression or elevation angle, it is necessary to perform further coordinate transformation using a projective transformation or the like, taking into account the tilt of the optical axis. However, even in this case, there is no essential difference from the explanation in FIG. 7(A).
[0073] While the stability of the posture is determined in steps S507 and S521 of FIG. 5, the destination may be identified from an image captured when the posture is stable. To do this, for example, the operator sends a signal to the control unit 15 while in a posture indicating the destination, which triggers the position identification unit 42 to acquire the target image. In this way, since the acquired image shows the operator indicating the destination, there is no need to wait for the posture to stabilize. The signal may be, for example, touching a predetermined button displayed by an application running on the mobile terminal or emitting a specific voice. In the latter case, when the specific voice is recognized by the mobile terminal, a signal is sent to the control unit 15 indicating that the destination may be identified.
[0074] As described above, according to this embodiment and the modified example, an operator outside the vehicle can specify a destination with a simple operation. Then, the vehicle can be driven automatically to the specified destination. In particular, when a distant destination is specified, the accuracy of the specification can be improved by setting a target object near the specified destination as the destination. Furthermore, by estimating the depth from the image, the three-dimensional position of the destination can be identified from the image.
[0075] The invention is not limited to the above-described embodiments, and various modifications and variations are possible within the scope of the invention. For example, when an operator specifies a destination, information about the destination target (e.g., the type and color of the target) may be provided by voice or text input along with the instruction. In this case, the target specified by the operator is estimated using the target information provided by the operator, thereby improving the accuracy of target identification. In this case, a microphone for detecting sound may be further provided as the external sensor 7, and information about the destination target may be provided based on the sound signal input from the microphone. Alternatively, information may be provided via a touch panel or the like of the HMI device 13.
[0076] The information about the target may be, for example, information indicating its position, direction, type, color, size, etc., or a combination thereof. For example, in the above embodiment, the depth direction of the arm is estimated based on the direction of the instructor's face, but it may be determined based on the recognized information by recognizing words such as "forward" or "backward." Alternatively, in the above embodiment, the target closest to the instruction line 314 is identified as the destination position from among the targets identified from the image, but it may be determined based on the recognized information by recognizing words such as "that red sign" or "blue vending machine." Of course, this is merely an example, and information about the target may be provided in other ways.
[0077] Furthermore, although the above embodiment has been described using a vehicle as an example, the present invention is not limited to vehicles and can be applied to other autonomously moving bodies. The moving body is not limited to a vehicle, but may include a small mobility that runs alongside a walking user to carry luggage or lead a person, and may also include other autonomously moving bodies (for example, walking robots).
[0078] Summary of embodiments The present embodiment described above can be summarized as follows.
[0079] (1) According to a first aspect of the present invention, there is provided a camera comprising: a photographing means for photographing an image; and a specifying means for specifying a destination position, which is a three-dimensional position of a destination, based on the image, the three-dimensional position is a position based on the position and imaging direction of the imaging means in three-dimensional space, The identification means Identifying a person from the image, and if the person is identified, estimating the three-dimensional positions of two key points of the person; If an intersection of a line connecting the two key points with the ground is within a predetermined range from the person, the intersection is identified as a destination position; If the intersection of the line connecting the two key points with the ground is not within the predetermined range from the person, the position of a target object identified from the image that exists within a predetermined distance from the line is identified as the destination position. A location system is provided, comprising: This configuration allows distant destinations to be set with high accuracy.
[0080] (2) According to the second aspect of the present invention, When the intersection of the line connecting the two key points and the ground is not within the predetermined range from the person, the specifying means specifies, as a destination position, the position of the target object closest to the line among the targets specified from the image. A location system is provided, comprising: This configuration allows distant destinations to be set with high accuracy.
[0081] (3) According to the third aspect of the present invention, the identification means further estimates the three-dimensional positions of the person's eyes and wrist as the two key points. A location system is provided, comprising: This configuration allows you to set your destination using gestures with your eyes and wrist.
[0082] (4) According to a fourth aspect of the present invention, the identification means, when a person is identified from the image, estimates the three-dimensional position of the person's feet based on the position of the feet in the image, and estimates the three-dimensional positions of the eyes and wrists based on the three-dimensional position of the feet. A location system is provided, comprising: With this configuration, the depth position of the person can be identified from the image.
[0083] (5) According to a fifth aspect of the present invention, the specifying means further comprises: a distance from the imaging means to the feet is estimated based on an image height of the feet in the image, and the estimated distance is used as a distance from the imaging means to the eyes to estimate a three-dimensional position of the eyes; Estimating the three-dimensional position of the wrist based on the position of the wrist in the image, the arm length based on the person's estimated height, and the apparent arm length in the image. A location system is provided, comprising: This configuration makes it possible to quickly and easily identify the depth position of the wrist from the person in the image.
[0084] (6) According to a sixth aspect of the present invention, the specifying means further comprises: a distance from the imaging means to the feet is estimated based on an image height of the feet in the image, and the estimated distance is used as a distance from the imaging means to the eyes to estimate a three-dimensional position of the eyes; The three-dimensional position of the wrist is estimated by shifting the position of the wrist in the image along the imaging direction to a position corresponding to the arm length based on the estimated height of the person. A location system is provided, comprising: With this configuration, the depth position of the wrist can be accurately identified from the person in the image. (7) According to a seventh aspect of the present invention, the specifying means estimates the position of the wrist in the photographing direction according to the direction of the person's face. A location system is provided, comprising: With this configuration, the direction indicated by the operator can be specified in accordance with the operator's intention. (8) According to an eighth aspect of the present invention, the device further comprises means for receiving an input, The specifying means estimates the position of the wrist in the imaging direction in response to an input by the input means. A location system is provided, comprising: With this configuration, the direction indicated by the operator can be specified in accordance with the operator's intention. (9) According to a ninth aspect of the present invention, there is further provided a means for receiving an input, The specifying means further specifies a predetermined target among the targets specified from the image based on the input as a destination position. A location system is provided, comprising: This allows the target position to be specified in accordance with the operator's intention. (10) According to a tenth aspect of the present invention, the target is located within a predetermined distance from the photographing means. A location system is provided, comprising: This configuration can prevent the setting of an inaccurate destination that is far away. (11) According to an eleventh aspect of the present invention, there is provided a mobile object equipped with any one of the above-described positioning systems. This configuration allows the destination of the vehicle to be set by gesture. (12) According to a twelfth aspect of the present invention, there is further provided a mobile body that is characterized in that a destination is set by the positioning system and that travels to the destination by automatic driving. With this configuration, the vehicle's destination can be set by gesture and the vehicle can travel there automatically.
[0085] The invention is not limited to the above-described embodiment, and various modifications and variations are possible within the scope of the gist of the invention. [Explanation of symbols]
[0086] 1: Vehicle control system, 2: Vehicle system, 3: Operation terminal, 7: External sensor, 13: Interface device (HMI device) 13, 15: Control device
Claims
1. An imaging means for capturing an image; and a specifying means for specifying a destination position, which is a three-dimensional position of a destination, based on the image, the three-dimensional position is a position based on the position and imaging direction of the imaging means in three-dimensional space, The identification means Identifying a person from the image, and if the person is identified, estimating the three-dimensional positions of two key points of the person; If an intersection of a line connecting the two key points with the ground is within a predetermined range from the person, the intersection is identified as a destination position; If the intersection of the line connecting the two key points with the ground is not within the predetermined range from the person, the position of a target object identified from the image that exists within a predetermined distance from the line is identified as the destination position. A location identification system characterized by:
2. 10. The location system of claim 1, When the intersection of the line connecting the two key points and the ground is not within the predetermined range from the person, the specifying means specifies the position of the target object closest to the line among the targets specified from the image as the destination position. A location identification system characterized by:
3. 3. The location system according to claim 1 or 2, The identification means estimates the three-dimensional positions of the person's eyes and wrist as the two key points. A location identification system characterized by:
4. 4. The location system of claim 3, When a person is identified from the image, the identification means estimates the three-dimensional position of the person's feet based on the position of the feet in the image, and estimates the three-dimensional positions of the eyes and wrists based on the three-dimensional position of the feet. A location identification system characterized by:
5. 5. The location system of claim 4, The identification means a distance from the imaging means to the feet is estimated based on an image height of the feet in the image, and the estimated distance is used as a distance from the imaging means to the eyes to estimate a three-dimensional position of the eyes; Estimating the three-dimensional position of the wrist based on the position of the wrist in the image, the arm length based on the person's estimated height, and the apparent arm length in the image. A location identification system characterized by:
6. 5. The location system of claim 4, The identification means a distance from the imaging means to the feet is estimated based on an image height of the feet in the image, and the estimated distance is used as a distance from the imaging means to the eyes to estimate a three-dimensional position of the eyes; The three-dimensional position of the wrist is estimated by shifting the position of the wrist in the image along the imaging direction to a position corresponding to the arm length based on the estimated height of the person. A location identification system characterized by:
7. 7. The location system according to claim 5 or 6, The identification means estimates the position of the wrist in the shooting direction according to the direction of the person's face. A location identification system characterized by:
8. 7. The location system according to claim 5 or 6, further comprising an input means for accepting an input, The specifying means estimates the position of the wrist in the imaging direction in response to an input by the input means. A location identification system characterized by:
9. 9. A location system according to any one of claims 1 to 8, comprising: further comprising means for accepting input; The specifying means further specifies a predetermined target among the targets specified from the image based on the input as a destination position. A location identification system characterized by:
10. 10. A location system according to any one of claims 1 to 9, comprising: The target is within a predetermined distance from the imaging means. A location identification system characterized by:
11. A mobile object equipped with the position specifying system according to any one of claims 1 to 10.
12. 11. The mobile body according to claim 10, wherein a destination is set by the position specifying system and the mobile body travels to the destination by automatic driving.
Citation Information
Patent Citations
Position teaching device and position teaching method to mover
JP2006051550A
Vehicle dispatch
JP2019530937A
Vehicle control device, vehicle control method, and program
JP2020163906A
Pointing input device
JP2020190941A
Vehicle control device
JP2021109530A