Estimation device, estimation method, estimation program, and recording medium

The estimation device uses an extended two-dimensional bounding box to accurately determine the orientation of motorcycles in captured images, addressing the challenge of erroneous yaw angle estimation in conventional techniques.

JP2025179563APending Publication Date: 2025-12-10DENSO CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024086397
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Conventional object detection techniques, particularly for motorcycles, struggle with accurately estimating the yaw angle, leading to erroneous orientation estimation of motorcycles relative to a host vehicle.

Method used

An estimation device and method that estimates the attitude of a motorcycle in a captured image by using an extended two-dimensional bounding box, including a circumscribing rectangle and ground contact points of the front and rear wheels, to accurately determine the motorcycle's orientation.

Benefits of technology

Accurately estimates the attitude of motorcycles, improving orientation detection and reducing errors in object detection systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025179563000001_ABST
    Figure 2025179563000001_ABST
Patent Text Reader

Abstract

To provide technology that enables accurate estimation of an attitude of a motorcycle with a rider as a target object.SOLUTION: An estimation device that estimates an attitude of a target object (B) in a captured image based on a captured image (Pg) of an own vehicle's destination includes a bounding box estimation unit and an attitude estimation unit. The bounding box estimation unit estimates an augmented two-dimensional bounding box (BB) within the captured image. The augmented two-dimensional bounding box includes at least a circumscribing rectangle (BBd) surrounding a motorcycle (B1) with a rider (B2) as the target object, and a front wheel contact point (P5) and a rear wheel contact point (P6) of the motorcycle. The attitude estimation unit estimates the attitude of the target object based on the augmented two-dimensional bounding box.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an estimation device, an estimation method, an estimation program, and a computer-readable, non-transient, physical recording medium on which such an estimation program is recorded, which estimate the attitude of a target object in a captured image based on the captured image of a destination of a vehicle. [Background technology]

[0002] BACKGROUND ART There is known a technique for detecting an object by setting a rectangular bounding box for an image captured by a camera installed in a vehicle (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-56717 Summary of the Invention [Problem to be solved by the invention]

[0004] In object detection, the orientation of the object can be an issue. For this reason, there has been a recent demand for accurate estimation of the orientation of an object. In particular, when the target object is a motorcycle, conventional techniques make it difficult to determine the yaw angle, which can lead to erroneous estimation that a motorcycle traveling ahead of the host vehicle or an oncoming motorcycle is oriented sideways. The present disclosure has been made in consideration of the circumstances exemplified above. [Means for solving the problem]

[0005] In one aspect of the present disclosure, an estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) of a destination of a host vehicle (V) based on the captured image, a bounding box estimation unit (602) that estimates, within the captured image, an extended two-dimensional bounding box (BB) that includes at least a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object and a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; an attitude estimation unit (604) that estimates the attitude of the target object based on the extended two-dimensional bounding box; It is equipped with: In another aspect of the present disclosure, a method for estimating an attitude of a target object (B) in a captured image (Pg) of a destination of a host vehicle (V) based on the captured image includes the following steps or processes: Within the captured image, an extended two-dimensional bounding box (BB) is estimated that includes at least a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object and a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; The pose of the target is estimated based on the extended two-dimensional bounding box. In yet another aspect of the present disclosure, an estimation program executed by an estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) of a destination of a host vehicle (V) based on the captured image includes, as processing executed by the estimation device, a process of estimating, within the captured image, a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object, and an extended two-dimensional bounding box (BB) including at least a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target object based on the extended 2D bounding box; Includes: In yet another aspect of the present disclosure, a computer-readable non-transient physical recording medium having recorded thereon an estimation program executed by an estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) of a destination of a host vehicle (V) based on the captured image includes, as a process included in the estimation program, a process of estimating, within the captured image, a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object, and an extended two-dimensional bounding box (BB) including at least a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target object based on the extended 2D bounding box; Includes:

[0006] In addition, in each section of the application documents, each element may be assigned a reference number in parentheses. However, such reference number merely indicates an example of the correspondence between the element and the specific means described in the embodiment described below. Therefore, the present disclosure is not limited in any way by the above-mentioned reference numbers. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a schematic diagram showing a vehicle to which the present disclosure is applied while traveling; [Figure 2] 2 is a block diagram showing a schematic device configuration of the in-vehicle system shown in FIG. 1. FIG. [Figure 3] 3 is a block diagram showing a schematic functional configuration realized in the target detection device shown in FIG. 2. FIG. [Figure 4] FIG. 4 is a diagram showing an outline of an example of the operation of the target detection device shown in FIGS. 2 and 3. [Figure 5] 4 is a flowchart showing an outline of an example of the operation of the target detection device shown in FIGS. 2 and 3. [Figure 6] 4 is a flowchart showing an outline of an example of the operation of the target detection device shown in FIGS. 2 and 3. [Figure 7] FIG. 4 is a conceptual diagram showing an outline of an example of operation of the target detection device shown in FIGS. 2 and 3. [Figure 8] FIG. 4 is a conceptual diagram showing an outline of an example of operation of the target detection device shown in FIGS. 2 and 3. DETAILED DESCRIPTION OF THE INVENTION

[0008] (Embodiment) Hereinafter, exemplary embodiments or specific examples of the present disclosure will be described with reference to the drawings as appropriate. Note that the following embodiments and their modifications, as well as the descriptions in the drawings related thereto, are schematic or simplified for the purpose of concisely explaining the contents of the present disclosure, and are not intended to limit the contents of the present disclosure in any way. Therefore, it goes without saying that the descriptions in the drawings do not necessarily coincide with the specific device configurations that are actually manufactured and sold. In other words, unless expressly limited by the applicant in the prosecution history of this application, it goes without saying that the present disclosure should not be interpreted as being limited by the descriptions in the drawings and the device configurations, functions, or operations described below corresponding thereto.

[0009] (In-vehicle system configuration) First, referring to FIG. 1 , an in-vehicle system 1 is mounted on a vehicle V and configured to execute various operations of the vehicle V. Hereinafter, the vehicle V equipped with the in-vehicle system 1 will be referred to as the "host vehicle." In this embodiment, the in-vehicle system 1 is equipped with a camera 2 that captures images of the surroundings of the host vehicle, and is configured to execute vehicle control operations of the host vehicle by performing image recognition of targets based on images captured by the camera 2. Note that in this specification, the term "target" includes "object." An "object" refers to a three-dimensional object such as a pedestrian or an obstacle. In addition to this, the term "target" includes two-dimensional detection targets such as road markings, as well as detection targets that are not necessarily related to a specific three-dimensional object, such as a step.

[0010] Vehicle control operations performed by the in-vehicle system 1 include operations for presenting information to occupants, driving control operations, and personal protection operations. "Information presentation" includes display and / or audio output. "Driving control" includes execution of longitudinal vehicle motion control subtasks and / or lateral vehicle motion control subtasks. The longitudinal vehicle motion control subtasks are starting, accelerating, decelerating, and stopping. The lateral vehicle motion control subtask is steering. "Personal protection operations" are operations for protecting the physical safety of vehicle occupants and pedestrians around the vehicle, and include operational control of protective devices such as airbags. Typically, the in-vehicle system 1 is configured as, for example, a so-called driving automation system, i.e., an automated driving system and / or a driving assistance system. Specifically, the in-vehicle system 1 is configured to, for example, detect an object B, such as a pedestrian or an obstacle, on a road surface Rd ahead of the vehicle using a camera 2, and perform various operations, such as collision avoidance operations and warning operations, based on the detection results.

[0011] The camera 2 is equipped with an image sensor such as a CCD or CMOS, and is mounted at a predetermined position on the host vehicle to capture images of the surroundings of the host vehicle. CCD stands for Charge Coupled Device. CMOS stands for Complementary Metal Oxide Semiconductor. In this embodiment, the host vehicle is equipped with at least a forward camera as the camera 2. Referring to FIG. 2, the in-vehicle system 1 is equipped with, in addition to the camera 2, an in-vehicle sensor 3, a locator 4, a communication device 5, a target detection device 6, an HMI device 7, and a vehicle control device 8. HMI stands for Human Machine Interface.

[0012] The camera 2, on-board sensor 3, locator 4, and communication device 5 are connected to the target detection device 6 via an on-board network so as to be able to exchange information or signals. The HMI device 7 and vehicle control device 8 are also connected to the target detection device 6 via the on-board network so as to be able to exchange information or signals. The on-board network is configured to comply with a predetermined communication standard such as CAN (international registered trademark: international registration number 1048262A). CAN (international registered trademark) is an abbreviation for Controller Area Network. Note that the on-board network may have, in addition to a main network conforming to CAN (international registered trademark), another main network or sub-network conforming to LIN, FlexRay, or the like. LIN is an abbreviation for Local Interconnect Network.

[0013] The on-board sensors 3 are configured to detect various quantities related to the driving state of the host vehicle. The "driving state" includes the driving operation state, driving behavior state, and driving environment state of the host vehicle. The "driving operation state" refers to the state related to the driving operation input of the host vehicle by the driver of the host vehicle or the vehicle control device 8 described later, and includes, for example, the steering amount, throttle opening, brake operation amount, shift range, etc. In other words, the on-board sensors 3 include an accelerator pedal sensor, a brake pedal sensor, a shift position sensor, a steering angle sensor, etc. The "driving behavior state" refers to the state related to the motion, i.e., physical behavior, of the host vehicle, and includes, for example, vehicle speed, acceleration, yaw rate, etc. In other words, the on-board sensors 3 include a vehicle speed sensor, a yaw rate sensor, an acceleration sensor, etc. The "driving environment state" refers to the environment around the host vehicle, and includes, for example, the illuminance, weather, outside temperature, road surface condition, the presence of objects such as pedestrians and other vehicles, etc. That is, the on-board sensors 3 include an illuminance sensor, a raindrop sensor, an outside air temperature sensor, a radar sensor, a laser radar sensor, a sonar sensor, etc. Among the on-board sensors 3, a sensor related to the presence state of a target is called an ADAS sensor. ADAS stands for Advanced Driver-Assistance Systems. The ADAS sensor may include a camera 2.

[0014] The locator 4 is configured to measure the position of the vehicle itself. Specifically, the locator 4 has at least a satellite positioning function that measures the position of the vehicle itself by receiving a positioning signal transmitted from a positioning satellite. Note that the locator 4 may also be configured to use an autonomous positioning function that uses an inertial sensor such as a gyro sensor or an acceleration sensor in order to improve the measurement accuracy of the vehicle's position in places where satellite radio waves are difficult to reach, such as inside a tunnel. Such an inertial sensor may be provided in the locator 4 or in the on-board sensor 3. For example, a commercially available locator 4 equipped with an inertial sensor is the "POSLV" positioning and orientation system for land vehicles manufactured by Applanix.

[0015] The communication device 5 is an in-vehicle communication module also referred to as DCM, and is configured to be able to communicate information with an external server Z via base stations around the vehicle using wireless communication compliant with communication standards such as LTE or 5G. DCM is an abbreviation for Data Communication Module. LTE is an abbreviation for Long Term Evolution. 5G is an abbreviation for 5th Generation. The communication device 5 is configured to be able to acquire various information such as road traffic information such as congestion information and the latest map information from the external server Z and output it to the target detection device 6, the HMI device 7, and the vehicle control device 8.

[0016] The target detection device 6 is configured to detect targets around the vehicle based on information and signals acquired from the camera 2, the on-board sensor 3, etc. Furthermore, the target detection device 6 generates and outputs signals required for information presentation in the HMI device 7 and driving control in the vehicle control device 8 based on the target detection results. In this way, the target detection device 6 has a configuration as an electronic circuit unit called an image processing ECU, a target recognition ECU, or a target detection ECU. ECU is an abbreviation for Electronic Control Unit. The configuration and functions of the target detection device 6 will be described in detail later.

[0017] The HMI device 7 includes a display device, an audio output device, and the like for presenting various types of information and warnings to the occupants of the vehicle. The display device may include a meter, a meter display, a center information display, a head-up display, an electronic mirror, and the like. The vehicle control device 8 is configured as a so-called driving ECU, which is an on-board computer that controls the driving force generation mechanism, driving force transmission mechanism, braking mechanism, steering mechanism, and the like of the vehicle. That is, the vehicle control device 8 is configured to execute longitudinal and / or lateral motion control of the vehicle. More specifically, the vehicle control device 8 is configured to be able to execute at least a part of the motion control of the vehicle, such as starting, acceleration / deceleration, braking, stopping, steering, and the like.

[0018] (Target detection device) In this embodiment, the target detection device 6 is configured as an on-board microcomputer including at least a processor 61 and a memory 62. The processor 61 includes at least one arithmetic unit having the functions or configuration of a CPU or MPU, and its peripheral circuits (e.g., a timer circuit, etc.). The memory 62 includes at least a RAM and a ROM or a nonvolatile rewritable memory among various non-transient physical storage media such as a ROM, a RAM, and a nonvolatile rewritable memory. The nonvolatile rewritable memory is a storage device that is rewritable when powered on but retains information in an unrewritable manner when powered off, such as a flash memory. The target detection device 6 is configured so that the processor 61 reads and executes a computer program from the memory 62 to realize a predetermined function for recognizing targets around the vehicle. The memory 62 stores the computer program as well as various data required to execute the program, such as initial values, maps, look-up tables, etc.

[0019] 3 shows an example of a functional configuration realized on the target detection device 6 serving as an on-board microcomputer by the processor 61 executing a computer program. That is, the target detection device 6 has, as such a functional configuration, an image acquisition unit 601, a bounding box estimation unit 602, a bounding box projection unit 603, and a position / orientation estimation unit 604. Each of these functional configurations will be described below. In the following description, unless otherwise noted, in this embodiment, the camera 2 refers to the front camera. However, it goes without saying that the present disclosure is not limited to such an embodiment.

[0020] The image acquisition unit 601 acquires image data captured by the camera 2. FIG. 4 shows an example of an image Pg captured by the camera 2, which is the target of acquisition by the image acquisition unit 601, in which a cyclist is captured as a target object B. The cyclist includes a motorcycle B1 and a rider B2, and may be referred to as a motorcycle B1 with rider B2. The motorcycle B1 includes a front wheel B11 and a rear wheel B12. In this embodiment, the image acquisition unit 601 receives image data corresponding to the captured image Pg captured using the camera 2 from the camera 2 and stores a certain amount of the image data in chronological order. In the following description, the "vertical direction" and "horizontal direction" of the captured image Pg are defined as follows. The "vertical direction" is the up-down direction in FIG. 4 and is the direction that determines the vertical resolution of the captured image Pg. The "horizontal direction" is the left-right direction in FIG. 4 and is the direction that determines the horizontal resolution of the captured image Pg.

[0021] The bounding box estimation unit 602 estimates a bounding box BB within the captured image Pg using machine learning (e.g., typically a deep neural network). In this embodiment, the bounding box BB is a so-called extended two-dimensional bounding box. The extended two-dimensional bounding box is a circumscribing rectangle BBd corresponding to a normal two-dimensional bounding box, with information on points other than the corner points of the circumscribing rectangle BBd added. The circumscribing rectangle BBd is a rectangle that circumscribes the target object B within the captured image Pg, and is a rectangle connecting the lower left point P1, the lower right point P2, the upper right point P3, and the upper left point P4 in this order. The bounding box BB includes at least the circumscribing rectangle BBd, a front wheel contact point P5, and a rear wheel contact point P6. The front wheel contact point P5 is the point where the front wheel B11 contacts the ground. The rear wheel contact point P6 is the point where the rear wheel B12 contacts the ground. The straight line connecting the front wheel ground contact point P5 and the rear wheel ground contact point P6 defines the attitude of the target object B, and is called the active line La.

[0022] The bounding box estimation unit 602 generates a bounding box BB based on a circumscribing rectangle BBd estimated using machine learning and the outer dimensions of the target object B. In this embodiment, the bounding box BB includes a circumscribing rectangle BBd, a front wheel ground contact point P5, a rear wheel ground contact point P6, and an upper end point P7. The upper end point P7 corresponds to the height of the vertex of the target object B in the captured image Pg. Specifically, in the example of FIG. 4 , the upper end point P7 is located at the midpoint (i.e., the center) between the front wheel ground contact point P5 and the rear wheel ground contact point P6 in the lateral direction and at the upper end position of the circumscribing rectangle BBd in the longitudinal direction (i.e., the same position as the upper right point P3 and the upper left point P4). The upper end point P7 may be located at the midpoint of the line segment connecting the upper side of the circumscribing rectangle BBd, i.e., the upper right point P3 and the upper left point P4, or may be located at the top of the head of the occupant B2. Furthermore, the outer dimensions of the target object B used to generate the bounding box BB may be estimated using machine learning, or the standard outer diameter of the two-wheeled vehicle B1 may be used as a predetermined value.

[0023] The bounding box projection unit 603 projects a bounding box BB onto the captured image Pg. The position / orientation estimation unit 604 estimates the position and orientation of the target object B based on the bounding box BB, i.e., based on projection result information by the bounding box projection unit 603. Details of the method for estimating the position and orientation will be described later. In this way, the target object detection device 6 is configured to function as an estimation device that estimates the orientation of the two-wheeled vehicle B1 with a rider B2, which is the target object B in the captured image Pg, based on the captured image Pg ahead of the host vehicle.

[0024] (Example of operation) An outline of the target detection operation by the target detection device 6 according to this embodiment will be described below. In the flowcharts shown in FIGS. 5 and 6, "S" is an abbreviation for "step." The target detection device 6 as an estimation device according to this embodiment, the estimation method and estimation program executed thereby, and a computer-readable, non-transient, tangible recording medium on which the estimation program is recorded will be collectively referred to as "this embodiment." This recording medium may be realized, for example, by a ROM, a non-volatile rewritable memory, a magnetic disk, an optical disk, or the like. Specifically, this recording medium may be realized in any format, for example, by an external server Z, a portable terminal device, an optical disk such as a CD-ROM, a memory card detachable from a computer device such as a terminal device, or the like.

[0025] When a predetermined target detection condition is met, the processor 61 provided in the target detection device 6 reads out the target detection routine shown in Fig. 5 from the memory 62 and repeatedly executes the routine at predetermined time intervals (e.g., 10 msec intervals) while the target detection condition is met. The target detection condition includes, for example, that the ignition switch is turned on, that the shift position is other than "P", etc. When the processor 61 starts the target detection routine shown in Fig. 5, it executes the processes of steps 101 to 104 in order.

[0026] In step 101, the processor 61 acquires a captured image Pg. The processing content of step 101 corresponds to the function of the image acquisition unit 601. In step 102, the processor 61 estimates a bounding box BB. At this time, the processor 61 may use the outer dimensions of the target object B used in estimating the bounding box BB estimated using machine learning, or may use the standard outer diameter dimensions of the motorcycle B1 as a predetermined value. Details of the process of estimating the bounding box BB will be described later. The processing content of step 102 corresponds to the function of the bounding box estimation unit 602. In step 103, the processor 61 projects the bounding box BB onto the captured image Pg. The processing content of step 103 corresponds to the function of the bounding box projection unit 603. In step 104, the processor 61 estimates the position and orientation of the target object B. The processing content of step 104 corresponds to the function of the position / orientation estimation unit 604.

[0027] (Bounding box estimation) In the process of estimating the bounding box BB in step 102, the processor 61 executes the processes of steps 201 to 206 shown in FIG.

[0028] In step 201, processor 61 generates a circumscribing rectangle BBd. In step 202, processor 61 selects a front wheel ground contact point P5 and a rear wheel ground contact point P6. In step 203, processor 61 selects an upper end point P7. In step 204, processor 61 selects corner points that form the circumscribing rectangle BBd. There are no particular limitations on the order in which steps 201 to 204 are performed, as long as there is no technical contradiction. That is, for example, the processing of step 204 may be performed simultaneously with or immediately after the processing of step 201.

[0029] In step 205, the processor 61 calculates the world coordinates of the front wheel ground contact point P5 and the rear wheel ground contact point P6. "World coordinates" refers to three-dimensional coordinates with a predetermined point (e.g., the center point of the vehicle) as the origin and coordinate axes in the vehicle length direction, vehicle width direction, and vehicle height direction, taking into account the dynamic angle of the vehicle (i.e., pitch, yaw, etc.). "WCS" in the figure stands for World Coordinate System. The bottom of the three-dimensional bounding box generated in the world coordinate system is formed by points r / 2 away from both ends of the active line La in its extension direction, and points W / 2 away on both sides in the width direction. r is the wheel diameter of the two-wheeled vehicle B1. W is the vehicle width of the two-wheeled vehicle B1 (e.g., the width of the handlebars), which corresponds to the width of the three-dimensional bounding box.

[0030] In step 206, the processor 61 converts each selected point from world coordinates to digital image coordinates. "Digital image coordinates" are coordinates on the captured image Pg, measured in units of pixels. "PCS" in the figure stands for Pixel Coordinate System or Program Coordinate System. The points used in the bounding box BB include the front wheel contact point P5, the rear wheel contact point P6, and the top end point P7.

[0031] The WCS coordinates of the front wheel contact point P5 are expressed as (0, H / 2, L / 2 - r / 2) when the center point of the three-dimensional bounding box is the origin. H is the height of the three-dimensional bounding box, and L is the length of the three-dimensional bounding box, i.e., the distance between the front end of the front wheel B11 and the rear end of the rear wheel B12 of the motorcycle B1. Similarly, the WCS coordinates of the rear wheel contact point P6 are expressed as (0, H / 2, -L / 2 + r / 2). Note that the leftmost point, rightmost point, and bottommost point may also be used for the bounding box BB. The leftmost point is the point located at the leftmost position in the three-dimensional bounding box. The rightmost point is the point located at the rightmost position in the three-dimensional bounding box. The bottommost point is the point located at the bottommost position in the three-dimensional bounding box.

[0032] As described above, in this embodiment, the bounding box estimation unit 602 estimates the bounding box BB as an extended two-dimensional bounding box based on the active line La, which is a straight line connecting the front wheel ground contact point P5 and the rear wheel ground contact point P6, and the external dimensions of the target object B. Note that the values ​​of W, H, L, and r, which are the external dimensions of the target object B, may be predetermined default values, or the results of image recognition or machine learning may be used.

[0033] 7 shows the relationship between the posture of a two-wheeled vehicle B1 with a rider B2 as the object B and an estimated example of a bounding box BB. As shown in FIG. 7, the front wheel ground contact point P5, the rear wheel ground contact point P6, and the top end point P7 can be accurately estimated regardless of the posture, i.e., orientation, of the object B. Furthermore, the positional relationship between the front wheel ground contact point P5 and the rear wheel ground contact point P6, i.e., the orientation of the active line La, correlates with the posture of the object B. As will be described later, the position / posture estimation unit 604 can estimate the posture based on the bounding box BB that does not include the corner points (i.e., the lower left point P1 to the upper left point P4) of the circumscribing rectangle BBd but does include the front wheel ground contact point P5, the rear wheel ground contact point P6, and the top end point P7.

[0034] (Position / Position Estimation) Details of the method for estimating the position and orientation of the target object B according to this embodiment will be explained using appropriate mathematical expressions. This estimation method is basically the same as the method described in Japanese Patent Application Laid-Open No. 2024-13021, a prior application of the applicant of the present application. Specifically, this estimation method estimates the position and orientation of the target object B based on the relationship that "each point of a three-dimensional bounding box in the camera coordinate system is constrained by two-dimensional information when projected onto the image coordinate system." The "two-dimensional information" referred to here is information represented by the bounding box BB. This constraint can be expressed by the following equation (1).

[0035]

number

[0036] In the above formula (1), the image coordinate x on the left side pcs =(x,y), and the camera coordinate X on the right side ccs = (X, Y, Z), and the right-hand side corresponds to the transformation of the camera coordinates into the image coordinate system. s is the scale factor of the projective transformation. R is the attitude of the target object B, and is expressed as R = (p, r, θ) using pitch angle p, roll angle r, and yaw angle θ. In this embodiment, p = r = 0. That is, this embodiment is concerned with only the yaw angle θ as the "direction" of the attitude of the target object B. T is the position of the target object B, and is expressed as T = (Tx, Ty, Tz). K is the internal parameter matrix of the imaging device 60, and can be expressed by the following equation (2). In this equation (2), fx and fy are the x and y components of the focal length of the camera, and cx and cy are the x and y components of the principal point of the camera.

[0037]

number

[0038] By modifying the above formula (1), the following formula (3) can be obtained.

number

[0039] From the above formulas (2) and (3), and xpcs=(x, y) and Xccs=(X, Y, Z), the following formula (4) can be obtained.

number

[0040] When the above formula (4) is rearranged for the four unknown variables θ and T = (Tx, Ty, Tz), two formulas can be obtained as shown in the following formula (5). Note that f1 and f2 in these two formulas are conveniently assigned ordinal numbers to the two obtained formulas, making them the "first formula" and the "second formula."

number

[0041] From the above formula (5), the following formulas (6) and (7) can be obtained.

number

[0042]

number

[0043] As described above, two simultaneous equations can be obtained by combining one point of the three-dimensional bounding box in the camera coordinate system with one point of the bounding box BB, which is an extended two-dimensional bounding box in the image coordinate system. These two simultaneous equations have sin θ, cos θ, Tx, Ty, and Tz as unknown variables. For example, since there are five unknown variables, if there are three pairs of points on the three-dimensional bounding box and points on the bounding box BB, six simultaneous equations can be obtained, and sin θ, cos θ, Tx, Ty, and Tz can be calculated. These can be calculated linearly. The simultaneous equations can be solved using, for example, the Moore-Penrose inverse matrix.

[0044] FIG. 8 is a diagram showing the relationship between each vertex of a three-dimensional bounding box in the camera coordinate system and the outer dimensions of the target B. As shown in FIG. 8, the center of the three-dimensional bounding box is set as the origin O, with the x-axis to the right, the y-axis downward, and the z-axis forward. For convenience, the outer shape of the target B is assumed to be a cube of W×H×L. In this case, the coordinates of the front upper left corner flt, front lower left corner flb, front upper right corner frt, front lower right corner frb, back upper left corner rlt, back lower left corner rlb, back upper right corner rrt, and back lower right corner rrb of the target B can be expressed as follows: That is, the camera coordinates Xccs=(X, Y, Z) can be calculated from the outer dimensions of the target B. flt=(-W / 2,-H / 2,L / 2) flb=(-W / 2,H / 2,L / 2) frt=(W / 2,-H / 2,L / 2) frb=(W / 2,H / 2,L / 2) rlt=(-W / 2,-H / 2,-L / 2) rlb=(-W / 2,H / 2,-L / 2) rrt=(W / 2,-H / 2,-L / 2) rrb=(W / 2,H / 2,-L / 2)

[0045] A system of simultaneous equations such as the above equation (5) can also be solved by a nonlinear method described below. For the combination of one point of the three-dimensional bounding box in the camera coordinate system and one point of the bounding box BB in the image coordinate system, M simultaneous equations are obtained by selecting M / 2 combinations according to the target type. In the case of a motorcycle B1 with a rider B2, six simultaneous equations are obtained using three points, as shown in the following equation (8).

number

[0046] M simultaneous equations are expressed as f1(X)~f M Let (X). For N unknown variables X = [θ, Tx, Ty, Tz] = [x1, x2, ..., xN], let its variation be Δx = (Δx1, Δx2, ..., ΔxN), and perform a first-order Taylor expansion to obtain the following equation (9). Note that "X" in each equation and explanation after equation (9) below refers to the unknown variable X = [θ, Tx, Ty, Tz], and is distinguished from the X coordinate value in Xccs = (X, Y, Z). The same applies to x1, x2, ..., xN and Δx1, Δx2, ..., ΔxN.

number

[0047] The following equation (10) is a determinant of the above equation (9).

number

[0048] The above equation (10) can be expressed by the following equation (11) using the Jacobian determinant J.

number

[0049] When M>N, the above equation (11) can be solved using the Moore-Penrose inverse, as shown in the following equation (12).

number

[0050] After calculating ΔX from the above equation (12), X = X - ΔX, and the procedure explained using the above equations (9) to (12) is repeated until ΔX converges. That is, it is repeated until the absolute value of ΔX becomes smaller than a predetermined value ε. Note that the initial value of X can be a value calculated by a linear solution.

[0051] (effect) As described above in detail, this embodiment estimates an extended two-dimensional bounding box that includes at least a circumscribing rectangle BBd that surrounds the motorcycle B1 with rider B2 as the object B, and the front wheel ground contact point P5 and the rear wheel ground contact point P6 of the motorcycle B1. Then, this embodiment estimates the attitude, i.e., orientation, of the object B based on the bounding box BB, which is the estimated extended two-dimensional bounding box. This makes it possible to accurately estimate the attitude of the motorcycle B1 with rider B2 as the object B.

[0052] (Variation) The present disclosure is not limited to the above-described embodiments and specific examples. Therefore, the above-described embodiments and the like can be modified as appropriate. Representative modifications will be described below. In the following description of the modifications, differences from the above-described embodiments and the like will be mainly described. Furthermore, the same reference numerals are used for parts that are identical or equivalent to each other in the above-described embodiments and the following modifications. Therefore, in the following description of the modifications, the explanations in the above-described embodiments and the like can be used as appropriate for components that have the same reference numerals as the above-described embodiments and the like, unless there is a technical contradiction or special additional explanation.

[0053] The present disclosure is not limited to the specific applications and device configurations shown in the above embodiments. For example, the host vehicle may be a so-called automobile or a motorcycle. There are no particular limitations on the type of automobile or motorcycle.

[0054] All or part of the target object detection device 6 may be configured to include a digital circuit, such as an ASIC or FPGA, configured to be able to realize the above-mentioned functions or operations. ASIC stands for Application Specific Integrated Circuit. FPGA stands for Field Programmable Gate Array. In other words, in the target object detection device 6, an on-board microcomputer portion and a digital circuit portion may coexist.

[0055] A computer program according to the present disclosure that enables the execution of various operations, procedures, or processes described in the above embodiments can be downloaded or upgraded via V2X communication using the communication device 5. V2X stands for Vehicle to X. Alternatively, such a computer program can be downloaded or upgraded via a terminal device installed in a vehicle manufacturing plant, a repair shop, a dealer, or the like. Such a computer program may be stored on a memory card, an optical disk, a magnetic disk, or the like.

[0056] In this way, each of the above functional configurations and processes may be realized by a special-purpose computer provided by configuring a processor 61 and memory 62 programmed to execute one or more functions embodied in a computer program. Alternatively, each of the above functional configurations and processes may be realized by a special-purpose computer provided by configuring a processor 61 with one or more dedicated hardware logic circuits. Alternatively, each of the above functional configurations and processes may be realized by one or more special-purpose computers configured by combining one or more processors 61 programmed to execute one or more functions and one or more memories 62 with one or more other processors 61 configured with one or more hardware logic circuits. Furthermore, a computer program may be stored in a computer-readable, non-transitory storage medium as instructions to be executed by a computer. In other words, each of the above functional configurations and processes may be expressed as a computer program including procedures for implementing the same, or as a non-transitory storage medium storing the computer program.

[0057] The present disclosure is not limited to the specific functions and operational modes described in the above embodiments. For example, the present disclosure may be effectively applied to target recognition in an image Pg captured by a camera other than a front camera (e.g., a rear camera).

[0058] It goes without saying that the elements constituting the above-described embodiments are not necessarily essential unless expressly stated as essential or clearly considered essential in principle. Furthermore, when numerical values ​​such as the number, value, amount, and range of components are mentioned, the present disclosure is not limited to those specific numbers unless expressly stated as essential or clearly limited to a specific number in principle. Similarly, when the shape, direction, positional relationship, etc. of components are mentioned, the present disclosure is not limited to those shapes, directions, positional relationships, etc. unless expressly stated as essential or clearly limited to a specific shape, direction, positional relationship, etc. in principle.

[0059] Similar expressions such as "acquire," "calculate," "estimate," "detect," and "sensing" may be substituted for each other as appropriate within the scope of technical inconsistency. Furthermore, "exceeding the threshold" and "above the threshold" may be substituted for each other as appropriate within the scope of technical inconsistency. The same applies to "below the threshold" and "below the threshold."

[0060] The variations are not limited to the above examples. For example, all or part of one of the variations may be combined with all or part of another, provided that no technical contradiction exists. Furthermore, all or part of the above specific example and all or part of the variations may be combined with each other, provided that no technical contradiction exists.

[0061] (Disclosure perspective) As is clear from the above description of the embodiments and modifications, this specification discloses at least the following matters.

[0062] [Perspective A] A method for estimating the attitude of a target object (B) in a captured image (Pg) based on a captured image (Pg) of a destination of a host vehicle (V), comprising: A captured image (Pg) is acquired using a camera (2) mounted on the vehicle; Within the acquired captured image, an extended two-dimensional bounding box (BB) is estimated that includes at least a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object and a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target based on the augmented two-dimensional bounding box; Estimation method.

[0063] [Perspective B1] An estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) of a destination of a host vehicle (V) based on the captured image, a bounding box estimation unit (602) that estimates, within the captured image, an extended two-dimensional bounding box (BB) that includes at least a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object and a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; an attitude estimation unit (604) that estimates the attitude of the target object based on the extended two-dimensional bounding box; An estimation device comprising: [Perspective B2] The extended two-dimensional bounding box further includes an upper end point (P7) corresponding to the height of a vertex of the target object in the captured image. An estimation device according to aspect B1. [Perspective B3] the bounding box estimation unit sets the upper end point to an intermediate position between the front wheel ground contact point and the rear wheel ground contact point in the horizontal direction of the captured image and to an upper end position of the circumscribing rectangle in the vertical direction of the captured image. The estimation device according to aspect B2. [Perspective B4] the posture estimation unit estimates the posture based on the extended two-dimensional bounding box that does not include corner points (P1, P2, P3, P4) in the circumscribing rectangle but includes the front wheel ground contact points, the rear wheel ground contact points, and the upper end point. The estimation device according to aspect B2 or B3. [Perspective B5] The bounding box estimation unit estimates the extended two-dimensional bounding box based on an active line (La) that is a straight line connecting the front wheel ground contact points and the rear wheel ground contact points. The estimation device according to any one of the aspects B1 to B4. [Perspective B6] the bounding box estimation unit estimates the front wheel contact points and the rear wheel contact points in the extended two-dimensional bounding box based on predetermined outer dimensions of the motorcycle or outer dimensions estimated by machine learning; The estimation device according to any one of the aspects B1 to B5.

[0064] [Perspective C1] An estimation method for estimating the attitude of a target object (B) in a captured image (Pg) of a destination of a host vehicle (V), based on the captured image, comprising: Within the captured image, an extended two-dimensional bounding box (BB) is estimated that includes at least a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object and a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target based on the augmented two-dimensional bounding box; Estimation method. [Perspective C2] The extended two-dimensional bounding box further includes an upper end point (P7) corresponding to the height of a vertex of the target object in the captured image. The estimation method according to aspect C1. [Perspective C3] In estimating the bounding box, the upper end point is set to an intermediate position between the front wheel ground contact point and the rear wheel ground contact point in the horizontal direction of the captured image and to an upper end position of the circumscribing rectangle in the vertical direction of the captured image. The estimation method according to aspect C2. [Perspective C4] The posture is estimated based on the extended two-dimensional bounding box that does not include the corner points (P1, P2, P3, P4) in the circumscribing rectangle but includes the front wheel contact points, the rear wheel contact points, and the upper end point. The estimation method according to aspect C2 or C3. [Perspective C5] The extended two-dimensional bounding box is estimated based on an active line (La) which is a straight line connecting the front wheel contact point and the rear wheel contact point. The estimation method according to any one of Aspects C1 to C4. [Perspective C6] estimating the front wheel contact points and the rear wheel contact points in the extended two-dimensional bounding box based on predetermined or machine learning estimated outer dimensions of the motorcycle; The estimation method according to any one of Aspects C1 to C5.

[0065] [Perspective D1] An estimation program executed by an estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) of a destination of a host vehicle (V) based on the captured image, The process executed by the estimation device includes: a process of estimating, within the captured image, a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object, and an extended two-dimensional bounding box (BB) including at least a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target object based on the extended 2D bounding box; Estimation programs including. [Perspective D2] The extended two-dimensional bounding box further includes an upper end point (P7) corresponding to the height of a vertex of the target object in the captured image. The estimation program according to aspect D1. [Perspective D3] In the process of estimating the bounding box, the upper end point is set to an intermediate position between the front wheel ground contact point and the rear wheel ground contact point in the horizontal direction of the captured image and to an upper end position of the circumscribing rectangle in the vertical direction of the captured image. The estimation program according to aspect D2. [Perspective D4] The posture is estimated based on the extended two-dimensional bounding box that does not include the corner points (P1, P2, P3, P4) in the circumscribing rectangle but includes the front wheel contact points, the rear wheel contact points, and the upper end point. The estimation program according to aspect D2 or D3. [Perspective D5] In the process of estimating the bounding box, the extended two-dimensional bounding box is estimated based on an active line (La) which is a straight line connecting the front wheel ground contact points and the rear wheel ground contact points. The estimation program according to any one of aspects D1 to D4. [Perspective D6] In the process of estimating the bounding box, the front wheel contact point and the rear wheel contact point in the extended two-dimensional bounding box are estimated based on predetermined outer dimensions of the motorcycle or outer dimensions estimated by machine learning. The estimation program according to any one of aspects D1 to D5.

[0066] [Perspective E1] A computer-readable non-transient tangible recording medium storing an estimation program executed by an estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) of a destination of a vehicle (V) based on the captured image, The processing included in the estimation program is a process of estimating, within the captured image, a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object, and an extended two-dimensional bounding box (BB) including at least a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target object based on the extended 2D bounding box; A recording medium including: [Perspective E2] The extended two-dimensional bounding box further includes an upper end point (P7) corresponding to the height of a vertex of the target object in the captured image. A recording medium according to aspect E1. [Perspective E3] In the process of estimating the bounding box, the upper end point is set to an intermediate position between the front wheel ground contact point and the rear wheel ground contact point in the horizontal direction of the captured image and to an upper end position of the circumscribing rectangle in the vertical direction of the captured image. A recording medium according to aspect E2. [Perspective E4] The posture is estimated based on the extended two-dimensional bounding box that does not include the corner points (P1, P2, P3, P4) in the circumscribing rectangle but includes the front wheel contact points, the rear wheel contact points, and the upper end point. A recording medium according to aspect E2 or E3. [Perspective E5] In the process of estimating the bounding box, the extended two-dimensional bounding box is estimated based on an active line (La) which is a straight line connecting the front wheel ground contact points and the rear wheel ground contact points. A recording medium according to any one of aspects E1 to E4. [Perspective E6] In the process of estimating the bounding box, the front wheel contact point and the rear wheel contact point in the extended two-dimensional bounding box are estimated based on predetermined outer dimensions of the motorcycle or outer dimensions estimated by machine learning. A recording medium according to any one of aspects E1 to E5. [Explanation of symbols]

[0067] 6 Target detection device (estimation device) 602 Bounding Box Estimation Unit 604 Position / orientation estimation unit B Target B1 Motorcycle BB Bounding Box (Extended 2D Bounding Box) P5 Front wheel contact point P6 Rear wheel contact point Pg Captured image V Vehicle

Claims

1. An estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) based on the captured image (Pg) of a destination of a host vehicle (V), a bounding box estimation unit (602) that estimates, within the captured image, an extended two-dimensional bounding box (BB) that includes at least a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object and a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; an attitude estimation unit (604) for estimating the attitude of the target object based on the extended two-dimensional bounding box; An estimation device comprising:

2. The extended two-dimensional bounding box further includes an upper end point (P7) corresponding to the height of a vertex of the target object in the captured image. The estimation device according to claim 1 .

3. the bounding box estimation unit sets the upper end point to an intermediate position between the front wheel ground contact point and the rear wheel ground contact point in the horizontal direction of the captured image and to an upper end position of the circumscribing rectangle in the vertical direction of the captured image. The estimation device according to claim 2 .

4. the posture estimation unit estimates the posture based on the extended two-dimensional bounding box that does not include corner points (P1, P2, P3, P4) of the circumscribing rectangle but includes the front wheel ground contact points, the rear wheel ground contact points, and the upper end point. The estimation device according to claim 2 .

5. The bounding box estimation unit estimates the extended two-dimensional bounding box based on an active line (La) that is a straight line connecting the front wheel ground contact points and the rear wheel ground contact points. The estimation device according to claim 1 .

6. the bounding box estimation unit estimates the front wheel contact points and the rear wheel contact points in the extended two-dimensional bounding box based on predetermined outer dimensions of the motorcycle or outer dimensions estimated by machine learning; The estimation device according to claim 1 .

7. An estimation method for estimating the attitude of a target object (B) in a captured image (Pg) of a destination of a host vehicle (V), based on the captured image, comprising: Within the captured image, an extended two-dimensional bounding box (BB) is estimated that includes at least a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object and a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target based on the augmented two-dimensional bounding box; Estimation method.

8. An estimation program executed by an estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) based on the captured image of a destination of a vehicle (V), comprising: The process executed by the estimation device includes: a process of estimating, within the captured image, a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object, and an extended two-dimensional bounding box (BB) including at least a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target object based on the extended 2D bounding box; Estimation programs including.

9. A computer-readable non-transient tangible recording medium storing an estimation program executed by an estimation device (6) that estimates the attitude of a target object (B) in a captured image (Pg) of a destination of a vehicle (V) based on the captured image, The processing included in the estimation program is a process of estimating, within the captured image, a circumscribing rectangle (BBd) surrounding the motorcycle (B1) with a rider (B2) as the target object, and an extended two-dimensional bounding box (BB) including at least a front wheel ground contact point (P5) and a rear wheel ground contact point (P6) of the motorcycle; estimating the pose of the target object based on the extended 2D bounding box; A recording medium including:

Citation Information

Patent Citations

  • Object detection device

    JP2021056717A