Image processing device

The image processing device uses transformed road surface 3D information to enhance monocular-based distance estimation, addressing inaccuracies in turning vehicles by integrating road surface 3D management and conversion units, ensuring reliable obstacle detection.

JP7727502B2Active Publication Date: 2025-08-21ASTEMO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021189408
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-22
Publication Date
2025-08-21
Estimated Expiration
2041-11-22

AI Technical Summary

Technical Problem

Existing image-based obstacle detection systems struggle to accurately calculate the distance to moving objects in monocular areas when the vehicle is turning at intersections, especially when there are differences in ground height relative to the roadway, due to changes in vehicle speed and road surface conditions, which can lead to inaccurate distance measurements and reduced reliability.

Method used

An image processing device that stores and transforms three-dimensional road surface information using vehicle movement and camera orientation to improve monocular-based distance estimation, incorporating road surface 3D information management and conversion units to maintain accuracy without additional sensors.

Benefits of technology

Enables accurate monocular-based distance measurements for obstacles during vehicle turns, enhancing system reliability by accounting for changes in vehicle speed and road surface conditions, particularly at intersections with sidewalk areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007727502000001
    Figure 0007727502000001
  • Figure 0007727502000002
    Figure 0007727502000002
  • Figure 0007727502000003
    Figure 0007727502000003
Patent Text Reader

Abstract

To provide an image processing apparatus which improves accuracy of distance measurement without requiring a special sensor, while maintaining accuracy of obstacle detection, using three-dimensional information from a road surface, especially, supports a case where a vehicle turns in an intersection which may include a sidewalk area (i.e., including a difference in a ground height with respect to a roadway area) that requires accurate monocular-based distance measurement for a detected obstacle.SOLUTION: An image processing device 110 is configured to: convert three-dimensional information (road surface 3D information) acquired through a road surface 3D information acquisition unit 131 and related to a road surface stored in a storage unit 181 of a road surface 3D information management unit 141, into road surface 3D information of current time by a coordinate conversion unit 191 of the road surface 3D information management unit 141 by using a motion of a vehicle; and determine a distance to an obstacle, by an obstacle ranging unit 161, using the converted road surface 3D information together with a monocular-based distance estimation process.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an on-board image processing device for image-based obstacle detection and recognition in the environment near the host vehicle, for example. [Background technology]

[0002] In recent years, image-based object detection devices have been used to detect nearby moving objects and obstacles. Such image-based object detection devices can be used in surveillance systems that detect intrusions or abnormalities, or in-vehicle systems that assist in safe driving of automobiles.

[0003] In vehicle applications, such devices are configured to display the surrounding environment to the driver and / or detect moving or static objects (obstacles) around the vehicle, notify the driver of a potential risk of a collision between the vehicle and the obstacle, and, based on a decision system, automatically stop the vehicle to avoid a collision between the vehicle and the obstacle.

[0004] Incidentally, there are two types of cameras used as sensors for monitoring the surroundings of a vehicle: monocular cameras and stereo cameras using multiple cameras. Stereo cameras can measure the distance to an object captured by using the parallax of an overlapping area (also called a stereo area) captured by two cameras spaced a predetermined distance apart. This allows for accurate assessment of the possibility of a collision with a surrounding object. On the other hand, in non-overlapping areas (also called monocular areas) captured by each camera alone (monocular camera), it is difficult to assess the distance to the object (a moving object such as a pedestrian) simply by detecting the captured object, making it difficult to assess the possibility of a collision. Therefore, one method for estimating the distance to a moving object detected in a monocular area is to apply road surface height information measured in a stereo area to the road surface in the monocular area.

[0005] However, in the above-mentioned conventional method, the road surface height in the stereo area is applied to the monocular area as it is, and therefore it is assumed that the road surface heights in the stereo area and the monocular area are the same, that is, that the road surface is flat. Therefore, when there is an incline or a step in the road surface in the monocular area (in other words, when the pedestrian detected in the monocular area is not at the same height as the road surface), it is difficult to estimate the distance accurately.

[0006] For example, the device disclosed in Patent Document 1 aims to address the above-mentioned problem by calculating with high accuracy the distance to a moving object detected in a monocular area even when the road surface heights in the overlapping area and the monocular area are different, and is equipped with a disparity information acquisition unit that acquires disparity information of the overlapping area of ​​multiple images taken by multiple cameras mounted on a vehicle, and an object distance calculation unit that calculates the distance between the vehicle and an object detected in a non-overlapping area other than the overlapping area in each of the images based on disparity information previously acquired in the overlapping area. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Japanese Patent Application Publication No. 2017-96777 Summary of the Invention [Problem to be solved by the invention]

[0008] The device described in the above-mentioned Patent Document 1 stores height information of the sidewalk surface calculated based on parallax information previously acquired in the overlapping area, and when the image area where the height of the sidewalk surface is measured moves into the non-overlapping area as the vehicle moves, the stored height information of the sidewalk surface is used to correct the position of the object in the image height direction and calculate the distance between the object and the vehicle. In other words, the device described in the above-mentioned Patent Document 1 is based on the premise that when the image area where the height of the sidewalk surface is measured moves into the non-overlapping area as the vehicle moves, the image area where the height of the sidewalk surface is measured is tracked from the overlapping area to the non-overlapping area as the vehicle moves.

[0009] Basic driving scenes of the host vehicle include straight driving and turning. In a scene in which the host vehicle is driving straight, the image area in the overlapping area moves to the non-overlapping area as the vehicle moves, and the image area can be tracked from the overlapping area to the non-overlapping area as the vehicle moves. Therefore, by applying the technology described in Patent Document 1 mentioned above, it is possible to calculate the distance to a moving object detected in the monocular area even if the road surface heights in the overlapping area and the monocular area are different.

[0010] On the other hand, in a scene where the host vehicle is turning, for example, at an intersection, a moving object (e.g., a pedestrian, a bicycle) in the sidewalk area (i.e., including the difference in ground height from the roadway area) may suddenly appear (be photographed) in the monocular region. As an example, when the vehicle turns and begins to cross the intersection, a moving object that was hidden behind the object (outside the camera's field of view) may appear in the monocular region. As another example, in a scene where the vehicle is traveling straight before turning, a moving object detected in the stereo region moves into the monocular region as the vehicle moves, and then moves outside the monocular region (i.e., out of the camera's field of view). At this point, the moving object is located to the side or rear of the host vehicle and can no longer be tracked in the camera's field of view. Later, when the vehicle turns and begins to cross the intersection, the moving object that was out of the camera's field of view may reappear in the monocular region (returns into the camera's field of view) as the vehicle turns.

[0011] However, as described above, the device described in Patent Document 1 is based on the premise that the image area in the overlapping area moves to the non-overlapping area as the vehicle moves, and that the image area is tracked from the overlapping area to the non-overlapping area as the vehicle moves. Therefore, for example, in a scene where the vehicle turns at an intersection that may include a sidewalk area, in a scene where a moving object in the sidewalk area suddenly appears in the monocular area (non-overlapping area), it is difficult to calculate the accurate distance of the moving object detected in the monocular area.

[0012] In addition, in general, in a scene where the vehicle is traveling straight, changes in vehicle speed and changes in current road surface conditions are considered to be relatively small, and therefore the degree to which changes in vehicle speed and current road surface conditions affect the spatial relationship between the on-board camera that generates changes in the initially set sensor attitude parameters (camera attitude parameters) and the road surface is considered to be small.

[0013] On the other hand, in a scene where the vehicle is turning, for example at an intersection, it is thought that the changes in vehicle speed and the changes in the current road surface conditions are relatively large, and therefore it is thought that the changes in vehicle speed and the current road surface conditions will have a large influence on the spatial relationship between the on-board camera that generates changes in the initially set sensor attitude parameters (camera attitude parameters) and the road surface in accordance with the turning movement of the vehicle.

[0014] However, the device described in Patent Document 1 does not consider the influence of changes in vehicle speed and current road surface conditions on the spatial relationship between the road surface and the on-board camera that generates changes in the initially set sensor attitude parameters (camera attitude parameters). Therefore, for this reason, it may not be possible to accurately calculate the distance to a moving object detected in the monocular area, for example, in a scene where the vehicle turns at an intersection that may include a sidewalk area.

[0015] Therefore, for example, when the vehicle is turning at an intersection that may include a sidewalk area (i.e., including differences in ground height relative to the roadway area), and at the same time, when changes in vehicle speed and current road conditions affect the spatial relationship between the onboard sensor and the road surface, generating changes in the initially set sensor attitude parameters, if the system is used in a scenario that requires supporting accurate monocular-based distance measurements for detected obstacles, the system may require additional sensors to correctly set the sensor attitude parameters and accurately define where the obstacle of interest contacts the road surface (which may vary based on road surface shape and sidewalk type), thereby increasing the overall system cost. On the other hand, if external changes related to sensor attitude changes and possible changes in road surface shape (caused by sidewalks and other changes in height above the road surface the vehicle is traversing) are not supported and correctly addressed, obstacle detection performance and distance measurement accuracy may be degraded, generating erroneous obstacle detection results, and reducing the reliability of the overall system.

[0016] The object of the present invention is to provide an image processing device that stores three-dimensional information (hereinafter also referred to as road surface 3D information) related to the road surface (generated from image data in the stereo region), transforms / converts the stored road surface 3D information to the current time using vehicle movement (e.g., calculated over time using speed and yaw rate from CAN data), and uses the three-dimensional information from the road surface together with monocular-based distance estimation processing to improve distance measurement accuracy without requiring special sensors while maintaining obstacle detection accuracy, and in particular can support cases where the vehicle is turning at an intersection that may include sidewalk areas (i.e., including differences in ground height relative to the roadway area) that require accurate monocular-based distance measurement for detected obstacles. [Means for solving the problem]

[0017] In order to achieve the above-mentioned object, the image processing device according to the present invention comprises a road surface 3D information detection unit that detects road surface 3D information including the three-dimensional structure of the road surface based on images obtained from multiple cameras mounted on a vehicle; a memory unit that stores the road surface 3D information acquired in chronological order; a coordinate conversion unit that converts the road surface 3D information based on the position of the vehicle and the orientation of the camera at a first time (t-1) and a second time (t) based on the amount of movement of the vehicle between the first time (t-1) and the orientation of the camera at the second time (t) into road surface 3D information based on the position of the vehicle and the orientation of the camera at the second time (t); and a ranging unit that calculates the distance to an object in the field of view of one of the multiple cameras based on the converted road surface 3D information acquired at the first time (t-1) and the road surface 3D information acquired at the second time (t). [Effects of the Invention]

[0018] According to the present invention, it is possible to perform accurate monocular-based distance measurements for detected objects when the vehicle is turning at an intersection that may include a sidewalk area where there may be a difference in ground height relative to the roadway area, thereby improving the reliability of the overall system.

[0019] Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a schematic configuration diagram of an image processing device according to a first embodiment of the present invention. [Figure 2] This is a scenario in which a vehicle is approaching an intersection and distance measurement of a specific pedestrian on the sidewalk is required when the vehicle is turning at the intersection. (a) is the earliest time (ta) at which the vehicle is turning left at the intersection, (b) is the time (tb) at which the vehicle is turning left at the intersection, and (c) is the current time (t) at which the vehicle is turning left at the intersection. [Figure 3]10 is a flowchart showing the processing executed by the road surface 3D information management unit 141 to store and update the calculated road surface 3D information. [Figure 4] FIG. 2 is an explanatory diagram of the relationship in three-dimensional space between the X axis, Y axis, Z axis, tilt angle (pitch angle), and roll angle. [Figure 5] 10 is a flowchart showing processing executed by a road surface 3D information management unit 141 for storing, updating, and deleting calculated road surface 3D information in an image processing device according to a second embodiment of the present invention. [Figure 6] 11 is a flowchart showing processing executed by a road surface 3D information management unit 141 for storing, updating, and deleting calculated road surface 3D information in an image processing device according to a third embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0021] Hereinafter, preferred embodiments of the image processing apparatus of the present invention will be described with reference to the drawings.

[0022] [First embodiment] The configuration and performance of the image processing device 110 of this embodiment will be described below with reference to Figures 1 to 4. Although not shown in the figures, the image processing device 110 has a configuration in which a CPU, RAM, ROM, etc. are connected via a bus, and the CPU controls the operation of the entire system by executing various control programs stored in the ROM.

[0023] It should be noted that in the configuration described below, two camera sensors (hereinafter sometimes simply referred to as cameras or sensors) are paired as a single in-vehicle stereo camera, and therefore correspond to the sensing unit 111. However, this does not limit the devices that can be used in other configurations in which a single monocular camera is used as the sensing unit 111.

[0024] 1 is a block diagram showing the configuration of an image processing device according to a first embodiment of the present invention. The image processing device 110 of this embodiment is mounted on, for example, a vehicle (host vehicle), and performs image processing (for example, affine transformation) on an image of the surroundings captured by a camera sensor (sensing unit 111), detects and recognizes surrounding objects shown in the image, and measures (calculates) the distance to the surrounding objects.

[0025] In FIG. 1, the image processing device 110 includes a sensing unit 111 including two camera sensors arranged at the same height, an image acquisition unit 121, a road surface 3D information acquisition unit 131, a road surface 3D information management unit 141, an obstacle detection unit 151, an obstacle ranging unit 161, and an alarm and control application unit 171.

[0026] (Image acquisition section) The image acquisition unit 121 processes images acquired by one or both camera sensors corresponding to the sensing unit 111 to adjust image characteristics for further processing. This processing may include, but is not limited to, image resolution adjustment, which can shrink or enlarge an input image to change the resulting image size, image region-of-interest selection, which can cut out (crop) a specific region of the input image from the original input for further processing, and image affine transformations such as rotation, scale, shear, and top-view image transformation, in which a flat ground plane is considered the reference. For affine transformations, geometric formulas or transformation tables can be pre-calculated or pre-adjusted. Parameters used for image resolution adjustment and image region-of-interest selection can be controlled based on the current driving environment and conditions (speed, turning speed, etc.). Additionally, stereo matching is performed using image signals received from both camera sensors corresponding to the sensing unit 111 to create a 3D range image of the scene ahead of the vehicle equipped with the image processing device. Stereo matching determines the unit area with the smallest difference between the image signals for each predetermined unit area in the two images being compared. In other words, areas that depict the same subject are detected. This forms a three-dimensional distance image, hereafter referred to as a disparity image. Note that disparity data is used to calculate distance values ​​from the camera to an object in a real environment. Furthermore, a V-disparity image is created by projecting vertical disparity data onto coordinate position V, obtaining disparity values ​​along the horizontal axis, and integrating the two-dimensional space to form a histogram image representing the frequency of each disparity value. In other words, a V-disparity image is an image whose vertical direction is represented by the vertical coordinate position V and whose horizontal direction is represented by the disparity value at coordinate position V.

[0027] In the case of top-view image transformation, the image acquisition unit 121 also has the function of calculating a difference image representing the discrepancy between at least two images acquired at different times by the sensing unit 111 and transformed into a top-view image. Known methods can be applied for this difference calculation, including, but not limited to, simple pixel-to-pixel difference calculation and filter-based image difference calculation.

[0028] (Road surface 3D information acquisition department) The road surface 3D information acquisition unit 131 has a function of estimating the three-dimensional surface characteristics / shape (height, distance, slope, etc.) of the road surface and its surroundings (sidewalk / step / road edge) based on the parallax image created by the image acquisition unit 121. Here, the surface height refers to the height position of the surface / road surface relative to a predefined plane, and refers to the height position estimated for each corresponding range in the depth direction for a vehicle equipped with the image processing device.

[0029] An example of the process for estimating surface characteristics (in this case, road surface height) is as follows.

[0030] First, a range (region) of the disparity image where the road surface normally exists (i.e., the road surface candidate) is set. Methods for setting the region may include, among others, using a pre-estimated road surface position, using the camera's infinity as a reference for setting the region, using a defined disparity range as valid road surface data, setting a trapezoid shape to define the road surface area, and using detected lane markings to define the road surface shape. The next step is to perform extraction of disparity values ​​that are valid and included in the road surface candidate region. During data extraction, available road surface disparity data is separated from possible obstacle data. Methods for separating road surface disparity data from the remaining disparity values ​​include, but are not limited to, using pre-detected obstacle position data to estimate the position closest to the road surface, and comparing the vertical disparity value with a pre-determined threshold to ensure that the disparity value gradually decreases vertically to the camera's infinity.

[0031] The next step is to project the disparity data of road surface candidates into the V-disparity space (vertical coordinate position V, horizontal disparity value) and create a histogram image that shows the frequency of disparity values ​​for each road surface candidate.

[0032] From this, a representative value can be selected for each disparity value in the vertical direction based on the amount of data available in a given range, reducing the amount of data describing the road surface.

[0033] The next step is to estimate straight lines (surface lines) passing through pixel neighborhoods with high histogram frequencies in the representative V-disparity data. Note that the estimated lines can also be curves, a set of straight lines, or a set of curves.

[0034] In the next step, the road surface line estimated in the previous step is received as input, and the road surface height position is calculated.

[0035] (Road surface 3D information management department) The road surface 3D information management unit 141 has a function of storing (memorizing) the three-dimensional surface characteristics / shape estimated by the road surface 3D information acquisition unit 131 in the form of road surface 3D information.

[0036] Depending on the device configuration, the method for storing road surface 3D information can be controlled based on the device configuration and / or the current state of the vehicle in which the image processing device is installed, by controlling the amount of data stored in each processing cycle (e.g., half the depth range), the distance range of the road surface 3D information to be stored (e.g., up to 25 m from the vehicle), and the interval at which the storage process is performed (e.g., every 50 ms when the vehicle speed is high). Other methods / configurations for storing road surface 3D information data can also be implemented.

[0037] The road surface 3D information management unit 141 also has a function of updating road surface 3D information stored in a previous period to the current time by converting (translating / rotating) the road surface 3D information stored in the previous period to the current time using the vehicle's movement over time (amount of movement) and the camera attitude (orientation) at each time when the road surface 3D information is stored. In this way, the converted 3D information is related to the position of the vehicle on which the image processing device is installed, so all stored road surface 3D information can be used at the current time.

[0038] That is, the road surface 3D information management unit 141 includes a memory unit 181 that stores the road surface 3D information acquired in chronological order by the road surface 3D information acquisition unit 131 (road surface 3D information based on the vehicle position and camera orientation at the time of acquisition), and a coordinate conversion unit 191 that converts the road surface 3D information based on the vehicle position and camera orientation at a previous time into road surface 3D information based on the vehicle position and camera orientation at the current time, based on the amount of movement of the vehicle between the previous time and the current time and the camera orientation at the current time (details will be explained later).

[0039] (Obstacle detection section) The obstacle detection unit 151 has a function of detecting three-dimensional objects in the image and calculating their positions by using the image created by the image acquisition unit 121, and by using, but not limited to, a difference image and a known clustering method that takes into account the distance between points (pixels) (e.g., the K-means algorithm) to create clusters of difference pixels that are close to each other and likely represent target obstacles on the road surface. In this specification, "obstacle detection" refers to processing that performs at least the following tasks: target object detection (position in image space), target object identification (e.g., automobile / vehicle, motorcycle, bicycle, pedestrian, pole, etc.).

[0040] (Obstacle distance measurement unit) The obstacle distance measurement unit 161 has the function of measuring the distance from the vehicle (on which the image processing device is installed) to one or more obstacles detected by the obstacle detection unit 151 and obtaining distance measurement values ​​from the vehicle to the target object in three-dimensional space, which enables calculation of the speed / velocity of the target object, by combining a monocular-based method that relies on geometric calculations using camera attitude parameters with a method that uses road surface 3D information managed by the road surface 3D information management unit 141 (more specifically, converted road surface 3D information acquired at a previous time and road surface 3D information acquired at the current time). Examples of use include (but are not limited to) an average distance measurement from the results of both methods, a weighted average of distance measurements from the results of both methods (weights for the results of each method are determined in advance or adjusted in real time based on the amount of road surface 3D information available) (default weights are equal for each method), distance measurement using a monocular-based method with an error rate calculation using road surface 3D information as a measure of reliability, and distance measurement using a monocular-based road surface 3D information method with an error rate calculation using a monocular-based method as a measure of reliability.

[0041] Furthermore, depending on the device configuration and / or current system state (processing stall / give up / error associated with any of the methods), the above methods can be used individually as needed.

[0042] (Alarm and Control Applications Division) The warning and control application unit 171 has the function of determining a warning routine (auditory or visual warning message to the driver) or control application to be executed by the vehicle in which the image processing device is installed, based on the obstacles recognized by the obstacle detection unit 151 and the results obtained by the obstacle distance measurement unit 161.

[0043] Here, the case where the image processing device 110 is applied as a system for monitoring the surroundings of a vehicle V1 will be described with reference to FIGS. 2(a), 2(b), and 2(c). The upper parts of FIGS. 2(a), 2(b), and 2(c) show a top view of a scenario in which a vehicle V1 turns left at an intersection, while FIGS. 2(a), 2(b), and 2(c) show the same vehicle V1 moving toward turning left at the intersection at different time periods. FIG. 2(a) shows the earliest time (ta) at which the vehicle V1 turns left at the intersection, FIG. 2(c) shows the current time (t) at which the vehicle V1 turns left at the intersection, and FIG. 2(b) shows a time (tb) between the time (ta) and the time (t) at which the vehicle V1 turns left at the intersection. The lower parts of Figures 2(a), 2(b) and 2(c) show the same scenario as seen from images acquired by the image acquisition unit 121, where the stereo region SR1 refers to the region where the image signals received from both camera sensors corresponding to the sensing unit 111 overlap and can therefore be used to calculate a disparity image and obtain road surface 3D information calculated by the road surface 3D information acquisition unit 131, and the monocular region MR1 and monocular region MR2 refer to the region where the image signals received from both camera sensors corresponding to the sensing unit 111 do not overlap and therefore only monocular images are available that can be used for obstacle detection by the obstacle detection unit 151.

[0044] For the period (ta) shown in Figure 2(a), vehicle V1 is crossing forward toward the intersection, and at this time, pedestrian P1 on the sidewalk is moving in the same direction as vehicle V1. At time (ta), road surface 3D information RI01 is acquired from stereo region SR1 by road surface 3D information acquisition unit 131 (with pedestrian P1's data removed), and then stored by road surface 3D information management unit 141 for further processing. At this time, pedestrian P2 is hidden behind object Q1 (outside the camera's field of view and not visible in the image).

[0045] For the period (tb) shown in FIG. 2(b), vehicle V1 is still crossing forward toward the intersection, and at this time, pedestrian P1 on the sidewalk is also still moving in the same direction as vehicle V1. At time (tb), road surface 3D information RI02 is acquired from the stereo region SR1 by road surface 3D information acquisition unit 131 (data of pedestrian P1 is removed) and stored by road surface 3D information management unit 141 for further processing. Next, the already stored road surface 3D information RI01 is transformed (translated / rotated) to the current time (tb) using the vehicle's movement (amount of movement) from time (ta) to time (tb) and the camera attitudes (orientations) at times (ta) and (tb). After the road surface 3D information is updated, all road surface 3D information stored by the image processing device is in the same time and space as the road surface 3D information stored at time (tb) and should therefore be related to the position of vehicle V1. At this point, pedestrian P2 is also still hidden behind object Q1 (outside the camera's field of view and not visible in the image).

[0046] For a period (t) shown in FIG. 2(c), vehicle V1 is making a left turn at an intersection. At this time, a target pedestrian P2 on the sidewalk is detected by the obstacle detection unit 151 in the monocular region MR1. The pedestrian P1 on the sidewalk continues to move away from vehicle V1. At time (t), road surface 3D information RI03 is acquired from the stereo region SR1 by the road surface 3D information acquisition unit 131 and stored by the road surface 3D information management unit 141 for further processing. Then, the already stored road surface 3D information RI01 and RI02 is transformed (translated / rotated) to the current time (t) using the vehicle's movement (amount of movement) from time (tb) to time (t) and the camera poses (orientations) at times (tb) and (t). After the road surface 3D information is updated, all road surface 3D information stored by the image processing device is in the same time and space as the road surface 3D information stored at time (t) and therefore should be related to the position of vehicle V1. During time period (t), the stored road surface 3D information RI01, RI02, and RI03 can be used by the obstacle ranging unit 161 to obtain distance measurements for the target pedestrian P2 (detected by the obstacle detection unit 151 within the monocular region MR1) in conjunction with a monocular-based method that relies on geometric calculations using camera pose parameters. A reference position in space of the road surface 3D information for the target pedestrian P2 can be calculated based on its position relative to the vehicle V1, in the form of an angle from the center of the image processing device to the center of the target pedestrian P2, and the calculated position can be used to access corresponding distance data from the road surface 3D information. In this way, high-precision distance measurements for the pedestrian P2 on the sidewalk can be achieved even when there is a difference in ground height relative to the roadway area (driving surface).

[0047] 3 is a flowchart showing an exemplary process executed by the road surface 3D information management unit 141 using the output of the road surface 3D information acquisition unit 131. Note that the current time is represented as (t) and the previous time as (tn), where n is a number from 1 to N.

[0048] First, in step S1 (a step of storing road surface 3D information at the current time (t)), the 3D surface characteristics / shape corresponding to the current time (t) estimated by the road surface 3D information acquisition unit 131 are searched for and stored (in the memory unit 181).

[0049] Next, step S2 (estimating camera pose based on road surface 3D information) estimates camera pose parameters based on the stored 3D surface characteristics / shape corresponding to the current time (t), and stores the results for further processing. An example of camera pose estimation (but is not limited to) is to use the 3D surface to calculate the slope of the road surface in front of the vehicle V1 in the X and Y axes, and then use the resulting road surface slope to calculate the pitch (tilt) and roll angles of the image processing device relative to the 3D surface in front of the vehicle V1 (see also FIG. 4).

[0050] Next, step S3 (step for checking the existence of previous data) checks the existence of previously stored road surface 3D information. If there is no road surface 3D information other than that stored at the current time (t), no additional processing is required in this step. If previous road surface 3D information has been stored (for example, at time (tn)), processing proceeds to step S4.

[0051] Next, in step S4 (a step of converting road surface 3D information to the current time (t)), the road surface 3D information stored in the previous period is updated to the current time (t).

[0052] This update can be performed by applying an affine transformation (also see Figure 4) in the form of X-axis and Z-axis translation and Y-axis rotation (yaw angle) based on the vehicle motion (difference in X and Z positions) (movement amount) occurring from a previous time (tn) to the current time (t), supplemented by an X-axis rotation (tilt angle) and Z-axis rotation (roll angle) based on the camera pose parameters at the current time (t).

[0053] Finally, the process is complete and passes to the obstacle ranging unit 161 for further processing.

[0054] By employing the above process, all stored road surface 3D information can be used for distance measurement in relation to the position of the vehicle at the current time (t). Furthermore, the update process updates all stored road surface 3D data to the latest current time (t) at the end of the process, so that the resulting transformation in future periods does not need to go back to each period in which each road surface 3D information was stored.

[0055] The configuration and operation of the image processing device according to the first embodiment have been described above. The image processing device according to the first embodiment enables accurate monocular-based distance measurement to be performed on a detected obstacle when the host vehicle is turning at an intersection that may include a sidewalk area where there may be a difference in ground height relative to the roadway area (driving surface), thereby improving the reliability of the entire system.

[0056] [Second embodiment] Next, an image processing device according to a second embodiment of the present invention will be described.

[0057] The basic configuration of the image processing apparatus according to the second embodiment differs from that of the first embodiment in the following respects.

[0058] The road surface 3D information management unit 141 of the second embodiment has an additional function of deleting stored road surface 3D information, as shown in FIG.

[0059] The road surface 3D information deletion process can be performed before or after the road surface 3D information update process, but in this example, the deletion process is performed after the update process, as shown in the flowchart of Fig. 5. In the flowchart of Fig. 5, step S5 (a step of deleting road surface 3D information) deletes (from the storage unit 181) the stored road surface 3D information to be deleted based on the device configuration. Examples of such a configuration in which road surface 3D information is to be deleted include (but are not limited to) road surface 3D information located a specific distance behind the vehicle (e.g., 30 meters behind the vehicle) and road surface 3D information stored at a specific time (e.g., 5 minutes ago) before the corresponding current time (t).

[0060] That is, the road surface 3D information management unit 141 (coordinate conversion unit 191) of this second embodiment deletes the road surface 3D information of the converted road surface 3D information (which may be the road surface 3D information before conversion) of the part that the vehicle has passed through at the current time (corresponding to a specific distance before the position at the current time, a specific time before the current time, etc.) (step S5).

[0061] By employing the above processing, it is possible to reduce the memory size required for the road surface 3D information stored in the image processing device while maintaining the same functional operation as in the first embodiment.

[0062] The configuration and operation of the image processing device according to the second embodiment have been described above. The image processing device according to the second embodiment enables accurate monocular-based distance measurement to be performed on detected obstacles when the host vehicle is turning at an intersection that may include a sidewalk area where there may be a difference in ground height relative to the roadway area (driving surface), improving the reliability of the entire system, reducing the amount of memory required for stored road surface 3D information, and maintaining the same functional operation as the first embodiment.

[0063] [Third embodiment] Next, an image processing device according to a third embodiment of the present invention will be described.

[0064] The basic configuration of the image processing apparatus according to the third embodiment differs from that of the first embodiment in the following respects.

[0065] The road surface 3D information management unit 141 in this third embodiment has the additional function of deleting stored road surface 3D information, and the road surface 3D information update process is divided into two steps (S4A, S4B) as shown in Figure 6.

[0066] Step S4A of the road surface 3D information update process updates the road surface 3D information stored in a previous time period to the current time (t) by applying an affine transformation in the form of (but not limited to) translation of the X and Z axes and rotation of the Y axis (yaw angle) based on the use of the vehicle movement (difference of X position and Z position) (movement amount) occurring from a previous time (tn) to the current time (t).

[0067] Step S4B of the road surface 3D information update process updates the road surface 3D information stored in the previous period (which remains in memory unit 181 after the deletion process) to the current time (t) by applying an affine transformation that complements the update performed by step S4A in the form of (but not limited to) an X-axis rotation (tilt angle) and a Z-axis rotation (roll angle) based on the camera attitude parameters at the current time (t).

[0068] The road surface 3D information deletion process is performed between step S4A of the road surface 3D information update process and step S4B of the road surface 3D information update process, as shown in the flowchart of Fig. 6. In the flowchart of Fig. 6, step S5 (a step of deleting road surface 3D information) deletes (from the storage unit 181) the stored road surface 3D information to be deleted based on the device configuration. Examples of such a configuration in which road surface 3D information is to be deleted include (but are not limited to) road surface 3D information located a specific distance behind the vehicle (e.g., 30 meters behind the vehicle) and road surface 3D information stored at a specific time (e.g., 5 minutes ago) before the corresponding current time (t).

[0069] That is, the road surface 3D information management unit 141 (coordinate conversion unit 191) of this third embodiment converts road surface 3D information based on the vehicle's position at a previous time (first time) into road surface 3D information based on the vehicle's position at a current time (step S4A) based on the amount of movement of the vehicle between a previous time (first time) and a current time (second time), deletes road surface 3D information for a portion of the converted road surface 3D information that the vehicle has passed through at the current time (corresponding to a specific distance before the position at the current time, a specific time before the current time, etc.) (step S5), and corrects the coordinate space of the road surface 3D information for a portion of the converted road surface 3D information that the vehicle has not passed through at the current time based on the camera orientation at the current time (step S4B).

[0070] By adopting the above processing, it is possible to reduce the amount of memory required for the road surface 3D information stored in the image processing device while maintaining the same functional operation as in the first embodiment, and to shorten the processing time required to update the road surface 3D information that remains stored in memory.

[0071] The configuration and operation of the image processing device according to the third embodiment are described above. The image processing device according to the third embodiment enables accurate monocular-based distance measurement to be performed on detected obstacles when the host vehicle is turning at an intersection that may include a sidewalk area where there may be a difference in ground height relative to the roadway area (driving road surface), improves the reliability of the overall system, reduces the amount of memory required for stored road surface 3D information, shortens the processing time required to update the road surface 3D information that remains stored in memory, and maintains the same functional operation as the first embodiment.

[0072] As described above, the image processing device 110 for obstacle detection and obstacle recognition according to this embodiment, for example as shown in FIG. 1, includes the following: A sensing unit 111 consisting of two sensing units (cameras) capable of capturing images of the scene in front of the device to which the apparatus is attached: an image acquisition unit 121 that processes the images acquired by the sensing unit 111 to adjust their characteristics (including but not limited to image size, image resolution, and image region of interest), performs 3D data generation to match the images acquired from both sensing units, and calculates the disparity for each pixel; A road surface 3D information acquisition unit 131 executes road surface shape estimation using the parallax information calculated by the image acquisition unit 121 and acquires road surface characteristics (height, distance, etc.) for each range in the depth direction. A road surface 3D information management unit 141 that executes a function (storage unit 181) of storing the road surface 3D information calculated by the road surface 3D information acquisition unit 131, and also executes a function (coordinate conversion unit 191) of updating the road surface 3D information by using the vehicle movement (movement amount) over time and transforming / converting the previously stored road surface 3D information to the current time: An obstacle detection unit 151 that performs object detection and object recognition using the images acquired by the image acquisition unit 121: an obstacle ranging unit 161 that performs monocular-based 3D distance measurement from the host vehicle to an obstacle detected by the obstacle detection unit 151 using a combination of geometric calculations using camera pose parameters and the use of available road surface 3D information updated at the current time by the road surface 3D information management unit 141; an alarm and control application unit 171 that determines an alarm routine or control application to be executed by the device to which the apparatus is attached based on current conditions, which may include outputs from at least the obstacle detection unit 151 and the obstacle ranging unit 161;

[0073] That is, the image processing device 110 according to this embodiment includes a road surface 3D information detection unit (road surface 3D information acquisition unit 131) that detects road surface 3D information including the three-dimensional structure of the road surface based on images obtained from a plurality of cameras mounted on a vehicle, a storage unit 181 that stores the road surface 3D information acquired in time series, and a storage unit 182 that calculates the position and time of the vehicle at a first time (t-1) based on the amount of movement of the vehicle between the first time (t) and a second time (t) and the orientation (attitude: tilt angle, roll angle) of the camera at the second time (t). a coordinate conversion unit 191 that converts (translates / rotates) road surface 3D information based on the position of the vehicle at the second time (t) and the orientation of the camera into road surface 3D information based on the position of the vehicle at the second time (t) and the orientation of the camera; and a distance measurement unit (obstacle distance measurement unit 161) that determines the distance to an object in the field of view (monocular area) of one of the multiple cameras based on the converted road surface 3D information acquired at the first time (t-1) and the road surface 3D information acquired at the second time (t).

[0074] In addition, based on the amount of movement of the vehicle between the first time (t-1) and the second time (t), the coordinate conversion unit 191 converts road surface 3D information based on the position of the vehicle at the first time (t-1) into road surface 3D information based on the position of the vehicle at the second time (t), deletes from the converted road surface 3D information the road surface 3D information of the portion through which the vehicle passed at the second time (t), and corrects the coordinate space of the road surface 3D information of the converted road surface 3D information of the portion through which the vehicle did not pass at the second time (t) based on the orientation of the camera at the second time (t).

[0075] By adopting this configuration, the obstacle ranging unit 161 performs monocular-based distance estimation processing, and by using road surface 3D information stored via the road surface 3D information acquisition unit 131 and converted to the current time by the road surface 3D information management unit 141, distance measurements can be performed reliably without adding a dedicated sensor, obstacle detection accuracy can be maintained, and support can be added for cases where the vehicle is turning at an intersection that may include a sidewalk area.

[0076] According to this embodiment, it is possible to perform accurate monocular-based distance measurements for detected objects when the vehicle is turning at an intersection that may include a sidewalk area where there may be a difference in ground height relative to the roadway area, thereby improving the reliability of the entire system.

[0077] While there have been described what are presently considered to be preferred embodiments of the invention, various modifications can be made thereto, and all modifications which come within the true spirit and scope of the invention are intended to be within the scope of the appended claims.

[0078] For example, in the above-described embodiment, a camera (a stereo camera using multiple cameras) is used as a sensor that monitors the surroundings of the vehicle and detects 3D road surface information including the three-dimensional structure of the road surface. However, a sensor such as a millimeter-wave radar or a laser radar may be used together with or instead of the camera. In other words, this embodiment may also be configured to use a monocular camera in combination with a sensor such as a millimeter-wave radar or a laser radar. Here, for example, the field of view (range) of the monocular camera may be wide-angle and exceed the measurement range (range) of the sensor. That is, the image processing device 110 according to this embodiment may be configured to include a road surface 3D information detection unit that detects road surface 3D information including the three-dimensional structure of the road surface based on information obtained from a sensor mounted on the vehicle; a memory unit that stores the road surface 3D information acquired in chronological order; a coordinate conversion unit that converts the road surface 3D information based on the position of the vehicle and the orientation of the sensor at a first time (t-1) and a second time (t) based on the amount of movement of the vehicle between the first time (t-1) and the orientation of the sensor at the second time (t) into road surface 3D information based on the position of the vehicle and the orientation of the sensor at the second time (t); and a ranging unit that calculates the distance to an object in the field of view of one camera mounted on the vehicle based on the converted road surface 3D information acquired at the first time (t-1) and the road surface 3D information acquired at the second time (t).

[0079] Furthermore, the present invention is not limited to the above-described embodiment, and includes various modifications. For example, the above-described embodiment has been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to an embodiment having all of the described configurations.

[0080] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a storage device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.

[0081] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]

[0082] 110 Image processing device 111 Sensing unit 121 Image acquisition unit 131 Road surface 3D information acquisition unit (road surface 3D information detection unit) 141 Road surface 3D information management department 151 Obstacle detection unit 161 Obstacle distance measuring section (distance measuring section) 171 Alarm and Control Applications Section 181 Storage section 191 Coordinate conversion section

Claims

1. a road surface 3D information detection unit that detects road surface 3D information including a three-dimensional structure of the road surface based on images obtained from a plurality of cameras mounted on the vehicle; a storage unit that stores the road surface 3D information acquired in time series; a coordinate conversion unit that converts road surface 3D information based on the position of the vehicle and the orientation of the camera at the first time into road surface 3D information based on the position of the vehicle and the orientation of the camera at the second time (t), based on the amount of movement of the vehicle between a first time (t-1) and a second time (t) and the orientation of the camera at the second time (t); An image processing device characterized by having a ranging unit that calculates the distance to an object in the field of view of one of the multiple cameras based on the converted road surface 3D information acquired at the first time (t-1) and the road surface 3D information acquired at the second time (t).

2. 2. The image processing device according to claim 1, The image processing device is characterized in that the coordinate conversion unit converts road surface 3D information based on the position of the vehicle at the first time (t-1) into road surface 3D information based on the position of the vehicle at the second time (t) based on the amount of movement of the vehicle between the first time (t-1) and the second time (t), deletes from the converted road surface 3D information the road surface 3D information of a portion through which the vehicle has passed at the second time (t), and corrects the coordinate space of the road surface 3D information of a portion through which the vehicle has not passed at the second time (t) based on the orientation of the camera at the second time (t).

3. 2. The image processing device according to claim 1, The image processing device is characterized in that the coordinate transformation unit deletes the road surface 3D information of the portion through which the vehicle passed at the second time (t) from the transformed road surface 3D information or the road surface 3D information before the transformation.

4. 2. The image processing device according to claim 1, The distance measurement unit calculates the distance to the object based on the converted road surface 3D information acquired at the first time (t-1), the road surface 3D information acquired at the second time (t), and the position detection result on the image of the object in the field of view of one of the multiple cameras.

5. a road surface 3D information detection unit that detects road surface 3D information including a three-dimensional structure of the road surface based on information obtained from a sensor mounted on the vehicle; a storage unit that stores the road surface 3D information acquired in time series; a coordinate conversion unit that converts road surface 3D information based on the position of the vehicle and the orientation of the sensor at the first time point into road surface 3D information based on the position of the vehicle and the orientation of the sensor at the second time point (t), based on the amount of movement of the vehicle between a first time point (t-1) and a second time point (t) and the orientation of the sensor at the second time point (t); An image processing device characterized by having a ranging unit that calculates the distance to an object in the field of view of one camera mounted on the vehicle based on the converted road surface 3D information acquired at the first time (t-1) and the road surface 3D information acquired at the second time (t).

Citation Information

Patent Citations

  • Image processing system

    JP2009085651A

  • Stereo camera system

    JP2017096777A

  • Stereo camera device

    WO2017090410A1

  • Object recognition device

    WO2021070537A1

  • Camera system

    WO2021124657A1