Image processing device and image processing method
A multi-camera system with advanced image processing generates seamless and accurate 360° point clouds, addressing the high cost of LiDAR by improving autonomous driving and ADAS systems with enhanced obstacle detection.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-02
AI Technical Summary
Existing autonomous driving technologies rely on expensive LiDAR for three-dimensional information acquisition, necessitating a low-cost and high-precision alternative.
A multi-camera system with a specific camera arrangement and image processing apparatus that generates and integrates point clouds using multiple cameras, incorporates noise reduction, vehicle motion estimation, and reliability calculation to create seamless and accurate 360° point cloud data around a vehicle.
Enables high-precision point cloud generation at a lower cost, improving autonomous driving and advanced driver assistance systems by enhancing accuracy and reliability of obstacle detection.
Smart Images

Figure JP2024034941_02042026_PF_FP_ABST
Abstract
Description
Image Processing Apparatus and Image Processing Method
[0001] The present invention relates to an image processing apparatus.
[0002] In order to reduce the driving burden of drivers and reduce traffic accidents, advanced driver assistance systems (ADAS: Advanced Driver Assistance Systems) at Level 1 and Level 2 that partially automate one or both of the operations of the accelerator, brake, and steering wheel operation have been developed and put into practical use. Furthermore, for the purpose of assisting the movement of the elderly and others, addressing the shortage of drivers in the transportation industry, and improving productivity, the development of autonomous driving technologies (AD: Autonomous Driving) at Level 3 and Level 4 is underway. A three-dimensional information acquisition device is essential for autonomous driving. As an essential three-dimensional information acquisition device for autonomous driving, mainly expensive LiDAR is used in existing technologies.
[0003] As the background art in this technical field, there is the following prior art. For example, in Patent Document 1 (Japanese Patent Application Laid-Open No. 2017-96777), a parallax information acquisition unit that acquires parallax information of an overlapping region of a plurality of images captured by a plurality of cameras mounted on a vehicle, and based on the parallax information acquired in the past in the overlapping region, calculates the distance between an object detected in a non-overlapping region other than the overlapping region in each image and the vehicle, and a stereo camera device characterized by including an object distance calculation unit are described.
[0004] Japanese Patent Application Laid-Open No. 2017-96777
[0005] There is a need for a low-cost and high-precision point cloud generation technology using a multi-camera system instead of LiDAR, which is a three-dimensional information acquisition device in existing technologies.
[0006] A representative example of the invention disclosed in this application is as follows: an image processing apparatus for processing images captured by multiple cameras, comprising: a point cloud generation unit that acquires multiple images captured by the multiple cameras while a vehicle is in motion, generates a disparity image from pairs of images whose field of view overlaps from the acquired multiple images, and converts the generated disparity image into a point cloud in the camera coordinate system; a point cloud integrating unit that converts the converted point cloud from the camera coordinate system to the vehicle coordinate system and integrates the point cloud in the vehicle coordinate system; and an estimation unit that converts the integrated point cloud at past time points using the momentum of the vehicle estimated from the integrated point cloud and generates an interpolated point cloud to add the converted point cloud to the point cloud at the most recent time point.
[0007] According to one aspect of the present invention, point cloud data of objects around a vehicle can be generated using a low-cost multi-camera system. Problems, configurations, and effects other than those described above will be clarified by the following description of the embodiments.
[0008] This figure shows an example of the camera arrangement of the camera system according to an embodiment of the present invention. This figure shows the stereo viewing area formed by the camera system according to an embodiment of the present invention. This figure shows an example of the hardware configuration of the image processing device according to an embodiment of the present invention. This figure shows an example of the logical configuration of the image processing device according to an embodiment of the present invention.
[0009] Figure 1 is a diagram showing an example of the camera arrangement of a camera system according to an embodiment of the present invention, where Figure 1(A) is a front view showing the front of the vehicle, Figure 1(B) is a rear view showing the rear of the vehicle, Figure 1(C) is a right side view showing the right side of the vehicle, and Figure 1(D) is a top view showing the top of the vehicle. Figure 2 is a diagram showing the stereo viewing area formed by the seven-camera system shown in Figure 1.
[0010] In embodiments of the present invention, the basic concept is to configure the camera system as follows. In this basic concept, firstly, as shown in Figure 1, a plurality of cameras C (CF, CRF, CRR, CLF, CLR, CB) are attached to the periphery of the vehicle 100 where reflections of the vehicle body are minimal. In this case, the camera C is a monocular camera consisting of a lens and one image sensor, or a camera consisting of a curved mirror such as a sphere, hyperboloid, cone, or paraboloid and one image sensor, and the periphery of the vehicle 100 to which the camera C is attached may be either inside or outside the vehicle. However, it is preferable to place the camera in a position where reflections of the body are minimal.
[0011] Secondly, as shown in Figure 2, multiple monocular cameras C are combined to form a stereo field of view around the entire perimeter. Each camera is positioned so that its shooting range (field of view) covers the entire perimeter of the vehicle 100, so that there are no gaps in the stereo field of view that exceed the width of the vehicle. In other words, the cameras C are positioned so that the edges of each stereo field of view touch or overlap, as shown in Figure 2.
[0012] Thirdly, as shown in Figure 2, it is desirable that the angle difference in the optical axis direction between the cameras forming the stereo pair be less than 90 degrees. This suppresses the decrease in distance measurement accuracy caused by the difference in optical axis angle.
[0013] Furthermore, the configuration example shown in Figure 1 is a camera system including seven cameras, with two forward cameras CF that capture images in front with a field of view of 120 degrees, a left front camera CLF with a field of view of 90 degrees and a right front camera CRF, a left rear camera CLR with a field of view of 120 degrees and a right rear camera CRR, and two rear cameras CB that capture images behind with a field of view of 60 degrees. The left and right cameras (CRF, CRR, CLF, CLR) should be positioned to capture images in diagonal directions in front of and behind the sides. With this camera arrangement, each direction (front, back, left, and right) is captured by two or more monocular cameras, enabling stereo vision in the overlapping field of view areas, and the distance to an object can be calculated from the parallax. Note that the camera arrangement does not have to be as described above, the field of view of the cameras can be widened (for example, with a fisheye lens), and the number of cameras may be more or fewer than that of the configuration shown in Figure 1, as long as a stereo vision area can be set up all around.
[0014] Furthermore, regarding the convention for the symbols assigned to the symbol C indicating a camera, the symbols F, B, R, and L, indicating front, back, left, and right, shall be added after the camera C. However, the front / back symbol shall be added after the left / right symbol. Therefore, for left and right cameras, those with the optical axis direction diagonally forward are distinguished by F, and those with the optical axis direction diagonally backward are distinguished by B. Also, when multiple cameras of the same type cannot be distinguished as described above, numbers shall be used as necessary. The same convention for assigning symbols also applies to the symbol RE, which indicates the field of view.
[0015] Following this convention, the "direction" of the optical axis and area of a camera in this specification shall be defined first in terms of front, rear, left, and right, as in front camera, rear camera, right camera, and left camera. However, left and right cameras may be collectively referred to as side cameras. Also, in front, rear, left, and right, left and right are described after front and rear. Therefore, left and right cameras whose optical axis is forward are called side-front cameras (right front camera, left front camera), and left and right cameras whose optical axis is rear are called side-rear cameras (right rear camera, left rear camera).
[0016] As shown in Figure 2, two 120-degree front cameras CF form a 120-degree stereo region REF in front, a 90-degree stereo region REL to the left side is formed by the combination of the left front camera CLF and the left rear camera CLR, a 90-degree stereo region RER to the right side is formed by the combination of the right front camera CRF and the right rear camera CRR, and a 60-degree (30 degrees and 30 degrees) stereo region REB to the rear is formed by the combination of the rear camera CB and the left rear camera CLB, and the rear camera CB and the right rear camera CRB. As a result, the all-around stereo region shown in Figure 2 is formed so that the edges of each stereo region touch or overlap, so there are no gaps larger than the width of the vehicle, and a seamless, continuous stereo region is formed around the entire circumference of the vehicle.
[0017] Figure 3 shows an example of the hardware configuration of the image processing apparatus 1 according to an embodiment of the present invention.
[0018] The image processing device 1 is an electronic control unit (ECU) having a communication device 3, an input / output device 4, a CPU (Central Processing Unit) 5, a memory 6, and a storage device 7. The communication device 3, the input / output device 4, the CPU 5, the memory 6, and the storage device 7 are connected to each other via an internal signal line 9 such as a bus.
[0019] The communication device 3 connects to other control devices via a network. The input / output device 4 is an interface for inputting and outputting data to the image processing device 1. The CPU 5 is an arithmetic unit that executes programs stored in the memory 6. By executing predetermined arithmetic processes, the CPU 5 operates as a functional block that provides various functions of the image processing device 1 (for example, a point cloud generation unit 10, a noise reduction unit 20, a point cloud integration unit 30, a vehicle motion estimation unit 40, and a reliability calculation unit 50). The memory 6 is a storage unit that has a volatile storage area for temporarily storing data used by the CPU 5 when executing programs. The storage device 7 is accessible by the CPU 5 and has a non-volatile storage area that includes a program area for storing programs executed by the CPU 5 and a data area for storing data used by the CPU 5 when executing programs.
[0020] Figure 4 shows an example of the logical configuration of the image processing apparatus 1 according to the present invention.
[0021] The image processing device 1 includes a point cloud generation unit 10, a noise reduction unit 20, a point cloud integration unit 30, a vehicle motion estimation unit 40, and a reliability calculation unit 50. Of these functional blocks, the point cloud generation unit 10, the noise reduction unit 20, and the point cloud integration unit 30 can be broadly classified into an automatic calibration point cloud integration unit, while the vehicle motion estimation unit 40 and the reliability calculation unit 50 can be broadly classified into a non-stereoscopic area point cloud interpolation processing unit.
[0022] The point cloud generation unit 10 treats two cameras with overlapping field of view as a single camera pair, generates a disparity image from time-stamped images (for example, moving images with known capture times for each frame) input from the two cameras constituting the camera pair (11, 12), converts the generated disparity image into a point cloud having three-dimensional coordinates in the camera coordinate system, and generates point clouds for multiple camera pairs (13, 14). If a continuous stereo image capturing the entire perimeter of a vehicle, as shown in Figure 2, is input to the point cloud generation unit 10, a highly accurate 360° point cloud of the entire perimeter of the vehicle can be generated. However, a stereo image capturing only a portion of the vehicle's perimeter may also be input to the point cloud generation unit 10 to generate a point cloud of that portion.
[0023] The disparity image generated by the point cloud generation unit 10 is a so-called disparity map, and the intensity of the color indicates the distance from the point of capture.
[0024] Furthermore, the following methods can be used to generate disparity images. First, an AI model for generating disparity images based on cost volume is used to estimate the disparity of each pixel in the overlapping (non-stereoscopic) region of images from two cameras with overlapping field of view. For example, using an AI model for generating disparity images, the following procedure can be followed. Step 1: A two-dimensional CNN (Convolutional Neural Network) is used to extract image features from time-stamped images input from two cameras. The two-dimensional CNN used for extracting image features is trained on the image and its high-dimensional image features, and uses deep learning to output the high-dimensional image features of the image as input. Step 2: For possible disparity values, the features of the two images calculated in Step 1 are concatenated to generate a four-dimensional cost volume consisting of height × width × disparity × features. Step 3: The three dimensions of height × width × disparity of the four-dimensional cost volume generated in Step 2 are normalized using 3D(de)convolution to generate a disparity distribution for each pixel.
[0025] Furthermore, point cloud generation from disparity images is performed using an algorithm based on the following formula, generating point clouds from the disparity image, focal length, pixel size, baseline length, and optical center.
[0026]
[0027] The point cloud generated by the point cloud generation unit 10 is input to the noise reduction unit 20.
[0028] The distance error of points in the point cloud generated by the point cloud generation unit 10 from parallax follows a squared relationship, proportional to the square of the actual distance of the point from the viewpoint. The noise reduction unit 20 clusters the point cloud, which is arranged in three-dimensional space using coordinate values, and estimates and removes points belonging to clusters with a number of points below a threshold as noise (21, 22). Since the noise reduction unit 20 removes points estimated to be noise, the accuracy of the point cloud can be further improved.
[0029] The noise reduction unit 20 places points in a logarithmic three-dimensional space using coordinate values, clusters the placed point cloud, and estimates and removes points belonging to clusters with a number of points below a threshold as noise. Due to the aforementioned squared relationship, general clustering methods are not accurate when applied to point clouds converted from disparity images. The logarithmic axis reduces the squared relationship between the distance error of points and the actual distance of those points, allowing for appropriate noise removal.
[0030] The point cloud from which the noise reduction unit 20 has removed noise is input to the point cloud integration unit 30.
[0031] The point cloud integration unit 30 integrates multiple point clouds generated from the parallax of multiple camera pairs into a unified vehicle coordinate system.
[0032] First, the point cloud integration unit 30 extracts image features from the timestamped images input from the two cameras (31). For image feature extraction, known methods such as SIFT (Scale-Invariant Feature Transform), ORB (Oriented FAST and Rotated BRIEF), and BRIEF (Binary Robust Independent Elementary Features) can be used.
[0033] The point cloud integration unit 30 then extracts point cloud feature points (32, 33). For example, known methods can be used to extract point cloud feature points, such as extracting plane points and edge points as point cloud feature points.
[0034] The point cloud integration unit 30 estimates a transformation matrix between the camera coordinate system and the vehicle coordinate system based on the extracted image feature points, the extracted point cloud feature points, and the prior probability of the transformation matrix (34). First, the point cloud integration unit 30 associates the image feature points of timestamped images input from the two cameras by feature point matching, creates image feature point pairs, and calculates the distance between image feature points for each image feature point pair. The point cloud integration unit 30 also associates the point cloud feature points of multiple point clouds generated from multiple disparity images by feature point matching, creates point cloud feature point pairs, and calculates the distance between point cloud feature points for each point cloud feature point pair. Then, using the distance between image feature points in the two-dimensional image, the distance between point cloud feature points in three-dimensional space, and the prior probability of the transformation matrix, the unit optimizes the following objective function to calculate a transformation matrix with a high posterior probability. The prior probability distribution of the transformation matrix is a multivariate normal distribution with the design value of the camera mounting position as the expected value and half of the allowable range of error in the mounting position as the standard deviation. Feature point matching can be performed using known methods such as ICP (Iterative Closest Point).
[0035]
[0036] The point cloud integration unit 30 then uses the calculated transformation matrix to convert the coordinates of the point cloud generated from the parallax of multiple camera pairs from the camera coordinate system to the vehicle coordinate system, and integrates the point cloud time by time in the stereo viewing area of the vehicle coordinate system (35). The point cloud integrated by the point cloud integration unit 30 is input to the vehicle motion estimation unit 40.
[0037] The vehicle motion estimation unit 40 estimates the momentum of the vehicle and interpolates the point cloud using the estimated momentum.
[0038] For example, the vehicle motion estimation unit 40 matches the integrated point cloud from past time periods with the integrated point cloud from the most recent time period (e.g., the current time period) to associate point cloud feature points, estimates the relative momentum of the vehicle between past and latest time periods, and estimates a transformation matrix between the vehicle coordinate system at the most recent time period and the vehicle coordinate system at past time periods (41). For point cloud matching, known methods such as ICP (Iterative Closest Point) can be used.
[0039] Then, the vehicle motion estimation unit 40 uses the estimated transformation matrix to transform the coordinates of the past point cloud into the coordinates at the latest time (42).
[0040] Then, the vehicle motion estimation unit 40 overlays the past point cloud transformed using the transformation matrix representing the momentum on the point cloud at the latest time to generate a point cloud to be interpolated (43). For example, when there are non-stereoscopic vision areas around the vehicle that are hidden by other objects and are impossible to be seen stereoscopically, or non-stereoscopic vision areas near the host vehicle where it is impossible to take pictures with a plurality of cameras, the point cloud may be interpolated in the non-stereoscopic vision area. Compared with the method of estimating vehicle motion based on the moving distance of the vehicle by the vehicle speed pulse and the direction by the steering angle proposed in Patent Document 1 (4 degrees of freedom of x, y, z, and yaw), the vehicle motion estimation unit 40 can estimate the vehicle momentum including 6 degrees of freedom of x, y, z, roll, pitch, and yaw, and interpolate the point cloud in the non-stereoscopic vision area based on the estimated vehicle momentum, so that the point cloud accuracy can be further improved.
[0041] The point cloud interpolated by the vehicle motion estimation unit 40 is input to the reliability calculation unit 50.
[0042] The reliability calculation unit 50 calculates the reliability of the points to be interpolated, and adopts the points whose calculated reliability is greater than or equal to the threshold value to generate a full-area point cloud.
[0043] For example, the reliability calculation unit 50 calculates the reliability of the points to be interpolated (51). Generally, the error of the estimated parallax value of a distant object is larger than the error of the estimated parallax value of a nearby object. Therefore, taking the reciprocal of the distance between the point to be interpolated and the origin of the vehicle coordinate system at the generation time as the reliability, the reliability N is calculated using the following formula. In the following formula, p is the coordinate of the point to be interpolated, C is the set of all points to be interpolated (regardless of reliability), N is the set of reliabilities of the points to be interpolated, and R is the set of points to be interpolated with high reliability.
[0044]
[0045] Then, the reliability calculation unit 50 adds points with high reliability among the points to be interpolated, that is, points whose reliability is greater than a predetermined threshold value when comparing the reliability of the points to be interpolated with the predetermined threshold value, to generate an entire region point group (52). By selecting points to be interpolated using reliability, it is possible to solve the so-called duplication problem in which a plurality of points representing one point of an object are arranged in a three-dimensional space. As a result, low-reliability targets can be removed and misrecognition can be suppressed.
[0046] Then, the reliability calculation unit 50 stores the generated point group in the memory 6 (53). The point group generated by the reliability calculation unit 50 can be acquired by the vehicle motion estimation unit 40 as a point group at a past time via the memory 6.
[0047] As described above, according to the image processing apparatus 1 of the embodiment of the present invention, even a low-cost multi-camera system can generate high-precision point group data of objects around the vehicle. In addition, since the integrated point group at the current time is interpolated from the integrated point group observed from the vehicle at a past time, high-precision point groups can be generated even in a non-stereoscopic viewing area. Further, since points to be interpolated are selected according to the reliability calculated based on the distance from the observation position at a past time, high-precision point group data can be generated in both the stereoscopic viewing area and the non-stereoscopic viewing area. For this reason, point group data of obstacles such as speed bumps, triangular cones, and curbs around the vehicle can be generated by an inexpensive multi-camera system, and the AD function and the ADAS function can be improved. For example, by using accurate point group data, the possibility of contact with obstacles can be reduced in automatic parking, and the suspension can be accurately controlled when crossing a speed bump.
[0048] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Also, the configuration of another embodiment may be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, addition, deletion, or replacement with another configuration may be performed.
[0049] Furthermore, each of the aforementioned configurations, functions, processing units, and processing means may be implemented in hardware, for example, by designing them as integrated circuits, or they may be implemented in software by having a processor interpret and execute programs that realize each function.
[0050] Information such as programs, tables, and files that implement each function can be stored in memory, hard disks, SSDs (Solid State Drives), or recording media such as IC cards, SD cards, and DVDs.
[0051] Furthermore, the control lines and information lines shown are those deemed necessary for explanation purposes and do not necessarily represent all control lines and information lines required for implementation. In reality, it can be assumed that almost all components are interconnected.
Claims
1. An image processing apparatus for processing images captured by multiple cameras, comprising: a point cloud generation unit that acquires multiple images captured by the multiple cameras while a vehicle is in motion, generates a disparity image from the acquired multiple images by pairs of images whose field of view overlaps, and converts the generated disparity image into a point cloud in the camera coordinate system; a point cloud integration unit that converts the converted point cloud from the camera coordinate system to the vehicle coordinate system and integrates the point cloud in the vehicle coordinate system; and an estimation unit that converts the integrated point cloud at past time points using the momentum of the vehicle estimated from the integrated point cloud and generates an interpolated point cloud to add the converted point cloud to the point cloud at the most recent time point.
2. An image processing apparatus according to claim 1, comprising a reliability calculation unit that calculates the reliability of points in the point cloud converted by the estimation unit and generates a point cloud by adding the points with the highest calculated reliability to the point cloud at the most recent time.
3. An image processing apparatus according to claim 1, characterized in that the plurality of cameras are arranged such that a stereoscopic viewable area is formed around the entire perimeter of the vehicle.
4. An image processing apparatus according to claim 2, characterized in that it has a storage unit for storing point cloud data generated by the reliability calculation unit.
5. An image processing apparatus according to claim 4, wherein the estimation unit acquires an integrated point cloud from the storage unit at the past time.
6. An image processing apparatus according to claim 2, wherein the confidence calculation unit calculates the confidence level of the point cloud converted by the estimation unit, and generates a point cloud by adding points whose calculated confidence level is greater than a predetermined threshold.
7. An image processing apparatus according to claim 1, characterized by comprising a noise reduction unit that removes points estimated to be noise from the converted point cloud of the camera coordinate system.
8. An image processing apparatus according to claim 7, wherein the point cloud integration unit extracts feature points of the point cloud in the camera coordinate system from a disparity image from which points estimated to be noise have been removed, associates the extracted feature points, calculates a first distance between the associated feature points, extracts feature points from a plurality of images captured by the plurality of cameras, associates the extracted feature points, calculates a second distance between the associated feature points, and integrates the point cloud in the vehicle coordinate system using a transformation matrix between the camera coordinate system and the vehicle coordinate system estimated using the calculated first distance and second distance.
9. An image processing apparatus according to claim 1, wherein the estimation unit associates points in the integrated point cloud at the most recent time with points in the integrated point cloud at past time, estimates the momentum of the vehicle from the positional difference of the associated points, transforms the integrated point cloud at past time with the estimated momentum, and interpolates the point cloud by adding the transformed point cloud to the point cloud at the current time.
10. An image processing method comprising an image processing device for processing images captured by a plurality of cameras, wherein the image processing device comprises a computing device and a storage device, and the image processing method comprises: a point cloud generation procedure in which the computing device acquires a plurality of images captured by the plurality of cameras while the vehicle is in motion, generates a disparity image from the acquired plurality of images with overlapping field of view areas, and converts the generated disparity image into a point cloud in the camera coordinate system; a point cloud integration procedure in which the computing device converts the converted point cloud from the camera coordinate system to the vehicle coordinate system and integrates the point cloud in the vehicle coordinate system; and an estimation procedure in which the computing device transforms the integrated point cloud at past time points using the momentum of the vehicle estimated from the integrated point cloud, and generates an interpolated point cloud to add the transformed point cloud to the point cloud at the most recent time point.
Citation Information
Patent Citations
Method for detecting feasible region based on camera and laser radar
CN113205604A
Camera scene analysis method, system and equipment for intelligent driving automobile and medium
CN118397588A
Vehicle operating system and vehicle operating method
JP2009292254A
Movement amount calculation apparatus
JP2022071779A
On-vehicle processing device, vehicle control device and self-position estimation method
JP2023039626A