Method for acquiring three-dimensional images with the aid of a stereo camera having two cameras and device for carrying out this method
The method addresses camera misalignment issues in stereo imaging by using multiple cameras with varying focal lengths and inertial measurement units, ensuring accurate 3D data acquisition for automated driving applications.
Patent Information
- Application Number
- JP2023543269
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-29
- Filing Date
- 2021-09-28
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2041-09-28
AI Technical Summary
Existing methods for acquiring three-dimensional images with stereo cameras face challenges in accurately compensating for camera misalignments, particularly for small objects and dynamic environments, leading to inaccuracies in automated driving applications.
A method that detects and compensates for misalignments between camera signatures by using multiple cameras with different focal lengths and incorporating inertial measurement units to adjust for relative positional changes, allowing for real-time triangulation and accurate 3D data acquisition.
Enables reliable and precise 3D data acquisition even in dynamic conditions, ensuring accurate triangulation and reduced errors in image acquisition, particularly suitable for automated driving systems.
Smart Images

Figure 0007813475000002 
Figure 0007813475000003 
Figure 0007813475000004
Abstract
Description
[Technical Field]
[0001] This patent application claims priority from German patent application DE 102020212285.7, the contents of which are incorporated herein by reference.
[0002] The present invention relates to a method for acquiring three-dimensional images with the aid of a stereo camera having two cameras. Furthermore, the present invention relates to a method for acquiring three-dimensional images with the aid of a stereo camera having two cameras. ,this It relates to an apparatus for carrying out the method. [Background technology]
[0003] An object detection device is known from US Pat. No. 5,629,999 and the references given therein. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] WO2013 / 020872A1 [Non-patent literature]
[0005] [Non-Patent Document 1] Karl Kraus, Photogrammetrie (= Photogrammetry), Volumes 1 and 2, Duemmler, Bonn, 1996 / 1997 [Non-patent document 2] Thomas Luhmann, Nahbereichsphotogrammetrie Grundlagen, Methoden und Anwendungen (= Basics, methods and applications of close-range photogrammetry), 3rd edition, Berlin / Offenbach 2010 [Non-patent document 3] Daniel Scharstein, Richard Szeliski: A taxonomy and evaluation of dense two-frame stereo correspondence algorithms, in International journal of computer vision, 47th ed, no. 1-3, 2002, pp. 7-42 [Non-patent document 4] Alex Kendall et al.: End-to-end learning of geometry and context for deep stereo regression. CoRR, vol. abs / 1703.04309, 2017 [Non-Patent Document 5] “James Stein estimator” (cf. the paper “Stein's estimation rule and its competitors - an empirical bayes approach” by B. Efron and C. Morris, Journal of the American Statistical Association 68 (341, pages 117 to 130, 1973) Summary of the Invention [Problem to be solved by the invention]
[0006] The object of the present invention is to provide a method for acquiring three-dimensional images for practical applications, in particular well suited for obtaining images for safeguarding automated driving (autonomous driving). [Means for solving the problem]
[0007] This object is achieved according to the invention by a method having the features of claim 1.
[0008] Misalignments between the signatures of scene objects resulting from camera misalignments (positional deviations) can be accurately detected and compensated for by this method.
[0009] A scene object may be a relatively small object whose image is, for example, a single pixel or less than 10 pixels in size for each camera. Alternatively, the scene object may have a larger amount and may encompass the entire image area. Examples of such larger scene objects, in the example of vehicle image acquisition, would be individual vehicle components or vehicle sections, or even the entire vehicle.
[0010] An example of a characteristic disparity determined in the image acquisition method is the image position displacement or position of the image of a scene object from the epipolar line of each camera. The displacement of each image position of the scene object along the epipolar line can also be taken into account in the determination method. These disparities are also referred to below as vertical disparity and horizontal disparity.
[0011] When filtering misregistration, it can be simply checked whether the misregistration is smaller than a specified misregistration tolerance value, which can be dynamically specified until the number of selected signature pairs is smaller than a specified limit value.
[0012] Instead of one stereo camera, two or more cameras may work together in an image acquisition method.
[0013] Each camera may be configured as an arrangement of several cameras that are mutually assigned. One of these mutually assigned cameras may be a fisheye camera. Such a fisheye camera may have a focal length of less than 20 mm. Another of these assigned cameras may be a telephoto camera that includes a telephoto lens with a focal length of at least 80 mm.
[0014] To determine the characteristic misalignment of the assigned signature pairs with respect to each other and to filter the misalignments, for example, each of the vertical misalignments can be summed to its square for all signature pairs, and the state variables of the stereo camera can be varied until this sum of squares is minimized.
[0015] A misregistration tolerance value in the form of a default sum of squares of vertical mismatch or standard deviation of vertical mismatch can also be used as a termination criterion to reduce the number of selected signature pairs below a predefined limit.
[0016] In particular, the relative motion between the cameras can be taken into account in the image acquisition method: an estimation of the relative positions of the cameras with respect to each other can be performed asynchronously to the triangulation calculation by determining the characteristic misalignment of the assigned signature pairs.
[0017] Position correction of the relative positions of the cameras to each other, particularly based on the relative position estimation, can also be performed immediately before each triangulation calculation. The camera position calibration is then available for each measurement process. Such rapid calibration allows reliable stereo measurements to be performed even when using image acquisition devices in which the relative positions of the cameras to each other are constantly changing. In particular, such position correction based on the camera relative position estimation can be performed using data from an inertial measurement unit to which the cameras are permanently connected. In particular, it is possible to use camera measurement devices with very large distances between the cameras (long-baseline stereoscopic imaging).
[0018] Known correction methods can be used for the triangulation calculations, and in this connection reference is made to the expert publications by [1] and [2].
[0019] In triangulation calculations, optical distance measurements are performed by measuring angles within a triangle defined within the method, which can be formed by one of the two cameras and two image points, or by two cameras and one image point.
[0020] When filtering the distance displacements to select the signature pairs to be assigned, those signature pairs that are more likely to belong to the same scene object in the 3D scene may be selected.
[0021] The filtering of the misalignment for selecting the signature pairs to be assigned may be performed using a filtering algorithm according to claim 2. The respective misalignment amounts that can be used for filtering are for example horizontal misalignment and / or vertical misalignment.
[0022] In the image acquisition method according to claim 4 or 5, even difficult-to-perceive three-dimensional scenes can be reliably assigned to scene objects and corresponding 3D data maps can be created without errors. Deviations in the relative orientation of the cameras of a stereo camera relative to one another can be accurately taken into account. The angular correction values identified in the calculation step can be used as state variables for the above-described transformations, e.g., for minimizing the sum of squares of the vertical discrepancy.
[0023] Angle correction values are characteristic angles for the relative positional relationships of the cameras with respect to each other, such as the baseline plane, the baseline tilt, the relative tilt of the camera image acquisition directions, or the tilt of the camera coordinate axes with respect to each other.
[0024] The determination method according to claim 6 results in particularly accurate 3D data values.
[0025] The method according to claim 7 allows for the inclusion of instantaneous changes in the position of the cameras relative to one another, for example due to vibrations. This data acquisition is therefore performed in real time. The time constant of this data acquisition can be less than 500 ms, in particular less than 100 ms. The data acquired by the inertial measurement unit provides raw information regarding the necessary correction of the relative position of the cameras relative to one another, which is thus optimized by the image acquisition. The corresponding misalignment detection can also be used for the calibration of the relative position of the cameras relative to one another. This can be done in particular before the respective triangulation steps are performed during the image acquisition method.
[0026] Another object of the present invention is to improve the reliability of data acquisition when photographing a measurement object.
[0027] This object, according to the present invention, Reference example below This is achieved by a method having the characteristics described in 1. A method for generating redundant images of a measurement object, comprising: - combining at least three cameras (74-76; 77-79) whose entrance pupil centers define a camera placement plane; -Place the object to be measured in the camera's field of view, - performing a triangulation measurement of at least one selected measurement point of the measurement object using at least three different camera pairs (74, 75; 74, 76; 75, 76; 77, 78; 77, 79; 78, 79) of the cameras (74 to 79); -Compare the results of triangulation measurements, A method comprising the steps.
[0028] According to the present invention, various cameras, for example provided in a camera device enabling an automatic method, can cooperate with each other for redundant photographing of a measurement object. In this case, the cameras are correspondingly coupled so that the acquisition results of a camera pair can be compared and checked using a third camera. Consequently, acquisition errors can be detected, and true redundancy can be created by comparing three independent acquisition results, and the acquired data is determined to be correct if at least two of the three acquisition results match each other. The measurement object can be arbitrary, as long as it has texture or structure and contains depth information about it.
[0029] Claim 8 The advantages of the device according to the invention correspond to those already explained above in connection with the method. The device may comprise at least one camera fixedly connected to an inertial measurement unit.
[0030] The three cameras may be arranged in the shape of a triangle, in particular in the shape of an isosceles triangle.
[0031] Claims including six cameras 9 The apparatus according to the present invention provides a further improved shooting redundancy. The six cameras may be arranged in the shape of a hexagon, in particular in the shape of a regular hexagon. The cameras may be arranged in a camera arrangement plane.
[0032] Claim 10The use of a further remote camera according to the formula (1) enables reliable matching of the results of image acquisition by three adjacently arranged cameras. The three cameras arranged adjacent to one another may be located in a camera arrangement plane. A distance factor characterizing the distance of the remotely arranged camera compared to the distance of the adjacently arranged cameras may be greater than 2, greater than 3, greater than 4, greater than 5, for example greater than 10. This distance factor may be selected so that close distances up to a close distance limit are covered by the adjacently arranged cameras, and the far distances (far ranges) of the camera's field of view are covered by adding distantly arranged cameras starting from this close distance limit. The device has at least three adjacently arranged cameras (74 to 76) and a camera (77) arranged at least twice as far away as the adjacently arranged cameras (74 to 76).
[0033] Reference example below The apparatus according to the present invention simplifies the alignment of triangulation measurements. In this device, a remotely located camera (77) is located within a camera placement plane (83) of three adjacent cameras (74-76).
[0034] Connecting at least one camera to an inertial measurement unit allows acceleration measurements of each camera and corresponding evaluation of the acquired acceleration values.
[0035] Exemplary embodiments of the invention are described in more detail below with reference to the drawings. [Brief explanation of the drawings]
[0036] [Figure 1] FIG. 1 is a plan view of an apparatus for calibrating the three-dimensional position of the center of a camera's entrance pupil, with additional calibration surfaces shown in both a neutral position outside the camera's field of view and an operational position within the camera's field of view. [Figure 2] FIG. 2 shows a view from direction II of FIG. 1 with an additional calibration surface in the neutral position. [Figure 3] FIG. 10 shows a diagrammatic representation to explain the positional relationships between the components of the calibration device. [Figure 4] FIG. 10 shows a more detailed view of the movable reference camera of the calibration apparatus, including a camera movement drive that moves the movable reference camera in multiple translational / rotational degrees of freedom. [Figure 5] 1A-1C are diagrams illustrating different orientations of a movable reference camera, namely eight orientation variations. [Figure 6] 1 shows a calibration panel having a calibration surface including calibration structures, which can be used as a primary calibration surface and / or as an additional calibration surface in a calibration device. [Figure 7] 1 shows, in a top view, the arrangement of an apparatus for determining the relative positions of the centers of the entrance pupils of at least two cameras mounted on a common support frame. [Figure 8] 1 is a diagrammatic representation of two cameras of a stereo camera acquiring a three-dimensional image, illustrating the coordinate and position parameters that determine the angular offset of the cameras relative to each other; [Figure 9] 9 shows again diagrammatically the two cameras of the stereo camera according to FIG. 8 capturing a scene object of a three-dimensional scene, with the misalignment parameters of the characteristic signature of the images captured by the cameras highlighted. [Figure 10] FIG. 10 is a block diagram for explaining a method for acquiring a three-dimensional image with the aid of a stereo camera according to FIGS. 8 and 9; [Figure 11] FIG. 1 shows an apparatus for carrying out a method for generating redundant images of a measurement object using, for example, two groups of three cameras each assigned common signal processing. [Figure 12] In a representation similar to that of Figure 8, the two cameras of a stereo camera for obtaining a three-dimensional image are again shown, with the coordinate and position parameters determining the corrections, in particular the angular corrections of the cameras relative to each other, illustrated. DETAILED DESCRIPTION OF THE INVENTION
[0037] The calibration device 1 serves to calibrate the three-dimensional position of the center of the entrance pupil of a camera 2 to be calibrated. The camera 2 to be calibrated is arranged in a cubic installation volume 3, which is highlighted by dashed lines in Figures 1 and 2. When the calibration method is performed, the camera 2 to be calibrated is firmly installed in the installation volume 3. A mount (holder) 4, which is only implied in Figure 1, serves for this purpose. The camera 2 to be calibrated is held by the mount 4 so that the camera 2 spans a predetermined calibration field of view 5, the boundaries of which are shown by dashed lines in the side view of the device 1 according to Figure 1.
[0038] The camera 2 to be calibrated may, for example, be a camera for a vehicle used to provide "self-driving" functionality.
[0039] In particular, to facilitate the description of the positional relationships of the cameras of apparatus 1 relative to each other and to field of view 5, an xyz coordinate system is depicted in each of Figures 1-3 unless otherwise indicated. In Figure 1, the x-axis extends perpendicular to and into the plane of the drawing. The y-axis extends upward in Figure 1. The z-axis extends to the right in Figure 1. In Figure 2, the x-axis extends to the right, the y-axis extends upward, and the z-axis extends perpendicular to and out of the plane of the drawing.
[0040] The entire viewing range of the calibration field of view 5 covers a detection angle of, for example, 100° in the xz plane. Other detection angles, for example between 10° and 180°, are also possible. In principle, it is also possible to calibrate the camera with a detection angle larger than 180°.
[0041] The mount 4 is fixed to a support frame 6 of the calibration device 1 .
[0042] The calibration device 1 has at least two, and in the version shown a total of four, fixed reference cameras 7, 8, 9 and 10 (see Fig. 3), of which only two, namely reference cameras 7 and 8, are visible in Fig. 1. The fixed reference cameras 7-10 have the function of recording the calibration field of view 5 from different directions.
[0043] FIG. 3 shows exemplary dimensional parameters that play a role in the calibration device 1.
[0044] The main lines of sight 11, 12, 13, and 14 of the fixed reference cameras 7 to 10 are indicated by dashed lines in FIG.
[0045] These main lines of sight 11 to 14 intersect at point C (see Figures 1 and 3). The coordinates of this intersection C are x c , y c and z c It is shown as follows.
[0046] The x-distance between reference cameras 7 and 10 on the one hand, and the x-distance between reference cameras 8 and 9 on the other hand, is dx in FIG. h The x coordinates of the fixed reference cameras 7 and 8 on the one hand and the x coordinates of the fixed reference cameras 9 and 10 on the other hand are respectively the same.
[0047] The y-distance between the fixed reference cameras 7 and 8 on the one hand, and the y-distance between the fixed reference cameras 9 and 10 on the other hand, are shown in Fig. 3 as dy h The y coordinates of the fixed reference cameras 7 and 10 on the one hand and the y coordinates of the fixed reference cameras 8 and 9 on the other hand are respectively the same.
[0048] The calibration device 1 further comprises at least one fixed main calibration surface, three main calibration surfaces 15, 16 and 17 in the illustrated embodiment, which are identified by corresponding calibration panels. The main calibration surface 15 is parallel to the xy plane and parallel to the z plane in the device according to Figures 1 and 2. c 1 and 2, two further lateral main calibration surfaces 16, 17 extend parallel to the yz plane on either side of the arrangement of the four fixed reference cameras 7-10. The main calibration surfaces 15-17 are furthermore fixedly mounted on the support frame 6.
[0049] The primary calibration surface has fixed primary calibration structures, examples of which are shown in FIG. 6 . At least some of these calibration structures are located within the calibration field of view 5. The primary calibration structures may have a regular pattern, for example, arranged in the shape of a grid. Corresponding grid points that are part of the calibration structures are shown at 18 in FIG. 6 . The calibration structures may have colorful pattern elements, as shown at 19 in FIG. 6 . Furthermore, the calibration structures may have various sizes. Pattern elements that are enlarged compared to the grid points 18 are highlighted at 20 in FIG. 6 as primary calibration structures. Furthermore, the primary calibration structures may include encrypted pattern elements, for example, a QR code 21 (see FIG. 6 ).
[0050] It is not necessary for the primary calibration surfaces 15-17 to be aligned in an xyz coordinate system according to Figures 1 and 2. Figure 3 shows an exemplary tilted arrangement of the primary calibration surface 15', for example tilted relative to the zy plane.
[0051] 3 further shows an XYZ external coordinate system, for example of a production hall, in which the calibration device 1 is accommodated. The xyz coordinate system of the coordinate system of the calibration device 1 on the one hand and the XYZ coordinate system of the production hall on the other hand are rotated by a tilt angle rot, as shown in FIG. Z may be inclined relative to each other by only
[0052] The primary calibration structures 15-17, 15' therefore lie in a primary calibration structure main plane (the xy plane of the device according to FIGS. 1 and 2) and also in a primary calibration structure angular plane (the yz plane of the device according to FIGS. 1 and 2). The primary calibration structure main plane xy is arranged at an angle greater than 5°, i.e., 90°, relative to the primary calibration structure angular plane yz. This angle relative to the primary calibration structure angular plane yz may be greater than 10°, greater than 20°, greater than 30°, greater than 45°, or greater than 60°, depending on the embodiment. Small angles, for example, in the range of 1° to 10°, can be used to approximate a curved calibration structure surface by the primary calibration structures. In the device according to FIGS. 1 and 2, and also in the device with the primary calibration surface 15' according to FIG. 3, more than two primary calibration structure surfaces 15-17, 15' can be arranged in different primary calibration structure planes.
[0053] The position of each primary calibration surface, e.g., primary calibration surface 15', relative to the xyz coordinate system can be defined via the position of the center of the primary calibration surface and two inclination angles of primary calibration surface 15' relative to the xyz coordinates. A further parameter characterizing each of primary calibration surfaces 15-17 or 15' is the grid spacing grid of grid points 18 of the calibration structure. Said grid spacing grid is shown in FIG. 6. The two grid values given horizontally and vertically for primary calibration surface 15' in FIG. 6 do not necessarily have to be equal to each other, but they must be fixedly predetermined and known.
[0054] Additionally, the positions of the colored pattern elements 19, enlarged pattern elements 20, and / or encrypted pattern elements 21 within the grid of grid points 18 are fixed for each primary calibration surface 15-17, 15'. These relative positions of the various pattern elements 18-21 to one another serve to identify each primary calibration surface and determine its absolute position in space. The enlarged pattern elements 20 can be used to assist in determining their respective positions. The different sizes of the pattern elements 18-20 and the encrypted pattern elements 21 allow for calibration measurements at near and far distances, and even measurements where the primary calibration surfaces 15-17, 15' are potentially highly tilted with respect to the xy plane.
[0055] Furthermore, the calibration device 1 has at least one, and in the illustrated embodiment three, additional calibration surfaces 22, 23, and 24, each including an additional calibration structure 25. The additional calibration surfaces 22-24 are implemented by cup-shaped calibration panels. The additional calibration structures 25 are each arranged on the additional calibration surfaces 22-24 in the form of a 3x3 grid. Each additional calibration structure 25 may have pattern elements of the type of pattern elements 18-21 described above in relation to the main calibration surface.
[0056] The additional calibration surfaces 22-24 are mounted together on a movable holding arm 26, which is rotatable about a rotation axis 28 extending parallel to the x-direction via a geared motor 27, i.e., a calibration surface movement drive. Via the geared motor 27, the additional calibration structures 22-24 can be moved between a neutral position and an operating position. The neutral position of the additional calibration structures 22-24, in which the additional calibration structures are arranged outside the calibration field 5, is shown in FIGS. 1 and 2, respectively. The operating position of the holding arm 26 and of the additional calibration surfaces 22-24 is shown in dashed lines in FIG. 1, where the operating position is rotated upward compared to the neutral position. In the operating position, the additional calibration surfaces 22-24 are arranged in the calibration field 5.
[0057] For example, as shown in FIGS. 1 and 2, in the operating position, the central additional calibration structure 25z (see also FIG. 3) is positioned parallel to the xy plane. Thus, the 3×3 grid arrangement of the additional calibration structures 25 is positioned in the operating position with three rows 251, 252, and 253 extending along the x direction and three columns extending parallel to the y direction. Adjacent rows and adjacent columns of the additional calibration structures 25 are inclined relative to each other by an inclination angle α, which ranges from 5° to 45°, e.g., 30°. For the four grid areas, each located at a corner of the 3×3 grid, this inclination angle α exists around two axes that are perpendicular to each other. This results in a cup-shaped basic structure of each additional calibration surface 22-24. The additional calibration structures 25 exist in a 3D configuration that deviates from a flat surface.
[0058] The calibration device 1 further comprises an evaluation unit 29 which processes the recorded camera data of the camera 2 to be configured and of the fixed reference cameras 7-10 and the state parameters of the device, i.e. in particular the positions of the additional calibration surfaces 22-24 and of the main calibration surfaces 15-17 and the positions and lines of sight of the reference cameras 7-10. The evaluation unit 29 may have a memory for image data.
[0059] The calibration device 1 further includes a movable reference camera 30 , which also functions to record the calibration field of view 5 .
[0060] FIG. 3 shows the degrees of freedom of movement of the movable reference camera 30: two tilt degrees of freedom and one translation degree of freedom.
[0061] Figure 4 shows details of the movable reference camera 30, which is movable by a camera movement drive 31 between a first field of view recording position and at least one other field of view recording position whose image acquisition direction differs from that of the first field of view recording position (see recording direction 32 in Figure 1).
[0062] The camera movement drive 31 includes a first rotation motor 33, a second rotation motor 34, and a linear movement motor 35. A camera head 36 of the movable reference camera 30 is mounted on a rotation part of the first rotation motor 33 via a holding plate 37. The camera head 36 can rotate around an axis parallel to the x-axis via the first rotation motor 33. The first rotation motor 33 is mounted on a rotation part of the second rotation motor 34 via another support plate 38. The camera head 36 can rotate around a rotation axis parallel to the y-axis via the second rotation motor 34.
[0063] The second rotary motor 34 is mounted on a linear movement unit 40 of the linear movement motor 35 via a holding bracket 39. The linear movement motor 35 enables linear movement of the camera head 36 parallel to the x-axis.
[0064] The camera movement drive 31 and the camera head 36 of the reference camera 30 are signal-connected to the evaluation unit 29. The position of the camera head 36 is accurately transmitted to the evaluation unit 29 depending on the positions of the motors 33-35 and on the mounting situation of the camera head 36 with respect to the first rotary motor 33.
[0065] The angular position of the camera head 36 that can be preset via the first rotary motor 33 is also called the pitch angle. Instead of the first rotary motor 33, the change in pitch angle can be performed by articulating the camera head 36 via a connecting shaft parallel to the x axis and a linear drive connected to the camera head 36 that is movable in the y direction by means of two stops that preset two different pitch angles. The angular position of the camera head 36 that can be preset via the second rotary motor 34 is also called the yaw angle.
[0066] FIG. 5 shows eight example variations of positioning the camera head 36 of the movable reference camera 30 using the three degrees of freedom of movement illustrated in FIG.
[0067] The image acquisition directions 32 are shown by dashed lines, each corresponding to a pitch angle ax and a yaw angle ay. In the top row of FIG. 5, the camera head is positioned at a small x-coordinate x min In comparison, in the bottom row of FIG. 5, the camera head 36 is positioned at a larger x-coordinate x max The eight image acquisition directions according to Figure 5 represent three sets of different parameters (positions x; ax; ay) with two discrete values for each of these three parameters.
[0068] In one variation of the calibration apparatus, the movable reference camera 30 may also be omitted.
[0069] To calibrate the three-dimensional position of the center of the entrance pupil of the camera 2 to be calibrated, the calibration device 1 is used as follows:
[0070] First, the camera 2 to be calibrated is held on a mount 4 .
[0071] A fixed primary calibration surface 15-17 or 15' is then acquired (recorded) with the camera 2 to be calibrated and the reference cameras 7-10 and 30. Additional calibration surfaces 22-24 are in a neutral position.
[0072] The additional calibration surfaces 22-24 are then moved between the neutral position and the working position by the calibration surface moving drive 27. The additional calibration surfaces 22-24 are then acquired by the camera 2 to be calibrated and by the reference cameras 7-10 and 30, the additional calibration structure 25 being in the working position. The recorded image data of the camera 2 to be calibrated and of the reference cameras 7-10 and 30 are then evaluated by the evaluation unit 29. This evaluation is performed by vector analysis of the recorded image data, taking into account the positions of the recorded calibration structures 18-21 and 25.
[0073] Once the main calibration surfaces 15-17 and the additional calibration surfaces 22-24 have been acquired, an initial acquisition (recording) of the main calibration surfaces 15-17, 15' on the one hand and the additional calibration surfaces 22-24 on the other hand can be performed by the movable camera 30 at a first field of view recording position and, after movement of the movable reference camera 30 with the camera movement drive 31, can be performed at at least one further field of view recording position. When evaluating the recorded image data, the image data of the movable reference camera 30 at at least two field of view recording positions are also taken into account.
[0074] The acquisition sequence for calibration surfaces 15-17 and 22-24 is as follows: First, primary calibration surfaces 15-17 are acquired (recorded) by movable camera 30 at a first field-of-view recording position. Next, additional calibration surfaces 22-24 are moved to the operating position and again acquired by movable camera 30 at the first field-of-view recording position. Movable reference camera 30 then moves to another field-of-view recording position, and additional calibration surfaces 22-24 remain in the operating position. Subsequently, additional calibration surfaces 22-24 are acquired by movable reference camera 30 at another field-of-view recording position. Additional calibration surfaces 22-24 are then moved to a neutral position, and further acquisitions (recordings) of primary calibration surfaces 15-17 are made by movable reference camera 30 at another field-of-view recording position. During this sequence, the primary calibration surfaces 15-17 may also be acquired by the fixed reference cameras 7-10 while the additional calibration surfaces 22-24 are in the neutral position, and when the additional calibration surfaces 22-24 are in the operating position, these additional calibration surfaces 22-24 may also be acquired by the fixed reference cameras 7-10.
[0075] With reference to FIG. 7, an apparatus 41 for determining the relative positions of the centers of the entrance pupils of at least two cameras 42, 43, 44 mounted on a common support frame 45 will now be described.
[0076] The cameras 42 to 44 may be pre-calibrated with respect to the positions of their respective entrance pupil centres with the aid of the calibration device 1 .
[0077] Once this relative position determination has been performed by the device 41, the nominal positions of the cameras 42-44 relative to the support frame 45, ie the target installation positions, are known.
[0078] The cameras 42-44 may be, for example, cameras on a vehicle used to provide "self-driving" functionality.
[0079] The apparatus 41 has a number of calibration structure-carrying parts 46, 47, 48, and 49. The calibration structure-carrying part 46 is a master part (main part) for specifying a master coordinate system xyz, in Fig. 7, the x-axis of this master coordinate system extends to the right, the y-axis extends upward, and the z-axis extends perpendicular to and outside the plane of the drawing.
[0080] For the calibration structures applied to the calibration structure carrying parts 46-49, what has been explained above for the calibration structures 18-21, in particular in relation to FIG. 6, applies.
[0081] The calibration structure carrying parts 46-49 are arranged around the support frame 45 in the operating position of the device 41, so that each of the cameras 42-44 captures at least two calibration structures of the calibration structure carrying parts 46-49. Such an arrangement is not necessary, so that at least some of the cameras 42-44 can acquire a calibration structure from exactly one of the calibration structure carrying parts 46-49. Furthermore, the arrangement of the calibration structure carrying parts 46-49 is such that at least one calibration structure of exactly one calibration structure carrying part 46 is acquired by two of the cameras 42-44. To ensure these conditions, the support frame 45 can optionally be moved relative to the calibration structure carrying parts 46-49, respectively, without changing their positions.
[0082] Figure 7 shows an example of the position of a support frame 45 with the actual positions of cameras 42, 43, 44 on the support frame, again not shown, where the field of view 50 of camera 42 captures the calibration structures of calibration structure carrying parts 46 and 47, while camera 43 with its field of view 51 captures the calibration structures of calibration structure carrying parts 47 and 48, while a further camera 44 with a field of view 52 captures the calibration structures of calibration structure carrying parts 48 and 49.
[0083] The relative positions of the calibration structure carrying parts 46 - 49 with respect to one another do not have to be precisely defined in advance, but must not change during the position determination method by the device 41 .
[0084] The device 41 comprises an evaluation unit 53 which processes the recorded camera data from the cameras 42 - 44 and possibly situation parameters during position determination, ie in particular the identification of the respective support frame 45 .
[0085] To determine the relative positions of the centers of the entrance pupils of the cameras 42-44, the device 41 is used as follows.
[0086] In a first preliminary step, the cameras 42-44 are mounted on a common support frame 45. In a further preliminary step, the calibration structure carrying parts 46-49 are arranged as a group of calibration structure carrying parts around the support frame 45. This can also be done by arranging the group of calibration structure carrying parts 46-49 in a preliminary step and then positioning the support frame relative to this group. In addition, an xyz coordinate system is defined by the alignment of the master part 46. The other calibration structure carrying parts 47-49 do not need to be aligned to this xyz coordinate system.
[0087] Now, the calibration structure-carrying parts 46-49 located in the field of view of the cameras 42-44 are recorded at the actual positions of the cameras 42-44, for example according to Fig. 7, and at a predetermined relative position of the support frame 45 with respect to the group of calibration structure-carrying parts 46-49. The recorded image data of the cameras 42-44 are then evaluated by the evaluation unit 53, so that the exact positions of the centers of the entrance pupils and the image acquisition directions of the cameras 42-44 are determined in the coordinate system of the master part 46. These actual positions are then transformed into the coordinates of the support frame 45 and matched with the nominal target positions. This can be done in terms of a best fit method.
[0088] In the determination method, the support frame can also be moved between different camera acquisition positions, so that at least one of the cameras whose relative position is to be determined records a calibration structure-carrying part not previously detected by that camera. This step of recording and moving the support frame can be repeated for all cameras whose relative positions to each other are to be determined until the condition is met that each of the cameras records at least a calibration structure of the two calibration structure-carrying parts, and at least one of the calibration structures is recorded by two cameras.
[0089] 8-10, a method for acquiring a three-dimensional image with the aid of a stereo camera 55a having two cameras 54, 55 will now be described. These cameras 54, 55 may have been calibrated in a preliminary step with the aid of a calibration device 1 and may further be measured with respect to their relative position with the aid of device 41. The cameras 54, 55 are again mounted on a support frame.
[0090] The camera 54 shown on the left in Fig. 8 is used as a master camera (parent camera) to define a master coordinate system xm, ym and zm, where zm is the image acquisition direction of the master camera 54. Therefore, the second camera 55 shown on the right in Fig. 8 is a slave camera (child camera).
[0091] The master camera 54 is permanently connected to an inertial master measurement unit 56 (IMU), which may be designated as a rate of rotation sensor, particularly in the form of a microelectromechanical system (MEMS). The master measurement unit 56 measures the pitch angle dax of the master camera 54. m , yaw angle day m and roll angle daz m This allows the change in angle of the master coordinate system to be measured, thereby making it possible to monitor the positional deviation of the master coordinate system in real time. The time constant for this real-time positional deviation detection may be greater than 500 ms, greater than 200 ms, or even greater than 100 ms.
[0092] The slave camera 55 is also rigidly connected to an associated inertial slave measurement unit 57, so that the pitch angle dax of the slave camera 55 s , yaw angle day s and roll angle daz sThe angular changes of the stereo camera 54a can be detected in real time, and thus the relative changes of the slave coordinate system xs, ys, zs with respect to the master coordinate system xm, ym, zm can be detected, again in real time. The relative motion of the cameras 54, 55 of the stereo camera 55a relative to one another can be detected in real time by the measuring devices 56, 57 and can be included in the method for acquiring the 3D image. The measuring devices 56, 57 can be used to predict changes in the relative positions of the cameras 54, 55 relative to one another. Image processing performed as part of the 3D image acquisition can then further improve this prediction of the relative positions. For example, even if the support frame on which the stereo camera 54a is mounted moves over an uneven surface and the cameras 54, 55 move continuously relative to one another, the results of the 3D image acquisition remain stable.
[0093] The line connecting the centers of the entrance pupils of cameras 54 and 55 is shown as 58 in FIG. 8 and represents the baseline of stereo camera 55a.
[0094] In the method of acquiring the three-dimensional image, the following angles related to the positional relationship of the slave camera 55 to the master camera 54 are captured:
[0095] -Angle by s , i.e., the tilt of the plane perpendicular to the plane xmzm and through which the base line 58 passes, about the tilt axis ym relative to the plane xmym; -bz s : tilt by about the tilt axis parallel to the slave coordinate axis zs relative to the plane xmzm of the base line 58 s The corresponding slope; -ax s : the tilt of the slave coordinate axis zs, i.e. the image acquisition direction of the slave camera 55, about the slave coordinate axis xs relative to the plane xmzm; -ay s :Tilt ax about slave coordinate axis ys relative to master plane xmym of slave coordinate axis xs s the corresponding slope; and -az s :The tilt ax of the slave coordinate axis ys about the slave coordinate axis zs relative to the master plane coordinates ymzm s ,ays The corresponding slope.
[0096] The following method is performed by measuring the angular change dax detected by the measuring devices 56, 57 on the one hand. m ,day m ,daz m ,dax s ,day s ,daz s Including these angles by s ,bz s ,ax s ,ay s ,az s is used to acquire a three-dimensional image with the aid of two cameras 54, 55, taking into account:
[0097] First, images of a three-dimensional scene with scene objects 59, 60, 61 (see FIG. 9) are acquired simultaneously by the two cameras 54, 55 of the stereo camera. The image acquisition of these images 62, 63 is performed simultaneously for both cameras 54, 55 in an acquisition step 64 (see FIG. 10).
[0098] Image acquisition may be integrated over multiple cycles of inertial measurement unit 56, 57 acquisition, particularly over a period that corresponds to multiple time constants for real-time displacement measurement sensing.
[0099] FIG. 9 shows diagrammatically the images 62, 63 of the cameras 54 and 55, respectively.
[0100] The image of the scene object 59 is captured by the master camera 54 in the image 62. M and the image of the scene object 60 is denoted by 60 M is shown.
[0101] The image of the scene object 59 is transferred to the image 63 of the slave camera 55. S The image of the scene object 61 is shown as 61 in the image 63 of the slave camera 55. S Furthermore, the image 59 of the master camera 54 is M ,60 Mcan be found at the corresponding x,y coordinates of the image frame in image 63 of slave camera 55.
[0102] Image position 59 M ,59 S The y-shift of the image position 59 is called the vertical misalignment or vertical misalignment VD perpendicular to the epipolar line of each camera. M ,59 S The x-displacement of is called the epipolar or horizontal disparity HD. In this connection, reference is made to the well-known terminology of epipolar geometry, where the parameter "center of the camera entrance pupil" is called the "saliency center" in this terminology. 2 images 60 M ,61 S show the same signature in images 62 and 63, and are therefore represented by the same image pattern in images 62 and 63, but actually arise from two different scene objects 60 and 61 in the three-dimensional scene.
[0103] Characteristic signatures of the scene objects 59-61 in the images are now determined separately for each of the two cameras 54, 55 in a determination step 65 (see Figure 10).
[0104] The signatures determined in step 65 are each summarized in a signature list, and in an assignment step 66, the signatures of the acquired images 62, 63 determined in step 65 are assigned in pairs. Identical signatures are thus assigned to each other with respect to the acquired scene objects.
[0105] Depending on the three-dimensional scene being acquired, the result of the assignment step 66 may be a very large number of assigned signatures, for example tens of thousands of assigned signatures and correspondingly tens of thousands of determined characteristic displacements.
[0106] In a further determination step 67, the characteristic misalignment of the assigned signature pairs with respect to each other, eg the vertical disparity VD and the horizontal disparity HD, is now determined.
[0107] For example, each determined vertical disparity VD is squared and summed for all assigned signature pairs. The angular parameter by s ,bz s ,ax s ,ay s ,az s This sum of squares can be minimized by varying the angles, which are dependent on these angles as explained above in connection with FIG.
[0108] In a next filtering step 68, the determined misregistrations are then filtered using a filter algorithm to select assigned signature pairs that are more likely to belong to the same scene object 59-61. The simplest variation of such a filter algorithm is selection by comparison with a predetermined tolerance value; only those signature pairs whose sum of squares is smaller than the predetermined tolerance value pass the filter. This default tolerance value can, for example, be increased until the number of selected signature pairs as a result of filtering is smaller than a predetermined limit value.
[0109] In one image algorithm variation, each misregistration VD, HD itself can be used to select an assigned signature pair, and then each misregistration is checked to see if it is less than a predetermined threshold.
[0110] In particular, multiple thresholds can be tested. For example, it can be checked for how many signature pairs the vertical mismatch is less than thresholds S1, S2, S3, and S4, where S1≦S2≦S3≦S4. This then results in four lists of accepted signature pair assignments (accepted correspondences) and rejected signature pair assignments (rejected correspondences). It is then examined how the number of accepted correspondences combined in each case depends on the threshold. The lowest threshold for which the number of correspondences remains approximately the same is used. This results in a heuristic method for filtering signature pairs so that "incorrect" signature pairs that do not belong to the same object are likely to be rejected.
[0111] As soon as the filtering results in a number of selected signature pairs being smaller than a predetermined limit, e.g., smaller than one-tenth of the signatures originally assigned in pairs, or, e.g., smaller than 500 signature pairs in absolute terms, a triangulation calculation is performed in step 69 to determine depth data for each scene object 59-61. In addition to the number of selected signature pairs, a default tolerance for the sum of squares of the characteristic misalignment, e.g., vertical misalignment VD, of the assigned signature pairs can serve as a termination criterion, in accordance with what has been explained above. The standard deviation of the characteristic misalignment, e.g., vertical misalignment VD, can also be used as a termination criterion.
[0112] Triangulation can be performed with each accepted signature pair, ie accepted correspondence.
[0113] As a result of this triangulation calculation, a 3D data map of the captured scene objects 59-61 in the captured image of the three-dimensional scene can be created and output as a result of a creating and outputting step 70. An example of such a 3D data map is a map of all points of each scene object, each with a respective value triple x representing the position of this respective scene point in Cartesian coordinates.i ,y i ,z i A 3D reconstruction of each scene object is possible with corresponding access to this 3D data map.
[0114] If filtering step 68 indicates that the number of selected signature pairs is still greater than the specified limit value, decision step 71 first determines angle correction values between the various selected assigned signature pairs and checks whether the photographed raw objects belonging to the various selected assigned signature pairs can be positioned in the correct position relative to each other in the three-dimensional scene. For this purpose, the angles described above in connection with Figure 8 are used, where these angles are corrected in real time for measurement monitoring using measurement devices 56, 57.
[0115] Based on the compensation calculations performed in decision step 71, it is possible to determine whether the scene objects 60, 61 have the same signature 60, e.g. M ,61 S Nevertheless, signature pairs that can be distinguished from each other in images 62, 63 and are therefore correspondingly assigned can be discarded as incorrectly assigned, thus correspondingly reducing the number of selected signature pairs.
[0116] After the angle correction has been performed, a comparison step 72 is performed, comparing the angle correction value determined for the signature pair with a predetermined correction value. If the comparison step 72 results in the angle values of the signature pair deviating from each other by more than the predetermined correction value, the filter algorithm used in the filtering step 68 is adapted in an adaptation step 73, so that after filtering with the adapted filter algorithm, a smaller number of selected signature pairs is produced than the number produced in the previous filtering step 68. This adjustment can be performed by eliminating signature pairs that differ in their discrepancy by more than a predetermined limit value. Similarly, the comparison criteria (from which the signatures of a potential signature pair are evaluated as equal and therefore assignable) can be set more critically in the adjustment 73.
[0117] Thus, this sequence of steps 73, 68, 71 and 72 is performed until the angular correction values of the remaining assigned signature pairs are found to deviate from each other by less than the specified correction value. Then, the triangulation calculation is performed again in step 69, which may include the angular correction values of the selected signature pairs, and the result obtained is generated and output, particularly in the form of a 3D data map.
[0118] FIG. 11 shows a method for generating redundant images of a measurement object. For this purpose, several cameras are coupled together, the centers of their entrance pupils defining a camera arrangement plane. FIG. 11 shows two groups, on the one hand, three cameras 74-76 (group 74a), and on the other hand, cameras 77, 78, 79 (group 77a). Group 74a on the one hand and group 77a on the other hand each have associated data processing units 80, 81 that process and evaluate the image data acquired by the associated cameras. The two data processing units 80, 81 are signal-connected to each other via signal line 82.
[0119] To capture a three-dimensional scene, for example, the cameras 74-76 of group 74a can be interconnected, thereby enabling, for example, 3D capture of this three-dimensional scene by the image acquisition method described above in connection with Figures 8-10. To create additional redundancy in this three-dimensional image acquisition, for example, the image acquisition results of a camera 77 of a further group 77a can be used, which are supplied to a data processing device 81 of group 77a and to a data processing device 80 of group 74a via a signal line 82. Due to the spatial distance from camera 77 to the cameras 74-76 of group 74a, there are significantly different viewing angles when photographing the three-dimensional scene, which improves the redundancy in the three-dimensional image acquisition.
[0120] A camera arrangement plane 83 defined by the cameras 74 to 76 of group 74a or the cameras 77 to 79 of group 77a is shown diagrammatically in FIG. 11 and is positioned at an angle relative to the drawing plane of FIG.
[0121] 3D image capture using just one group 74a, 77a of cameras is also called intra-image capture. 3D image capture involving at least two groups of cameras is also called inter-image capture.
[0122] Triangulation can be performed independently using, for example, stereo cameras 78, 79, cameras 79, 77, and cameras 77, 78. The triangulation positions of these three devices must all coincide.
[0123] Camera groups such as groups 74a, 77a can be arranged in the shape of a triangle, in particular in the shape of an isosceles triangle. A hexagonal six-camera arrangement is also possible.
[0124] Compared to the distance between the cameras of one group 74a, 77a, the cameras of the other group are at least twice as far away. The distance between cameras 76 and 77 is therefore at least twice the distance between cameras 75 and 76 or between cameras 74 and 76. This distance factor may be larger, for example, greater than 3, greater than 4, greater than 5, or even greater than 10. The near range of the cameras covered by each group 74a, 77a may be in the range of, for example, 80 cm to 2.5 m. By adding at least one camera from each other group, a far range (wide area) can also be captured by the image capture device beyond the near range limit.
[0125] In the following, a method for obtaining a three-dimensional image will be described, which should be understood as a supplement to the above description.
[0126] For 3D image acquisition, a system of equations (simultaneous equations) is established to calculate the (u,v) coordinates of, for example, cameras 54, 55 with known focal lengths f. 右 and a known image point (u,v) 左 This system of equations is solved for the image point correspondences:
[0127]
number
[0128] Equation 1 is explained in more detail below in connection with FIG.
[0129] In equation 1: u l ,v l : Scene objects in the left image 59,59 l (=59 M ) means the image coordinates of the scene feature under consideration, using the example of u r ,v r : Scene objects in the right image 59,59 r (=59 S ) means the image coordinates of the features, using the example;
[0130] f l ,f l : means the focal length of the left and right cameras 54, 55; λ l ,λ r : means the control variable along the ray from the center of the left or right camera 54, 55 through the image point of the left or right camera. At λ=0, the center of the left and right cameras 54, 55, respectively, and at λ=1, the position 59 of the scene object 59. l ,59 r are on the respective images 62, 63 of the respective cameras 54, 55. For λ>1, the point lies on a further ray in front of the left or right camera 54, 55 relative to the respective image plane 62, 63 up to and beyond the intersection of the scene object 59.
[0131] t x ,t y ,t z : It means the coordinate of the center position of the right camera 55 in the coordinate system of the left camera 54. The length of this vector t is the baseline length (blen) or the baseline 58 (see Figure 8). This length of the vector t can be determined from the installation situation of the two cameras 54, 55. After the length of t is known, the three coordinates t x ,t y and t z This leaves two degrees of freedom with
[0132] R XX ,...,R ZZ : means the parameters of the rotation matrix to describe the rotational position of the right camera 55 in the coordinate system of the left camera 54. This rotation has three rotational degrees of freedom.
[0133] The optical axes of the two cameras 54, 55 are shown by dashed lines in FIG. 12, again comparable to FIG.
[0134] Equation 1 above is the triple u l (Equation 1.1), v l (Equation 1.2) and f lIt can be written as three equations 1, 1, 1.2 and 1.3 for (Equation 1.3). Therefore, it is a system of equations with three equations (Equations 1.1-1.3) and two unknowns (λ1, λ2). These equations can be converted into one equation by eliminating the unknowns (λ1, λ2).
[0135] A correspondence is the agreement of feature points when the same scene object is captured by different cameras 54, 55 (left l, right r) of the image capture device. If a feature point of a scene object is captured by camera 54 at image coordinates ui, vi, and the same feature point is captured by camera 55 at image coordinates uj, vj, this is a (positive) correspondence. Each correspondence can be assigned a vertical disparity VD. Therefore, at least five correspondences are required per camera pair, and the number can be somewhat higher, but is advantageously significantly higher for greater statistical stability.
[0136] The image point correspondences may be stereo correspondences, for which see the technical articles [3] and [4].
[0137] In Equation 1, u, v are the Cartesian coordinates of the respective image points, for example the respective Cartesian coordinates x, y or coordinates in the direction of the respective horizontal and vertical disparities HD, VD.
[0138] The vector (u,v,f) describes the position of a ray from the origin of the camera 54, 55 with control parameter λ.
[0139] Equation 1 describes the position change from the right camera (eg, camera 55) to the left camera (eg, camera 54) in terms of three position parameters and a rotation matrix R via a translation vector t, i.e., six degrees of freedom.
[0140] Since the length of the displacement vector t (baseline length = baseline 58 in Figure 8 = blen) cannot be estimated (estimated) without scalar information in the 3D scene, but is measured or otherwise known in advance, there are five degrees of freedom: two displacements normalized by blen perpendicular to blen (which can also be interpreted as rotations of blen), and three rotations of the right camera to the left (i.e., five rotations).
[0141] An example of five degrees of freedom is the aforementioned angle .times. ... s ,bz s ,ax s ,ay s ,az s is.
[0142] The system of equations GL1 has two unknowns, namely the two control parameters λ l ,λ r It consists of three equations with five degrees of freedom: the transformation matrix T and the rotation matrix R.
[0143] From these three equations, an equation can be formed that no longer has any unknown variables, so that a suitable transformation creates exactly one equation without the λ parameter, i.e., depending only on the 5 degrees of freedom.
[0144] With at least five equations, the degrees of freedom can be calculated, for example, by an estimator, which solves a system of equations that is usually overdetermined by minimizing the residuals. Such estimators are known from the literature as [5].
[0145] The correspondences, i.e., image point correspondences according to equation 1 above, must be linear and independent of each other.
[0146] With multiple correspondences, false positives can be filtered out by an algorithm based on weighting of features with respect to the size of the distance to the epipolar line of each camera 54, 55.
[0147] False positive correspondences are usually distributed to a first approximation. True correspondences have a characteristic accumulation, i.e., they are not usually distributed. The filtering results in false positive correspondences are removed, so that true correspondences remain within the framework of the convergence algorithm. Positives allow the estimator to converge to the actual values. This estimator also works for squint cameras (compare the above description for the misalignment between zm and zs, especially Figure 8), and converges faster for fisheye lenses.
[0148] If the image capture device includes two cameras 54, 55, then five degrees of freedom must be determined. If there are three cameras that can be configured into two camera pairs, then there are ten degrees of freedom to determine.
[0149] Generally, n cameras are provided in the image capture device. For one camera pair, 5 degrees of freedom can be estimated. Each camera beyond the first camera pair, which has 5 degrees of freedom, contributes 6 additional degrees of freedom. With n cameras, this results in 6n-1 estimable degrees of freedom.
[0150] An IMU (inertial sensor for rotation rate) is integrated into both cameras 54, 55 of the image capture device. From the rotation rate in one cycle (monoperiod frame), the rotational misalignment is estimated, and therefore the rotational misalignment of the cameras relative to each other. The residual error is compensated by applying the method. When using an IMU, the IMU translation can also be used. When using an IMU, acceleration values can be obtained, which helps to improve the prediction accuracy.
[0151] The prediction from the IMU data may be integrated over multiple cycles for stabilization, and for this purpose the method is also calculated over these cycles.
[0152] The number of cameras can also be increased to three or more cameras for joint calculation instead of two. Even when the cameras are aligned in a row, various viewpoints improve the estimation results.
[0153] As cameras spread out over the area, the estimates are refined again.
[0154] The estimate improves again if the camera is more oblique and therefore opens up to a larger overall field of view. In this sense, a fisheye can be a better camera 54, 55 than a standard lens.
[0155] Each scene object may be a corresponding larger image region (blob), e.g., a part of an automobile, and the correspondence of larger blobs may be determined more accurately than the correspondence of smaller features.
[0156] The same 3D feature can be found in three images by three cameras. If it is found three times, it is considered plausible. A corresponding plausibility check may be included in the misalignment filtering of the image acquisition method. If a particular feature is acquired by two cameras, its position can be predicted when acquired by a third camera. If it is actually captured during the image capture by the third camera, the 3D feature in question is considered plausible. Otherwise, a misassignment is assumed.
[0157] An increasing number of cameras increases the likelihood and therefore the rejection of the false positive correspondences mentioned above.
[0158] Each estimate weights the correspondence via a distance weight to the epipolar line of the respective camera 54, 55. Before each estimation step, each epipolar line is calculated. In the case of multiple estimates, the epipolar lines change and so do the weights. Up to a threshold, for example, a weight of "1" can be assumed. Up to twice the threshold, the weights can decrease linearly to 0.
[0159] The weight changes during multiple estimations provide information about the likelihood of false positives, which leads to new weightings. The magnitude and value of each weight is derived by comparing the corresponding numbers when using four different thresholds, S1–S4.
[0160] The weighting is described, for example, via a cost function that depends on the amount of the threshold.
[0161] The correspondences should be distributed as evenly as possible in the image. For this purpose, the image is divided into regions (e.g., 9 evenly divided regions). The number of features in a region is ideally constant. A ZDF (region partition function) describes the distribution of features in the image. If the distribution is unfavorable, the estimate is rejected.
[0162] As long as there are correspondences that are evenly distributed across each image, the estimator will be more robust in terms of the results obtained. When subdividing a region, it can be guaranteed that a minimum number of correspondences is present in every region. This minimum number of correspondences must be present in the region at least during the final estimation performed. For example, correspondences from regions with a large number of correspondences compared to the average number of correspondences in the subdivided regions can then be weighted less by a smaller weighting factor.
[0163] Irregularly distributed correspondences can be distributed evenly by removing them in overcrowded regions as an alternative to the reweighting described above.
[0164] Irregularly distributed correspondences can be identified by weighting in overcrowded regions.
[0165] The ZDF distribution can also be adjusted from the vicinity of the epipolar line via weights, which are followed by a product of weights that depend on the vicinity of the respective epipolar line on the one hand and on the other hand on the respective region.
[0166] Correspondences are weighted by their distance to the current epipolar line. In principle, a large number of correspondences, i.e., still with a large distance to the respective epipolar line, can be considered. However, it is better to have a relatively large number of correspondences with a small distance to the respective epipolar line. For this purpose, the curve of the number of correspondences is considered with the maximum epipolar distance. This generally s-shaped curve showing the number of correspondences as a function of epipolar distance (first a small increase, then a large decrease, then a weak increase again) is analyzed, and the position of the maximum gradient is determined (inflection point). A small change in small distances results in a statistically almost linear increase in rejected correspondences. Beyond the actual error, this increase remains nominally constant.
[0167] This position is a compromise between as many correspondences as possible and not too many correspondences with large distances.
[0168] Near the actual value, the correspondences accumulate, i.e., increase in significance. Far from the actual value, the correspondences are usually statistically distributed (background noise).
[0169] The result of the estimation can be averaged over a period of time. Thus, several images can be recorded in succession and subjected to the corresponding estimation. This results in temporal filtering.
[0170] The estimates that are averaged are only averaged up to the maximum past. Thus, a moving average is calculated in real time. Correspondingly long disturbances in the past are again forgotten and no longer interfere with the latest estimates.
[0171] The relative positions of the cameras used in the image acquisition device can be estimated using the image acquisition method. The estimated results of different image acquisition scenarios can be compared. If they differ too much from each other, they can be rejected and deactivated (fail-safe), or one of the two results can be confirmed by a third result and the mission can continue (fail-operational). The measurements from these realigned cameras can also be confirmed by a match between two of the three measurements.
[0172] External calibration using image acquisition methods is based on the intrinsic calibration of each camera (distortion, focal length, etc.).
[0173] The model error of the intrinsic calibration can be integrated into the weight of the distance to the epipolar line. Depending on the epipolar line distance, i.e., the epipolar curve, a prediction of the coordinates of the scene objects 59-61 in the image of the second camera 55 can be made directly taking into account the distortion description. The weight of such a description can be increased depending on the expected residual error of the distortion correction.
[0174] To capture a wide field of view using as few cameras as possible, the cameras can be equipped with fisheye lenses with focal lengths of less than 20 mm, particularly less than 10 mm. At larger distances, the distance resolution is lower due to the smaller focal lengths than when using standard or telephoto lenses. Estimation using standard or telephoto lenses is more stable than estimation using fisheye lenses. If a telephoto camera with a focal length of at least 80 mm (particularly at least 150 mm or at least 200 mm) is mounted or assigned to the fisheye camera in a shape-stably manner to optimize the aforementioned cameras 54, 55, 74-76, 77-79, the estimation of a fisheye camera can be converted to a telephoto camera. When two fisheye cameras are estimated, a new estimation is automatically generated for the associated telephoto camera. However, the higher resolution of the telephoto camera contributes to the estimation of the system of at least two fisheye cameras and the associated telephoto camera. Thus, the image acquisition device no longer only has a fisheye camera, but also an additional associated telephoto camera. For example, two fisheye cameras may form a stereo pair (see the above example of cameras 54 and 55). Each fisheye camera may have a telephoto camera permanently or fixedly assigned to it. From the correspondence of the fisheye cameras on the one hand and the correspondence of the telephoto camera on the other hand, an augmented system of equations corresponding to equation 1 above may be constructed and then solved.
[0175] The estimate can be averaged over multiple frames. The averaging can be done by IMU-value-normalized estimate. Larger movements are less disturbing.
[0176] The averaging of the IMU-value-normalized estimates can be smoothed and thus stabilized by IMU-detected jerk movements.
[0177] A motion model can support the estimation. The prediction of angles with their probability balanced by the current measurement and measurement error can support dynamic measurements. The motion model is based on the stiffness assumption of the camera's support structure or support frame. The movement must be slower than the exposure time, but can be faster than the cycle time, i.e., the image refresh time.
[0178] The general movement of the superstructure can be recorded and an AI (artificial intelligence) model can be derived from this training set, which allows for better predictions. For example, motor vibrations introduced into the camera support structure can be compensated for.
[0179] The AI model can detect sudden movements (such as jerks) and help reject the estimate.
[0180] The estimation of the relative camera positions may be asynchronous with the measurement using triangulation, which is based on classical stereo algorithms of rectified images.
[0181] Image acquisition methods use native features to form correspondences.
[0182] Classical stereo matching methods compare gray value differences and look for the epipolar curve with the smallest gray value difference, typically along a horizontal line. When the gray values are equal, a signature exists. A "distinctive" feature is converted into a signature from the gray value variations of the feature environment. Thus, a comparison of gray values within the feature environment occurs in the camera image. Only features whose signatures occur relatively rarely in the image, for example, less than 10 times, less than 8 times, or less than 7 times, are used.
[0183] Correspondences can also be used for distance measurement using triangulation. Features are extracted and used for triangulation (specific stereo). The same features are used for estimation in conjunction with the image acquisition method. Once the estimation is done, new results can be added asynchronously to the measurement. While measurements are done several times, the estimation is done only once. This is done due to the slowly changing relative positions of the cameras.
[0184] Image acquisition methods can quickly detect strong distortions in coarse (binned) images, i.e., images in which many pixels are grouped together. When these are detected, the exact changes can be determined in finer images and made available for measurement. Estimates with smaller deviations are not provided because computer-intensive calculation of distortion parameters for correction images should be avoided. The distortion parameters are parameters for describing the desired correction based on the distortion description, i.e., based on the image error description of the cameras 54, 55.
[0185] The rejected deviation must, of course, be smaller than the deviation that can be tolerated by DenseStereo, i.e., by the classical stereo matching method described above. Linear errors of one pixel are usually acceptable; this depends on the size of the environments being compared. However, such linear errors, even when they are small, will result in errors in the distance measurement.
[0186] In case of highly deformable (soft) mechanical structures between the respective cameras under consideration, e.g., cameras 54 and 55, an estimation must be performed before each measurement. Distortion parameters are calculated from the estimation data, then a correction image is calculated for line fidelity, and the distance is determined by triangulation from the DenseStereo correspondence, i.e., from the position of minimum distance.
[0187] Calculating distortion parameters for DenseStereo is computationally intensive. For idiosyncratic stereo, i.e., assignment of idiosyncratic features without line fidelity and gray value comparisons, only comparisons of signatures are performed, allowing triangulation to be calculated directly from the correspondence of image acquisition methods.
[0188] The image acquisition method also provides a confidence measure for the correspondence through a weighting of the distance to the epipolar line, which is on the one hand the probability of a correspondence across more than two cameras, and on the other hand an estimate of the magnitude of the expected noise, ultimately the signal-to-noise ratio (SNR).
[0189] Therefore, a synchronized sequence of triangulation and estimation using characteristic features is efficient and gives a measure for the reliability of the measurement points.
[0190] A specific triangulation method can measure the distance from the correspondence, where the center of the distance between the oblique rays generated from the two corresponding image points of the two stereo cameras 54, 55 is considered. This distance must be below a predetermined threshold.
[0191] However, if three or more cameras are used, there can be more than three correspondences per feature. The distance measurement becomes more accurate and reliable. However, the number of measurement points is usually reduced. The intersection of the successful triangulations between camera pair 1 / 2 and camera pair 2 / 3 is smaller than the union of the two successful triangulations.
[0192] If multiple corresponding signatures are found in a camera pair, multiple correspondences are created. From these correspondences of the same signature, the one with the lowest cost function relative to the epipolar line is selected. Thus, the signature closest to the epipolar curve is selected, not necessarily the signature closest in space.
[0193] Multiple may be selected, but each selected signature may only be used once. Once a correspondence is selected, the two locations in both images should not be used for further correspondence. [Explanation of symbols]
[0194] 54,55 Camera 55a Stereo Camera 59~61 Scenery Objects 62,63 images
Claims
1. A method for acquiring a three-dimensional image with the aid of a stereo camera (55a) having two cameras (54, 55), comprising: - acquiring (64) images (62, 63) of a three-dimensional scene simultaneously by the two cameras (54, 55) of the stereo camera (55a), - determining (65) characteristic signatures of scene objects (59-61) in each acquired image (62, 63); - assigning (66) the signatures of said acquired images (62, 63) to each other in pairs, - determining (67) the characteristic misregistrations HD, VD of the assigned signature pairs relative to each other, - filtering (68) said misregistrations to select assigned signature pairs; - performing triangulation calculations (69) to determine depth data for each of the scene objects (59-61) based on the selected signature pairs; - creating (70) a 3D data map of the captured scene objects (59-61) within the captured image of the three-dimensional scene; A method having the steps.
2. 6. The method of claim 1, further comprising: filtering (68) the misregistrations using a filtering algorithm to select assigned signature pairs that belong to the same scene object in the three-dimensional scene.
3. 3. The method according to claim 1, wherein the triangulation calculation (69) is performed as soon as the number of selected signature pairs is less than a predetermined limit value.
4. As long as the number of selected signature pairs is greater than a predetermined limit, - determining (71) angular correction values between different selected and assigned signature pairs in order to check whether the photographed live objects (59-61) belonging to the different selected and assigned signature pairs can be correctly positioned within the 3D scene, - comparing (72) each of the angular corrections determined for the signature pair with a predetermined correction value; - creating (70) a 3D data map of the captured scene objects (59-61) within the captured image of the three-dimensional scene; 3. The method of claim 2, further comprising the steps of:
5. As long as the number of selected signature pairs is greater than a predetermined limit, - insofar as the angular corrections of the signature pairs deviate from each other by more than the predetermined correction value; adjusting (73) said filtering algorithm, so that after filtering (68) with said adjusted filtering algorithm a number of selected signature pairs is obtained that is smaller than the number obtained in the previous filtering step, and repeating said comparison (72); - insofar as the angular correction values of the signature pairs deviate from each other by at most the predetermined correction value, --performing triangulation calculations (69) to determine depth data for each of said scene objects (59-61); - creating (70) a 3D data map of the captured scene objects (59-61) within the captured image of the three-dimensional scene; 5. The method of claim 4, further comprising the steps of:
6. 6. The method of claim 4 or 5, wherein the triangulation calculation (69) includes the angle correction values of the selected signature pairs.
7. 7. The method according to claim 1, wherein during the acquisition (64) of the images (62, 63), data is simultaneously acquired from inertial measurement units (56, 57) to which the cameras (54, 55) are fixedly connected.
8. Apparatus for carrying out the method according to any one of claims 1 to 7.
9. 9. The device according to claim 8, characterized in that it comprises six cameras (74-79).
10. 10. The device according to claim 8 or 9, characterized in that it has at least three adjacently arranged cameras (74-76) and a camera (77) arranged at least twice as far away as the adjacently arranged cameras (74-76).
11. 11. Apparatus according to claim 9 or 10, characterized in that at least one of the cameras (54, 55; 74-76; 77-79) is fixedly connected to an inertial measurement unit (56, 57).
Citation Information
Patent Citations
Image processor, image processing method and program
JP2018032144A
Systems and Methods for Estimating and Refining Depth Maps
US20180027224A1
Multiple camera calibration chart
US20180367681A1
Systems and methods for multi-camera placement
US20190364206A1
Object detection device for a vehicle, vehicle with such an object detection device and method for determining a relative positional relationship of stereo cameras with respect to one another
WO2013020872A1