System comprising a stereo camera and method for calculating confidence values and / or error values

DE102016104378B4Active Publication Date: 2025-08-14ROBOCEPTION GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE102016104378
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-03-04
Filing Date
2016-03-10
Publication Date
2025-08-14
Estimated Expiration
2036-03-10

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

System, comprising: - at least one stereo camera (20) with a first optic (21) and a second optic (22); - at least one processing device (14, 24) which is in communicative connection with the stereo camera (20) in order to receive from the stereo camera (20) first image data (ImgL) which are assigned to the first optics (21) and second image data (ImgR) which are assigned to the second optics (22); - at least one memory (13) with instructions, where the instructions cause the processing device (14, 24): - to calculate at least one disparity image (DM) with disparity values ​​(d(p)) for a plurality of pixels (p) based on the first and second image data (ImgL, ImgR); and - to calculate a confidence value (k(p)) for at least some of the image points (p) based on the image data (ImgL, ImgR), characterized in that the instructions cause the processing device (14, 24): - to determine at least one first candidate disparity value (d(p), d'(p), d''(p)) and at least one second candidate disparity value (d(p), d'(p), d''(p)) for at least one pixel (p); - to determine a confidence value (k(p)) for the pixel (p) by comparing the candidate disparity values ​​(d(p), d'(p), d''(p)), wherein the confidence value (k(p)) indicates a reliability for the pixel (p); and - to select and / or calculate a single disparity value (d(p)) as the resultant disparity value (d(p)) for the disparity image (DM) for the at least one pixel (p).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a system in which confidence and / or error values ​​are calculated and to a method for calculating confidence and / or error values.

[0002] Cameras, especially stereo cameras, are used in a wide variety of industrial applications. One application is the motion planning and navigation of devices such as robots, robot arms, or vehicles. Cameras assist a robot in motion planning and navigation within a given environment and enable it to locate, grasp, and manipulate objects. Stereo cameras are particularly advantageous in this context because, thanks to two lenses arranged side by side, they enable the capture of the environment with depth information.

[0003] Typically, disparity images are generated from the image data acquired with the stereo camera, for example, images. These images indicate the distance from the camera and / or an optical center to the respective point in the scene / environment for each pixel. For example, if one considers a left and a right image, the disparity image indicates the displacement distance for the pixel in the left image to the right image: d(p)=x L (p)-x R (p). The spatial distance Z is calculated using the following formula: Z=b⋅F / d

[0004] The spatial distance Z is therefore inversely proportional to disparity d, where b is the so-called base distance between the two optics, and F is a focal length factor.

[0005] Numerous methods for determining disparity images are known. One method gaining increasing acceptance in industry is semi-global matching, as described, for example, in a paper by Heiko Hirschmüller titled "Accurate and Efficient Stereo Processing by Semi-Global Matching and Mutual Information" (IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA, June 20-26, 2005).

[0006] Paper D1 ("Exploiting the Power of Stereo Confidences", Pfeiffer, D; Gehrig, S.; Schneider, N.; in 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, pp. 297-304) investigates the use of stereo confidence features in stereo vision applications widely used in gaming, robotics, and driver assistance systems. While conventional methods apply a threshold to confidence values, resulting in a thinned disparity map, this does not fully exploit the available information. Instead, the authors propose propagating confidence values ​​along with disparities in a Bayesian framework. Based on annotated video data, a mapping of confidence values ​​to probabilities for disparity outliers is created.The study extends the Stixel-World representation to integrate stereo confidence features into the underlying sensor modeling process of maximum a posteriori estimation. Evaluations on a large real-world traffic database show that the use of stereo confidence features significantly reduces the number of false object detections while maintaining a high detection rate.

[0007] Document D2 ("Stereo Processing by Semiglobal Matching and Mutual Information", Hirschmüller, H.; in: IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. 2, pp. 328-341, Feb. 2008) describes the semiglobal matching (SGM) stereo method. It uses a pixel-wise matching cost calculation based on mutual information (MI) to compensate for radiometric differences between input images. It also covers methods for occlusion detection, subpixel optimization, and multibaseline matching. Furthermore, post-processing steps for removing outliers, addressing specific challenges in structured environments, and interpolating gaps are presented. Finally, strategies for processing almost arbitrarily large images and for fusing disparity images using orthographic projection are proposed.

[0008] Document D3 ("A Quantitative Evaluation of Confidence Measures for Stereo Vision," in: IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 34, no. 11, pp. 2121-2133, Nov. 2012) presents a comprehensive evaluation of confidence measures for stereo matching, comparing both widely used methods and newly proposed techniques. The authors categorize these measures based on their approach to stereo cost estimation and analyze their strengths and weaknesses. The evaluation, conducted on binocular and multibaseline ground truth datasets, assesses the ability of each method to rank depth estimates, detect occlusions, and generate accurate depth maps. The study closes a research gap in the field of stereo vision and aims to provide valuable insights for researchers in binocular and multiview stereo applications.

[0009] Fundamentally, such methods must be highly performant, as current disparity images must be available virtually in real time for navigation. On the other hand, there is a demand for highly reliable data to rule out failures or errors in safety-relevant applications, such as autonomous driving. Ultimately, these are conflicting interests, as reliability requires more data to be acquired and processed, while performance requires as few computational operations as possible.

[0010] Based on this prior art, the present invention aims to provide an improved system for obtaining and processing environmental information. Furthermore, a corresponding method is to be provided.

[0011] The object is achieved by a system according to claim 1, a method according to claim 11, and a computer-readable memory according to claim 12. Further advantageous embodiments of the invention are set forth in the dependent claims.

[0012] In particular, the task is solved by a system that includes: - at least one stereo camera with a first lens and a second lens; - at least one processing device which is in communicative connection with the stereo camera in order to receive from the stereo camera first image data which are associated with the first optics and second image data which are associated with the second optics; - at least one memory with instructions; the instructions causing the processing device to: - to calculate at least one disparity image with disparity values ​​for a plurality of pixels based on the first and second image data; and - to calculate a confidence value and / or an error value for at least some of the pixels based on the image data characterized in that the instructions instruct the processing facility: - to determine at least one first candidate disparity value and at least one second candidate disparity value for at least one pixel; - to determine a confidence value for the pixel by comparing the candidate disparity values; and - to select and / or calculate a single disparity value for the at least one pixel as the result disparity value for the disparity image.

[0013] A key aspect of the present invention is that, in addition to the disparity image, additional information is provided that provides information about the reliability of the information obtained (e.g., confidence value and / or error value). Unlike prior art, confidence values ​​and / or error values ​​are not specified statically, but are calculated individually for each disparity image.

[0014] Preferably, the confidence values ​​and / or error values ​​are calculated individually for some or all pixels of the disparity image.

[0015] Through this individual calculation of the confidence values ​​and / or error values, a program processing the disparity image is able to make decisions based on the reliability of the obtained disparity image.

[0016] If, for example, a robot or a robot arm is controlled, the corresponding program can determine that the acquired disparity images, particularly in a highly relevant area, contain disparity values ​​that are not sufficiently reliable. In this case, the program can decide to move the robot or robot arm so that further image information, for example in the form of the first and second image data, is obtained. This approach can produce a more reliable and / or less error-prone disparity image. In the automotive sector, for example, it would be conceivable to resort to other sensors, such as lasers and / or ultrasound, for unreliable disparity images obtained by image processing. In a (further) embodiment, a robot and / or vehicle can decide, based on the confidence, whether a certain area is safe to enter.High-confidence measurements can indicate that an area has been reliably measured as clear. The robot or vehicle can then successively travel a specified distance.

[0017] Ultimately, the system according to the invention enables program sequences to be designed reliably. Furthermore, very fast and resource-efficient algorithms can be used to obtain disparity images, allowing for a more detailed analysis under certain conditions. This saves computing resources.

[0018] In one embodiment, the first image data and the second image data are rectified image data obtained based on calibration data from the stereo camera. The image data can be a grayscale image. Color images can also be used as image data, with intensity values ​​for different color channels (e.g., three) being stored in one embodiment.

[0019] The disparity image can be a disparity matrix. Regardless of the chosen data structure, the disparity values ​​can be stored in 1-bit, 2-bit, 3-bit, ..., 8-bit data fields. Preferably, 4-bit values ​​are used. In one embodiment, the spatial distance Z is inversely proportional to the respective disparity value. In this case, the value zero can indicate an infinite distance.

[0020] According to the invention, the term "pixels" can be understood to indicate preferably two-dimensional coordinates in a grid. For example, a pixel can indicate an x- and a y-value for the corresponding x- and y-axes.

[0021] Both the error values ​​and the confidence values ​​can be stored in corresponding matrices. For example, an error matrix can be provided for the error values ​​and / or a confidence matrix for the confidence values. In one embodiment, the data structure of the error values ​​and / or the confidence values ​​has a corresponding entry (error value or confidence value) for each entry in the disparity image.

[0022] In one embodiment, the disparity image has the same number of rows and columns as the error matrix and / or the confidence matrix. This high resolution regarding the errors and / or confidence values ​​allows for a very precise evaluation of the disparity image.

[0023] The disparity image, the confidence values, and / or the error values ​​can be saved in well-known graphic formats, such as GIF or TIFF, that support multiple image layers (multilayers). For example, the disparity values ​​can be stored in a first layer, and / or the confidence values ​​in a second layer, and / or the error values ​​in a third layer. Using well-known formats allows for computational operations to be performed on the data using established hardware. For example, modern graphics cards allow for very fast, partially parallelized processing of corresponding image formats.

[0024] The instructions may cause the processing device to determine at least one first candidate disparity value and at least one second candidate disparity value. For example, in the semi-global matching method, candidate disparity values ​​are calculated via eight or four independent paths before selecting a specific disparity value as the result disparity value for the corresponding pixel.

[0025] This means that for each pixel, eight or four candidate disparity values ​​are determined, at least partially independently of the other candidate disparity values. Semi-global matching can consider eight or four paths from eight or four different directions. To determine the disparity, the eight or four paths are combined at each pixel by summing the path costs and then searching for the minimum. Finally, in this step, after the individual path costs have been determined, a new search for a minimum of the sum of the objective functions over each path can begin, although preferably the "intermediate results" for the candidate search on the individual paths are used to determine this semi-global minimum. Preferably, the disparity value that provides the semi-global minimum is searched for in the entire possible result space.According to the invention, this disparity value can be defined as the result disparity value. The result disparity value will often correspond to one of the candidate disparity values. However, a disparity value that does not correspond to any of the candidate disparity values ​​can also be determined as the result disparity value. A semi-global minimum does not necessarily have to be determined to determine the result disparity value. According to the invention, instead of determining the semi-global minimum, it is possible to select one of the candidate disparity values ​​as the result disparity value, for example, the candidate disparity value with the lowest cost on the associated path.

[0026] The candidate disparity values ​​and their similarity to each other can be used to determine a corresponding confidence value. For example, a variance can be determined, which can then be converted into a scaled confidence value.

[0027] In one embodiment, for each pixel, the number of candidate disparity values ​​is determined that match the selected result disparity value or deviate only very slightly, e.g., by an amount of less than one. The threshold can be specified as an absolute value, e.g., 1 or 2, or as a relative value, e.g., 10% or 20%. The number of slightly deviating candidate disparity values ​​can be divided by the total number of available candidate disparity values ​​for the respective pixel. The result can serve as a starting value to calculate a confidence value for the respective pixel. For example, a basic confidence G of 0.5 can be assumed, with the result of the calculation described above being added to this basic confidence G. Preferably, this result of the calculation described above is multiplied by the factor F, e.g., 0.5.

[0028] In general, the confidence value can be calculated according to the following formula: k(p)=G+F∑iT[|di−d(p)|≤1]n

[0029] The factor F can be between 0.2 and 0.6. Preferably, the factor is 0.5, as stated above. The value n can specify the number of candidate disparity values ​​for the pixel p. The variable di can be the i-th candidate disparity value for the pixel p.

[0030] If multiple disparity values ​​are considered when determining a result disparity value, the result disparity value is preferably selected using an objective function. The objective function can assign an assessment value or value to each disparity value. In one embodiment, the disparity value is selected for which the applied objective function yields the lowest possible value. Ultimately, this is an optimization problem in which the objective function is minimized.

[0031] In one embodiment, the objective function is chosen such that cost and / or smoothness criteria are taken into account. According to the invention, cost criteria can be criteria that provide information about the similarity of specific intensity values ​​and / or color values ​​within the image data, e.g.: C(p,d(p))

[0032] Stereo matching attempts to match pixels from the first image data (first lens) to pixels from the second image data (second lens), allowing a displacement—the disparity—between the pixels to be determined. Due to varying lighting conditions and the lens used, the respective pixels may have different intensity and / or color values. The objective function penalizes such differences.

[0033] Furthermore, stereo matching typically assumes that neighboring pixels in the image data have an identical or similar distance from the optics. This means that similar disparity values ​​are expected to be determined. A preferred objective function penalizes disparity differences in the neighborhood Np. An objective function according to the invention can be structured as follows: E(M)=∑p∈M(C(p,d(p))+∑q∈NPg1⋅T[|d(p)−d(q)|=1]+∑q∈Npg2⋅T[|d(p)−d(q)|>1])

[0034] The function T preferably returns the value 1 if the condition specified in the square brackets is met. Otherwise, the function T returns the value zero. Np can be the set of pixels or disparity values ​​in the immediate vicinity of a pixel p.

[0035] In one embodiment, the error values ​​are calculated using a hierarchical approach. The instructions may cause the processing device to: - to calculate at least a first disparity value of the disparity image based on the first image data and second image data and to assign a first constant error value, e.g. 0.5, to the first disparity value, wherein the first and second image data are present at a first (high) resolution; - compressing the first image data into third image data and the second image data into fourth image data, wherein the third image data and fourth image data have a second resolution which is lower than the first resolution; - to calculate at least a second disparity value of the disparity image based on the third and fourth image data and to assign a second constant error value, e.g. 1, to the first disparity value.

[0036] The determination of the measurement error according to the invention is preferably based on a hierarchical approach, which can be extended to virtually any stereo matching method. For example, disparity images calculated at full image resolution often contain holes. A hole refers to an image point within the disparity image for which no permissible disparity value can be determined.

[0037] In a hierarchical approach, one attempts to fill these holes by performing a corresponding analysis on image data that has a lower resolution. For example, using known methods, two or four neighboring pixels in the first image data can be combined to form a single pixel, for example in the third image data. Similarly, two or four pixels from the second image data can be combined to form one pixel in the fourth image data. Ultimately, an appropriate method increases image sharpness. With increased image sharpness, it is to be expected that a corresponding disparity image will have fewer holes. According to the invention, disparity values ​​determined at the coarser hierarchy level can be transferred to a higher hierarchy level. In this way, for example, holes in a disparity image based on the full image resolution can be filled.When transferring to a higher hierarchy level, an adjustment of the determined disparity value may be necessary. Preferably, the adjustment factor to be multiplied is the inverse element of the scaling factor used for scaling. For example, if you are working with half the image resolution at the next higher hierarchy level (scaling factor = 1 / 2), the obtained disparity value must be multiplied by a factor of two when transferring (adjustment factor = 2).

[0038] In one embodiment, the hierarchical method halves the image resolution. It is possible to work with multiple hierarchical levels: 1. Hierarchy level: full image resolution; 2. Hierarchy level: half image resolution; 3. Hierarchy level: ¼ image resolution; n. Hierarchy level: 12n−1 Image resolution

[0039] In one embodiment of the invention, each hierarchy level is assigned a specific, for example, constant, error value. Depending on which hierarchy level a particular disparity value is taken from, the corresponding error value can be obtained according to the following formula: 0.5⋅2n−1

[0040] As already mentioned, the system can comprise a robot and / or a robot arm to which the stereo camera is attached. The system can comprise a control unit for controlling the robot and / or robot arm, wherein the control unit decides, depending on the confidence values ​​and / or error values, whether the robot and / or robot arm is moved in order to capture a specific object and / or a specific section of the environment using the stereo camera from a further position and / or orientation. Alternatively or additionally, the controller can only cause a robot or a vehicle to travel through a specific area if the associated disparity image indicates that the area is free of obstacles and the confidence for the relevant range of the disparity value is high, e.g., greater than 50%.

[0041] The problem described at the outset is further solved by a method according to claim 11.

[0042] In particular, the problem is solved by a method comprising the following steps: - capturing or receiving first and second image data from a stereo camera; - Calculating, based on the first and second image data, at least one disparity image with disparity values ​​for a plurality of pixels; and - Calculate a confidence value and / or an error value for at least some of the pixels - determining at least one first candidate disparity value and at least one second candidate disparity value for at least one pixel; - determining a confidence value for the pixel by comparing the candidate disparity values, the confidence value indicating a reliability for the pixel; and - Selecting and / or calculating for the at least one pixel a single disparity value as the result disparity value for the disparity image.

[0043] The method has similar advantages to those already described in connection with the system according to the invention. The method can be implemented by the system according to the invention, in particular by the processing device described therein.

[0044] The process may further implement some or all of the procedural features described in connection with the system.

[0045] The object mentioned at the outset is further achieved by a computer-readable memory with instructions for implementing the described method, wherein the instructions are designed to cause a computing unit to implement this method when the instructions are executed on the computing unit.

[0046] Further advantageous embodiments emerge from the subclaims.

[0047] The invention is described below using several exemplary embodiments, which are explained in more detail with reference to the figures. Herein: Fig. 1 a robot with a stereo camera that is communicatively connected to a computer (front view); Fig. 2 the robot Fig. 1 with the stereo camera in a schematic side view; Fig. 3 an embodiment of the stereo camera; Fig. 4 is a schematic representation of a processing method in which a disparity image, an error matrix and a confidence matrix are determined from a left and a right image; Fig. 5 a schematic representation of a method in which error values ​​are determined in a hierarchical approach; Fig. 6 a resulting image with several layers.

[0048] In the following description, the same reference numbers are used for identical and equivalent parts.

[0049] Fig. Figure 1 shows a schematic view of a system within which the inventive method can be implemented. The system comprises a computer 10, a stereo camera 20, and a robot 30. The computer 10 has a processing device 14, a display device 12, and a computer-readable memory 13. The processing device 14 executes instructions that cause the computer 10 to implement the inventive method described below.

[0050] The stereo camera 20 is as shown in the Fig. 1 and Fig.2, is attached to the robot 30. The robot 30 has a robot arm 31, at the end of which the stereo camera 20 is attached. The robot arm 31 also has a robot tool 32. Purely by way of example, it can be assumed that the robot tool 32 is a gripper for grasping an object. The robot arm 31 has a joint so that the stereo camera 20 can be pivoted together with the robot arm 31.

[0051] The stereo camera 20, in turn, has a first lens 21 and a second lens 22. The stereo camera 20 is suitable for capturing a series of left images ImgL and right images ImgR (cf. Fig.4). The stereo camera 20 is aligned such that its viewing direction runs along the alignment of the robot tool 32. According to the invention, the stereo camera 20 can also be attached to a different location on the robot 30. For example, the stereo camera 20 can be aligned such that it captures the robot arm 31 and the robot tool 32. In this exemplary embodiment, it may be necessary to move the entire robot 30 in order to appropriately align the stereo camera 20. In the case described above, it is sufficient to move the robot arm 31, which is connected to the stereo camera 20, in order to capture an object to be grasped from a different direction and / or position.

[0052] The described system is suitable for detecting an object in an unknown environment and grasping it using the robot tool 32.

[0053] For this purpose, it is necessary that the computer 10 or the stereo camera 20 processes the left and right images ImgL, ImgR.

[0054] Fig. Figure 3 illustrates that the stereo camera 20, in addition to the first and second optics 21, 22, has its own processing device 24. This processing device 24 can perform some or all of the steps of the method according to the invention. In the following description, some essential method steps are performed by the processing device 24. In another embodiment, it is possible for the processing device 14 of the computer 10 to perform these processing steps.

[0055] According to the invention, a disparity image DM, an error matrix FM, and a confidence matrix KM are generated by the processing device 24 based on a left image ImgL and a right image ImgR. In this context, images ImgL, ImgR, disparity images DM, error matrices FM, and confidence matrices are common, which have several hundred entries. Thus, the disparity image DM can specify disparity values ​​d(p) for several hundred pixels p. The error matrix FM can provide error values ​​f(p) for each of these pixels p. Likewise, the confidence matrix KM can specify a confidence value k(p) for each of the pixels p. For the Fig.In the embodiment described in Figure 4, for illustrative purposes, it is assumed that the respective data structure has only four entries. Thus, for the disparity image DM, a first disparity value d(p1), a second disparity value d(p2), a third disparity value d(p3), and a fourth disparity value d(p4) result. The respective disparity values ​​d(p1), d(p2), d(p3), and d(p4) are assigned to a first pixel p1, a second pixel p2, a third pixel p3, and a fourth pixel p4.

[0056] Accordingly, after applying the method according to the invention, the confidence matrix KM provides a confidence value k(p1), k(p2), k(p3) or k(p4) for each of the pixels p1, p2, p3, p4.

[0057] Likewise, the error matrix FM provides an error value for each pixel p1, p2, p3, p4, namely the error values ​​f(p1), f(p2), f(p3) and f(p4), respectively.

[0058] In the described embodiment, semi-global matching is carried out to determine the disparity image DM, as explained in the article by Heiko Hirschmüller already mentioned (see “Accurate and Efficient Stereoprocessing by Semi-Global Matching and Mutual Information” (IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA, June 20-26, 2005).

[0059] A particular advantage of semi-global matching is that eight candidate disparity values ​​are available for each pixel p1, p2, p3, p4. Due to the fact that only four pixels are considered in the described embodiment, there are only three candidate disparity values ​​for each disparity value d(p1), d(p2), d(p3), d(p4). Purely for the sake of example, it is assumed that for the first pixel p1, the following candidate disparity values ​​are available after minimizing the following objective function: E(M)=∑p∈M(C(p,d(p))+∑q∈NPg1⋅T[|d(p)−d(q)|=1]+∑q∈Npg2⋅T[|d(p)−d(q)|>1]) available: d(p1), d'(p1), d''(p1).

[0060] The semi-global matching algorithm selects the disparity value that yields the smallest value when applying the objective function E(M). Each candidate disparity value is assigned a different evaluation set M of pixels according to the semi-global matching algorithm, e.g., M1 for d(p1), M2 for d'(p1), M3 for d''(p1). The evaluation set M1 can include the pixels p1, p2, the evaluation set M2 the pixels p1, p4, and the evaluation set M3 p1, p3.

[0061] For the sake of example, it is assumed that in this case the result disparity value is the candidate disparity value d(p1). Therefore, the method according to the invention would enter d(p1) as the result disparity value in the disparity image DM.

[0062] To illustrate the invention, it is assumed that the following candidate disparity values ​​were calculated for the first pixel p1: d(p1)=1; d'(p1)=2; d''(p1)=1.

[0063] To determine the confidence value k(p1), the following formula is used: k(p)=0.5+0.5∑iT[|di−d(p)|≤1]n

[0064] The function T[] is 1 if its argument is true and 0 otherwise. n specifies the number of candidate disparity values ​​d(p1), d'(p1), d''(p1). In this example, n=3. The result for k(p1)=0.5+0.51+0+13≈83%

[0065] The algorithm according to the invention would assign a probability of approximately 83% to the disparity value d(p1) = 1. This means that the reliability of the corresponding value is relatively high.

[0066] In a similar way, confidence values ​​d(p2), d(p3), d(p4) can be calculated for the other pixels p2, p3, p4. The method according to the invention can also be readily applied to much larger matrices.

[0067] The determination of the error matrix FM is based on a hierarchical approach. Assuming that in the disparity image DM the Fig. 4 a value can be determined for all disparity values ​​d(P1), d(p2), d(p3), d(p4), then in the error matrix FM all error values ​​f(p1), f(p2), f(p3), f(p4) would be assigned an identical value, for example 0.5.

[0068] In reality, due to the image quality of the left image ImgL and the right image ImgR, it is often not possible to determine all disparity values ​​d(p1), d(p2), d(p3), d(p4). Gaps or holes arise that can be filled using a hierarchical approach. In this hierarchical approach, pixels from the images ImgL, ImgR are combined to form new pixels. In an embodiment, as exemplified by the Fig. As explained in section 5, the image resolution is halved at each hierarchy level.

[0069] The Fig.Figure 5 shows a first left image ImgL and a first right image ImgR with a 4x4 resolution, i.e., each image ImgL, ImgR has 16 pixels p (full image resolution in the 1st hierarchy level). Assuming that, for example, the upper left value of the disparity image DM (e.g., d(p1)) cannot be calculated, a mapping function is applied to the first left image ImgL and the first right image ImgR according to the invention. This mapping function results in the resolution of the images being halved (= half image resolution in the 2nd hierarchy level). This results in a second left image ImgL' and a second right image ImgR. Starting from the second left image ImgL' and the second right image ImgR', a second disparity image DM' can be fully determined. This means that the upper left pixel can also be determined. In the Fig.In the example shown in Figure 5, it is assumed that the value 4 is determined for this disparity value d(p1'). This disparity value can be transferred to the higher hierarchy level (cf. disparity image DM, which is assigned to the 1st hierarchy level), whereby the value is multiplied by a predetermined adjustment factor - in the exemplary embodiment, the adjustment factor is 2. For the 1st hierarchy level, the disparity value 8 results. This means that the previous hole in the disparity image DM at the location of the first pixel p1 is filled with the disparity value that was determined for the second disparity image DM'.

[0070] According to the invention, it is assumed that an error value F(p) increases when disparity values ​​have to be transferred from correspondingly higher hierarchy levels. In general, an error value estimate according to the invention results from f(p)=0.5*2l−1 where l indicates the hierarchy level from which the corresponding disparity value was taken or adopted.

[0071] For the described embodiment, the error value for the first pixel p1 would be F(p1) = 0.5 × 2 1 = 1. Assuming that the lower right pixel of the disparity image DM of the Fig. 5 in the 1st hierarchy level, the corresponding error value is F(p4) = 0.5 × 2 0 = 0.5. According to the invention, all error values ​​for the error matrix FM can be calculated in a corresponding manner. According to the invention, disparity values ​​d(p) can also be adopted from higher hierarchy levels, which leads to correspondingly higher error values ​​f(p).

[0072] The data obtained can be saved in common image formats. For example, the Fig.Figure 6 shows an exemplary representation in which the acquired data is stored in three different levels. The first level stores the disparity image DM, the second level stores the confidence matrix KM, and the third level stores the error matrix FM.

[0073] In the previous exemplary embodiment, both a confidence matrix determination and an error matrix determination were performed. According to the invention, it is possible to calculate a disparity image DM only in conjunction with the inventive error matrix FM or only in conjunction with the inventive confidence matrix KM. Depending on the application, the determination of the confidence matrix KM or the determination of the error matrix FM can be omitted.

[0074] Furthermore, in the preceding embodiment, the semi-global matching algorithm was used to determine the confidence values. Theoretically, it is possible to apply the general inventive concept for determining confidence values ​​to various methods in which multiple candidate disparity values ​​are determined. Likewise, in the described embodiment, it was assumed that the semi-global matching algorithm typically determines eight candidate disparity values ​​for a pixel in addition to the result disparity value. However, the method can also be modified according to the invention so that only four or even only two candidate disparity values ​​are determined per pixel p. Corresponding confidence matrices can also be calculated here.

[0075] At this point, it should be noted that all parts described above, viewed individually and in any combination, particularly the details shown in the drawings, are claimed as essential to the invention. The same applies to the method steps explained. List of reference symbols ImgL, ImgR, ImgL', ImgR' Image DM disparity image d(p), d(p1), d(p2), d(p3), d(p4) disparity value d(p), d'(p), d''(p) candidate disparity value FM error matrix f(p), f(p1), f(p2), f(p3), f(p4) error value KM confidence matrix k(p), k(p1), k(p2), k(p3), k(p4) confidence value Erg result image E(M) objective function M, M1, M2, M3 evaluation set (subset of pixels) 10 computers 12 Display device 13 Computer-readable memory 14 Processing facility 20 stereo camera 21 First Optics 22 Second optics 24 processing facility 30 robots 31 Robot arm 32 robot tools 33 Robot processing device

Claims

[1] System comprising: - at least one stereo camera (20) with a first optic (21) and a second optic (22); - at least one processing device (14, 24) which is in communicative connection with the stereo camera (20) in order to receive from the stereo camera (20) first image data (ImgL) which are assigned to the first optics (21) and second image data (ImgR) which are assigned to the second optics (22); - at least one memory (13) with instructions, where the instructions cause the processing device (14, 24): - to calculate at least one disparity image (DM) with disparity values ​​(d(p)) for a plurality of pixels (p) based on the first and second image data (ImgL, ImgR); and - to calculate a confidence value (k(p)) for at least some of the image points (p) based on the image data (ImgL, ImgR), characterized by , that the instructions cause the processing device (14, 24): - to determine at least one first candidate disparity value (d(p), d'(p), d''(p)) and at least one second candidate disparity value (d(p), d'(p), d''(p)) for at least one pixel (p); - to determine a confidence value (k(p)) for the pixel (p) by comparing the candidate disparity values ​​(d(p), d'(p), d''(p)), wherein the confidence value (k(p)) indicates a reliability for the pixel (p); and - to select and / or calculate a single disparity value (d(p)) as the resultant disparity value (d(p)) for the disparity image (DM) for the at least one pixel (p). [2] System according to claim 1, characterized by that the instructions cause the processing device (14, 24) to calculate an error matrix (FM) with error values ​​(f(p)) and / or confidence matrix (KM) with confidence values ​​(k(p)). [3] System according to claim 2 characterized by that the instructions cause the processing device (14, 24) to output a result image (Erg) with several levels, for example in GIF or TIFF format, wherein a first level comprises the disparity values ​​(d(p)) and / or a second level comprises the confidence values ​​(k(p)) and / or a third level comprises the error values ​​(f(p)). [4] System according to one of the preceding claims, characterized bythat the instructions cause the processing device (14, 24): to determine at least one first candidate disparity value (d(p), d'(p), d''(p)) and at least one second candidate disparity value (d(p), d'(p), d''(p)) for at least one pixel (p), wherein to determine the candidate disparity values ​​(d(p), d'(p), d''(p)) an objective function (E(M)) is optimized, in particular by means of dynamic programming, wherein during the optimization of the objective function for the first candidate disparity value (d(p), d'(p), d''(p)) the objective function (E(M)) is based on a first selection (M1) of pixels (p) and during the optimization of the objective function (E(M)) for the second candidate disparity value (d(p), d'(p), d''(p)) the objective function (E(M)) is based on a a second, at least partially different selection (M2) of pixels (p) is applied. [5] System according to one of the preceding claims, in particular according to claim 4, characterized bythat the objective function (E(M)) is chosen such that the objective function (E(M)) is composed of at least one cost and at least one smoothness criterion and is preferably structured as follows: E(M)=∑p∈M(C(p,d(p))+∑q∈NPg1⋅T[|d(p)−d(q)|=1]+∑q∈Npg2⋅T[|d(p)−d(q)|>1]) where (C(p, d(p))) is a cost function and the function T is equal to 1 if the condition is met and T=0 if the condition is not met and Np specifies the neighboring pixels to the pixel p. [6] System according to one of the preceding claims, characterized by that the instructions cause the processing device (14, 24): - to determine a plurality (n) of candidate disparity values ​​(d(p), d'(p), d''(p)) for a plurality of pixels (p); - to calculate a result disparity value (d(p)) for each of the pixels (p) of the disparity image (DM); and - to calculate a confidence value (k(p)) for each of the image points (p), wherein the respective confidence value (k(p)) is greater than the number of candidate disparity values ​​(d(p), d'(p), d''(p)) corresponding to the result disparity value (d(p)) divided by the total number of candidate disparity values ​​(d(p), d'(p), d''(p)) for the respective image point (p). [7] System according to one of the preceding claims, in particular according to claim 6, characterized by , that the confidence value (k(p)) for a pixel (p) is calculated according to the following formula: k(p)=G+F∑iT[|di−d(p)|≤1]n , where F is a factor between 0.2 and 0.6, in particular between 0.4 and 0.55, where G=1 - F, where n is the number of candidate disparity values ​​for the pixel (p) and d i is the i-th candidate disparity value for the pixel (p). [8] System according to one of the preceding claims, characterized bythat the instructions cause the processing device (14, 24) to calculate error values ​​(f(p)) for at least one pixel (p) in a hierarchical approach. [9] System according to one of the preceding claims characterized by that the instructions cause the processing device (14, 24): - to calculate at least a first disparity value (d(p1)) of the disparity image (DM) based on the first image data (ImgL) and second image data (ImgR) and to assign a first (constant) error value (f(p1)), e.g. 0.5, to the first disparity value (d(p1)), wherein the first and second image data are present with a first (high) resolution; - to process the first image data (ImgL) into third image data (ImgL') and the second image data (ImgR) into fourth image data (ImgR'), wherein the third image data (ImgL') and fourth image data (ImgR') have a second resolution which is lower than the first resolution; - to calculate at least a second disparity value (d(p2)) of the disparity image (DM) based on the third and fourth image data (ImgL', ImgR') and to assign a second constant error value, e.g. 1, to the first disparity value. [10] System according to one of claims 2, 8 or 9 characterized by : a robot and / or a robot arm to which the stereo camera (20) is attached, a) wherein a control unit decides, depending on the confidence value (k(p)) and / or the error value (f(p)), whether the robot and / or robot arm is moved in order to capture a specific object and / or a specific section of the environment by means of the stereo camera (20) from a further position and / or with a different orientation and / or b) wherein a control unit decides whether the robot and / or robot arm can be moved depending on the confidence value (k(p)) and / or the error value (f(p)). [11] Procedure, characterized by the steps: - capturing or receiving first and second image data (ImgL, ImgR) by means of a stereo camera (20); - Calculating, based on the first and second image data (ImgL, ImgR), at least one disparity image (DM) with disparity values ​​(d(p)) for a plurality of pixels (p); and - Calculate a confidence value (k(p)) and / or an error value (f(p)) for at least some of the image points (p) - determining at least one first candidate disparity value (d(p), d'(p), d''(p)) and at least one second candidate disparity value (d(p), d'(p), d''(p)) for at least one pixel (p); - determining a confidence value (k(p)) for the pixel (p) by comparing the candidate disparity values ​​(d(p), d'(p), d''(p)), wherein the confidence value (k(p)) indicates a reliability for the pixel (p); and - selecting and / or calculating for the at least one pixel (p) a single disparity value (d(p)) as the resultant disparity value (d(p)) for the disparity image (DM). [12] Computer readable memory with instructions for implementing the method according to claim 11 when the instructions are executed on a computing unit.

Citation Information

Patent Citations

  • Method for monitoring of stereo camera arrangement used for detecting environment of e.g. vehicle during manufacturing vehicle, involves determining confidence measure indicating efficiency and / or reliability of arrangement

    DE102011108995A1

  • Spatio-temporal confidence maps

    US20140153784A1

  • Method and apparatus for performing depth estimation

    US20150178936A1