Detection, 3D reconstruction, and tracking of multiple rigid objects moving relative to each other
Through the multi-object beam adjustment method of direct three-dimensional image alignment and photometric measurement, the problem of insufficient detection and tracking accuracy in multi-object scenes in the prior art is solved, and higher robustness and accuracy are achieved, and multiple moving objects in complex dynamic scenes can be effectively processed.
Patent Information
- Application Number
- CN202080043551.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-05
- Filing Date
- 2020-05-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-05-28
AI Technical Summary
The prior art has problems of insufficient accuracy and robustness when detecting, reconstructing and tracking multiple rigid objects moving relative to each other from image data of a single camera, especially in multi-object scenes, resulting in increased blur areas and incorrect object clustering.
The multi-object beam adjustment method using direct three-dimensional image alignment and photometric measurement is used to detect and track multiple rigid objects through sparse pixel selection and joint optimization parameters. This method combines the method in direct sparse odometer (DSO), extends the beam adjustment of photometric measurement, and adapts to the processing of multi-object clusters.
It improves the accuracy and robustness of object detection and tracking, eliminates ambiguity in object clusters, enhances the three-dimensional reconstruction ability in complex dynamic scenarios, and can effectively identify and segment multiple moving objects.
Smart Images

Figure CN113994376B_ABST
Abstract
Description
[0001] The present invention relates to a method and device for detecting, three-dimensionally reconstructing and tracking multiple rigid objects moving relative to each other from an image sequence of at least one camera, and can be used in particular for assisted driving or automatic driving in the framework of a camera-based environment detection system.
[0002] The following methods are known for detecting, 3D reconstructing and tracking objects from image data of a (single) camera:
[0003] Motion reconstruction (SFM):
[0004] A common way to extract 3D structure from video data is to use an indirect approach: As a preprocessing step, image correspondences from multiple camera images are identified. Only in a subsequent step are the epipolar geometry, 3D structure and the determination of the relative motion of the cameras detected. The term indirect approach describes the two-stage process of first computing the optical flow and then computing the 3D structure from the optical flow (Reconstruction from Motion (SFM)).
[0005] Bundle Adjustment:
[0006] Bundle Adjustment (German: ) is used to optimize structural and motion parameters with the aid of multiple images. Geometric errors in point or line correspondences, such as backprojection errors, are minimized.
[0007] Photometric bundle adjustment:
[0008] Photometric bundle adjustment (German: Photometrischer ) optimizes structure and motion using image intensity and image gradient based on a probabilistic photometric error model:
[0009] Alismail et al., Photometric Bündle Adjustment for Vision-Based SLAM, arXiv:1608.02026v1[cs.CV], 5. August 2016, published on arXiv on August 5, 2016.
[0010] Photometric bundle adjustment is applied to single object problems (e.g., moving camera + rigid, stationary environment), which corresponds to the problem of visual odometry (VO) or self-localization and mapping (SLAM).
[0011] Direct Sparse Odometry (DSO) by Engel et al. (Alismail et al., Photometric Bündle Adjustment for Vision-Based SLAM, arXiv:1608.02026v1[cs.CV], 5. August 2016) is a method that combines a direct probabilistic model (minimization of photometric error) with a consistent joint optimization of all model parameters, including the structure geometry as inverse depth of the reference image midpoint, the camera trajectory, and the affine sensor characteristic curves, focal length and principal point for each image. Tracking and photometric bundle adjustment with direct 3D image alignment / registration are used to implement visual odometry, where a static scene is assumed. For a one-time initialization, a coarse-to-fine bundle adjustment is used based on two camera images. DSO does not use keypoint correspondences, but rather uses a single camera or a stereo camera system.
[0012] SFM with multiple objects:
[0013] Known methods for 3D reconstruction of multiple objects are, for example, keypoint-based methods, in which a sparse flow field is pre-calculated, and methods based on pre-calculated dense optical flow fields are also known.
[0014] Ranftl et al. in “Dense Monocular Depth Estimation in Complex Dynamic Scenes” (DOI: 10.1109 / CVPR.2016.440) presents the reconstruction of moving objects together with their environment. To this end, motion segmentation is performed by assigning each pixel to a different motion model based on pre-computed dense optical flow.
[0015] The object of the present invention is to provide improved object detection, three-dimensional reconstruction and tracking of a plurality of objects moving relative to one another based on images from a camera or based on images from a plurality of rigidly connected cameras.
[0016] The following considerations are a starting point:
[0017] In some domains and scenarios, indirect methods are inferior to direct photometric methods in terms of accuracy and robustness. In multi-object SFM methods, the reduction in measurement accuracy leads to an increase in ambiguous areas, which in turn leads to incorrect object clustering. For example, objects that only move slightly differently cannot be identified as two objects. Ultimately, the quality of object clustering / identification of moving objects is limited by the uncertainty associated with the indirect methods that previously determined the optical flow error distribution, and sparse optical flow is also limited by the low density of the keypoint set. This leads to:
[0018] 1. The limited minimum solid angle of each motion model (→ large minimum object size / small maximum object distance),
[0019] 2. Increased the minimum deviation in the direction of motion detectable by the method, and
[0020] 3. Limited applicability in situations where there are few key points.
[0021] First, the viewpoints of the present invention and its design variations are described below:
[0022] 1. Detection and tracking of multiple rigid objects with the help of direct 3D image alignment and photometric multi-object bundle adjustment as an online method, based on the selection of key cycles and pixel selection (sparseness)
[0023] The invention extends the methods used in particular for direct sparse odometry (DSO) so that a method for object clustering, for identifying all different moving rigid objects in a camera video (the entire rigid static environment can be referred to as an object here) is combined in an adapted form with an extended photometric bundle adjustment. The result includes not only the determination of the trajectory and structure of the objects themselves in motion, but also the determination of the movement of the camera system relative to the stationary environment, as well as the determination of the structure of the stationary environment.
[0024] Online Method:
[0025] Although all parameters are jointly optimized based on image data from multiple time points, the method is suitable for simultaneous application during data detection (as opposed to using bundle adjustment as a batch processing method after data acquisition). The method is also suitable for detecting objects that are only temporarily visible.
[0026] Sparse, no regularization:
[0027] To reduce the computational effort of the photometric bundle adjustment, only those pixels are selected that are expected to have a relevant contribution or relevant constraints to solving all object trajectory estimates. Typically, these pixels are several orders of magnitude fewer than the pixels in the input image. No regularization term is needed to regularize the depth estimate, and the potential systematic errors that come with it are thus avoided.
[0028] The core of the method is the joint optimization (maximum a posteriori probability estimation) of the following parameters:
[0029] - Depth of multiple selected points of multiple selected images, representing multiple objects by inverse depth (1 parameter per point and object)
[0030] - Optional: Normal vector for each selected point of multiple objects (2 parameters per point and object)
[0031] -Number of sports models
[0032] - Trajectory of each kinematic model (pose or 3D position and 3D rotation for each key cycle)
[0033] - Assign the selected points to the motion model (1 parameter per point and motion model, implemented with soft or hard assignment)
[0034] - an estimate of the (e.g. affine) sensor characteristic curve for each image (not indicated below for readability; see e.g. Engel et al. DSO Kapitel 2.1 Calibration), and
[0035] - Estimated focal length and principal point (not indicated in the following text for readability; see for example Engel et al. DSO Kapitel 2.1 Calibration).
[0036] Minimize the error function
[0037] E:=Ephoto+Ecomp+Egeo
[0038] Among them, Ephoto is the photometric error of the selected set of uncovered pixels, Ecomp is a priori term assuming the synthesis of multiple motion model scenes, and Egeo is the prior assumption on the geometry of each object.
[0039] The photometric error term is defined as
[0040]
[0041] The photometric error of the observations in image j of point p or motion model m
[0042]
[0043] Here, m is the set of motion models, g mis an optional weight based on the prior model of the camera model geometric error, which has different degrees of influence depending on the size of the object, F is the set of all images in the dynamic bundle adjustment window, and P i is the set of all active points in image i, and obs(p) is the set of all observations of other images and point p. n is the weight of the pattern point n (the neighborhood N around p p ), I i and I j Represents the grayscale values of the two images, Project a point n into camera image j with the help of motion model m and assign inverse depth represents the probability of a point being associated with motion model m, where:
[0044]
[0045] ||.|| γ represents the Huber norm.
[0046] Since the number of motion models is usually unobservable, the minimum number should be preferred. To this end, if necessary, a prior term E is defined based on the parameter comp , and assume a probability distribution of the number of objects. For example, E comp It can be a strictly monotonically increasing function of the number of objects, or it can be based on a criterion of minimum description length.
[0047] Prior term E geo Geometric assumptions, such as compactness requirements of objects, can be expressed to resolve ambiguities / fuzziness in clustering. This is, for example, modeling probabilities for mutually different object correspondences (i.e. object boundaries) when observing each pair of adjacent points. Thus, it is preferred to perform segmentation of objects with as few object boundaries as possible. The term can also be omitted, for example, in application scenarios with little ambiguity.
[0048] To determine the observability or quantity obs(p), projections outside the image edge or with negative depth (in the target camera) are first removed. To determine the coverage caused by other structures, for example, the photometric error of each projection is evaluated analytically, or a coverage analysis is used (see "Coverage").
[0049] optimization
[0050] To optimize the error function, the Levenberg–Marquardt method is used alternately for trajectory parameters and structure parameters when the object assignment is fixed (this corresponds to a bundle adjustment of the photometric measurements of each object), and then the correspondence is optimized using, for example, the interior point method in the case of fixed geometry and fixed number of objects (when using soft assignment), or using, for example, image cutting. For this purpose, the depth parameter of each selected point is required for the external object; if it has not been optimized during the bundle adjustment, it can be optimized in advance.
[0051] In the upper optimization loop, the structure, trajectory and object assignment are first optimized in the described alternation and repetition until convergence is achieved, and then, to optimize the number of motion models, new configuration hypotheses (objects and their point assignments) are constructed in such a way that the total error is expected to be reduced. The new configuration hypotheses are evaluated analytically according to the initialization method. [See also Figure 4 ]
[0052] Key cycle management
[0053] The optimal selection of bundle adjustment images from the image data stream can be object specific. An exemplary strategy is: one object is almost stationary, → select a very low critical cycle frequency, another object is moving rapidly, → select a high critical cycle frequency.
[0054] Possible issues:
[0055] The cluster parameters cannot be optimized because the photometric error terms cannot be determined for all objects for the union of all critical cycles, since the object poses in the bundle adjustment are only determined in object-specific critical cycles.
[0056] Possible solutions:
[0057] For all objects, poses are determined for all external (not object-specific) key cycles using direct image alignment. Here, poses are determined only for these points in time without optimizing the structure. Photometric error terms for each point and each motion model can be determined for each internal and external key cycle, which are required to optimize the assignment of points to the motion model.
[0058] If the selection of the key loop of the object does not change during the loop, no new data is generated for the subsequent optimization. In this case, the loop is reduced to a pure tracking or estimation of the pose of the object by means of direct image alignment.
[0059] Hypothesis Formulation (1): Detection of an Alternative Motion Model
[0060] Detect another Motion Modelor other moving objects: a hypothesis of the object configuration (hypothesis H is a specific set or hypothesis of all model parameter values) can be constructed based on (photometric) error analysis of additional dense scattering points with optimal depth (these additional points do not participate in the bundle adjustment of the photometric measurements). Determination of fully optimal hypothesis H old The locations and times of high errors in the configuration (e.g. the last iteration) accumulate, and new hypotheses H are defined if necessary. new , where the new hypothesis contains other objects in the measured error accumulation region. The criteria for building and evaluating new hypotheses can be defined as follows:
[0061] E comp (H new )+E geo (H new )+C photo (H new ) <E(H old )
[0062] Among them, C photo (H new ) is a heuristic estimate of the expected photometric error of the new hypothesis, for example based on prior assumptions and measured error accumulation. First, only the heuristic estimate C is used photo (H new ), since the structure and trajectory of the new object are not yet exactly known. During the initial optimization, the assumptions are evaluated (and thus E is determined). photo (H new ) and E(H new )). If the total error becomes larger than that of another fully optimized hypothesis (e.g., the configuration of the last iteration), the hypothesis is discarded during the coarse-to-fine initialization. The final hypothesis that is not discarded becomes the new configuration of the current cycle.
[0063] Coverage modeling, described below, is important when forming hypotheses to avoid false positive detections due to coverage.
[0064] Hypothesis formation (2): Eliminating motion model
[0065] If it is determined that there are too many hypothesized objects, that is, the presence of some motion models will increase the overall error, then these motion models and their associated parameters will be removed from the error function. The method steps for determining whether the presence of a motion model will increase the overall error are as follows:
[0066] For each object, based on the previous configuration assumption H old Periodically construct a new configuration hypothesis H that no longer contains the object in question new . newThe optimization is performed and the overall error is determined at the same time. In general, the assumption of distant objects is usually expected to be E comp (H new ) <E comp (H old ) and E photo (H new ) <E photo (H old Then check whether E(H new ) <E(H old ), that is, whether the overall error of the new hypothesis is smaller than the overall error of the original hypothesis. If it is smaller than the overall error of the original hypothesis, the new hypothesis is adopted, that is, the motion model is eliminated.
[0067] Instead of a comprehensive optimization of the new hypothesis (i.e. a joint optimization of all model parameters), the following simplification for determining an upper bound on the overall error can be performed: Only the point assignments assigned to the removed objects are optimized and all structural and trajectory parameters are retained. This process is very fast.
[0068] Initialization of new point depths in a known motion model
[0069] The new point depths can be optimized using a one-dimensional brute force search on the discretized depth values followed by a Levenberg-Marquard optimization method. The discretization distance is adjusted in such a way that it matches the expected convergence radius of the optimization (approximately 1 pixel spacing in projection). Alternatively, a combination of a coarse-to-fine approach and a brute force search can be used to reduce runtime:
[0070] An image pyramid may be generated for an image, where, for example, pyramid level 0 corresponds to the original image (with full pixel resolution), pyramid level 1 corresponds to the image with half-pixel resolution (along each image axis), pyramid level 2 corresponds to the image with quarter-pixel resolution, and so on.
[0071] Starting from a coarse pyramid level (reduced pixel resolution), after a brute force search, point depth regions are excluded with high error values using discrete depth values (fit to the pyramid resolution). After changing to a finer pyramid level, only the point depth regions that have not been excluded are re-evaluated by brute force search. Afterwards, after the finest pyramid level is completed, the best depth hypothesis is refined with the help of the Levenberg-Marquard optimization method. Other remaining hypotheses can be marked to indicate the ambiguity of the corresponding point depth.
[0072] During initialization, coverage and other non-modeled effects have to be taken into account, e.g. by the methods in Section “Coverage”, by removing outlier projections and / or weighting projections, e.g. based on a priori assumptions that the probability of coverage depends on the time interval.
[0073] Initialization of the new motion model and its point depth
[0074] The structural and trajectory parameters of the new motion model are initially unknown and must be initialized within the convergence radius of the non-convex optimization problem.
[0075] The generation of correspondences in an image sequence (sparse or dense optical flow) is computationally intensive and prone to errors. The present invention also solves the problem of initializing a local motion model without having to explicitly compute optical flow between critical cycles.
[0076] Possible issues:
[0077] 1. The convergence region of a photometric bundle adjustment can be roughly estimated by the region in parameter space where all projections in all images are no more than about 1 pixel away from the correct projection. Therefore, all parameters of the trajectory (multiple images!) and point depths must be initialized well enough so that as many point projections as possible are within 1 pixel of the correct solution before the Levenberg–Marquardt optimization is performed.
[0078] 2. No correspondences or optical flow are available to generate initial motion estimates
[0079] A possible solution is: instead of DSO's global 2-frame coarse-to-fine approach, use a new local "multi-frame near-to-far / coarse-to-fine approach":
[0080] The local structure parameters of all critical loops are initialized with 1 and the trajectory parameters with 0. (As an alternative, a priori assumptions and superior brute force searches can be incorporated into the initial values as listed below).
[0081] first
[0082] a) points at only a coarsely selected pyramid level, and
[0083] b) Only observations that are close in time / location to the respective owner image are evaluated (eg for a point in the fifth image, observations are first evaluated only in the fourth and sixth images).
[0084] During the bundle adjustment optimization process, the resolution is now gradually increased as increasingly distant observations are evaluated. At the latest in the last iteration, the maximum resolving pyramid level and all observations are used.
[0085] The combination of a) and b) aims to significantly expand the convergence region in parameter space: Thus, only those terms are evaluated for which the current state (usually) lies in a region where the linearized minimum is a good approximation of the actual minimum.
[0086] During the local multi-frame near-to-far / coarse-to-fine initialization, structure+trajectory and point correspondences are optimized alternately. Object clustering is more accurate as the resolution increases.
[0087] Since convergence to the global minimum is not always guaranteed with the described method, a coarse-to-fine brute-force search can also be used, similar to the above method for initializing the point depths: different initial value hypotheses are started by optimizing using coarse pyramid levels and are continuously selected through error checking, so that ideally only correctly configured hypotheses are fully optimized to the finest pyramid levels.
[0088] The discretization starting values required for the coarse-to-fine brute-force search can be derived from a priori object models that propose approximate areas of typical object trajectories and depths of convex shapes, where the camera's native motion relative to the rigid background can be "subtracted". The depths of the initial points of the new motion model can also be derived from the optimized depths of old object clusters with fewer motion models.
[0089] advantage:
[0090] Besides the need to initialize all parameters for multiple frames, the advantage over the 2-frame coarse-to-fine approach of DSO is that the trilinear constraints (>=3 frames) are implicitly used in the first iteration, which first makes the constraints explicit for the linear feature class points. As a result, incorrectly assigned points are already more reliably identified as "foreign to the model" in the first iteration. In addition, the coarse-to-fine brute force search is supplemented to reduce the risk of converging to local minima (the photometric bundle adjustment problem is severely non-convex, i.e. it includes local minima).
[0091] cover
[0092] Possible issues:
[0093] Overlays are not modeled in bundle adjustment errors and lead to potential erroneous object clustering or wrong assumptions.
[0094] Modeling of coverage is difficult due to the “sparse” approach.
[0095] Possible solutions:
[0096] A very dense distribution of points used to form the hypothesis can be used to geometrically predict mutual coverage of the points. If coverage of observations is determined, these observations are removed from the error function.
[0097] In the case of multiple objects, the approximate relative scales for coverage modeling between the objects must be known, which can be estimated using specific domain model assumptions, for example, if stereo information is not available. The relative scales of the different objects can also be determined with the aid of additional coverage detection or determination of the depth order of the objects. This can be achieved, for example, by determining which point or object is in the foreground based on the photometric error of two points of two objects when there is a predicted collision or overlap.
[0098] Point selection
[0099] The points used for the (sparse) bundle adjustment are chosen in such a way that, even for small objects, constraints present in all images are used as much as possible. For example, a fixed number of points is selected for each object. This enables a very dense point selection for very small objects and thus effectively uses almost all available relevant image information for the image part of the mapped object.
[0100] This results in a heterogeneous / non-uniform point density when observing the entire solid angle of the camera image, whereas a uniform density distribution results for individual objects.
[0101] 2. Extending the method to multi-camera systems
[0102] The present invention extends the multi-object approach from point 1 to multi-camera systems: videos from (one or) multiple rigidly connected synchronized cameras are processed in a joint optimization process with potentially different intrinsic properties (e.g. focus / distortion) and detection areas.
[0103] In the context of multi-camera systems, the term "key cycle" includes the set of all camera images acquired during a camera cycle or capture time.
[0104] The error function is adjusted so that
[0105] a) Different camera models and (previously known) relative camera positions are given by different projection functions π j m Modeling
[0106] b) for each time period and motion model, estimate the position parameters (rotation and translation) relative to the camera system reference point (rather than relative to the camera center), and
[0107] c) F represents the set of all images of all cameras of the selected key cycle, and obs(p) represents the set of all images of point p observed in all cameras and key cycles (optionally, redundant observations can be removed to save computation time). Point p can be selected from all images in F.
[0108] The formulation or method uses all available constraints across all images from all cameras and makes no assumptions about the camera system configuration. It can therefore be used for any baseline, camera orientation, any overlapping or non-overlapping detection regions, and for highly heterogeneous intrinsic properties (such as telephoto optics and fisheye optics). A possible application is a camera system with wide-angle cameras calibrated in all (sky) directions and some telephoto cameras (or stereo cameras) calibrated in critical spatial directions.
[0109] Tracking with direct image alignment is extended to direct image alignment of multiple cameras. This results in the same changes as for bundle adjustment of multi-camera photometry:
[0110] The trajectory optimization is performed about a reference point of the camera system (not about the camera center) while minimizing the sum of the photometric errors in all cameras. Here, all available constraints are used, including the photometric errors of the projection between the cameras. Here, too, the projection function must be specifically adapted to the respective camera model and relative position in the camera system.
[0111] initialization:
[0112] Since minimization of the new error function is also part of the initialization of the configuration hypothesis, all available constraints for all cameras are also used in the initialization phase. For example, the scale of objects in the overlapping area is automatically determined in this way. If the estimated scale deviates significantly from the correct value, objects that were initialized in one camera and then come into the field of view of the second camera must be reinitialized if necessary.
[0113] 3. Visual odometry with higher accuracy and stronger scaling capabilities
[0114] Both the segmentation of the rigid background and the use of multi-camera optimization improve the accuracy and robustness of visual odometry relative to DSO, especially in difficult scenes where most of the image consists of moving objects or in scenes with little structure in a single camera.
[0115] In camera systems with static or dynamic overlapping regions, the absolute scale of the visual odometry can be determined based on an evaluation of inter-camera observations of a point when the scale of the relative positions of the cameras is known.
[0116] 4. Automatic calibration of intrinsic photometric parameters, intrinsic geometric parameters and external parameters
[0117] Vignetting can be approximated or modeled parametrically. The same applies to the model of the sensor characteristic curve. The parameters of each camera that is desired can be optimized in the above-mentioned direct multi-object bundle adjustment. Due to the high accuracy of the structure and trajectory estimation and due to the modeling of objects that are themselves moving, the accuracy of the model optimization is expected to increase, for example compared to a combination of pure visual odometry.
[0118] Distortion modeling and determination of intrinsic geometric parameters: The parameters of each camera that one wishes to derive can be optimized in the above-described direct multi-object bundle adjustment. Due to the high accuracy of the structure and trajectory estimates, and due to the modeling of inherently moving objects, the accuracy of the model optimization is expected to be improved, for example compared to a combination of pure visual odometry.
[0119] Estimation of extrinsic parameters: The relative positions of the cameras to each other can be optimized in the above-described direct multi-object bundle adjustment. Due to the high accuracy of the structure and trajectory estimates, and due to the modeling of intrinsically moving objects, the accuracy of the model optimization is expected to be improved, e.g. compared to a purely visual multi-camera odometry combination.
[0120] It should be noted here that if a metric reconstruction is to be carried out, at least one distance between the two cameras must be determined as an absolute metric reference in order to avoid scaling drifts.
[0121] The initial values for all parameters of the camera calibration must be determined in advance and provided to the specified method. It must be ensured that the parameter vector is within the convergence range of the error function for the coarsest pyramid level due to sufficient accuracy of the initial values. The initial values can also be incorporated into the error function together with a prior distribution to prevent ambiguities that depend on the application. In addition, the constraints for the calibration parameters that are deleted when discarding / replacing critical cycles can be retained in linearized form using the marginalization methods used in the DSO.
[0122] 5. Fusion with other sensors and methods
[0123] a. Fusion with other object recognition methods such as pattern recognition (deep neural networks, etc.) has high potential because the error distributions of the two methods are largely uncorrelated.
[0124] Exemplary application cases are object detection, 3D reconstruction and tracking for integration with pattern recognition-based systems in autonomous vehicles with stereo cameras and surround view camera systems.
[0125] b. Fusion with inertial sensor systems and vehicle odometry promises high potential for solving native motion estimation (== 3D reconstruction of the static environment of the “object”) in critical scenarios and determining absolute scaling.
[0126] c. Fusion with other environment detection sensors, especially radar and / or lidar.
[0127] 6. Application
[0128] Applications from points 1 to 5 are used for the detection and tracking of moving traffic participants, the reconstruction of the rigid, stationary vehicle environment, and the estimation of the vehicle's own motion by driver assistance (ADAS) systems or automated driving (AD, full name: automated driving) systems.
[0129] The applications of points 1 to 5 are used for environmental detection and support for self-positioning in autonomous systems such as robots or drones, support for self-positioning of virtual reality (VR) glasses or smartphones, and 3D reconstruction of moving objects in monitoring (such as fixed cameras for traffic monitoring).
[0130] Advantages of the present invention and variations of its design
[0131] 1. The proposed method does not require local correspondence retrieval as a preprocessing step, which is a complex, error-prone and time-consuming task.
[0132] 2. The proposed method can significantly improve the accuracy of all estimates in some critical cases compared to indirect methods. In multi-object clustering, the improved accuracy of motion estimation can eliminate ambiguities, such as separating / identifying two objects with almost the same motion or with almost the same direction of motion in a camera image.
[0133] 3. The locking behavior of the direct photometric method facilitates the convergence of the solution of a single object problem to the primary motion model (as opposed to convergence to an erroneous "compromise" solution) in the case of multiple simultaneous motion models, and the secondary motion model can then be identified as a simultaneous model. This behavior has a beneficial effect on distinguishing motion models and improves the convergence of the multi-object problem to the correct overall solution.
[0134] This feature does not exist in the commonly used indirect methods.
[0135] 4. Improved visual odometry by identifying moving objects: Moving objects are a disruptive factor in traditional methods (such as DSO). In the new method, moving objects are automatically identified and removed from their own motion estimation based on the static environment.
[0136] 5. The described method allows sampling of pixels from contrast-affected areas with an almost arbitrary density. Together with a relatively high accuracy of motion and structure estimation, this enables detection and in particular tracking of objects with relatively small solid angles and relatively low resolutions.
[0137] This feature is also not available in the commonly used indirect methods.
[0138] 6. By using multi-camera expansion, the detection area can be expanded and / or the resolution can be increased within a certain angle range, which will increase the robustness and accuracy of the overall solution accordingly. In addition:
[0139] a. A high degree of accuracy and robustness of native motion estimation is achieved by using cameras with a combined detection solid angle that is as large as possible (e.g. a multi-camera system that covers a full 360 degrees horizontally).
[0140] b. The combination of a) with one or more cameras with higher coverage / higher resolution (telephoto cameras) also allows for more accurate measurement of the trajectories of distant objects, which are distinguished from the robust, accurate estimation of their own motion (or relative motion to a static environment) achieved by a).
[0141] c. If the relative position of the two cameras is known, there is an overlap region of the field of view of the two cameras, where the absolute scaling of the structure can be observed. By using the concepts of a) and b), absolute distance estimation can also be performed based on the overlap region of highly heterogeneous cameras, such as telephoto cameras and fisheye cameras.
[0142] d. The stereo depth information present in the overlapping area significantly simplifies the identification of moving objects and can also be used in situations that are ambiguous for the monocular case, for example, it can be used in situations where objects have the same direction of movement but different speeds, which is not uncommon in road traffic.
[0143] 7. Dynamic estimation of relevant camera parameters: Compared with a one-time calibration, automatic calibration of intrinsic photometric parameters, intrinsic geometric parameters and extrinsic parameters significantly improves the accuracy of the calibrated parameters.
[0144] The method according to the invention (implemented by a computer) for detecting, three-dimensionally reconstructing and tracking a plurality of objects moving relative to each other from an image sequence of at least one camera comprises the following steps:
[0145] a) selecting images of specific recording time points (=critical cycles) from the image sequence of the at least one camera,
[0146] b) based on sparsely selected pixels in the key loop, jointly optimizing all parameters of a model for describing the rigid objects moving relative to each other based on the image of the key loop, wherein the model includes parameters for describing the pose, number, and three-dimensional structure of the rigid objects in the key loop, and parameters for describing the correspondence between the selected pixels and the rigid objects, for which
[0147] c) minimization of an error function (S20), wherein the error function comprises a photometric error E associated with the image intensity of a plurality of critical cycles photo , and a first a priori energy term E related to the number of rigid objects comp ,as well as
[0148] d) Recursively output the number, 3D structure and trajectory of the (currently) detected rigid objects in the image sequence.
[0149] At least one camera may be a single monocular camera or a multi-camera system. The camera or multi-camera system may be arranged in particular in a vehicle for detecting the vehicle environment during driving of the vehicle. In the case of a vehicle-mounted multi-camera system, the system may be in particular a stereo camera system or a surround view camera system, wherein, for example, four satellite cameras are mounted on four sides of the vehicle and have a large opening angle to ensure 360-degree detection of the vehicle environment, or may be a combination of the two camera systems, a stereo camera system and a surround view camera system.
[0150] Typically, the entire stationary background is selected as one of a plurality of rigid objects that move relative to one another. In addition to the rigid, stationary environment, at least one further rigid object that moves itself is detected, reconstructed in three dimensions, and tracked (traced). Thus, the rigid object that moves itself moves relative to the stationary "background object". If at least one camera also moves during the recording of the image sequence, the stationary background object moves relative to the camera, and typically the rigid object that moves itself also moves relative to the camera.
[0151] The optimization in step a) is performed based on sparsely selected pixels or on sparse sets of pixels, i.e., not based on all pixels of an image or image portion ("dense"), nor based on partially dense ("semi-dense") selection of image regions. For example, J.Engel et al. proposed a "semi-dense" depth map method in a paper entitled "LSD-SLAM: Large-Scale Direct Monocular SLAM (LSD-SLAM: Large-Scale Monocular Instant Localization and Mapping Method Based on Direct Method)" published at the European Conference on Computer Vision (ECCV) in September 2014. In particular, each corresponding pixel that contributes to the reconstruction of the motion can be selected, for example, by the minimum spacing from other points, and stands out from their direct environment in a characteristic manner, so that they can be well identified in the figure below. (Assumption) The three-dimensional structure of an object corresponds to the spatial geometry of the object. The posture of an object corresponds to the position and orientation of the object in three-dimensional space. The time course of the object's posture corresponds to the trajectory of the object. The output of parameters for determining the number of objects, the three-dimensional structure and the trajectory can preferably be performed in a cyclical cycle, in particular "online", meaning in real time or continuously while receiving new images from at least one camera. Images can be processed "as fast as new images appear".
[0152] According to another preferred embodiment of the method, the error function includes a second prior energy term E geo , the second prior energy term is related to the three-dimensional structure of the rigid object.
[0153] The error function preferably comprises the following (model) parameters:
[0154] Inverse depth for each selected pixel in each single motion model;
[0155] the number of motion models, where a motion model is assigned to each currently assumed rigid object;
[0156] The pose (3D position and 3D rotation, i.e. 6 parameters) of each single motion model and activated key loop; and
[0157] The probability of correspondence between each selected pixel and each motion model. After optimization, the probability of the selected pixel being assigned is equal to 1 for the motion model and equal to 0 for other motion models.
[0158] Optionally, the normal vector of each selected pixel of each motion model is considered as an additional parameter.
[0159] The error function preferably includes the following additional (model) parameters:
[0160] The sensor characteristic curve for each image, and
[0161] The focal length and the principal point of each camera (see for example DSO chapter 2.1 "Calibration" by Engel et al.), so that the joint optimization of all parameters results in an automatic calibration of the at least one camera.
[0162] Preferably, direct image alignment is performed using one or more image pyramid levels to track the individual objects. Here, the relative 3D position and 3D rotation (pose) of the objects visible in the loop can be estimated by 3D image registration based on the images and depth estimates of other loops and optionally by a coarse-to-fine approach.
[0163] In an advantageous manner, for the optimization of the error function, a photometric bundle adjustment is performed alternately with an optimization of trajectory parameters and structure parameters according to object-specific key cycles (per motion model and the pose of the key cycle) and an optimization of the correspondence between pixels and motion models. The selection of the key cycles selected from the image sequence for use in the photometric bundle adjustment can be made motion model-specific. For example, the frequency (of the selected images) can be adapted to the relative motion of the object.
[0164] Preferably, the number of motion models is subsequently optimized, wherein, in case of adding a motion model to the error function or removing a motion model from the error function, the selected pixels are reassigned to the motion model and the optimization of the error function is restarted.
[0165] Preferably, the at least one camera is moved relative to the object against a stationary, rigid background.
[0166] In a preferred embodiment of the method, a plurality of image sequences are captured by means of a camera system comprising a plurality of synchronized (on-board) cameras and the image sequences are provided as input data to the method. All parameters are jointly optimized to minimize the resulting error function. The model parameters include the pose of each object relative to the camera system (rather than relative to one camera). The pixels may be selected from the images of all cameras. The pixels are selected from the images of the critical cycle of at least one camera.
[0167] For the selected pixels, the observations from at least one camera and at least one critical cycle are accepted as energy terms for the photometric error. In this case, preferably individual geometric camera models and photometric camera models as well as the relative positions of the cameras to one another are taken into account.
[0168] Furthermore, it is preferred to perform multi-camera direct image alignment for tracking of a single object using one or more pyramid levels. For this purpose, the following images are preferably used:
[0169] a) All images for all loops whose pose is known and the selected points have known depth. These images are repeatedly combined and warped here (building a prediction of the expected image for each camera in the loop that retrieves the pose).
[0170] b) All images of the cycle of gestures are retrieved. These images are each repeatedly compared with the predicted image from a) for this camera combination.
[0171] The model parameters preferably include other intrinsic photometric parameters, other intrinsic geometrical parameters and / or extrinsic camera parameters of the at least one camera, so that a joint optimization of all parameters is applied to the automatic calibration of the at least one camera. In other words, the camera intrinsic photometric parameters (e.g. vignetting and sensor characteristic curve), intrinsic geometrical parameters (e.g. focal length, principal point and distortion) and / or extrinsic model parameters (e.g. relative position of the cameras to one another) are automatically calibrated / automatically optimized. Vignetting, sensor characteristic curve and distortion are preferably approximated parameterized. Thus, all new model parameters can be determined together when the error function is minimized (in a process).
[0172] Another subject of the invention relates to a device for detecting, three-dimensionally reconstructing and tracking a plurality of rigid objects moving relative to one another from a sequence of images of at least one (onboard) camera received by an input unit. The device comprises an input unit, a selection unit, an optimization unit and an output unit.
[0173] The selection unit is configured to select images of a plurality of acquisition times (=key cycles) (determined by the selection unit) from the image sequence.
[0174] Optimize cell configuration for,
[0175] a) based on sparsely selected pixels in the key loop, all model parameters of a model for describing rigid objects moving relative to each other are jointly optimized according to the image of the key loop, wherein the model includes parameters for describing the pose, number, and three-dimensional structure of the rigid objects in the key loop, and parameters for describing the correspondence between the selected pixels and the rigid objects, for which
[0176] b) Minimize the error function, where the error function includes the photometric error E photo and the first a priori energy term E related to the number of rigid objects comp .
[0177] The output unit is configured to cyclically output the number, the three-dimensional structure (geometry) and the trajectory of the rigid objects moving relative to one another detected by the optimization unit from the image sequence.
[0178] The device may include, in particular, a microcontroller or microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC (application-specific integrated circuit), an FPGA (field programmable gate array) and other similar components, interfaces (input units and output units), and software for executing the corresponding method steps.
[0179] Thus, the invention may be implemented in digital electronic circuitry, computer hardware, firmware, or software.
[0180] The following is a more detailed description of the embodiments and drawings.
[0181] Figure 1 a shows a series of five images from the camera on the left side of the vehicle;
[0182] Figure 1 b shows the three-dimensional reconstruction of the vehicle environment;
[0183] Figure 1 c shows the 3D reconstruction of the first (self-moving) rigid object;
[0184] Figure 1 d shows the 3D reconstruction of the second (stationary) rigid object corresponding to the stationary background;
[0185] Figure 2 Four schematic camera images showing the panoramic (surround view) system of the host vehicle (bottom) and a 3D point reconstruction of the host vehicle's environment (top);
[0186] Figure 3 A schematic flow chart of a method for a multi-camera system is shown;
[0187] Figure 4 A schematic flow chart showing a method for data selection and minimization of the error function for each image cycle; and
[0188] Figure 5 A host vehicle is shown with a surround view camera system, a front telephoto camera and a device for detecting, three-dimensionally reconstructing and tracking a plurality of rigid objects moving relative to each other.
[0189] Figure 1a shows a series of five images (L0, L1, ..., L4), which were taken by the left camera of the vehicle at shooting time points t0, ..., t4 while the vehicle was driving. In images L0, ..., L4, a vehicle 19 driving in the overtaking lane on the left side of the vehicle can be seen. The edge of the left lane is defined by the wall 11. Behind it, at one-third of the image, the trees next to the lane can be glimpsed. The wall 11, the trees, the lane and the lane markings are part of the stationary environment of the vehicle. The entire stationary environment of the vehicle is regarded as a rigid object. The imaged vehicle 19 is a rigid object that moves relative to the first object (of the stationary environment). The imaged vehicle 19 is driving faster than the vehicle, that is, it is overtaking the vehicle.
[0190] If the method is to be implemented based on only one camera, each image corresponds to a (recording) cycle. If these five images are five key cycles for imaging the vehicle (=object to which the motion model is assigned), these key cycles are referred to as motion model-specific key cycles.
[0191] Figure 1 b shows the three-dimensional reconstruction of the scene achieved according to an embodiment of the method. However, for this three-dimensional reconstruction, not only Figure 1 The key loop captured by the left camera is shown in the summary, and the key loop captured by the rear camera, front camera and right camera synchronized by the camera system at the same shooting time points t0, ..., t4 is used. This will be combined later Figure 2 Further explanation. Figure 1 In b, we can see points that do not fully reflect the three-dimensional relationship but are roughly identifiable (sparse). Figure 1 The 3D reconstruction is performed from a perspective approximately from above, with different camera orientations. (Other) vehicles that can be well estimated in spatial form are imaged, and aspects of the rigid environment, in particular walls, are imaged as two parallel lines behind or above the vehicle. There are individual points on the roadway.
[0192] Figure 1 cOnly show the Figure 1 3D reconstruction of another vehicle 29 of b. The vehicle 29 is a moving rigid object. The method enables reliable tracking of the vehicle 19 from the image series L0, ..., L4. In addition to the three-dimensional position and dimensions, the trajectory of the vehicle 19 can also be determined from the tracking, i.e., in particular the speed and rotation in all three spatial directions.
[0193] Figure 1 d shows only Figure 1b) 3D reconstruction of the stationary (immobile) rigid environment of the vehicle. The stationary rigid environment of the vehicle is also considered as a (relatively) moving rigid object. Here, the position of the vehicle in this environment is given directly. The relative motion of the reconstructed environment determined by this method is the same as the inverse of the vehicle's own motion. Figure 1 The three-dimensional reconstruction of the wall 11 of a is identified as two parallel lines 29.
[0194] Figure 2 Four schematic camera images L10 , F10 , R10 , H10 of the vehicle's panoramic (surround view) system are shown below, and the three-dimensional point reconstruction of the vehicle's environment is shown above.
[0195] At the bottom left, the corrected image L10 of the onboard camera looking to the left can be seen. Also shown next to it are the corrected images F10, R10, H10 of the onboard cameras looking forward, right and backward. It can be seen that all four images L10, F10, R10, H10 have a black road surface with white lane markings 12, 13, 15, 16 in the respective field of view. Another vehicle 19 is driving diagonally to the left in front of the vehicle. The rear of the other vehicle 19 is detected by the image L10 of the left camera and the front by the image F10 of the front camera. The imaged vehicle 19 is a rigid object that moves itself. In the image L10 of the left camera, the wall 11 can be identified as the lane edge boundary between the lane and the surrounding landscape (trees, hillside) of the lane. Below the wall 11, the lane limit marking (line) 12 with a solid line is imaged, which defines the left lane edge of the three-lane carriageway. In the image F10 of the panoramic system's front camera, the dashed lane markings on the left 13 and right 15 are imaged, which delimit the left and right edges of the middle lane in which the vehicle is currently traveling. The right edge of the lane is identified by another solid lane limit marking 16. In the image R10 of the right camera, the guardrail 17 is imaged as a lane edge limit, below which the right lane limit marking 16 can be identified. From the image H10 of the rear camera, it can also be seen that the vehicle is traveling in the middle lane of the three-lane system, where the right lane marking 15 of the vehicle on the left in the image and the left lane marking 13 of the vehicle on the right in the image can also be identified as dashed lines between the two solid lane limit markings (not numbered in the image R10). In the upper part of all four images, the sky can be seen. The wall 11, the lane markings 12, 13, 15, 16 and the guardrail 17 are part of the stationary environment of the vehicle. The entire stationary environment of the vehicle is regarded as a rigid object.
[0196] Each of the four cameras records a series of images (videos) while the vehicle is moving. From these image series, a three-dimensional reconstruction of the scene is achieved according to the method embodiment with multiple (synchronous) cameras. Figure 2The points representing the three-dimensional relationship can be seen above. The visualization is done from a bird's eye view (top view). The trajectory of the vehicle so far is represented by a solid line 24. This line is not part of the three-dimensional structure, but is used to visualize the reconstructed trajectory of the camera system of the vehicle. The right end 28 of the line 24 corresponds to the current position of the vehicle, which is Figure 2 Not shown. On the left side in front of your vehicle (or on the top Figure 2 In the upper right of the center, the outline of another vehicle 29 can be seen. Moving objects can be tracked robustly and very accurately, so that their properties can be determined for the driver assistance system or the automatic driving system of the host vehicle. The three-dimensional reconstruction can include the following as components of the stationary background (from top to bottom): the marking wall (left lane edge limit) as a slightly dense and slightly extended line 21 (which is composed of points), the left solid lane limit marking 22, the left dashed lane marking 23 of the host vehicle, the right dashed lane marking 25 of the host vehicle, the right solid lane limit marking 26 and the line 27 with individual guardrail posts again slightly denser and more extended. The emergency lane of the carriageway is located between the right solid lane dividing marking 26 and the guardrail "line" 27.
[0197] Figure 3 The exemplary embodiment of the method for a multi-camera system is shown as an example. A similar method can also be used for various variants on a monocular camera system.
[0198] In a first step S12, the parameters of the error function are initialized. The error function is used to calculate the error of each image according to the parameters. Thus, minimization of the error function provides the parameters of the model that best matches each image. The parameters include:
[0199] - Depth parameters for multiple points in multiple images of multiple objects
[0200] - Optional: Normal vector for each selected point (2 parameters per point)
[0201] -Number of sports models
[0202] - Multiple motion models (3+3 parameters for position and rotation per time step respectively), where a motion model is assigned to an object. A rigid background (i.e. a stationary environment in real space) is also considered as an object. Background objects are also assigned a motion model.
[0203] - Assignment of points to motion models (1 parameter per point and motion model, implemented by soft assignment or optionally hard assignment)
[0204] - estimate the sensor characteristic curve, and
[0205] - Estimate focal length and principal point.
[0206] The parameters can be initialized by selecting the number of motion models as 1, initializing the trajectory with 0, initializing the inverse depth with 1, and thus implementing a coarse-to-fine initialization.
[0207] In step S14, new individual images of a cycle are obtained from a plurality of synchronous cameras. A cycle describes a set of images created by the synchronous cameras in a shooting cycle (corresponding to a shooting time point). The new individual images are provided to the method or system, for example, by cameras, storage or similar products.
[0208] In the following step S16, a direct image alignment of multiple cameras is performed for each currently existing motion model (corresponding to the currently assumed object or the current object hypothesis) in order to determine the motion parameters in the current cycle (and the new individual images). For example, it can be assumed that a moving rigid object currently moves relative to a stationary rigid background. Since the stationary background is also considered as a moving rigid object, this is the simplest case for multiple, i.e. two, rigid objects moving in different ways. The camera can perform a movement relative to the stationary background, so that in each image sequence, the background is not stationary in the camera system coordinate system, but performs a relative movement. Each currently assumed object is described by a motion model. The (pose) parameters of each object of the new (i.e. current) cycle are determined with the help of direct image alignment of multiple cameras.
[0209] Direct image alignment is different from bundle adjustment, but it has something in common with photometric bundle adjustment: the photometric error function to be minimized is the same. In direct image alignment, the depth is not optimized, but assumed to be known, while the new pose is only estimated during the minimization of the photometric error (grey value difference). Here, a prediction of the new cycle image is iteratively generated with the help of image warping or similar 3D rendering (based on old images, known structure, trajectory) and the latest object pose is matched until the prediction is most similar to the new image. For more information on single-camera direct image alignment based on holography, see, for example: https: / / sites.google.com / site / imagealignment / tutorials / feature-based-vs-direct-image-alignment (Call time: March 12, 2019).
[0210] Then, in step S20, data (key cycles, pixels) are selected and the error function is minimized. The details of this aspect will be explained in more detail below. The parameters obtained in this way are output in the next step S22. Then step 14 can be continued, that is, each new image of the new cycle is obtained.
[0211] Figure 4 A schematic diagram of a method flow is shown for data selection and error function minimization for each image loop ( Figure 3S20 in the flow) and for subsequent parameter output (S22).
[0212] In a first step S200, a key cycle of each motion model (which corresponds to an object) is selected from the set of all camera cycles.
[0213] In step S201 , points are selected in the images of key cycles of all motion models.
[0214] In step S202, new parameters of the error function for describing other point depths and point correspondences are initialized.
[0215] In step S203, the motion parameters and structural parameters of each object are optimized according to the key cycles specific to the object by means of photometric bundle adjustment.
[0216] In step S204 , multi-camera direct image alignment is performed for the key loop outside the object.
[0217] In step S205, the correspondence between pixels and objects or motion models is optimized.
[0218] In the subsequent step S206, it is checked whether (sufficient) convergence has been achieved. If sufficient convergence has not (yet) been achieved because the point correspondence has been changed, the process continues with step S200.
[0219] If convergence has been achieved, then in the subsequent step S207 the number of motion models (objects) and the correspondence between pixels and motion models are optimized.
[0220] In the following step S208, it is checked whether (sufficient) convergence of the correlation has been achieved.
[0221] If the numbers do not match, the number of motion models is checked in the subsequent step S209.
[0222] If the number is too high, the motion model and its associated parameters are removed in step S210 and the method is continued with step S200. This can be done in the following way:
[0223] For each object, a new configuration hypothesis is regularly evaluated that no longer contains the object. A check is performed to see whether the total error is reduced as a result. If this reduces the total error, the configuration is accepted or the specified object is deleted.
[0224] This new upper limit on the total error can be determined by optimizing only the point assignments of the relevant points and retaining all structural and trajectory parameters. This method is very fast (compared to a complete optimization of such a new hypothesis with missing objects). For this, please also refer to the above section "Hypothesis Formation (2): Eliminating the Motion Model".
[0225] If the number is too small, new parameters for describing other motion models (objects) of the error function are initialized in step S211 (see above: "Formulation of hypotheses (1): Detection of other motion models") and the method is continued with step S200.
[0226] As long as the quantities match, ie convergence has been reached in step S208 , the parameters are output in step S22 .
[0227] Figure 5 The vehicle 1 is shown, which has a panoramic surround camera system, a front telephoto camera and a device 2 for detecting, three-dimensionally reconstructing and tracking multiple rigid objects moving relative to each other. The detection areas of the four single cameras of the panoramic surround system are illustrated by four triangular areas (L, F, R, H) around the vehicle 1. The triangular area L (F, R or H) on the left side (front, right or rear) of the vehicle corresponds to the detection area of the left side (front, right or rear) camera of the panoramic surround camera system. In the windshield area of the vehicle 1, a telephoto camera is arranged, and its detection area T is indicated by a dotted triangle. The telephoto camera can be, for example, a stereo camera. The camera is connected to the device 2 and transmits the captured image or image series to the device 2.
Claims
1. A method for detecting, three-dimensionally reconstructing and tracking a plurality of rigid objects (11, 13, 15, 16, 17; 19) moving relative to one another from an image sequence of at least one camera, the method comprising the following steps: a) selecting an image at a specific capture time point from the image sequence of the at least one camera, the key cycle comprising a collection of all camera images acquired at the capture time point, b) jointly optimizing all parameters of the model used to describe the rigid object (11, 13, 15, 16, 17; 19) based on the image of the key loop, based on sparsely selected pixels in the key loop, in, The model includes parameters for describing the pose, number, and three-dimensional structure of rigid objects (11, 13, 15, 16, 17; 19) in the critical loop, and parameters for describing the correspondence between selected pixels and rigid objects (11, 13, 15, 16, 17; 19), for which c) minimizing an error function (S20), wherein the error function comprises a photometric error E associated with the intensity of the image of the plurality of critical cycles photo , and a first a priori energy term E related to the number of rigid objects (11, 13, 15, 16, 17; 19) comp ,as well as d) Output the number, 3D structure and trajectory of rigid objects (11, 13, 15, 16, 17; 19) detected from the image sequence.
2. The method according to claim 1, It is characterized in that The error function includes a second prior energy term E geo , the second prior energy term is related to the three-dimensional structure of the rigid object (11, 13, 15, 16, 17; 19).
3. The method according to claim 1 or 2, It is characterized in that The error function includes the following model parameters: Inverse depth for each selected pixel in each single motion model; the number of motion models, wherein a motion model is assigned to each currently assumed rigid object (11, 13, 15, 16, 17; 19); The poses for each individual motion model and key cycles activated; and The probability of correspondence of each selected pixel with each motion model.
4. The method according to claim 3, It is characterized in that The error function also includes the following model parameters: The sensor characteristic curve for each image, and The focal length and the principal point of each camera, and thus the joint optimization of all parameters, results in an automatic calibration of the at least one camera.
5. The method according to claim 1 or 2, It is characterized in that Direct image alignment is performed using one or more image pyramid levels to track individual objects (11, 13, 15, 16, 17; 19).
6. The method according to claim 1 or 2, It is characterized in that To optimize the error function, bundle adjustment by photometry is performed alternately for optimization of trajectory parameters and structural parameters according to object-specific key cycles ( S203 ) and optimization of the correspondence between pixels and motion models ( S205 ).
7. The method according to claim 6, It is characterized in that The number of motion models is then optimized (S207), wherein, in case a motion model is added to the error function or removed from the error function, the selected pixels are reassigned to the motion model and the optimization of the error function is restarted.
8. The method according to claim 1 or 2, It is characterized in that The at least one camera is moved relative to an object (11, 13, 15, 16, 17) relative to a stationary, rigid background.
9. The method according to claim 1 or 2, It is characterized in that A plurality of image sequences are captured by means of a camera system comprising a plurality of synchronized cameras, wherein the model parameters comprise the pose of each object (11, 13, 15, 16, 17; 19) relative to the camera system, wherein pixels can be selected from all cameras, wherein pixels are selected from the image of a critical cycle of at least one camera, wherein for the selected pixels the observations in at least one camera and at least one critical cycle are used as energy terms of the photometric error, and a joint optimization of all parameters is performed in order to minimize the resulting error function.
10. The method according to claim 9, It is characterized in that To track the individual objects (11, 13, 15, 16, 17; 19), a direct image alignment of multiple cameras is performed using one or more image pyramid levels, wherein a joint optimization of all model parameters is performed to minimize the resulting error function, wherein the model parameters include the pose of each object (11, 13, 15, 16, 17; 19) relative to the camera system, wherein for a selected pixel, the observation in at least one camera of the cycle to be optimized is used as an energy term of the photometric error.
11. The method according to claim 4, It is characterized in that The model parameters also include other intrinsic photometric parameters, other intrinsic geometrical camera parameters and / or extrinsic camera parameters of the at least one camera, such that a joint optimization of all parameters results in an automatic calibration of the at least one camera.
12. A device (2) for detecting, three-dimensionally reconstructing and tracking a plurality of rigid objects (11, 13, 15, 16, 17; 19) moving relative to one another from a sequence of images of at least one camera received by an input unit, the device comprising the input unit, a selection unit, an optimization unit and an output unit, in, Select the unit configuration for, a) selecting an image at a specific shooting time point from an image sequence of at least one camera, the key cycle including a set of all camera images acquired at the shooting time point; Optimized unit configuration for b) based on sparsely selected pixels in the key cycle, all model parameters of a model for describing rigid objects (11, 13, 15, 16, 17; 19) moving relative to each other are jointly optimized according to the image of the key cycle, wherein the model includes parameters for describing the posture, number, and three-dimensional structure of the rigid objects (11, 13, 15, 16, 17; 19) in the key cycle, and parameters for describing the correspondence between the selected pixels and the rigid objects (11, 13, 15, 16, 17; 19), for this purpose c) Minimize the error function, where the error function includes the photometric error E photo and the first a priori energy term E related to the number of rigid objects (11, 13, 15, 16, 17; 19) comp , Output unit configuration for d) Output the number, 3D structure and trajectory of rigid objects (11, 13, 15, 16, 17; 19) detected from the image sequence.
Citation Information
Patent Citations
Method and device for determining correspondence, preferably for the three-dimensional reconstruction of a scene
CN101443817A
Real-time camera tracking method for dynamically-changed scene
CN103646391A