Methods for evaluating and / or improving a depth map of a monitoring area, as well as the arrangement of depth maps for implementing the method
By employing monocular depth estimation and camera parameter scaling, the method generates precise and user-friendly 3D representations of monitored areas, addressing the challenges of creating accurate three-dimensional surveillance models with multiple cameras.
Patent Information
- Application Number
- DE102024206367
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-05
- Publication Date
- 2026-01-08
AI Technical Summary
Existing surveillance systems struggle to create accurate and intuitive three-dimensional representations of monitored areas, especially when using multiple cameras, as traditional methods lack precision and require extensive knowledge of camera setups and monitored areas.
A method involving monocular depth estimation followed by scaling using intrinsic and extrinsic camera parameters, along with object detection and trajectory analysis, to generate a comprehensive and accurate depth map, which is then fused with image information to create an intuitive 3D visualization model.
This approach produces a highly accurate and user-friendly 3D representation of monitored areas, allowing for improved depth map evaluation and intuitive understanding, even when using multiple cameras with varying setups.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for evaluating and / or improving a depth map of a monitoring area and a depth map arrangement for implementing the method. State of the art
[0002] Camera surveillance is used to monitor public or private spaces. It is possible to move the cameras or zoom in on specific areas. In the field of technical image processing, as well as in surveillance applications, depth information from the monitored areas is also used. In this way, the two-dimensionally mapped surveillance area is transformed into a three-dimensional surveillance space.
[0003] Patent application US 10,346,996 B2 discloses techniques and systems for determining image depth from semantic labels. In one or more implementations, a digital media environment includes one or more computing devices to control a determination of depth within an image. Regions of the image are semantically labeled by the one or more computing devices. At least one of the semantically labeled regions is decomposed into a plurality of segments, formed as planes that generally extend perpendicular to a base plane of the image. The depth of one or more of the multiple segments is then derived based on the relationships between the respective segments and their respective positions on the base plane of the image.A depth map is created that describes the depth for the at least one semantically labeled area, at least partially, based on the derived depths for the one or more segments of the multitude of segments. Disclosure of the invention
[0004] The invention relates to a method for evaluating and / or improving a depth map of a monitoring area with the features of claim 1, a depth map arrangement with the features of claim 12, a computer program with the features of claim 13, and a computer-readable data carrier with the features of claim 14. Preferred advantageous embodiments of the invention are described in the dependent claims, the following description, and the accompanying figures.
[0005] The invention relates to a method for evaluating and / or improving a depth map of a monitoring area.
[0006] The monitored area can be located in a commercial, private, and / or public area. The monitored area can be continuous, partially continuous, or segmented with open spaces between the segments.
[0007] A comprehensive depth map containing depth information for the monitored area is provided. This depth map can be displayed as a matrix, with the depth information encoded in the matrix points. From a data processing perspective, the depth information is encoded in the same way as color information in an image, particularly in raster graphics. For example, the depth information is encoded in shades of gray, where each shade of gray corresponds to a depth. The "depth" is expressed either as a radial distance, i.e., a radial distance to the camera, or as an actual or standard depth (in the narrower sense), i.e., a distance in the camera's main line of sight. Both distance specifications (radial or along a single axis / depth (in the narrower sense)) are common in the literature. However, the distance along a single axis is most frequently specified.The depth map is specifically designed as a scaled depth map in which the depth information is encoded in absolute values, e.g. in meters.
[0008] The overall depth map can also be represented as a 3D point cloud, 3D map, voxel grid, 2D elevation map, or 3D mesh. In particular, the overall depth map can be the result of displaying a multitude of depth maps in a common coordinate system. In this case, the depth information is plotted in the common coordinate system.
[0009] At least one object is detected within the monitored area. Preferably, multiple surveillance cameras are arranged within the monitored area. In particular, more than five surveillance cameras, and specifically more than ten surveillance cameras, are arranged within the monitored area. The surveillance cameras can be stationary. Alternatively, they can be mobile and / or pan-tilt cameras, such as PTZ cameras. In a very small-scale version, only one or two surveillance cameras can be arranged within the monitored area.
[0010] The object is specifically designed as a mobile and / or moving object. Detection can be achieved, for example, via digital image processing or AI. In particular, the object's position is also detected.
[0011] Within the scope of the invention, it is proposed that the depth map be evaluated and / or improved based on the detected object, in particular its position. During the evaluation, it is specifically checked whether the detected object, especially its position, and the depth map correspond plausibly. It is generally assumed that the detected object rests on a base surface of the depth map. If the detected object is lifted from the base surface and thus "floating in the air," it can be assumed that the depth map is incorrect at the corresponding object position.
[0012] Alternatively or additionally, the overall depth map can be improved based on the detected object. For example, in the aforementioned case, the base area can be varied so that the detected object, particularly at its location, rests on the base area. Specifically, improving the overall depth map involves modifying individual points or areas while leaving other points or areas unchanged. These improvements are implemented as local enhancements to the overall depth map, not as global enhancements. Specifically, the overall depth map is improved only, or at least primarily, at the object's location.
[0013] One aspect of the invention is that the detected object moves across the monitored area, essentially acting as a scanning element, and thus scans the area. This scanning can be recorded and used to evaluate and / or improve the overall depth map.
[0014] It is particularly preferred that object recognition of the detected object is performed. After object recognition, at least the object type of the detected object is known. The object type can be, for example, a vehicle, in particular a bus, passenger car, truck, bicycle, motorcycle, person, animal, etc. Object recognition can be performed, for example, by pattern matching within the framework of digital image processing and / or using AI.
[0015] Based on the object type, object information can be estimated, and in particular determined. This object information provides a further basis, especially a priori knowledge such as object size or orientation, for evaluating and / or improving the overall depth map.
[0016] Knowing the object type allows, for example, conclusions to be drawn about a typical object size or – if the exact object type, such as vehicle type, etc., is known – about an exact object size, whereby the overall depth map can be evaluated and / or improved with knowledge of the exact object size of the detected object.
[0017] Knowing the object's size, it's possible, for example, to estimate the distance between the object and the surveillance camera that captured the underlying image. This distance can then be compared with the depth information in the overall depth map to evaluate and potentially improve it. This assumes that the surveillance camera is positioned within the shared coordinate system of the overall depth map or that the distance can at least be represented within that system.
[0018] The object's orientation allows direct conclusions to be drawn about the local plane normal in the overall depth map, so that it can be evaluated and / or improved.
[0019] With a preferred level of detail, the detected object can be identified as a person, and personal information can be estimated as object information. Typically, a person's height of approximately 1.80 m is assumed, and an object orientation, i.e., a person orientation, can be estimated. Based on this object information, the overall depth map can be evaluated and / or improved.
[0020] Alternatively or additionally, the detected object is identified as a vehicle as the object type. Vehicle information is then estimated and, in particular, determined. Specifically, it is possible to identify a specific vehicle type or model, with very precise vehicle information available from databases, for example. Furthermore, it is particularly easy to determine the object orientation, such as driving orientation or direction, especially for vehicles. Based on this vehicle information, the overall depth map can be evaluated and / or improved.
[0021] A polyhedron is preferably derived for the detected and / or recognized object. The polyhedron represents the object, and the overall depth map is evaluated and / or improved based on the polyhedron. The object's position preferably corresponds to the polyhedron's position. Using polyhedra makes further data processing of the detected and / or recognized object particularly easy.
[0022] In particular, coordinates will be derived from images of the monitored area, defining projected vertices of a polyhedron in the respective image. These coordinates are specifically defined as image coordinates and / or 2D coordinates and represent points in the image on or within the image plane. Specifically, eight vertices are determined, defining six faces (front, back, top, bottom, left, right for vehicles in the direction of travel). Depending on the object's geometry, right angles are not necessarily required. The polyhedron preferably encloses the object in such a way that neither parts of the object protrude beyond the polyhedron, nor is it too large, creating a gap between the polyhedron and the object. In particular, the polyhedron has six faces, specifically as a cuboid, and most preferably as a rectangular cuboid. The polyhedron represents the object.
[0023] Based on the polyhedron, a polyhedron position can be defined as the object's position. The polyhedron position can be, for example, a center of gravity, a center of gravity projected onto the base, or the center point of the polyhedron; preferably, the polyhedron position is defined as a vertex, center point, or foot of the polyhedron.
[0024] The image is used to derive an object orientation, particularly of the vehicle. Specifically, the image is used to determine the front and back of the object, especially in relation to the polyhedron. The combination of the object orientation and the polyhedron position can optionally be referred to as the pose of the polyhedron and / or the object. The use of the polyhedron in conjunction with vehicles as objects is particularly advantageous.
[0025] A further consideration is that by detecting the coordinates of the projected vertices of the polyhedron, the detection is independent of the specific object type and can therefore be used in a wide variety of cases. It is also simpler from a data processing perspective to process only the pose of a polyhedron rather than using a multitude of the object's vertices.
[0026] In a preferred implementation of the invention, the coordinates of the polyhedron are represented as p3D coordinates. These are the projection of the eight (3D) vertices of the object-enclosing polyhedron onto the image. Such p3D coordinates can be easily derived from the image using a variety of methods.
[0027] The p3D coordinates, as a frame projection, are the projection onto the image of a (real-world) polyhedron, particularly a cuboid, that directly surrounds the object, forming a 3D frame. The 3D frame and / or the polyhedron is formed, in particular, by straight line segments. "Directly surrounding" means that a surface defined by the imagined 3D frame or polyhedron (e.g., a polyhedron defined by the frame) surrounds the object as closely as possible, for example, with the smallest possible volume. The respective surfaces and edges of the surface spanned by the 3D frame or the polyhedron touch the surface of the object. In other words, a so-called "3D bounding box" is determined as the 3D frame or polyhedron. The 3D frame is an enclosing body, preferably in the form of a rectangular cuboid.
[0028] In a preferred embodiment of the invention, a trajectory of the detected object is detected within the monitoring area. In particular, the detected object moves along the trajectory within the monitoring area, with control points of the trajectory having different timestamps, especially consecutive timestamps. It is provided that the overall depth map is evaluated and / or improved based on the detected trajectory of the object.While a single detection of the object can already provide a basis for evaluating and / or improving the overall depth map, the detection of an object's trajectory provides a further basis for evaluating and / or improving the overall depth map: It can be assumed that the object cannot undergo sudden changes in altitude along the trajectory, so that the individual control points of the trajectory have a higher degree of significance, provided they follow a plausible trajectory.
[0029] In a preferred implementation, an uncertainty map is generated as the evaluation tool. This map displays the evaluation results of the overall depth map based on the detected object. For example, a positive evaluation result can be entered if the detected object is consistent and / or plausible with the overall depth map. For instance, the detected object is located on a base area within the overall depth map. A negative evaluation result can be entered if the detected object is inconsistent and / or implausible with the overall depth map. For example, the detected object is detached from a base area within the overall depth map ("floating in mid-air"). A neutral evaluation result can be entered if certain areas of the overall depth map have not yet been evaluated with a detected object.
[0030] In the case of an alternative or further training, an inconsistency map is generated as an evaluation, displaying inconsistent areas of the overall depth map. Such inconsistent areas can arise, for example, if the detected object travels along its trajectory through a base area in the overall depth map or exhibits other implausible behavior.
[0031] In an alternative or advanced training scenario, a boundary map is generated as part of the assessment. This boundary map displays elevation changes and / or mechanical boundaries of the overall depth map. Mechanical boundaries can, for example, be road boundaries. By utilizing a large number of detected objects and, in particular, their trajectories, the boundary map can be created to be highly informative.
[0032] In a preferred advanced training, the overall depth map is fused in a fusion step with image information, particularly color information, from the images of the monitored area into a common visualization model of the monitored area. Knowledge of the overall depth map allows the common visualization model to be built and enriched with image information, especially color information, resulting in a 2D or 3D model as a visualization model of the monitored area.
[0033] The visualization model, and thus the monitoring area, can be displayed and / or monitored by monitoring personnel via appropriate output devices. It may be possible to display the uncertainty map, the inconsistency map, and / or the boundary map to enable an intuitive and easy understanding of the monitoring area and the significance of the overall depth map.
[0034] Alternatively, the uncertainty map, the inconsistency map and / or the boundary map can be fed back into the evaluation and / or improvement module to improve the overall depth map based on this data.
[0035] A further consideration is that the traditional distribution of surveillance camera images across different screens is not intuitively accessible to surveillance personnel. This form of display therefore has significant disadvantages, as it requires, on the one hand, that the surveillance personnel are familiar with the camera system's setup and the monitored area in order to spatially locate the displayed sections, and on the other hand, that they understand which areas are not monitored in order to understand, for example, where a person who is no longer visible might have gone. This problem is exacerbated, however, if, for example, a service provider who manages several of a client's buildings connects to the video system of a monitored area to investigate an alarm. In this case, it cannot be assumed that the operator is familiar with the system or the monitored area.
[0036] This method generates a 3D representation of the monitored area that is significantly more intuitive to understand and also offers several new possibilities that a traditional representation does not allow or only allows with difficulty. The goal / result is therefore to display larger, interconnected camera arrays in a single, fused view.
[0037] When specifying the provision of the overall depth map in the monitored area, the majority of monitoring cameras are arranged, with each monitoring camera being able to capture an image of a sub-area of the respective monitoring camera.
[0038] Some or all of the monitored areas depicted in the respective image may overlap. The surveillance cameras can also capture multiple images; in particular, they can record image sequences or streams comprising a multiple set of images.
[0039] In a depth detection step, a depth map containing depth information for the monitored area is created for each surveillance camera image. The depth map can be represented as a matrix, with the depth information encoded in the matrix points. From a data processing perspective, the depth information is encoded in the same way as color information in the image, particularly in the form of a raster graphic. Specifically, the resolution of the depth map corresponds to the resolution of the underlying image.
[0040] The depth information in the depth detection step is unscaled and / or represented as relative information in the depth map. Specifically, the depth information is not displayed metrically, i.e., in meters, etc., but rather in unscaled values.
[0041] In a scaling step, the majority of the depth maps are scaled to a common coordinate system and merged into the overall depth map. The common coordinate system can, for example, be a world coordinate system. Alternatively, the common coordinate system can be a common relative coordinate system, so that a common relative scaling is implemented in the scaling step. Preferably, the common coordinate system is a common absolute coordinate system, which is metrically scaled, so that the depth information or other distances are available in metric units, such as meters, for the absolute coordinate system.
[0042] In a preferred embodiment of the invention, the depth map or depth maps are implemented in the depth detection step using a monocular depth estimation (mono-depth) method. A wide variety of such methods are publicly available. In particular, an aluminum method is used.
[0043] In principle, the depth detection step and the scaling step can be implemented in a single step and / or in a shared algorithm, so that the depth information is available in absolute, metric, or relative quantities within the same coordinate system. An example of this is "ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth" (https: / / arxiv.org / abs / 2302.12288). This network (theoretically) predicts metric (correct) outputs. In practice, however, it has been found that the information is too imprecise for the given use case, and rescaling is required. Other examples include the image processing by Tesla or "Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D" (https: / / arxiv.org / abs / 2008.05711). However, these projects sometimes make very strong assumptions about the environment or consistently use the same / similar camera views.Therefore, these methods probably do not generalize very well to a surveillance camera scenario.
[0044] In a preferred implementation of the invention, the individual depth maps are temporarily stored as an intermediate result. This intermediate storage can take place on a temporary or permanent storage medium. Thus, the depth maps are fully available before the scaling step. This implementation emphasizes that the depth detection step and the scaling step are not implemented in a single step and / or in a single computer. Rather, the advantages and accuracy of a method for creating an unscaled depth map are first utilized, followed by the advantages of a method for scaling the unscaled depth map.
[0045] In a preferred embodiment of the invention, scaling is performed based on at least one scaling information. This scaling information contains data that enables the depth maps to be scaled. In particular, this information is selected from intrinsic camera parameters of at least one surveillance camera, extrinsic camera parameters (especially location) of at least two surveillance cameras with overlapping surveillance areas, and / or overlap information from surveillance areas.
[0046] By dividing the process into a depth detection step and a scaling step with scaling information, the visualization model can be created with exceptional accuracy. Relative depth detection methods are highly precise and yield reliable detection results. By using scaling information in the scaling step, the relative scaling of the individual depth maps can be transferred very accurately and reliably to the common coordinate system, resulting in a robust foundation for the overall visualization model and leading to a meaningful and high-quality visualization. Specifically, scaling is performed using metric units, such as meters.
[0047] Typically, 3D reconstruction / depth estimation requires multiple camera views. This paper proposes estimating the depth map from a single image. AI-based methods, such as Vision Transformers for Dense Prediction (DPT), are suggested for implementation, as they estimate depth information from a single image. These methods are typically not as accurate as those based on multiple cameras. Some of these mono-depth methods also estimate an absolute depth / distance. However, this is very challenging and therefore error-prone, as the method must distinguish, for example, between a telephoto view of a vehicle and a wide-angle view of the same vehicle. In both cases, the vehicle might appear the same size in the image, but the camera could be over a hundred meters away in the case of a telephoto camera and only a few meters away in the case of a wide-angle camera.
[0048] Methods like the proposed DPT therefore proceed differently and estimate an "unscaled" depth without direct metric meaning. The area estimated to be the nearest point in the depth image is represented, for example, as white (maximum value), and the furthest point is represented, for example, as black (minimum value). In other words, distances or depths are simply specified between the boundaries of "nearest" and "farthest." The scaling in between typically corresponds to an inverse depth as depth information, since the methods are often trained on stereo data, and the inverse depth represents a disparity.
[0049] A possible implementation of DPT can be downloaded from https: / / github.com / isl-orgfDPT. The publication "Vision Transformers for Dense Prediction: René Ranftl, Alexey Bochkovskiy, Vladlen Koltun" (arXiv:2103.13413) contains a theoretical description of the implementation of the method, downloadable at: https: / / arxiv.org / pdf / 2103.13413.pdf.
[0050] For practical application, the depth (relative or absolute, especially metric) for the common coordinate system must be recovered when using methods with "non-metric" depths. This is implemented in the scaling step based on the scaling information. To convert the depth maps (e.g., from DPT) into a true / metric or relative scale in the common coordinate system, two values can typically be determined: an absolute offset and a scaling factor. Depending on the method, there may be more, fewer, or different parameters. In principle, if one of the surveillance cameras has a scaled depth map that has been scaled using the scaling information, the scaling can be transferred to another surveillance camera that has an overlapping surveillance area. This is achieved by
[0051] In a preferred embodiment of the invention, the scaling information comprises extrinsic and / or intrinsic camera parameters.
[0052] The extrinsic camera parameters can, for example, be defined as the positions of the surveillance cameras within the common coordinate system. If the surveillance areas overlap, the scaling can be implemented based on this scaling information.
[0053] The intrinsic camera parameters, such as the focal length of the surveillance camera, can be used to scale the depth map of the image from the surveillance camera with the known focal length. If the intrinsic camera parameters are known for only one surveillance camera, and the monitoring areas of the other surveillance cameras overlap with the monitoring area of the camera with the known focal length, the scaling can be applied to the images of the other surveillance cameras.
[0054] Alternatively or additionally, the scaling information can be determined by an additional measuring device or by at least one surveillance camera acting as a measuring device. For example, a laser scanner can be used, which measures absolute distances to objects in the monitored area as scaling information. In the scaling step, the absolute distances are assigned to the corresponding areas in the depth maps, thereby scaling the depth map. It can also be provided that a known measuring object is positioned in the monitored area. By capturing the measuring object with the surveillance camera and knowing its dimensions, the depth map is scaled. A vehicle or person with known or typical dimensions that happens to be present in the monitored area can also be used as the measuring object.
[0055] In a preferred embodiment of the invention, a ground plane of the monitored area or sub-area is determined in the image and / or the depth map. This determination can be achieved, for example, by semantic segmentation, a technique known from digital image processing. By incorporating scaling information, such as a known calibration, extrinsic and / or intrinsic camera parameters, the position of the monitoring camera relative to the ground plane is known, allowing the ground plane to be scaled in the image and / or the depth map. If the sub-areas of the monitoring share a common ground plane and overlap, the scaling can be performed without further information. The scaled ground plane thus constitutes the scaling information, or a derived scaling information, in the scaling step.
[0056] Therefore, it is proposed to use the scaled ground plane as scaling information. The ground plane should preferably be determined automatically or semi-automatically to simplify setup for the user. For example, semantic segmentation can be used, which can optionally be determined by DPT. All image regions identified as "roads" are assumed to be the ground plane. By adding a known (extrinsic) calibration, the camera's position relative to the ground plane is known. By comparing the depth values in the "road" region with the expected depth values based on the plane assumption, the two unknown parameters can be determined, and the depth map can be scaled.
[0057] For a particularly simple implementation, it is proposed that the depth map contain map points with depth information, where the map points correspond to or are defined as corresponding points in the associated image. Thus, the map points with the depth information form a point cloud, which is then transformed into the overall depth map via the scaling step within the common coordinate system. Due to the correspondence between the map points and the image points, the point cloud can be easily enriched with color information from the images during the fusion step.
[0058] In a preferred embodiment of the invention, the visualization model can, for example, be displayed on a 3D screen, allowing a user or monitoring personnel to move within the visualization model using a virtual camera. Alternatively or additionally, the visualization model can be displayed using VR glasses, allowing the user or monitoring personnel to change their view by moving their head. It is also possible to define one or more virtual cameras, with the field of view of the virtual cameras displayed on multiple screens or screen sections. This has the advantage that the virtual cameras can define fields of view that are intuitively understandable for the user and / or monitoring personnel.
[0059] A further aspect of the invention is a depth mapping arrangement, which is specifically designed to implement the method described above. The depth mapping arrangement is specifically configured as a digital data processing device, such as a computer, a server, or a cloud. The depth mapping arrangement has an interface for receiving the depth mapping and / or images from the surveillance cameras. The images each show a sub-area of the surveillance area. Optionally, the depth mapping arrangement includes the majority of the surveillance cameras located in the surveillance area. Alternatively, the depth mapping arrangement includes the images of the sub-areas of the surveillance area.
[0060] The depth map arrangement includes a detection module for detecting at least one object in the monitored area and an evaluation and / or improvement module for evaluating and / or improving the depth map based on the at least one detected object.
[0061] The overall depth map assembly optionally includes a depth detection module, which is configured for the depth detection step. Furthermore, the overall depth map assembly includes a scaling module, which is configured for the scaling step. Additionally, the overall depth map assembly optionally includes a fusion module, which is configured for implementing the fusion step.
[0062] Optionally, the depth map arrangement includes an output device designed for outputting, in particular visualizing, the visualization model. The output device can be, for example, a 3D screen, VR glasses, or a screen for displaying the fields of view of virtual cameras within the visualization model.
[0063] A further aspect of the invention is a computer program comprising commands which, when the program is executed by a computer or visualization device, cause it to execute the method / steps of the method according to the invention. A further aspect of the invention is a computer-readable, in particular non-volatile, data carrier on which the computer program is stored.
[0064] Further features, advantages, and effects of the invention will become apparent from the following description of preferred embodiments and the accompanying figures. These show: Fig. 1 a schematic block diagram a depth map arrangement in general form as an embodiment of the invention; Fig. 2 a schematic block diagram a depth map device in general form for the depth map arrangement in the Fig. 1; Fig. 3 a schematic block diagram a concretization of the depth mapping device in the Fig. 2; Fig. 4 a schematic representation of a visualization model as a result of the depth mapping device of the preceding figures; Fig. 5. A camera view from a virtual camera in the visualization model in the Fig. 4.
[0065] The Fig. Figure 1 shows a schematic block diagram of a total depth map arrangement 50 as an embodiment of the invention and to describe an embodiment of the method,
[0066] The depth map arrangement 50 has an interface 51 for receiving a depth map 52 from a monitoring area 53. The depth map 52 can be obtained in particular from the depth map device of the Fig. 2 and Fig. 3. The surveillance area 53 is in particular recorded with one or a plurality of surveillance cameras 54 and can form a single surveillance area 53, a continuous surveillance area 54 or a segmented surveillance area 54.
[0067] The depth map 52 represents the monitoring area 53, but instead of color values or the like, depth information is entered. The depth map can be, for example, a 2D depth map, a 3D point cloud, a voxel grid, a 2D elevation map, or a 3D mesh. In particular, the depth map 52 is scaled, specifically metrically escalated.
[0068] The depth mapping system 50 includes a detection module 55, the detection module 55 being configured to detect an object 56, in particular a moving and / or mobile object 56, within the monitored area 53. Detection is performed, for example, using images from the surveillance camera 54. Detection can be carried out via digital image processing or other algorithms.
[0069] The depth map arrangement 50 includes an evaluation and / or improvement module 57, which is designed to evaluate and / or improve the depth map 52 based on the detected object 56.
[0070] To improve and evaluate this overall depth map 52, and in particular to validate it, observations will be used. One possibility here is P3D detections of vehicles as objects 56, assuming a vehicle size (vehicle-specific via model, vehicle class, or default size) as object information.
[0071] The evaluation and / or improvement module 57 checks whether the detection—or temporal series of detections—corresponds to the overall depth map 52. If this is not the case, the map is adjusted and thereby improved. In the simplest case, the depth values below (and weighted in the vicinity of) the vehicle (object 56) are corrected to the corresponding height / distance. The correction should be made in small increments. An extension of this is to evaluate the entire vehicle trajectory, since improving the overall depth map 52 by considering the trajectory prevents any sudden changes in elevation within the map.
[0072] Similarly, when using people as objects, 56 person detections (2D and projected 3D, i.e., body pose) can be used.
[0073] In general, several criteria can be used for estimation / correction: The size and orientation of a vehicle / person as object 56 allows direct inferences about the local plane normal in the overall depth map 56. The condition is that persons / vehicles as object 56 cannot instantaneously change their size, orientation, and speed. Here, temporal consistency is exploited. Further features can include a speed estimate and its consistency over time for vehicles as object 56, and step consistency for persons as object 56.
[0074] All these criteria can also be implicitly estimated using a neural network or a filter as the basis for the evaluation and improvement module 57. To further improve the process, semantic segmentation can optionally be used. This allows, for example, inferences about meaning / semantics. In this way, the overall depth map 52 could also be augmented, resulting in an improved semantic description alongside the overall depth map 52. This could, for example, describe the areas where vehicles drive or where only pedestrians move.
[0075] For the procedure, it can be very advantageous to use a representation of the uncertainty in the relevant areas in addition to the overall depth map 52. Areas with no observations could be marked as uncertain, and areas with contradictory detections could be marked as such and treated specifically during the correction process. These could be an uncertainty map 58 and an inconsistency map 59. The latter can be particularly helpful when repeated dynamic changes occur (e.g., opening and closing a bridge). Furthermore, it can be very advantageous to include (and correct / update) a representation of depth discontinuities or object boundaries as a boundary map 60. Ideally, a correction, especially an improvement of the overall depth map 52 by an observed vehicle as object 56, should only have an effect up to the corresponding boundary.This boundary map 60 (edge map) could be initialized and / or corrected from observations, image edges, depth jumps or semantic segmentation.
[0076] Finally, it could be advantageous to run the process continuously. The result would be constantly updated as new observations (or a buffered set of them) become available. Simultaneously, it could then be useful to mark the additional maps (Edge, Uncertainty / Inconsistency, Semantic) as less certain over time if no new observations are available for an extended period. This is primarily intended to prevent the map from being classified as increasingly reliable, thus ensuring that, for example, structural changes go unnoticed.
[0077] To initialize the process, a total depth map 52 is generated (or, for example, a default map from level-based calibration is used). If a mono-depth method is used, it might be useful to merge several time-shifted mono-depth results here to exclude dynamic changes from the total depth map 52.
[0078] Simultaneously, potentially dynamic objects, such as vehicles, people, etc., could also be detected through semantic segmentation or object detection, and the corresponding entries in the depth map, uncertainty map, and inconsistency map could be initialized accordingly. The simplest concrete implementation of the method (filter) could be in the form of a Kalman filter.
[0079] The depth map arrangement 50 has a fusion module 61 for implementing a fusion step, wherein the fusion module 61 is configured to fuse the scaled depth map 52 with image information from the monitoring sub-areas 4 or the monitoring area 53 in a common visualization model 11 of the monitoring area 53.
[0080] The visualization model 11 can subsequently be displayed in an output device 62. The visualization model 11 is specifically designed as a 3D model of the surveillance area 53, which is realistically represented by the enrichment with image information and which can be displayed by a user or surveillance personnel in any way, for example by means of a 3D screen, VR glasses or via one or more screens, each as an output device 62, whereby any view can be displayed in the visualization model, for example by using virtual cameras.
[0081] This method generates a 3D representation of the scene in surveillance area 53, which is significantly more intuitive to understand and also offers several new possibilities that a classic representation does not allow or only allows with difficulty. The goal / result is to display larger, contiguous camera arrays in a single, fused view.
[0082] When evaluating the overall depth map 52, the uncertainty map 58, the inconsistency map 59 and / or the boundary map 60 are generated and displayed together in a display unit 61. In this way, a user is informed about the evaluation status of the overall depth map 52.
[0083] When the overall depth map 52 is improved, the overall depth map 52 is corrected immediately, as described. Alternatively or additionally, the uncertainty map 58, the inconsistency map 59 and / or the boundary map 60 are fed back into the evaluation and improvement module 57 to improve the overall depth map 52.
[0084] As previously explained, the depth map 52 can be designed as a two-dimensional map whose structure resembles an underlying image and is therefore also designed as a matrix, with the depth information represented as pixels. However, it is also possible for the depth map 52 to be designed as a 3D point cloud, with multiple depth maps plotted in a common coordinate system.
[0085] The Fig. Figure 2 shows in a schematic block diagram a depth map generation device 1, which is configured to generate the depth map 52, in particular as a 3D point cloud.
[0086] In the surveillance area 53, a plurality of surveillance cameras 54 are arranged, which are located in the Fig. 1 and Fig. 2 are shown as a block. Each surveillance camera 54 records an image of a surveillance sub-area 4, which lies within the field of view of the corresponding surveillance camera 3. The surveillance sub-areas 4 can be non-overlapping, as schematically shown in the lower part of the surveillance area 53. The surveillance sub-areas 4 can also overlap, forming overlapping areas 5, as schematically shown in the upper part of the surveillance area 2.
[0087] The depth mapping device 1 has an interface 6 for receiving images from the surveillance cameras 3. Optionally, the surveillance cameras 3 form part of the depth mapping device 1.
[0088] The depth map generation device 1 has a depth detection module 7, wherein the depth detection module 7 is configured to create, in a depth detection step 100, for each image of the surveillance cameras 3, a depth map with depth information of the surveillance sub-area 4 associated with the image. The depth information in the respective depth maps is unscaled and / or presented as relative information.
[0089] The depth map generation device 1 has a scaling module 8, which is configured to perform a scaling of the depth maps from all surveillance cameras 3 in a common coordinate system in a single scaling step. Specifically, the depth maps are brought into a common metric of the common coordinate system by the scaling module 8 and / or in the scaling step. In particular, the depth information of the different depth maps is directly comparable. In principle, the scaled depth information can already be configured as absolute depth information, so that the depth information of the depth maps is configured as metric depth information in a common coordinate system.Alternatively, the scaled depth information can be designed as relative depth information, which is comparable in the common coordinate system, but is not expressed in absolute values, such as meters.
[0090] The representation of the depth maps in the common coordinate system forms the overall depth map 52.
[0091] In principle, the depth detection module 7 and the scaling module 8 can be implemented as a single, combined module that performs depth detection step 100 and scaling step 200, thus generating the scaled depth map in a single step. As previously described, this is a comparatively difficult and complex task for the corresponding evaluation methods.
[0092] Accordingly, in this embodiment, it is proposed that the depth detection step 100 is first performed, for example, based on the previously described DPT, and subsequently the scaling is performed based on at least one scaling information. The scaling information contains information that enables the depth maps to be scaled. This information is specifically selected from intrinsic camera parameters of at least one surveillance camera, extrinsic camera parameters (especially location) of at least two surveillance cameras with overlapping surveillance areas, and / or overlap information from surveillance areas. By introducing the scaling information as a priori knowledge, the estimation of the scaled depth maps can be significantly improved, thus making the method and the visualization device 1 more robust.
[0093] Overall, a completely different type of representation is proposed, combining methods for monocular depth estimation (Mono-Depth) and (camera) calibration.
[0094] The Fig. Figure 3 shows an embodiment of the invention in the form of a more detailed configuration of the depth mapping device 1 in the Fig. 2. Identical or corresponding components are provided with the same or corresponding reference numerals.
[0095] The Fig. Figure 3 shows an exemplary implementation of the method for generating the overall depth map 52 based on Vision Transformers for Dense Prediction (DPT), whereby several individual images (one image per surveillance camera 54) are used to predict the non-metric depth and semantics. An intrinsic and extrinsic calibration, determined in advance (only) based on the (here 24) individual images, is used to rescale the depth maps, determine the lines of sight corresponding to the image points, and convert the relative position of the surveillance cameras 54, and thus the 3D point clouds, into a common 3D view. Filtering was also performed for improved visualization.
[0096] It should be noted that all steps shown here (except for determining the calibration) can be carried out independently for each surveillance camera 54, and therefore in parallel. Not shown are • Live processing, • Anonymization, a • More sophisticated fusion, which, for example, corrects (incorrect) depth jumps in overlapping views, and • a color correction that corrects the color balance of the individual surveillance cameras 3.
[0097] In the exemplary embodiment in the Fig. Images from 24 different surveillance cameras were received. First, a calibration was determined by manually selecting corresponding points in the images as scaling information. From this, the intrinsic (distortion, focal length) and extrinsic (relative position of the cameras to each other and the ground plane) parameters could be determined. The calibration process could also be carried out in other ways (e.g., using known lens data and autocalibration methods).
[0098] First, a depth map and a semantic segmentation are generated using DPT in the depth detection module 7. Basic level regions (here, the "Road" class) are used from the semantic segmentation. It was assumed for scaling purposes that this is indeed a (basic) level. Alternative assumptions and methods are discussed in the next section.
[0099] Based on the calibration, the required depth / distance for each pixel can be predicted if it were a point on the ground plane. By comparing the corresponding "road" regions of the depth map with this data, the two scaling parameters (in this case) can be determined, enabling the depth map to be converted into metric depths / distances. Finally, the intrinsic calibration of surveillance camera 3 is used again to calculate (colored) 3D point clouds from the now metric depths.
[0100] Many monocular depth estimation methods "smooth out" depth jumps, resulting in unsightly artifacts in the depth reconstruction. To filter these out, an edge detector was applied to the depth map. A better filtering method would apply this not to the depth map (which represents the inverse depth), but to the actual depths (or a mixture or combination thereof), since the method shown here becomes less sensitive at greater distances (the inverse depth values become very small due to the increased distance, and consequently, so do the jumps). The choice and type of filtering can, of course, be modified (see also the next section). Finally, the point clouds from the individual surveillance cameras are combined. In the simplest case, these are simply displayed together and independently. A minor adjustment was made here.Only the nearest points from each surveillance camera 54 are displayed. This prevents gross miscalculations at greater distances from obscuring or overwriting the good results of a local surveillance camera 54. It should be noted that this is based solely on calibration information and is therefore independent of the actual content (3D point clouds) of the other surveillance cameras 54. This ensures that fully parallel processing remains possible. Example results:
[0101] In Fig. Figure 4 shows the overall result of the 24 camera images as visualization model 11. The image in Fig. Figure 4 is a screenshot from the CloudCompare software, which can display 3D point clouds. In some areas, especially those with many occlusions, the 3D reconstruction is currently inaccurate and somewhat difficult to interpret. It should be noted that DPT is a method from 2021, and better results can likely be expected with more recent methods. Furthermore, the method was probably not trained on corresponding images from this (monitoring) area 53. However, if a virtual camera is moved closer to the actual camera positions in the point cloud visualization model 11, the potential becomes clearer.
[0102] This is clearly evident in Fig. 5 to see where a partial excerpt 12 from Fig. Figure 4 shows that the images also exhibit a similar color balance, which is why the transition between the two point clouds on the ground is invisible. The output device 62 can therefore be a screen that displays the partial view 12 as the field of view of a virtual camera. Alternative or improved presentation:
[0103] In a further and improved form of representation, a simple anaglyph method is used to create a true 3D image on a screen as output device 62, using red-cyan glasses. This representation significantly enhances scene perception, as the content can be perceived directly, not just through texture and depth differences achieved by (virtually) rotating / moving the camera / scene. An even better representation is achieved using a VR headset, such as an Oculus Rift, as navigation (rotating and moving) and depth perception are significantly easier / more intuitive, and a larger field of view is utilized. Another alternative is a 3D monitor as output device 62 with appropriate glasses. Live View:
[0104] The figures only describe a static image / point cloud. For one application, for example, (A) a current image / point cloud is generated every few minutes / seconds. Alternatively, this could also be done live (for each or nearly each) image. Since this can be very computationally intensive (despite the possibility of independent processing for each surveillance camera 3), it could be advantageous (B) to perform the depth determination less frequently and only update the current texture (color image). The depth map could, for example, be a "background" depth image generated from several depth maps taken at different times, so that moving objects are not included. Moving objects could then be displayed separately, for example, in the form of avatars or simple shapes (cylinders, cuboids) onto which the texture (color image) is projected.The detection of these objects typically takes place in the surveillance camera 54, so the processing should be significantly faster. Anonymization:
[0105] If, as in the previous point ("Live View"), a background is generated and people, vehicles, etc., are drawn separately, this also represents a possible approach to anonymization. A further step would be to replace the background texture (workshop in the described example) with a different one. Then the user or monitoring personnel would only see the 3D structure and avatars, etc., and wouldn't even know what they were actually seeing. This represents (almost) perfect anonymization. Calibration tracking
[0106] For pan-tilt (PTZ) surveillance cameras, the time-dependent calibration must be taken into account. If the calibration change is known, it can be easily integrated. However, this requires recalculating the depth map. Alternatively, if a complete depth background map for a pan-tilt camera has been pre-generated, no recalculation is necessary; only a corresponding adjustment to the current field of view is required. Improved Fusion
[0107] In the described method, the 3D point clouds from each surveillance camera (54) are simply displayed together. However, it frequently happens that parts of the camera views overlap, and errors in depth estimation or calibration lead to duplicate structures. In these cases, the same structure appears multiple times in different locations. A method could be used to detect and merge these duplicate structures.
[0108] On the one hand, this could be an averaging of the depth and texture information at the appropriate location, or a clever stitching technique. Color correction (color balancing)
[0109] In Fig. In image 4, for example, it is clearly visible on the ground surface that no color balancing was performed between the images. Fig. In contrast, image 5 clearly shows how much a good (here random) color balancing improves the illusion of a single result compared to a superposition of many individual results. Rescaling methods and Ground Plane Selection:
[0110] In the described method, the ground plane and automatic recognition of the corresponding image areas via semantic segmentation were used to introduce metric information into the "non-metric" depth images. If a continuous ground plane actually exists, this has the advantage that it can be done fully automatically. Of course, this cannot be applied in many situations. Furthermore, the reconstruction in the described method is metric, but not "in meters." This means that lengths can be compared, but lengths cannot be directly specified in meters. For this, absolute scale information, such as the height of a person, the height of the surveillance camera (54 cm above the ground), or similar, must be introduced.
[0111] Alternatively, metric information can be introduced, for example, via image correspondences in the overlap area of two camera views / monitoring sub-areas 4. In this way, depth could be determined via triangulation (as in classic stereo methods). The correct scale could, for example, be derived from the known distance between two surveillance cameras 54 or the camera height. Alternatively, with surveillance cameras 54, such as the Bosch Flexidome Multi, it could be exploited that the surveillance cameras 54 can only be moved within a ring whose radius is known. Alternatively, the correct scale can also be determined by a few direct distance measurements, e.g., with a laser distance meter.
[0112] Selecting a suitable base plane or other plane could also be implemented, for example, using methods like Segment-Anything, where the operator can select appropriate areas with just a few mouse clicks. Removing implausible depth measurements:
[0113] As demonstrated in the procedure, implausible depth measurements can be discarded through filtering. This can be done using simple methods. However, it can also be useful, for example, to identify whether meaningful structures emerge in the overlap area of several surveillance cameras (see also "Improved Fusion"). Finally, timestamps should also be removed (automatically), as no meaningful depth can be attributed to them. Parallel processing
[0114] As mentioned previously, depending on the application, processing can be completely parallel, except for the visualization. This means that, for example, even individual surveillance cameras can generate three 3D point clouds or a depth map in a suitable format and send them to a central server or alarm monitoring center. For information on the amount of data that needs to be transferred, see also "Internal 3D Representation". Rectification:
[0115] Many mono-depth methods, such as DPT, are not designed for image data with significant distortion, which is often used. Therefore, the images are first rectified, resulting in significant improvements. A processing step right at the beginning of the process can thus be to rectify the image or even split it into several rectified images and then recombine the depth results of the individual images. This can be particularly useful with very wide-angle cameras (e.g., fisheye). Internal 3D representation:
[0116] Representing the depth map as a point cloud (several million individual points) allows for a high, but unnecessary, level of detail. Alternatively, a 3D mesh could be used, which enables very fast texture rendering. This method of representing results is common in computer games, etc., and is both fast to process and requires less storage space and bandwidth. See-through:
[0117] The combination of 3D information and multiple views also allows for functions such as See-Through, where, for example, people who are hidden in the current (virtual) view are schematically displayed. Improved tracking:
[0118] The 3D information allows for improved tracking of people, since it is known that tracked people will, for example, disappear behind a car if they continue their walking movement. Other scenarios:
[0119] The described method used an example with many surveillance cameras (54) and a large overlap area. However, the method does not inherently require this. For example, another scenario would be monitoring the fence of a substation. Here, many surveillance cameras (54) are typically used, but with little or no overlap. In this case, a satellite or topographic map could be enhanced with 3D point clouds / meshes. The visualization could be as described previously. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 10,346,996 B2
[0003] Zitierte Nicht-Patentliteratur
[0000] ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth“ (https: / / arxiv.org / abs / 2302.12288
[0043] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D“ (https: / / arxiv.org / abs / 2008.05711
[0043] Vision Transformers for Dense Prediction: René Ranftl, Alexey Bochkovskiy, Vladlen Koltun (arXiv:2103.13413
[0049] https: / / arxiv.org / pdf / 2103.13413.pdf
[0049]
Claims
[1] Methods for evaluating and / or improving a depth map (52) of a monitoring area (53), where the overall depth map (52) of the monitoring area (53) is provided, wherein at least one object (56) is detected in the monitoring area (53), where the overall depth map (52) is evaluated and / or improved based on the detected object (56). [2] Method according to claim 1, characterized by , that the detected object (56) is recognized, whereby, based on the recognized object (56), an object size and / or object orientation is estimated as object information, whereby the overall depth map (52) is evaluated and / or improved based on the object information. [3] Method according to claim 1 or 2, characterized by, that the detected object (56) is recognized as a person, whereby person information is estimated, whereby the overall depth map (52) is evaluated and / or improved based on the person information [4] Method according to any one of the preceding claims, characterized by , that the detected object (56) is recognized as a vehicle, whereby vehicle information is estimated based on the recognized vehicle, and the overall depth map (52) is evaluated and / or improved based on the vehicle information [5] Method according to any of the preceding claims, wherein a polyhedron is derived for the detected and / or recognized object which represents the detected object (56), wherein the overall depth map (52) is evaluated and / or improved on the basis of the polyhedron [6] Method according to any one of the preceding claims, characterized by, that a trajectory of the detected object (56) is detected in the monitoring area (53), whereby the overall depth map (52) is evaluated and / or improved based on the detected trajectory of the object (56). [7] Method according to any one of the preceding claims, characterized by , that an uncertainty map (58) is generated as an evaluation, wherein the uncertainty map (58) displays evaluation results of the depth map (52) based on the detected object (56). [8] Method according to any one of the preceding claims, characterized by , that an inconsistency map (59) is generated as an evaluation, in which inconsistent areas of the overall depth map (52) are displayed on the basis of the detected object (56) in the inconsistency map (59). [9] Method according to any one of the preceding claims, characterized by, that a boundary map (60) is generated as an assessment, in which height jumps and / or mechanical limits of the overall depth map (52) are displayed in the boundary map (60). [10] Method according to any one of the preceding claims, characterized by , that in a fusion step the depth map (52) is fused with image information from images of the monitored area (53) in a visualization model (11). [11] Method according to one of the preceding claims, wherein a plurality of surveillance cameras (54) are arranged in the surveillance area (53) to provide the overall depth map (52), wherein each surveillance camera (54) can take an image of a surveillance sub-area (4) of the respective surveillance camera (54), wherein in a depth detection step (100) a depth map with depth information of the monitored sub-area (4) is created for each of the images from the surveillance cameras (54), wherein in a scaling step (200) a scaling of the depth maps for a common coordinate system of the monitoring sub-areas (4) is carried out in order to form the overall depth map (52). [12] Depth map arrangement (1), in particular for implementing the method according to any one of the preceding claims, with an interface (51) for taking over a depth map (52) of a monitoring area (53), with a detection module (55) for detecting at least one object (56) in the monitored area, with an evaluation and / or improvement module (57) for evaluating and / or improving the overall depth map (52) based on the detected object (56). [13] Computer program comprising instructions which, when the program is executed by a computer or the depth map arrangement (50) according to claim 12, cause it to execute the method / steps of the method according to any one of claims 1 to 11. [14] Computer-readable data carrier on which the computer program according to claim 13 is stored.
Citation Information
Patent Citations
US10,346,996B2