Method for visualizing a surveillance area, visualization device, computer program and computer-readable data carrier
The method integrates depth detection and scaling of unscaled depth maps from multiple cameras into a common coordinate system, creating an intuitive 3D model that addresses the challenges of non-intuitive split-screen displays in monitoring systems, enhancing monitoring efficiency and understanding.
Patent Information
- Application Number
- DE102024203186
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2025-10-09
AI Technical Summary
Existing monitoring systems require multiple cameras with non-intuitive split-screen displays, necessitating operator familiarity with camera systems and monitored areas, and are challenging for service providers unfamiliar with the system to understand and navigate.
A method that combines depth detection and scaling of unscaled depth maps from multiple cameras into a common coordinate system, creating a 3D visualization model with accurate metric scaling, using AI-based methods like monocular depth estimation and camera calibration to merge images into a single intuitive view.
Enables a clear, intuitive 3D representation of monitored areas, allowing seamless navigation and understanding without prior knowledge of the camera system, and offering enhanced monitoring capabilities.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for visualizing a surveillance area, a visualization device, a computer program and a computer-readable data carrier. State of the art
[0002] To effectively monitor larger areas, buildings, or facilities, a large number of cameras often need to be installed. The video streams from the cameras are presented to an operator on numerous monitors.
[0003] The document US 9,706,264 B2 discloses a system comprising an image sensor configured to receive a selection of a plurality of fields of view from a user, capture an image associated with each of the plurality of fields of view, combine the image associated with each of the plurality of fields of view into a single stream, and transmit the single stream to at least one computing device. Disclosure of the invention
[0004] The invention relates to a method for visualizing a surveillance area having the features of claim 1, a visualization device having the features of claim 10, a computer program having the features of claim 12 and a computer-readable data carrier having the features of claim 13.
[0005] Preferred or advantageous embodiments of the invention emerge from the subclaims, the following description and the attached figures.
[0006] The subject of the invention is a method for visualizing a surveillance area. The surveillance area can be located in a commercial, private, and / or public area. The surveillance area can be continuous, alternatively, it can be partially continuous or segmented with open intermediate areas between the segments.
[0007] A plurality of surveillance cameras are arranged in the surveillance area. In particular, more than five surveillance cameras, and more specifically more than ten surveillance cameras, are arranged in the surveillance area. The surveillance cameras can be stationary cameras. Alternatively, they can be mobile and / or movable cameras, such as PTZ cameras. In a very small version, two surveillance cameras can also be arranged in the surveillance area.
[0008] The surveillance cameras each record an image of a monitoring sub-area of the surveillance area that lies within the field of view of the respective surveillance camera. Some or all of the monitoring sub-areas depicted in the respective image can be overlapping. The surveillance cameras can also record multiple images; in particular, the surveillance cameras can record image sequences or streams comprising a plurality of images.
[0009] In a depth detection step, a depth map containing depth information of the monitored sub-area is created for each of the surveillance camera images. The depth map can be configured as a matrix, with the depth information encoded in the matrix points. From a data-technical perspective, the depth information is encoded in the same way as color information in or within the image, particularly as a raster graphic. In particular, the resolution of the depth map corresponds to the resolution of the underlying image. For example, the depth information is encoded in grayscale, with the grayscale corresponding to a depth. The "depth" is configured, in particular, as a radial distance, i.e., a radial distance from the camera, or an actual or standard depth (in the narrower sense), i.e., a distance in the camera's main line of sight. Both distance specifications (radial or along only one axis / depth (in the narrower sense)) are common in the literature.However, most of the time the distance is given along only one axis.
[0010] In the depth detection step, the depth information is presented unscaled and / or as relative information in the depth map. In particular, the depth information is not represented in metric units, such as meters, but in unscaled values.
[0011] In a scaling step, the majority of depth maps are scaled for a common coordinate system. The common coordinate system can be configured, for example, as a world coordinate system. It can be provided that the common coordinate system is configured as a common relative coordinate system, so that a common, relative scaling is implemented in the scaling step. Preferably, the common coordinate system is configured as a common absolute coordinate system, which is scaled metrically, so that the depth information or other distances for the absolute coordinate system are available in metric units, for example, in meters, etc.
[0012] In a fusion step, the scaled depth maps are merged with image information, particularly color information, from the images of the monitored sub-areas to create a common visualization model of the monitored area. With knowledge of the scaled depth maps in the common coordinate system, the common visualization model can be constructed and enriched with image information, particularly color information, resulting in a 3D model as a visualization model of the monitored area. The visualization model, and thus the monitored area, can be monitored, for example, by surveillance personnel using appropriate output devices.
[0013] The scaled depth maps are positioned in the common coordinate system in the scaling step. Alternatively, the scaled depth maps are positioned in the visualization model in the fusion step.
[0014] The invention is based on the idea that the classic division of surveillance camera images across different screens is not intuitive for surveillance personnel. This form of display therefore has significant disadvantages. On the one hand, it requires surveillance personnel to be familiar with the structure of the camera system and the monitored surveillance area in order to be able to spatially classify the displayed surveillance sub-areas. On the other hand, it requires an understanding of which areas are not monitored in order to understand where, for example, a person who is no longer visible might have gone. This problem is exacerbated, however, if, for example, a service provider who looks after several of a customer's buildings connects to the video system of a surveillance area, for example to investigate an alarm. In this case, it cannot be assumed that the operator is familiar with the system or the surveillance area.
[0015] This process creates a 3D representation of the surveillance area that is significantly more intuitive to understand and also offers several new possibilities that a conventional representation cannot provide or only allows with difficulty. The goal / result is thus that larger, interconnected camera arrays are displayed in a single fused view.
[0016] In a preferred embodiment of the invention, the depth map(s) are converted in the depth detection step using a monocular depth estimation method (mono-depth). A wide range of such methods are publicly available. In particular, this is an AI method.
[0017] In principle, the depth detection step and the scaling step can be implemented in a single step and / or in a single AI, so that the depth information is available in absolute, metric quantities or in relative quantities within the shared coordinate system. One example is "ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth" (https: / / arxiv.org / abs / 2302.12288). This network (theoretically) predicts metric (correct) outputs. In practice, however, it can be found that the information is too imprecise for the current use case, and rescaling is required. Further examples are Telsa's image processing or "Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D" (https: / / arxiv.org / abs / 2008.05711). However, these projects sometimes make very strong assumptions about the environment or always use the same / similar camera views.Therefore, these methods probably generalize very poorly to a surveillance camera scenario.
[0018] In a preferred implementation of the invention, the individual depth maps are cached as an intermediate result. The cache can be stored on a temporary or permanent storage medium. Thus, the depth maps are available in their entirety before the scaling step. This implementation emphasizes that the depth detection and scaling steps are not implemented in a single step and / or in a single AI. Rather, the advantages and accuracy of a method for creating an unscaled depth map are first utilized, followed by the advantages of a method for scaling the unscaled depth map.
[0019] In a preferred embodiment of the invention, the scaling is performed based on at least one piece of scaling information. The scaling information has information content that enables the depth maps to be scaled. In particular, this information is selected from intrinsic camera parameters of at least one surveillance camera, extrinsic camera parameters (in particular location) of at least two surveillance cameras with overlapping surveillance sub-areas, and / or overlap information of surveillance sub-areas.
[0020] By dividing the process into a depth detection step and a scaling step with the scaling information, the visualization model can be created with particular precision. Relative depth detection methods are very accurate and lead to reliable detection results. By using scaling information in the scaling step, the relative scaling of the individual depth maps can be transferred very accurately and reliably into the common coordinate system, thus providing a very robust basis for the common visualization model and leading to a meaningful and high-quality visualization model. In particular, scaling is performed in metric units, such as meters.
[0021] Typically, 3D reconstruction / depth estimation requires multiple camera views. Here, it is proposed to estimate the depth map from a single image. AI-based methods, such as Vision Transformers for Dense Prediction (DPT), are proposed for implementation. These methods are typically not as accurate as methods based on multiple cameras. Some of these mono-depth methods also estimate absolute depth / distance. However, this is very challenging and therefore error-prone, as the method must distinguish, for example, between a telephoto view of a vehicle and a wide-angle view of the same vehicle. In both cases, the vehicle could appear the same size in the image, but the camera could be over a hundred meters away in the case of a telephoto camera and only a few meters away in the case of a wide-angle camera.
[0022] Methods such as the proposed DPT therefore take a different approach and estimate an "unscaled" depth without direct metric meaning. The area estimated as the nearest point in the depth image is represented, for example, as white (maximum value), and the farthest point is represented, for example, as black (minimum value). In other words, distances or depths are simply specified between the "nearest" and "farthest" boundaries. The scaling in between typically corresponds to an inverse depth as depth information, since the methods are often trained on stereo data, and the inverse depth corresponds to a disparity.
[0023] A possible implementation of DPT can be downloaded from https: / / github.com / isl-org / DPT. The publication "Vision Transformers for Dense Prediction: René Ranftl, Alexey Bochkovskiy, Vladlen Koltun" (arXiv:2103.13413) contains a theoretical description for the implementation of the method, available for download at: https: / / arxiv.org / pdf / 2103.13413.pdf.
[0024] For a meaningful application, the depth (relative or absolute, especially metric) for the shared coordinate system must be recovered when using methods with "non-metric" depths. This is implemented in the scaling step based on the scaling information. To convert the depth maps (e.g. from DPT) into a true / metric or relative scale in the shared coordinate system, two values can typically be determined: an absolute offset and a scaling factor. Depending on the method, there may be more, fewer, or different parameters. In principle, if one of the surveillance cameras has a scaled depth map that has been scaled using the scaling information, the scaling can be transferred to another surveillance camera that has an overlapping surveillance area. This is achieved by
[0025] In a preferred embodiment of the invention, the scaling information comprises extrinsic and / or intrinsic camera parameters.
[0026] The extrinsic camera parameters can, for example, be represented by the positions of the surveillance cameras in the shared coordinate system. If the surveillance areas of the surveillance cameras with a known position overlap, the scaling can be implemented based on this scaling information.
[0027] The intrinsic camera parameters, such as the focal length of the surveillance camera, can be used to scale the depth map of the image from the surveillance camera with the known focal length. If the intrinsic camera parameters are only known for one surveillance camera and the surveillance areas of the other surveillance cameras overlap with the surveillance area of the surveillance camera with the known focal length, the scaling can be applied to the images of the other surveillance cameras.
[0028] Alternatively or additionally, the scaling information can be determined by an additional measuring device or by at least one surveillance camera as the measuring device. For example, a laser scanner can be used, which measures absolute distances to objects in the surveillance sub-area as scaling information. In the scaling step, the absolute distances are assigned to the corresponding areas in the depth maps, the depth map being scaled by this assignment. It can also be provided that a known measuring body is positioned in the surveillance sub-area, wherein the depth map is scaled by detecting the measuring body with the surveillance camera and knowing the dimensions of the known measuring body. A vehicle or a person with known or usual dimensions that happens to be present in the surveillance sub-area can also be used as the measuring body.
[0029] In a preferred embodiment of the invention, a ground plane of the surveillance area or the surveillance sub-area is determined in the image and / or in the depth map. The determination can be made, for example, by semantic segmentation, known from digital image processing. By adding the scaling information, such as a known calibration, extrinsic and / or intrinsic camera parameters, the position of the surveillance camera relative to the ground plane is known, so that the ground plane in the image and / or in the depth map can be scaled. In the event that the surveillance sub-areas have a common ground plane and the surveillance sub-areas overlap, the scaling can be performed without further information. The scaled ground plane thus forms the scaling information or a derived scaling information in the scaling step.
[0030] It is therefore proposed to use the scaled ground plane as scaling information. The ground plane should preferably be determined automatically or semi-automatically to simplify setup for the user. For example, semantic segmentation can be used, which can optionally be determined by DPT. All image regions identified as "road" are assumed to be the ground plane. By adding a known (extrinsic) calibration, the camera's position relative to the ground plane is known. By comparing the depth values in the "road" region with the expected depth values based on the plane assumption, the two unknown parameters can be determined and the depth map scaled.
[0031] For a particularly simple implementation, it is proposed that the depth map comprise map points with depth information, where the map points correspond to the pixels in the associated image, i.e., are configured as correspondence points. Thus, the map points with the depth information form a point cloud, which is transferred to the common coordinate system via the scaling step. Due to the correspondence of the map points to the pixels, the point cloud of the map points can be easily enriched with color information from the images in the fusion step.
[0032] In a preferred embodiment of the invention, the visualization model can be displayed, for example, on a 3D screen, wherein a user or the monitoring personnel can move within the visualization model using a virtual camera. Alternatively or additionally, the visualization model can be displayed using VR glasses, wherein the user or the monitoring personnel can change their own view by moving their head. It is also possible to define one or more virtual cameras, wherein the field of view of the virtual cameras is displayed on multiple screens or screen sections. This has the advantage that the virtual cameras can be used to define fields of view that are intuitively understandable for the user and / or the monitoring personnel.
[0033] A further object of the invention is formed by a visualization device, which is designed in particular to implement the method as described above. The visualization device is designed in particular as a digital data processing device, such as a computer, a server, or a cloud. The visualization device has an interface for receiving images from the surveillance cameras. The images each show a surveillance sub-area of a surveillance area. Optionally, the visualization device comprises the majority of surveillance cameras arranged in the surveillance area. Alternatively, the visualization device comprises the images of the surveillance sub-areas of the surveillance area.
[0034] The visualization device comprises a depth detection module configured for the depth detection step. Furthermore, the visualization device comprises a scaling module configured for the scaling step. Furthermore, the visualization device comprises a fusion module configured to implement the fusion step.
[0035] Optionally, the visualization device comprises an output device configured to output, in particular visualize, the visualization model. The output device can be configured, for example, as a 3D screen, VR glasses, or a screen for displaying fields of view of virtual cameras in the visualization model.
[0036] A further subject matter of the invention is a computer program comprising instructions that, when executed by a computer or the visualization device, cause the computer or the visualization device to carry out the method(s) of the method according to the invention. A further subject matter of the invention is a computer-readable data carrier, in particular a non-volatile, machine-readable memory, on which the computer program is stored.
[0037] Further features, advantages, and effects of the invention will become apparent from the following description of preferred embodiments and the accompanying figures. These show: Fig. 1 is a schematic block diagram of a visualization device in general form as an embodiment of the invention; Fig. 2 a schematic block diagram a concretization of the visualization device in the Fig. 1; Fig. 3 a schematic representation of a visualization model as a result of the visualization device of the preceding figures; Fig. 4 a camera section of a virtual camera in the visualization model in the Fig. 3.
[0038] The Fig. 1 shows a schematic block diagram of a visualization device 1 which is designed to implement a method for visualizing a surveillance area 2.
[0039] In the surveillance area 2, a plurality of surveillance cameras 3 are arranged, which are in the Fig. 1 are shown summarized as a block. Each surveillance camera 3 records an image of a surveillance sub-area 4 located within the field of view of the corresponding surveillance camera 3. The surveillance sub-areas 4 can be non-overlapping, as schematically shown in the lower area of the surveillance area 2. The surveillance sub-areas 4 can also overlap, forming overlapping areas 5, as schematically shown in the upper area of the surveillance area 2.
[0040] The visualization device 1 has an interface 6 for receiving images from the surveillance cameras 3. Optionally, the surveillance cameras 3 form a component of the visualization device 1.
[0041] The visualization device 1 has a depth detection module 7, wherein the depth detection module 7 is configured to create, in a depth detection step 100, a depth map for each image from the surveillance cameras 3 with depth information of the surveillance sub-area 4 associated with the image. The depth information in the respective depth maps is unscaled and / or embodied as relative information.
[0042] The visualization device 1 has a scaling module 8, wherein the scaling module 8 is configured to perform a scaling step of scaling the depth maps of all surveillance cameras 3 in one or the common coordinate system. In particular, the depth maps are brought into a common metric of the common coordinate system by the scaling module 8 and / or in the scaling step. In particular, the depth information of the various depth maps is directly comparable with one another. In principle, the scaled depth information can already be configured as absolute depth information, so that the depth information of the depth maps is configured as metric depth information in a common coordinate system.Alternatively, the scaled depth information may be formed as relative depth information, which is comparable in the common coordinate system but is not expressed in absolute values, such as meters.
[0043] The visualization device 1 has a fusion module 9 for implementing a fusion step 300, wherein the fusion module 9 is designed to fuse the scaled depth maps with image information of the surveillance sub-areas in a common visualization model of the surveillance area 2.
[0044] The visualization model can subsequently be displayed in an output device 10. The visualization model is designed, in particular, as a 3D model of the surveillance area 2, which is realistically displayed by the enhancement with image information and which can be displayed by a user or surveillance personnel in any manner, for example, using a 3D screen, VR glasses, or via one or more screens, each as an output device 10. Any view can be displayed in the visualization model, for example, by using virtual cameras.
[0045] The process creates a 3D representation of the scene in surveillance area 2 that is significantly more intuitive to understand and also offers several new possibilities that a conventional representation does not allow or only allows with difficulty. The goal / result is that larger, connected camera arrays are displayed in a single fused view.
[0046] In principle, the depth detection module 7 and the scaling module 8 can be implemented as a single module that performs the depth detection step 100 and the scaling step 200, so that the scaled depth map is generated in a single step. As previously described, this is a comparatively difficult and complex task for the corresponding evaluation methods.
[0047] Accordingly, in the exemplary embodiment, it is proposed that the depth detection step 100 is first performed, e.g., based on the previously described DPT, and subsequently the scaling is performed based on at least one piece of scaling information. The scaling information has information content that enables the depth maps to be scaled. In particular, it is information selected from intrinsic camera parameters of at least one surveillance camera, extrinsic camera parameters (in particular location) of at least two surveillance cameras with overlapping surveillance sub-areas, and / or overlap information of surveillance sub-areas. By introducing the scaling information as a priori knowledge, the estimation of the scaled depth maps can be significantly improved, so that the method and the visualization device 1 are designed more robustly.
[0048] Overall, a completely different type of representation is proposed, combining methods for monocular depth estimation (mono-depth) and (camera) calibration.
[0049] The Fig. 2 shows an embodiment of the invention in the form of a concrete embodiment of the visualization device 1 in the Fig. 1. Identical or corresponding components are provided with the same or corresponding reference symbols.
[0050] The Fig. Figure 2 shows an exemplary implementation of the method based on Vision Transformers for Dense Prediction (DPT), where multiple individual images (one image per surveillance camera 3) are used to predict the non-metric depth and semantics, and a previously determined intrinsic and extrinsic calibration based (only) on the (here 24) individual images is used to rescale the depth maps, determine the lines of sight corresponding to the pixels, and convert the relative position of the surveillance cameras 3 and thus 3D point clouds into a common 3D view. Filtering was also performed for improved display. It should be noted that all steps shown here (except for determining the calibration) can be performed independently for each surveillance camera 3, and thus in parallel. Not shown are: • Live processing, • Anonymization, a • more clever fusion, which corrects, for example, (incorrect) depth jumps in overlapping views, and • a color correction that corrects the color balance of the individual surveillance cameras 3.
[0051] In the example shown in the Fig. 2, images from 24 different surveillance cameras 3 were received. First, a calibration was determined, whereby corresponding points in the images were manually selected as scaling information. From this, the intrinsic (distortion, focal length) and extrinsic (relative position of the cameras to each other and the ground plane) parameters could be determined. The calibration process could also be performed in other ways (e.g., using known lens data and autocalibration methods).
[0052] First, both a depth map and a semantic segmentation are generated using DPT in the depth detection module 7. Ground plane regions (here, the "Road" class) from the semantic segmentation are used. As scaling information, it was assumed that this is indeed a (ground) plane. Alternative assumptions and methods are discussed in the next section.
[0053] Based on the calibration, the required depth / distance for each pixel can be predicted if it is a point on the ground plane. By comparing the corresponding "road" regions of the depth map with this data, the scaling parameters – in this case two – can be determined, which allow the depth map to be converted into metric depths / distances. Finally, the intrinsic calibration of surveillance camera 3 is used again to calculate (colored) 3D point clouds from the now metric depths.
[0054] In many methods for monocular depth estimation, depth jumps are "smoothed," resulting in unsightly artifacts in the depth reconstruction. To filter these out, an edge detector was applied to the depth map. A better type of filtering would do this not on the depth map (which represents the inverse depth), but on the actual depths (or a mixture or combination), since the method shown here becomes less sensitive at greater distances (the inverse depth values become very small due to the greater distance, and thus also the jumps). The choice and type of filtering can, of course, be different (see also the next section). Finally, the point clouds of the individual surveillance cameras 3 are combined. In the simplest case, these are simply displayed together and independently of each other. A small adjustment was made here.Only the nearest points from each surveillance camera 3 are displayed. This prevents gross miscalculations at greater distances from overshadowing / overwriting good results from a local surveillance camera 3. It should be noted that this is done only based on the calibration information and thus independent of the actual content (3D point clouds) of the other surveillance cameras 3. This allows for fully parallel processing. Example results:
[0055] In Fig. 3 shows the overall result of the 24 camera images as visualization model 11. The image in Fig. Figure 3 is a screenshot from the CloudCompare software, which can be used to display 3D point clouds. In some areas, especially those with a high level of occlusion, the 3D reconstruction is currently incorrect and somewhat difficult to interpret. It must be noted that DPT is a method from 2021, and better results can probably be expected with more recent methods. Furthermore, the method was probably not trained on corresponding images from this (surveillance) area 2. However, if a virtual camera is moved in the visualization model 11 as a point cloud visualization closer to the real camera positions, the potential becomes more apparent. This is clearly shown in Fig. 4, where a section 12 from Fig. 3. Here, the images also have a similar color balance, which is why the transition between the two point clouds on the ground is invisible. The output device 10 can thus be a screen that displays the partial section 12 as the field of view of a virtual camera. Alternative or improved representation:
[0056] A further and better form of representation uses a simple anaglyph process to generate an actual 3D representation on a screen as output device 10 with the help of red-cyan color glasses. This representation significantly aids scene perception, as the content can be perceived not only through texture differences and depth differences by (virtual) rotating / moving the camera / scene, but directly. An even better representation results from the use of a VR headset, such as an Oculus Rift, as navigation (rotating and moving) and depth perception are significantly easier / more intuitive and a larger field of view is utilized. Another alternative is a 3D monitor as output device 10 with appropriate glasses. Live View:
[0057] The figures only describe a static image / point cloud. For one application, for example, (A) a current image / point cloud is generated every few minutes / seconds. Alternatively, this could also happen live (for each or almost each) image. Since this can be very computationally intensive (despite the possibility of independent processing for each surveillance camera 3), it could be advantageous (B) to determine the depths less frequently and only update the current texture (color image). The depth map could, for example, be a "background" depth image generated from several, temporally offset, depth maps so that moving objects are not included. Moving objects could then also be displayed separately, for example in the form of avatars or as simple shapes (cylinders, cuboids) onto which the texture (color image) is projected.The detection of these objects typically already occurs in surveillance camera 3, so processing should be significantly faster. Anonymization:
[0058] If, as in the previous point ("Live View"), a background is created and people, vehicles, etc., are drawn separately, this also represents a possible approach to anonymization. A further step would be to replace the background texture (workshop in the case of the described example) with a different one. Then the user or surveillance personnel would only see the 3D structure and avatars, etc., and would not even know what is actually being seen. This represents (almost) perfect anonymization. Calibration tracking
[0059] For moving (PTZ) surveillance cameras, the time-varying calibration must be taken into account. If the calibration change is known, it can be easily integrated. However, this requires recalculating the depth map. Alternatively, if a complete depth background map was previously generated for a moving camera as surveillance camera 3, no recalculation is necessary; only a corresponding conversion to the current field of view is required. Improved Fusion
[0060] In the described method, the 3D point clouds of each surveillance camera are simply displayed together. It often happens that parts of the camera views overlap, and errors in depth estimation or calibration lead to duplicate structures. The same structure appears multiple times at different locations. A method could be used to detect and fuse these structures. This could be achieved by averaging the depth and texture information at the corresponding location or by clever stitching. Color correction (color balancing)
[0061] In Fig. 3, for example, it is clearly visible on the floor surface that no color balancing was performed between the images. Fig.4, however, clearly shows how a good (here random) color balancing improves the illusion of a single result compared to a superposition of many individual results. Rescaling procedure and ground plane selection:
[0062] In the described method, the ground plane and automatic detection of the corresponding image areas via semantic segmentation were used to incorporate metric information into the "non-metric" depth images. If a continuous ground plane is actually present, this has the advantage that this can be done fully automatically. Of course, this cannot be applied in many situations. Furthermore, the reconstruction in the described method is metric, but not "in meters." This means that lengths can be compared, but lengths cannot be directly specified in meters. For this, absolute scale information, such as the height of a person, the height of surveillance camera 3 above the ground, or similar, must be incorporated.
[0063] Alternatively, metric information can be incorporated, for example, via image correspondences in the overlap area of two camera views / surveillance sub-areas 4. For example, depth could be determined via triangulation (as in classic stereo methods). The correct scale could be derived, for example, from the known distance between two surveillance cameras 3 or the camera height. Alternatively, surveillance cameras 3, such as the Flexidome Multi from Bosch, could take advantage of the fact that the surveillance cameras 3 can only be moved within a ring whose radius is known. Alternatively, the correct scale can also be determined through a few direct distance measurements, e.g., with a laser rangefinder.
[0064] A selection of a suitable ground plane or other plane could also be implemented using methods such as Segment-Anything, where the operator can select corresponding areas with just a few mouse clicks. Removing implausible depth measurements:
[0065] As shown in the procedure, implausible depth measurements can be discarded by filtering. This can be achieved using simple methods. However, it can also be useful, for example, to detect whether meaningful structures emerge in the overlap area of multiple surveillance cameras (see also "Improved Fusion"). Finally, timestamps should also be (automatically) removed, since no meaningful depth can be assigned to them. Parallel processing
[0066] As mentioned previously, depending on the application, completely parallel processing can be performed, except for the visualization. This means that, for example, even individual surveillance cameras can generate three 3D point clouds or a depth map in a suitable format and send them to a central server or alarm monitoring center. For the amount of data that must be transmitted, see also "Internal 3D Representation." Rectification:
[0067] Many monodepth methods, such as DPT, are not designed for image data with significant distortion, as is often the case. Therefore, the images are first rectified, which yields significant improvements. A processing step right at the beginning of the process can therefore also be to rectify the image or even split it into several rectified images and then combine the depth results of the individual images. This can be particularly useful with very wide-angle cameras (e.g., fisheye). Internal 3D representation:
[0068] Representing the data as a point cloud (several million individual points) allows for a high, but unnecessary, level of detail. Alternatively, a 3D mesh could be used, which allows for very fast rendering of the texture. This representation of the results is common in computer games, etc., and is quick to process while requiring less storage space and transmission bandwidth. See-through:
[0069] The combination of 3D information and multiple views also allows for functions such as see-through, where people who are hidden in the current (virtual) view are schematically displayed. Improved tracking:
[0070] The 3D information allows for improved tracking of people, since it is known that tracked people will, for example, disappear behind a car if they continue walking. Other scenarios:
[0071] In the described method, an example with many surveillance cameras (3) and a large overlap area was shown. However, the method does not essentially require this. Another scenario would be monitoring the fence of a substation. Typically, many surveillance cameras (3) are used here, but with little or no overlap area. For example, a satellite or topographic map could be enhanced with 3D point clouds / meshes. The representation could be as described above. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] US 9,706,264 B2
[0003] Zitierte Nicht-Patentliteratur
[0000] ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth“ (https: / / arxiv.org / abs / 2302.12288
[0017] Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D“ (https: / / arxiv.org / abs / 2008.05711
[0017] https: / / github.com / isl-org / DPT
[0023] https: / / arxiv.org / pdf / 21 03.13413.pdf
[0023]
Claims
[1] Method for visualising a surveillance area (2), wherein a plurality of surveillance cameras (3) are arranged in the surveillance area (2), wherein each surveillance camera (3) can record an image of a surveillance sub-area (4) of the respective surveillance camera (3), wherein in a depth detection step (100) a depth map with depth information of the surveillance sub-area (4) is created for the images of the surveillance cameras (3), wherein in a scaling step (200) a scaling of the depth maps is carried out for a common coordinate system of the monitoring sub-areas (4), wherein in a fusion step (300) the scaled depth maps are fused with image information of the images of the surveillance sub-areas (4) in a common visualization model (11) of the surveillance area (2). [2] Method according to claim 1, characterized bythat in the depth detection step (100) the depth map is implemented using a method for monocular depth estimation. [3] Method according to claim 1 or 2, characterized by that the respective depth map is temporarily stored as an intermediate result before the scaling step (200). [4] Method according to one of the preceding claims, characterized by that in the scaling step (200) the scaling is carried out on the basis of at least one scaling information. [5] Method according to claim 4, characterized by that the scaling information comprises extrinsic and / or intrinsic camera parameters of the surveillance cameras (3). [6] Method according to one of the preceding claims 4 or 5, characterized by that the scaling information is determined by an additional measuring device or by at least one surveillance camera (3) as a measuring device. [7] Method according to one of the preceding claims 4 to 6, characterized by that in the image and / or in the depth map a base plane of the surveillance area (2) or of the surveillance sub-area (4) is determined and scaled, wherein the scaling of the depth maps is carried out on the basis of the scaled base plane as scaling information. [8] Method according to one of the preceding claims, characterized by that the depth map has map points with depth information, where the map points correspond to or match the image points in the associated image. [9] Method according to one of the preceding claims, characterized by that the visualization model (11) is displayed as an output device (10) via a 3D screen, VR glasses and / or as a camera image of a virtual camera in the visualization model (11). [10] Visualization device (1), in particular for implementing the method according to one of the preceding claims, with an interface (6) for receiving images from a plurality of surveillance cameras (3), wherein the images each show a surveillance sub-area (4) of a surveillance area (2), with a depth detection module (7), wherein the depth detection module (7) is designed to create a depth map with depth information of the surveillance sub-area (4) for the images of the surveillance camera (3), wherein the depth information is unscaled and / or formed as relative information in the depth map, with a scaling module (8), wherein the scaling module (8) is designed to perform a scaling of the depth maps in a common coordinate system for the monitoring sub-areas (4), with a fusion module (9), wherein the fusion module (9) is designed to fuse the scaled depth maps with image information of the surveillance sub-areas in a common visualization model (11) of the surveillance area (2). [11] Visualization device (1) according to claim 10, characterized by an output device (10), wherein the output device (10) is designed to display the visualization model (11). [12] Computer program comprising instructions which, when the program is executed by a computer or the visualization device (1) according to claim 10, cause the computer or the visualization device (1) to carry out the method / steps of the method according to any one of claims 1 to 9. [13] A computer-readable data carrier on which the computer program according to claim 12 is stored.
Citation Information
Patent Citations
Multiple field-of-view video streaming
US9706264B2