PTZ occlusion control

By comparing the camera's field of view and dynamically adjusting the occlusion area, the problem of low-authority users having difficulty achieving context awareness in multi-camera monitoring systems is solved, achieving a balance between visibility of objects of interest and privacy protection in scene overview.

CN116962882BActive Publication Date: 2026-01-02AXIS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310359044.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-26
Filing Date
2023-04-06
Publication Date
2026-01-02
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

In multi-camera surveillance systems, users with lower authorization levels find it difficult to gain contextual awareness through scene overviews, especially when using PTZ cameras to track objects of interest, where privacy occlusion in the scene overview hinders the understanding of the object's location.

Method used

By comparing the fields of view of different cameras, the occluded areas in the output video stream are dynamically adjusted to ensure that objects of interest are visible in the scene overview. Pixelation or de-occlusion techniques are used, combined with object detection and tracking algorithms, to dynamically adjust the occlusion level to improve context awareness.

Benefits of technology

It improves the ability of users with lower authorization levels to perceive the location of objects of interest in a scene, while maintaining privacy protection and enhancing the system's context awareness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116962882B_ABST
    Figure CN116962882B_ABST
Patent Text Reader

Abstract

PTZ occlusion control is provided. A method of generating an output video stream, comprising: receiving a first input video stream from at least one first camera capturing a scene; receiving a second input video stream from a second camera capturing only a part of the scene; and generating the output video stream from the first input video stream, including suppressing occlusion of a particular region of the output video stream or at least reducing a level of occlusion of the particular region of the output video stream in response to determining that the particular region of the output video stream depicts a part of the scene depicted in the second input video stream. Corresponding apparatus, camera monitoring system, computer program and computer program product are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to (privacy) obscuring in video streams generated in a camera surveillance system. In particular, the present disclosure relates to such obscuring when multiple video streams are received from multiple cameras capturing the same scene. BACKGROUND

[0002] In modern camera surveillance systems, multiple cameras can be used to capture the same scene. For example, one or more cameras can be used to generate a scene overview, while one or more other cameras can be used to generate a more detailed view of the scene. A camera used to generate a scene overview can for example be a camera with a larger field-of-view (FOV) such that it can capture a larger part of the scene. Other examples can include combining video streams from multiple cameras, each camera having a smaller FOV and each camera capturing a different part of the scene, in order to form a combined video stream that still captures a larger part of the scene. In order to generate a more detailed view of the scene, so-called pan-tilt-zoom (PTZ) cameras can preferably be used, which PTZ cameras are able to for example zoom in on an object or area of interest in the scene and are able to for example provide more details about such object or area than are available from a scene overview only. Subsequently, a user of the camera surveillance system can for example use the scene overview to identify the presence of an object or area of interest and subsequently use the PTZ camera to study this object or area in more detail (e.g. by zooming in to get closer to the object of interest, and / or by using the pan and / or tilt functionality of the PTZ camera to track the object of interest if it moves in the scene).

[0003] In such camera surveillance systems, it is not uncommon that different users have different levels of authorization, and that viewing a more detailed view of the scene can require a higher level of authorization than is required for viewing a scene overview only. This is because a scene overview provides less details about individual objects (e.g. persons) than a PTZ camera. In particular, in order to further reduce the level of authorization required for viewing a scene overview, (privacy) obscuring can be applied in the scene overview such that for example the identity of a person, a license plate or similar information is hidden and cannot be recovered by a low level of authorization user observing the scene overview. On the other hand, for a PTZ camera, such obscuring is typically not performed such that if an object of interest is found, the full potential of the PTZ camera can be exploited and thus the level of authorization required for viewing a more detailed view of the scene generated by the PTZ camera is typically higher.

[0004] If a PTZ camera is used, for example, to track an on-site event (e.g. to track the movement of an object of interest), it can also be desirable to use the scene overview to improve situational awareness. For example, if the PTZ camera has zoomed in on an object of interest, it can be difficult to understand the exact location of the object in the scene if only the more detailed view of the scene is used. Moreover, by using the scene overview, it can be easier to locate the object within the scene. However, if a privacy obstruction is performed in the scene overview, it can be difficult to exploit the full potential of the scene overview for situational awareness. SUMMARY

[0005] To improve the above-described situation in which the use of a (privacy) obstruction of an object and / or area in the scene overview is used to lower the level of authorization required to view the scene overview, but in which it is also desirable to use the scene overview to improve situational awareness, the present disclosure provides an improved method of generating an output video stream and corresponding apparatus, camera surveillance system, computer program and computer program product, as defined in the attached independent claims. Various embodiments of the improved method, apparatus, camera surveillance system, and computer program and computer program product are defined in the attached dependent claims.

[0006] According to a first aspect, a method of generating an output video stream is provided. The method comprises receiving a first input video stream from at least one first camera that captures a scene. The method further comprises receiving a second input video stream from a second camera that only partially captures the scene. The method also comprises generating the output video stream from the first input video stream, including suppressing an obstruction of, or at least reducing the level of an obstruction of, a particular region of the output video stream in response to determining that the particular region of the output video stream depicts a part of the scene that is depicted in the second input video stream. In some embodiments, it is designed that the obstruction can be suppressed (or reduced) in a region of the output video stream that corresponds to the part of the scene that is depicted in the second input video stream, and that the obstruction can be performed as usual (i.e. not suppressed or reduced) in one or more other regions of the output video stream that do not correspond to the part of the scene that is depicted in the second input video stream. It is noted that “suppressing the obstruction” or “reducing the level of the obstruction” means that such a region would indeed be obstructed in the output video stream if the designed method that takes the content of the second input video stream into account is not used.

[0007] The designed method improves the currently available techniques and solutions as it takes into account what is currently seen in the second input video stream from the second camera and then prevents (or reduces the level of) occlusion of the corresponding area in the scene overview. By doing so, the situational awareness is improved as, for example, the object of interest currently tracked in the more detailed view of the scene becomes visible also in the scene overview and specifically this "de-occlusion" in the scene overview is performed dynamically and adapted to the changes in the FOV of the second camera.

[0008] In some embodiments of the method, determining that the specific area of the output video stream depicts a portion of the scene depicted in the second input video stream can comprise comparing the field of view (FOV) of the at least one first camera with the field of view of the second camera. As used herein, the FOV of a camera can comprise, for example, the position of the camera relative to the scene or at least relative to another camera, the zoom level (optical or digital) of the camera, the tilt / pan / rotation of the camera, etc. By comparing the FOV of the two cameras, it can thus be determined whether the two cameras capture the same scene portion and specifically which area of the image captured by one camera corresponds to a portion of the scene captured by the other camera. In other embodiments, it is designed that other techniques can also be used to achieve the same effect and are based on, for example, image analysis and feature recognition rather than on the FOV parameters of the cameras.

[0009] In some embodiments of the method, the method can comprise generating an output video stream from a plurality of first input video streams from a plurality of first cameras, each first camera capturing a portion of a scene. The specific area of the output video stream can depict the entire scene as captured by at least one of the plurality of first cameras (e.g., as depicted in at least one of the plurality of first input video streams). For example, it is designed that two or more cameras can be arranged such that they are aligned in different directions and thus capture different portions of the scene (which can have some overlap, but not necessarily). The input video streams from these cameras can then be combined into a single composite video stream on which the video output stream is based. The video streams from each camera can be, for example, stitched together to create the illusion of receiving a single input video stream from a camera with a very large FOV. In other embodiments, the input video streams from each first camera can instead be shown next to each other in the output video stream, such as in, for example, a horizontal array, a vertical array, and / or in a grid or similar manner.

[0010] In some embodiments of the method, the obscuring can comprise pixelating a specific region of the output video stream. Reducing the level of obscuring can comprise reducing a block size of the pixelation for the specific region of the output video stream. This can be advantageous in that, for example, the resolution of the object of interest can be the same in both the output video stream and the second output video stream showing the second input video stream.

[0011] In some embodiments of the method, the method can further comprise determining, based on the first input video stream, that an extended region of the output video stream comprising the specific region is a region to be obscured. The method can comprise obscuring at least a portion of the extended region in the output video stream other than the specific region. In other words, it can be designed that still a portion of a larger obscuring object that is otherwise desired can be retained, while only the specific region of such larger obscuring object that corresponds to a portion of the scene that is also captured by the second camera is "de-obscured".

[0012] In some embodiments of the method, the extended region of the output video stream can be a predefined region of the first input video stream, or, if in other words, the extended region of the output video stream corresponds to a predefined portion of the scene. Such predefined region / portion of the scene can for example be manually set when configuring the camera surveillance system in order to, for example, statically obscure one or more features or regions of the scene that are a priori known, for example, where people usually appear or reside. Such features and / or regions can be, for example, windows of a residential building, windows of a coffee shop or restaurant, a day care center, a playground, an outdoor service area, a parking lot, a bus or train station, etc.

[0013] In some embodiments of the method, the portion of the scene depicted in the second input video stream can be defined as a portion of the scene comprising at least one object of interest. In other words, the otherwise desired obscuring of the first input video stream and the output video stream can be more dynamic than static and, for example, tracks people or other objects as they move within the scene, and wherein the obscuring is first performed only after it is also determined that the people and / or objects should be obscured otherwise (if not inhibited from obscuring them according to the designed method). For example, obscuring of an object can only be required if it is first determined (by using, for example, object detection or similar methods) that the object belongs to a class to be obscured (e.g. people, license plates, etc.).

[0014] In some embodiments of the method, the part of the scene depicted in the second input video stream can be defined as the entire scene depicted in a predefined area of the second input video stream, such as for example the entire area of the second input video stream. In other embodiments, the predefined area of the second input video stream can correspond to less than the entire area of the second input video stream. For example, it can be designed to detect in which area of the second input video stream, for example, the object of interest is located, and to only de-occlude (or more precisely, suppress occlusion) in the output video stream the part of the first input video stream corresponding to that area, thus "de-occluding" a smaller area of the scene in the output video stream than the area of the scene captured by the second camera.

[0015] In some embodiments of the method, the method can be executed in a camera arrangement. The camera arrangement can comprise the second camera and the at least one first camera. In other embodiments, it is designed that the second camera and the at least one first camera are provided as separate cameras, not forming part of the same camera arrangement, and that the method is then executed in any of these cameras, or for example jointly by all cameras (assuming that these cameras have some way of communication with each other).

[0016] According to a second aspect of the present disclosure, a device for generating an output video stream is provided. The device comprises a processor. The device further comprises a memory storing instructions which, when executed by the processor, cause the device to: receive a first input video stream from at least one first camera capturing a scene; receive a second input video stream from a second camera capturing only a part of the scene; and generate an output video stream from the first input video stream, including suppressing occlusion or at least reducing a level of occlusion of a particular area of the output video stream in response to determining that the particular area of the output video stream depicts a part of the scene depicted in the second input video stream.

[0017] Thus, the device according to the second aspect is configured to perform the corresponding steps of the method of the first aspect.

[0018] In some embodiments of the device, the device is further configured (i.e. the instructions, when executed by the processor, cause the device to) perform any of the embodiments of the method described herein.

[0019] In some embodiments of the device, the device is one of the at least one first camera and the second camera. In other embodiments of the device, the device is a camera arrangement comprising all of the at least one first camera and the second camera.

[0020] According to a third aspect of the present disclosure, a camera surveillance system is provided. The camera surveillance system comprises at least a first camera, a second camera, and a device as designed herein according to the second aspect (or any embodiment thereof).

[0021] According to a fourth aspect of the present disclosure, a computer program for generating an output video stream is provided. The computer program product is configured to, when executed by a processor of a device (e.g. a device according to the second aspect or any embodiment thereof), cause the device to: receive a first input video stream from at least one first camera capturing a scene; receive a second input video stream from a second camera capturing only a part of the scene; and generate the output video stream from the first input video stream, including suppressing an occlusion of a particular region of the output video stream or at least reducing a level of an occlusion of the particular region of the output video stream in response to determining that the particular region of the output video stream depicts a part of the scene that is depicted in the second input video stream.

[0022] Hence, the computer program according to the fourth aspect is configured to cause the device to perform the method according to the first aspect.

[0023] In some embodiments of the computer program, the computer program is further configured to cause the device to perform any embodiment of the method of the first aspect as described herein.

[0024] According to a fifth aspect of the present disclosure, a computer program product is provided. The computer program product comprises a computer-readable storage medium storing a computer program according to the fourth aspect (or any embodiment thereof as described herein). The computer-readable storage medium may, for example, be non-transitory and be provided as, for example, a hard disk drive (HDD), a solid state drive (SDD), a USB flash drive, an SD card, a CD / DVD, and / or any other storage medium capable of non-transitorily storing data.

[0025] Other objects and advantages of the present disclosure will be apparent from the following detailed description, drawings and claims. Within the scope of the present disclosure, it is designed that all features and advantages described with reference to, for example, the method of the first aspect are relevant, applicable to, and can also be used in conjunction with any feature and advantage described with reference to the device of the second aspect, the surveillance camera system of the third aspect, the computer program of the fourth aspect, and / or the computer program product of the fifth aspect, and vice versa. BRIEF DESCRIPTION OF DRAWINGS

[0026] Exemplary embodiments will be described below with reference to the accompanying drawings, in which:

[0027] Figure 1A general camera surveillance system is schematically illustrated;

[0028] Figure 2A An embodiment of generating an output video stream according to the present disclosure is schematically illustrated;

[0029] Figure 2B A flowchart of an embodiment of a method according to the present disclosure is schematically illustrated;

[0030] Figure 2C An alternative embodiment of a method according to the present disclosure is schematically illustrated; and

[0031] Figure 3 Various embodiments of an apparatus for generating an output video stream according to the present disclosure are schematically illustrated.

[0032] In the drawings, like reference numerals will be used to refer to like elements throughout. Unless explicitly stated otherwise, the drawings represent essentially non-scale views. Unless specifically stated otherwise, the drawings are not to scale, and, for the sake of clarity, other elements can be omitted or suggested only. As illustrated in the drawings, the (absolute or relative) dimensions of elements and regions can be exaggerated or belittled relative to their true values for the purpose of illustration, and thus, are provided to illustrate the general structure of the embodiments. DETAILED DESCRIPTION

[0033] Figure 1 A general camera surveillance system 100 is schematically illustrated, comprising camera devices 110, an apparatus 150 and a user terminal 160. The camera devices 110 comprise second cameras 130 and four first cameras 120 (of which only three cameras 120a, 120b and 120c are visible in the Figure 1 The first cameras 120 are overview cameras mounted in a circular pattern and each aimed in a different direction so that they together capture a 360-degree horizontal view of e.g. a scene. The second cameras 130 are pan-tilt-zoom (PTZ) cameras and are mounted below the first cameras 120 and (if viewed e.g. from above or below) in the center of the circular pattern.

[0034] The first cameras 120 generate first input video streams 140 (i.e. four input video streams 140a to 140d if there are four first cameras 120), while the second cameras generate second input video streams 142. The input video streams 140 and 142 are provided to the apparatus 150. The apparatus 150 is configured to receive the input video streams 140 and 142 and to generate corresponding output video streams 152 and 154.

[0035] The output video streams 150 and 152 are provided by the device 150 to the user terminal 160, wherein the first output video stream 152 is displayed on a first monitor 161 of the terminal 160, and wherein the second output video stream 154 is displayed on a second monitor 162 of the terminal 160. In Figure 1 In the particular example of the system 100 illustrated in Fig. 1, there are four first input video streams 140a to 140d (one provided by each camera 120) provided to the device 150, and the device 150 in turn combines the first input video streams 140a to 140d into a single combined first output video stream 152. For example, the device 150 can arrange the input video streams 140a to 140d in a 2x2 grid such that each quadrant of the 2x2 grid displays one first input video stream from a corresponding first camera. Of course, other alternative arrangements of converting such multiple first input video streams into a single combined first output video stream can also be devised. The various first input video streams can for example also be stitched together to create a panoramic first output video stream or the like.

[0036] In this context, it can be devised that the device 150 is a separate device from the camera arrangement 110 and the user terminal 160. In other embodiments, the device 150 can alternatively be an integrated part of the camera arrangement 110 or the user terminal 160.

[0037] The user of the terminal 160 can use the first output video stream 152 displayed on the first monitor 161 for situational awareness, while the second output video stream 154 displayed on the second monitor 162 can be used to provide more details about a particular part of the scene. For example, the user can detect an object or area of interest by watching the first monitor 161, and subsequently adjust the field of view (FOV) of the second camera 130 such that the second camera 130 captures the object or area of interest in more detail. For example, the FOV of the second camera 130 can be adjusted, e.g. by scaling, panning, rotating and / or tilting the second camera 130 accordingly, such that the second output video stream 154 displayed on the second monitor 162 provides a zoomed-in view of the particular object or area of interest. For situational awareness, e.g. to still keep track of where the zoomed-in object is currently located in the scene, the first output video stream 152 displayed on the first monitor 161 can be used.

[0038] The FOV adjustment of the second camera 130 can also be automatic, such that if an object or region of interest is detected in the first input video stream 140 and / or in the corresponding first output video stream 152 (e.g., by using object detection and / or tracking algorithms), the FOV of the second camera 130 is adjusted so that the second input video stream 142 (and thus also the corresponding second output video stream 154) provides a close-up view of the detected object or region of interest when displayed on the second monitor 162. Alternatively, it can be designed to use, for example, digital zoom, without altering the second input video stream 142, instead of physical / optical zoom (i.e., by adjusting the lens of the second camera 130).

[0039] For privacy reasons, it may be necessary (e.g., according to local laws) to obscure one or more "sensitive" parts of a scene so that, for example, the identity or other personal characteristics of objects present in these parts of the scene, vehicles, or the like, cannot be determined by viewing the first monitor 161. Such sensitive parts may, for example, correspond to a shop window, a window of a residential building, a bus or train station, a restaurant, a coffee shop, or similar places where, for example, people may be present. Therefore, the camera system 100 and device 150 are thus configured such that one or more privacy obscuring objects are applied when the first output video stream 152 is generated. Figure 1 In the specific example illustrated, three such privacy screens 170a, 170b, and 170c are applied. The first privacy screen 170a is located in front of a portion of the scene including a shop, the second privacy screen 170b is located in front of a portion of the scene including a bus stop, and the third privacy screen 170c is located in front of a residential unit on the ground floor / first floor of a building. Figure 1 In this context, these privacy occlusions are displayed as shadow lines, but they could also be completely opaque, making it impossible to obtain detailed information about people or similar objects behind these occlusions by viewing the first monitor 161. In other embodiments, privacy occlusions 170a to 170c can be created, for example, by pixelating the relevant areas, or even by making the objects to be occluded transparent (e.g., if an image of the scene without these objects is available previously). Other methods of occluding objects have also been devised.

[0040] exist Figure 1In the specific example illustrated, a particular object of interest 180 exists in the first input video stream 140a and consists of a person currently attempting to break into the store. Person 180 appears to have smashed a shop window, and therefore the user of terminal 160 is interested in tracking the progress of this intrusion, allowing them to monitor the person 180's movements in real time. For example, by adjusting the FOV of the second camera 130 accordingly, a close-up view of person 180 is provided in the second output video stream 154 displayed on the second monitor 162. No privacy masking is applied in the second output video stream 154.

[0041] As previously described herein, system 100 is typically configured such that the authorization level required to view the first output video stream 152 is lower than the authorization level required to view the second output video stream 154. This is because the scene overview provided by the first output video stream 152 contains less detail about individual objects in the scene, and, that is, because privacy occlusions 170a to 170c have been applied before the user can view the first output video stream 152 on the first monitor 161. However, in situations such as Figure 1 In the scenario illustrated, the user of terminal 160 expects to use the first output video stream 152 for context awareness, making it easier for the user to know where person 180 is currently located in the scene. For example, if person 180 begins to move, it may be difficult to know where person 180 is currently located in the scene simply by looking at the second terminal 162. For example, knowing the current location of person 180 can be important if law enforcement or similar personnel are being called. However, this context awareness is hindered because person 180 is hidden in the first output video stream 152 by privacy obscurations 170a to 170c, and therefore it is more difficult to use the first output video stream 152 displayed on the first monitor 161 for this purpose.

[0042] Now will also refer to Figure 2A to Figure 2C This paper will describe in more detail how the methods designed in this paper can help improve the situation described above.

[0043] Figure 2A The schematic map illustrates a design embodiment of method 200 for generating an output video stream, while Figure 2BA flow chart schematically illustrates the steps included in such a method 200. In a first step S201, a first input video stream 140 is received from at least one first camera (e.g., one or more of the cameras 120a-c), wherein the first input video stream 140 captures / depicts a scene. In a second step (S202, which can also be performed before or simultaneously with step S201), a second input video stream 142 is received from a second camera (e.g., the camera 130). The second input video stream 142 only captures / depicts a part of the scene. For example, the FOV of the second camera 130 is smaller than the FOV (or combined FOV) of the first camera (or cameras) 120. In other words, the second input video stream 142 provides a close-up view of an object or region of the scene. As discussed earlier herein, such a result can be obtained, for example, by mechanical / optical zooming of the second camera 130 (i.e., by moving the lens of the second camera 130) or, for example, by digital zooming. In the latter case, even if it is not necessary to subsequently change the second input video stream 142 generated by the second camera 130 in order to generate the second output video stream 154, for the sake of explaining the designed improved method, it can anyway be assumed that the second input video stream 142 thus only partially captures the scene.

[0044] The designed method 200 then continues in a third step S203 with generating an output video stream 220 from the first input video stream 140, as will now be described. In order to generate the output video stream 220, it is first determined whether a particular region 212 of the resulting output video stream 220 depicts a part 210 of the scene depicted in the second input video stream 142. If so, the generation of the output video stream 220 suppresses or at least reduces the level of occlusion of the particular region 212 of the output video stream.

[0045] In Figure 2AIn the example illustrated herein, the second input video stream 142 includes a close-up view of person 180 and depicts a portion of the scene also found in the first input video stream 140, as illustrated by the dashed rectangle 210. To determine the position and size of rectangle 210, it is designed to use, for example, information about the current FOV of the second camera 130 and information about the current FOV of one or more first cameras 120, and compare them. In the designed method 200, it can then be confirmed that a specific region 212 of the corresponding output video stream 220 depicts the same portion of the scene as the second input video stream 142, i.e., as illustrated by the dashed rectangle 210. As used herein, knowing information about the FOV of a camera includes, for example, knowing the camera's position relative to the scene (and / or relative to one or more other cameras), the camera's zoom level (optical / mechanical and / or digital), the camera's orientation (e.g., camera pan, tilt, and / or roll), and similar information. By comparing the FOV of the two cameras, it is possible to calculate, for example, whether the first camera depicted the same part of the scene as depicted by the second camera, and specifically, which region of the video stream generated by the first camera corresponds to that same part of the scene.

[0046] When applying one or more privacy censors 230a to 230c (as already referenced) Figure 1 When occlusion is applied to objects such as shops, bus stops, and residential units (as discussed), the designed method 200 does not apply occlusion in a specific area 212 of the output video stream 220. As a result, the contextual awareness of the scene obtainable from the output video stream 220 is thus enhanced because there is no obstruction covering the person 180. If the person 180 begins to move, for example, the person 180 will not be occluded in the output video stream 220 as long as the person 180 remains within the FOV of the second camera 130 that generates the second input video stream 142. It can be designed, for example, that the second camera 130 tracks the person 180 (either manually controlled by the user or, for example, using object detection and / or tracking), and the specific area 212 of the output video stream 220 is thus updated as the FOV of the second camera 130 changes, so that the person 180 remains unoccluded in the output video stream 220. If the designed method is used to improve… Figure 1 The general-purpose camera monitoring system 100 illustrated herein may be configured, for example, to perform the improved method 200 and, for example, to replace the first output video stream 152 with the output video stream 220, thereby displaying the output video stream 220 instead of the first output video stream 152 on the first monitor 161.

[0047] exist Figure 2AIn the example illustrated in FIG. 2, the particular region 212 of the output video stream 220 in which occlusions are suppressed corresponds to the entirety of the portion of the scene captured / depicted in the second input video stream 142. In other words, the portion of the scene depicted in the second input video stream 142 is defined as the entire scene depicted in the predefined region of the second input video stream 142, such as, for example, the entire region of the second input video stream 142. In other embodiments, of course, the particular region 212 of the output video stream 220 in which occlusions are suppressed can be made to correspond to only a portion of the scene captured / depicted in the second input video stream 142. For example, the suppression can be designed to occlude only 90% of the portion of the scene captured / depicted in the second input video stream 142, only 80% of the portion, only 75% of the portion, etc. In such cases, the predefined region of the second input video stream 142 can be smaller than the entire region of the second input video stream 142. In any embodiment, it is important that occlusions be suppressed in a sufficient region of the output video stream 220 to allow objects and / or regions of the scene of interest to be visible in the output video stream 220.

[0048] Figure 2C Another embodiment of the designed method is schematically illustrated, showing only the resulting output video stream 220. In the output video stream 220, suppression of occlusions is instead performed in the entire scene captured by one of the first input cameras 120, i.e., in a region of the output video stream 220 that potentially corresponds to a larger portion of the scene than captured by the second camera 130. Here, the particular region 212 of the output video stream in which occlusions are suppressed is a region that depicts the entire scene captured by one of the cameras 120, even though the portion of the scene depicted in the second input video stream 142 corresponds only to the region marked by the dashed rectangle 210. In other words, in this embodiment, the first input video stream 140 is a composite video stream from multiple first cameras 120, e.g., cameras 120a-c, each capturing a portion of the scene, and the particular region 212 of the output video stream 220 depicts the entire scene captured by at least one of the multiple first cameras 120. In other cases, where, for example, the dashed rectangle 210 would span input video streams from more than one first camera 120, the suppression of occlusions can be designed, for example, in all portions of the scene captured by the corresponding first cameras 120.

[0049] In contrast to that illustrated in FIG. 2, and as already referenced with respect to, for example, FIG. 1, the particular region 212 of the output video stream 220 in which occlusions are suppressed can be made to correspond to only a portion of the scene captured / depicted in the second input video stream 142. For example, the suppression can be designed to occlude only 90% of the portion of the scene captured / depicted in the second input video stream 142, only 80% of the portion, only 75% of the portion, etc. In such cases, the predefined region of the second input video stream 142 can be smaller than the entire region of the second input video stream 142. In any embodiment, it is important that occlusions be suppressed in a sufficient region of the output video stream 220 to allow objects and / or regions of the scene of interest to be visible in the output video stream 220. Figure 2C Figure 2A ​As explained, the present disclosure is also designed to output a specific region 212 of the output video stream 220 is a part of a larger extended region of the output video stream 220 in which it is determined that occlusion should be applied otherwise. In other words, after occlusion is suppressed in the specific region 212, there can be other regions in the output video stream 220 that still occlude, such that at least a part of the extended region of the output video stream 220 other than the specific region 212 is still occluded, for example. Therefore, the proposed and designed method 200 has the advantage that it allows for improved context awareness without performing de-occlusion of regions in the scene that are not relevant to the object of interest that is currently tracked, for example, i.e., such that as long as the location of the object in the scene does not coincide with the tracked object, their privacy can still be respected.

[0050] As used herein, the extended region of the output video stream 220 can be, for example, a predefined region of the first input video stream 140. The extended region can be statically defined, for example, using a video management system (VMS) or similar system, for example, by a user manually drawing / marking one or more polygons in which occlusion is to be applied. This can be useful, for example, if the first camera 120 is static and always captures the same scene portion. In other embodiments, alternatively, the extended region in which occlusion is to be performed (in addition to the specific region) is dynamic and can change over time, and various input video streams are analyzed based on object detection and / or tracking algorithms, for example, to detect / track objects that are classified as sensitive and for which occlusion should generally be applied. Therefore, if the second camera 130 is currently tracking the object of interest, the designed method is able to overrule such occlusion decisions, such that the object of interest remains unoccluded in the scene overview provided by the output image stream 220 (and displayed on the first monitor 162 of the user terminal 160, for example).

[0051] In another embodiment of the method 200, instead of completely suppressing the occlusion in the certain area 212, the occlusion level is changed, i.e. reduced. In such an embodiment, the occlusion is typically performed by increasing the pixelization of the object. For example, by dividing the input images of the input image stream into blocks of MxM pixels (where M is an integer), each pixel in such a block can be replaced by the average pixel value in that block. By increasing the size of the blocks, i.e. by increasing M, the level of detail in the image can thus be reduced, since a pixel block of size M=1 would correspond to performing no occlusion / pixelization at all, while a pixel block of M equal to the full width and / or height of the image would render all the image with the same color, corresponding to the average color of all the original pixels. In the method designed herein, it can be assumed that the occlusion is performed by performing such MxM block averaging at least in the areas of the output video stream 220 where there are sensitive areas or objects. For example, the occlusion can be performed with a pixelization block size of 8x8. Thus, in the certain area 212 of the output video stream 220, the occlusion level can be obtained by reducing the block size used for the pixelization to e.g. 4x4 or 2x2 (or even to 1x1, which would equal to completely removing the occlusion). This can be especially useful if, for example, the resolution of the first camera 120 is larger than the resolution of the second camera 130, since the occlusion level of the output video stream 220 is then adapted so that it matches the resolution level of the second camera 130 (on the first monitor 161). By doing so, the user of the terminal 160 will see the object of interest 180 on the first monitor 161 and on the second monitor 162 with the same resolution. As an example, using Figure 1The camera surveillance system 100 illustrated in the figure can assume that each of the four first cameras 120 has a 4K resolution (e.g. 3840 x 2160), while the second camera 130 has a 1080p resolution (e.g. 1920 x 1080). The first cameras 120 can for example be arranged to each cover a horizontal FOV of 90 degrees (thus together generating a 360 degree view of the scene, without taking any overlap between the first input video streams 140 into account). In other embodiments, there can only be a single first camera 120 of 4K resolution, arranged to cover a horizontal FOV of 90 degrees. The second camera 130 can be such that it covers a horizontal FOV of 90 degrees when fully zoomed out. As a result, the pixel density of the second camera 130 will correspond to a 2 x 2 pixelized block size in the output video stream 220. If the second camera 130 is zoomed in, the corresponding pixelized block size in the output video stream 220 will instead be 1 x 1, e.g. no pixelization / occlusion. In case the second camera 130 is fully zoomed out and occlusion is typically performed by averaging over e.g. 4 x 4, 8 x 8, 16 x 16 or similar pixel block sizes, the level of occlusion in a particular area 212 of the output video stream 220 can thus be adjusted by reducing the block size (e.g. to 2 x 2) to match the block size of the second camera 130. Thus, the pixel density of the person 180 will be the same when viewed on both the first monitor 161 and the second monitor 162 of the terminal 160.

[0052] In this document, the method 200 is designed to be performed by e.g. a camera surveillance system (e.g. the camera surveillance system 100 illustrated in Figure 1 the figure, but with the device 150 reconfigured / replaced to perform the method 200) (or in a camera surveillance system), and / or by a device which will now be described in more detail with reference to Figure 3 The present disclosure provides such a camera surveillance system.

[0053] Figure 3A schematic diagram illustrates an apparatus 300 for generating an output video stream according to one or more embodiments of the present disclosure. The apparatus 300 includes at least a processor (or “processing circuitry”) 310 and a memory 312. As used herein, a “processor” or “processing circuitry” can be, for example, any combination of one or more of a suitable central processing unit (CPU), multiprocessor, microcontroller (μC), digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), etc., capable of executing software instructions stored in the memory 312. The memory 312 can be external to the processor 310 or internal to the processor 310. As used herein, “memory” can be any combination of random-access memory (RAM) and read-only memory (ROM), or any other type of memory capable of storing instructions. Memory 312 contains (i.e., stores) instructions that, when executed by processor 310, cause device 300 to perform the methods described herein (i.e., method 200 or any embodiment thereof). Device 300 may further include one or more additional items 314, which in some cases may be necessary for performing the method. In some example embodiments, device 300 may be, for example, as shown below. Figure 1The camera arrangement as illustrated in Fig. 1 1 is convenient and practical in that all cameras are provided as a single unit, and the FOVs of the cameras can be more easily compared, as the relative positions of the cameras are already known from the factory. It is also designed that the camera arrangement can comprise the circuitry required to perform the method 200, e.g. the device 300 as described above, and the camera arrangement can output the output video stream 220, e.g. and can also output the second output video stream 152, so that the camera arrangement can be directly connected to a user terminal without the need for any intermediate device / processing system. In other embodiments, the device for performing the method 200 as designed herein can be a separate entity that can be provided between the camera arrangement and the user terminal. The present disclosure also provides for a camera arrangement comprising the functionality for performing the method 200 as designed.

[0054] As Figure 1 The camera arrangement as illustrated in Fig. 1 1 is convenient and practical in that all cameras are provided as a single unit, and the FOVs of the cameras can be more easily compared, as the relative positions of the cameras are already known from the factory. It is also designed that the camera arrangement can comprise the circuitry required to perform the method 200, e.g. the device 300 as described above, and the camera arrangement can output the output video stream 220, e.g. and can also output the second output video stream 152, so that the camera arrangement can be directly connected to a user terminal without the need for any intermediate device / processing system. In other embodiments, the device for performing the method 200 as designed herein can be a separate entity that can be provided between the camera arrangement and the user terminal. The present disclosure also provides for a camera arrangement comprising the functionality for performing the method 200 as designed.

[0055] It is also designed herein that, for example, the at least one first camera 120 can be provided as its own entity / unit, and the second camera 130 can also be provided as its own different entity / unit. For example, a fisheye camera can be used to provide an overview of a scene, and be mounted in, for example, a ceiling of a building or similar, while a PTZ camera can be mounted elsewhere (e.g. on a wall), and can be used to provide a more detailed view of the scene. In such an arrangement, the device as designed herein can for example be provided as part of the one or more first cameras for generating an overview of the scene, or as part of the second camera for generating a more detailed view of the scene. In other designed embodiments, the device is instead provided as a separate entity or as part of a user terminal, and is configured to receive respective input video streams from the respective cameras and generate an output video stream.

[0056] The present disclosure is also designed to provide a computer program and a corresponding computer program product, as previously described herein.

[0057] In overview of the embodiments presented herein, the present disclosure provides an improved way of generating an output video stream containing an overview of a scene in a system in which there is also an additional second camera for generating a more detailed view of a part of the scene. By suppressing occlusions in the output video stream containing the overview of the scene based on a part of the scene currently captured by the second camera, or by at least reducing the level of occlusions in the output video stream containing the overview of the scene, the use of the output video stream providing an overview of the scene for context awareness is improved, and the output video stream can thus be used, for example, to better keep track of the position or other whereabouts of an object or region of interest captured by the second camera when, for example, the object moves in the scene.

[0058] Although the above can describe features and elements in particular combinations, each feature or element can be used alone without the other features and elements or in various combinations with or without other features and elements. Accordingly, the present disclosure is not limited to the specific combinations of features and elements as described, but is intended to encompass various combinations of features and elements in various combinations and sub-combinations, as specifically set forth herein or as specifically realized by persons skilled in the art, in practice of the claimed application.

[0059] In the claims, the words "comprising" and "containing" do not exclude other elements and the wording "a" or "an" does not exclude a plurality. The mere fact that certain features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

[0060] List of reference signs

[0061] 100 camera surveillance system

[0062] 110 camera device

[0063] 120a-c first camera

[0064] 130 second camera

[0065] 140a-d first input video stream

[0066] 142 second input video stream

[0067] 150 apparatus

[0068] 152 first output video stream

[0069] 154 second output video stream

[0070] 160 user terminal

[0071] 161 first monitor

[0072] 162 second monitor

[0073] 170a-c (privacy) obscuring applied in the first output video stream

[0074] 180 object of interest

[0075] 200 improved method

[0076] S201-S203 steps of the improved method

[0077] 210 scene portion captured by the second camera

[0078] 212 specific area of the output video stream

[0079] 220 output video stream

[0080] 230a-c (privacy) obscuring applied in the improved output video stream

[0081] 300 improved apparatus

[0082] 310 processor

[0083] 312 memory

[0084] 314 optional add-on

[0085] 316 optional network interface

[0086] 318 communication bus

Claims

1. A method of generating an output video stream, comprising: receiving a first input video stream from at least one first camera capturing a scene; receiving a second input video stream from a second camera capturing only a part of the scene; and generating an output video stream from the first input video stream, characterized in that: in response to determining that a particular region of the output video stream depicts a part of the scene depicted in the second input video stream, occlusion to the particular region of the output video stream is suppressed or at least the level of occlusion to the particular region of the output video stream is reduced, and occlusion is performed as usual in one or more other regions of the output video stream not corresponding to the part of the scene depicted in the second input video stream, and wherein the part of the scene depicted in the second input video stream is a predefined region of the second input video stream and / or a part of the scene comprising at least one object of interest. Determining that the particular region of the output video stream depicts the part of the scene depicted in the second input video stream comprises:

2. The method of claim 1, wherein, comparing a field of view of the at least one first camera with a field of view of the second camera.

3. The method of claim 1, comprising: generating the output video stream from a plurality of first input video streams from a plurality of first cameras, each of the first cameras capturing a part of the scene, and wherein the particular region of the output video stream depicts all of the scene captured by at least one of the plurality of first cameras. Occlusion comprises:

4. The method of claim 1, wherein, pixellating the particular region of the output video stream, and wherein reducing the level of occlusion comprises reducing a block size of the pixellation for the particular region of the output video stream. Determining that an extended region of the output video stream comprising the particular region is a region to be occluded based on the first input video stream comprises occluding at least a part of the extended region in the output video stream other than the particular region.

5. The method of claim 1, comprising:

6. The method of claim 5, wherein the extended region of the output video stream is a predefined region of the first input video stream.

7. The method of claim 1, wherein a predefined region of the second input video stream is an entire region of the second input video stream.

8. The method of claim 1, wherein the method is performed in a camera device comprising the at least one first camera and the second camera.

9. An apparatus for generating an output video stream, comprising: a processor, and a memory storing instructions that, when executed by the processor, cause the apparatus to: receive a first input video stream from at least one first camera capturing a scene; receive a second input video stream from a second camera capturing only a part of the scene; and generate an output video stream from the first input video stream, characterized in that: ​ inhibiting or at least reducing a level of occlusion of a particular region of the output video stream in response to determining that the particular region of the output video stream depicts a part of the scene depicted in the second input video stream, and performing occlusion as usual in one or more other regions of the output video stream that do not correspond to the part of the scene depicted in the second input video stream, and wherein the part of the scene depicted in the second input video stream is a predefined region of the second input video stream and / or a part of the scene comprising at least one object of interest.

10. The device of claim 9, wherein the device is or comprises the at least one first camera and one of the second camera.

11. A camera surveillance system comprising: at least one first camera, a second camera, and the device of claim 9.

12. A non-transitory computer-readable storage medium storing a computer program for generating an output video stream, the computer program configured to, when executed by a processor of a device, cause the device to: receive a first input video stream from at least one first camera capturing a scene; receive a second input video stream from a second camera capturing only a part of the scene; and generate an output video stream from the first input video stream, characterized in that: inhibiting or at least reducing a level of occlusion of a particular region of the output video stream in response to determining that the particular region of the output video stream depicts a part of the scene depicted in the second input video stream, and performing occlusion as usual in one or more other regions of the output video stream that do not correspond to the part of the scene depicted in the second input video stream, and wherein the part of the scene depicted in the second input video stream is a predefined region of the second input video stream and / or a part of the scene comprising at least one object of interest.

13. A computer program product comprising a computer program configured to, when executed by a processor of a device, cause the device to perform the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Automatically expanding the zoom capability of a wide-angle video camera

    US20060056056A1

  • Alarm dependent video surveillance

    US20200126383A1