PTZ Masking Control
By dynamically adjusting privacy masking in camera surveillance systems based on the fields of view of multiple cameras, the method improves situation awareness by ensuring tracked objects remain visible in the overview video stream.
Patent Information
- Application Number
- JP2023069712
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-26
- Filing Date
- 2023-04-21
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2043-04-21
AI Technical Summary
In camera surveillance systems, privacy masking in overview video streams can hinder situation awareness when a PTZ camera is tracking an event, as it becomes difficult to accurately locate objects in the scene without unmasking relevant areas.
A method is introduced to dynamically adjust privacy masking in the overview video stream by comparing the fields of view of multiple cameras. Specifically, when a second camera captures a portion of the scene, the masking in the corresponding regions of the overview video stream is reduced or aborted, allowing objects tracked by the PTZ camera to be visible.
This approach enhances situation awareness by ensuring that objects being tracked by a PTZ camera remain visible in the overview video stream, improving the user's ability to understand the object's location within the scene.
Smart Images

Figure 0007699167000001 
Figure 0007699167000002 
Figure 0007699167000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to (privacy) masking in video streams generated in a camera surveillance system. In particular, the present disclosure relates to such masking when receiving multiple video streams from multiple cameras that capture the same scene.
Background Art
[0002] In modern camera surveillance systems, multiple cameras can be used to capture the same scene. For example, one or more cameras can be used to generate an overview of the scene, and one or more other cameras can be used to generate a more detailed view of the scene. The camera used to generate the overview of the scene may be, for example, a camera with a larger field of view (FOV) so that it can capture a larger portion of the scene. As another example, video streams from multiple cameras, each having a smaller FOV and each capturing a different portion of the scene, can be combined to form a combined video stream that still captures a larger portion of the scene. To generate a more detailed view of the scene, preferably, a so-called pan-tilt-zoom (PTZ) camera can be used, which can, for example, zoom in on an object or area of interest within the scene and provide more details about such object or area than may be available from the overview of the scene alone. The user of the camera surveillance system can then, for example, use the overview of the scene to identify the presence of an object or area of interest and then use the PTZ camera to investigate the object or area in more detail (e.g., by zooming in closer to the object of interest and / or by using the pan and / or tilt functions of the PTZ camera to follow the object of interest as it moves across the scene).
[0003] In such a camera monitoring system, it is not uncommon for different users to have different levels of authority, and to view a more detailed view of a scene, a higher level of authority may be required than that required to view only an overview of the scene. This is because the overview of the scene does not provide more details about individual objects (such as people) than a PTZ camera. In particular, in order to further reduce the level of authority required to view the overview of the scene, for example, (privacy) masking can be applied to the overview of the scene such that information such as a person's identity or license plate number is hidden and cannot be restored by a user with a low level of authority viewing the overview of the scene. On the other hand, in the case of a PTZ camera, when a target object is found, all the capabilities of the PTZ camera can be utilized, and thus, usually, a higher level of authority is required to view a more detailed view of the scene generated by the PTZ camera, and such masking is not usually performed.
[0004] When a PTZ camera is used, for example, to follow an event that is currently occurring (e.g., to track the movement of a target object), it may be desirable to also use the overview of the scene to improve situation awareness. For example, when zooming in on a target object using a PTZ camera, if only a more detailed view of the scene is used, it can be difficult to accurately recognize where in the scene the object is currently located. By also using the overview of the scene, such difficulties in locating objects in the scene can be reduced. However, when privacy masking is performed in the overview of the scene, it can be difficult to utilize all the capabilities of the overview of the scene for situation awareness. SUMMARY OF THE INVENTION
[0005] The (privacy) masking of objects and / or regions within an overview of a scene is used to reduce the required level of permissions needed to view the overview of the scene, but in cases where it is also desirable to use the overview of the scene to improve situation awareness, the present disclosure provides an improved method of generating an output video stream, as well as a corresponding device, camera monitoring system, computer program, and computer program product as defined in the appended independent claims. The various embodiments of the improved method, device, camera monitoring system, as well as the computer program and computer program product are defined in the appended dependent claims.
[0006] According to a first aspect, a method of generating an output video stream is provided. The method includes receiving a first input video stream from at least one first camera that captures a scene. The method further includes receiving a second input video stream from a second camera that captures only a portion of the scene. The method also includes generating an output video stream from the first input video stream, including aborting or at least reducing the masking of a particular region of the output video stream in response to determining that the particular region of the output video stream depicts a portion of the scene that is depicted in the second input video stream. In some embodiments, it is contemplated that masking can be aborted (or reduced) in regions of the output video stream corresponding to portions of the scene depicted in the second input video stream, and normal (i.e., not aborted or reduced) masking can be performed in one or more other regions of the output video stream that do not correspond to portions of the scene depicted in the second input video stream. It should be noted that "aborting the masking" or "reducing the level of masking" implies that such regions would actually be masked in the output video stream if the assumed method of taking into account the content of the second input video stream were not used.
[0007] The proposed method improves on the currently available techniques and solutions in that it takes into account what is currently seen in the second input video stream from the second camera and then prevents (or reduces the level of) masking of the corresponding areas in the overview of the scene. By doing so, for example, an object that is currently being tracked in a more detailed view of the scene becomes visible also in the overview of the scene, and in particular, such "unmasking" in the overview of the scene is done dynamically and adapts to changes in the FOV of the second camera, thus improving the situation awareness.
[0008] In some embodiments of the method, determining that a particular region of the output video stream depicts a portion of the scene depicted in the second input video stream may include comparing the field of view (FOV) of at least one first camera with the FOV of the second camera. As used herein, the FOV of a camera may include, for example, the position of the camera relative to the scene or at least another camera, the zoom level (optical or digital) of the camera, or the tilt / pan / rotation of the camera. Thus, by comparing the FOVs of the two cameras, it is possible to determine whether the two cameras are capturing the same portion of the scene, and in particular, which regions of the image captured by one camera correspond to the portion of the scene captured by the other camera. In other embodiments, other techniques may be used to achieve the same effect, and it is envisioned that other techniques may be based on image analysis and feature identification, for example, instead of the FOV parameters of the camera.
[0009] In some embodiments of the method, the method may include generating an output video stream from a plurality of first input video streams from a plurality of first cameras, each capturing a portion of the scene. A particular region of the output video stream may depict all of the scene captured by at least one of the plurality of first cameras (such as being depicted in at least one of the plurality of first input video streams). For example, it is envisioned that two or more cameras may be arranged (optionally with some overlap, but not necessarily) such that they are directed in different directions and thus capture different portions of the scene. The input video streams from these cameras may then be combined into a single composite video stream upon which the video output stream is based. The video streams from each camera may be stitched together, for example, to create the illusion of receiving only a single input video stream from a very large FOV camera. In other embodiments, the input video streams from each first camera may instead be shown adjacent to each other in the output video stream, for example, in a horizontal array, a vertical array, and / or a grid, etc.
[0010] In some embodiments of the method, masking may include pixelating a particular region of the output video stream. Reducing the level of masking may include reducing the block size used for pixelating a particular region of the output video stream. This may be advantageous in that, for example, the resolution of the object of interest may be the same in both the output video stream and a second output video stream representing a second input video stream.
[0011] In some embodiments of the present method, the method may further include determining, based on a first input video stream, that an extended area of an output video stream including a specific area is an area to be masked. The method may include masking at least a portion of the extended area outside the specific area in the output video stream. In other words, it is assumed that a part of a larger mask that would otherwise be desired is still retained, and only a specific area of such a larger mask corresponding to a part of the scene captured by the second camera can be "unmasked".
[0012] In some embodiments of the present method, the extended area of the output video stream may be a predetermined area of the first input video stream, that is, in other words, the extended area of the output video stream corresponds to a predetermined part of the scene. For example, such a predetermined area / part of the scene can be manually set when configuring a camera monitoring system to statically mask one or more features or areas of the scene where people usually are or are known in advance to reside. Such features and / or areas can be, for example, windows of residential buildings, windows of coffee shops or restaurants, daycare centers, playgrounds, outside service areas, parking lots, bus or train stops, etc.
[0013] In some embodiments of the method, the portion of the scene depicted in the second input video stream can be defined as the portion of the scene that includes at least one object of interest. In other words, the otherwise desired masking of the first input video stream and the output video stream can be made more dynamic and less static, for example, it can follow when a person or other object moves within the scene, and the masking is only performed after it is further determined that the person and / or object should be masked (if they should not be stopped from being masked in the assumed way). For example, the masking of an object may only be desired if it is first determined that the object belongs to a class (such as a person, license plate, etc.) for which the object should be masked (for example, by using object detection, etc.).
[0014] In some embodiments of the method, the portion of the scene depicted in the second input video stream can be defined as all of the scene depicted within a predetermined region of the second input video stream, such as the entire region of the second input video stream. In other embodiments, the predetermined region of the second input video stream can correspond to a region smaller than the entire region of the second input video stream. For example, it is envisioned that the region of the second input video stream where, for example, an object of interest is located can be detected, and only the portion of the first input video stream in the output video stream corresponding to this region is unmasked (or rather, the masking is stopped), thus "unmasking" a region of the scene in the output video stream that is smaller than the region captured by the second camera.
[0015] In some embodiments of the present method, the method can be performed with a single camera configuration. This camera configuration can include at least one first camera and a second camera. In other embodiments, at least one first camera and a second camera are provided as separate cameras that do not form part of the same camera configuration, and the method is then assumed to be performed by any of the cameras, or jointly by, for example, all the cameras (assuming the cameras have some communication method with each other).
[0016] According to a second aspect of the present disclosure, a device for generating an output video stream is provided. The device includes a processor. The device further includes a memory for storing instructions, which, when executed by the processor, cause the device to receive a first input video stream from at least one first camera that captures a scene, receive a second input video stream from a second camera that captures only a portion of the scene, and generate an output video stream from the first input video stream, including aborting the masking of a specific region of the output video stream or at least reducing the level of masking in response to determining that a specific region of the output video stream depicts a portion of the scene that is projected onto the second input video stream.
[0017] Thus, the device according to the second aspect is configured to perform the corresponding steps of the method according to the first aspect.
[0018] In some embodiments of the device, the device is further configured to perform any of the embodiments of the method described herein (i.e., such that the instructions, when executed by the processor, cause it to do so).
[0019] In some embodiments of the device, the device is one of at least one first camera and a second camera. In other embodiments of the device, the device is a camera configuration including all of at least one first camera and a second camera.
[0020] According to a third aspect of the present disclosure, a camera monitoring system is provided. The camera monitoring system includes at least a first camera, a second camera, and a device as contemplated herein according to the second aspect (or any embodiment thereof).
[0021] According to a fourth aspect of the present disclosure, a computer program for generating an output video stream is provided. When executed by a processor of a device (e.g., a device according to the second aspect or any embodiment thereof), the computer program product causes the device to receive a first input video stream from at least one first camera that captures a scene, receive a second input video stream from a second camera that captures only a portion of the scene, and generate an output video stream from the first input video stream, wherein a specific region of the output video stream is masked or at least the level of masking is reduced in response to determining that the specific region of the output video stream depicts a portion of the scene that is projected onto the second input video stream. Thus, the computer program according to the fourth aspect is configured to cause the device to perform the method according to the first aspect.
[0022] Thus, the computer program according to the fourth aspect is configured to cause the device to perform the method according to the first aspect.
[0023] In some embodiments of the computer program, the computer program is further configured to cause the device to perform any embodiment of the method of the first aspect described herein.
[0024] According to a fifth aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer-readable storage medium storing a computer program according to the fourth aspect (or any embodiment thereof described herein). The computer-readable storage medium may be, for example, non-transitory and may be provided as, for example, a hard disk drive (HDD), a solid state drive (SSD), a USB flash drive, an SD card, a CD / DVD, and / or any other storage medium capable of non-transitory storage of data.
[0025] Other objects and advantages of the present disclosure will become apparent from the following detailed description, the drawings, and the claims. It is assumed that all features and advantages described with reference to, for example, the method of the first aspect are related to, applicable to, and may be used in combination with any features and advantages described with reference to the device of the second aspect, the surveillance camera system of the third aspect, the computer program of the fourth aspect, and / or the computer program product of the fifth aspect within the scope of the present disclosure, and vice versa.
[0026] Exemplary embodiments will be described below with reference to the accompanying drawings.
Brief Description of the Drawings
[0027]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 3
Modes for Carrying Out the Invention
[0028] In the drawings, unless otherwise specified, the same reference numerals are used for the same elements. Unless the contrary is explicitly stated, the drawings show only the elements necessary to illustrate the exemplary embodiments, but other elements may be omitted or merely proposed for clarity. As shown in the figures, the (absolute or relative) sizes of the elements and regions may be enlarged or reduced with respect to their true values for illustrative purposes, and are thus provided to show the general structure of the embodiments.
[0029] FIG. 1 schematically shows a general camera monitoring system 100 including a camera configuration 110, a device 150, and a user terminal 160. The camera configuration 110 includes four first cameras 120 (only three of which, cameras 120a, 120b, and 120c, are visible in FIG. 1) and a second camera 130. The first cameras 120 are overview cameras mounted in a circular pattern, each facing a different direction so as to together capture, for example, a 360-degree horizontal view of a scene. The second camera 130 is a pan-tilt-zoom (PTZ) camera and is mounted below the first cameras 120 at the center of the circular pattern (when viewed, for example, from above or below).
[0030] The first cameras 120 generate a first input video stream 140 (i.e., four first input video streams 140a - d if there are four first cameras 120), and the second camera generates a second input video stream 142. The input video streams 140, 142 are provided to the device 150. The device 150 is configured to receive the input video streams 140, 142 and generate corresponding output video streams 152, 154.
[0031] The output video streams 150, 152 are provided by the device 150 to the user terminal 160. The first output video stream 152 is displayed on the first monitor 161 of the terminal 160, and the second output video stream 154 is displayed on the second monitor 162 of the terminal 160. In a specific example of the system 100 shown in FIG. 1, four first input video streams 140a - d (one from each camera 120) are provided to the device 150, and the device 150 combines the first input video streams 140a - d into a single combined first output video stream 152. For example, the device 150 can arrange the input video streams 140a - d in a 2×2 grid such that each quadrant of the 2×2 grid displays one first input video stream from the corresponding first camera. Of course, other alternative configurations for combining such a plurality of first input video streams into a single combined first output video stream are also envisioned. For example, various first input video streams can be stitched together to create a panoramic first output video stream and the like.
[0032] In this specification, it can be assumed that the device 150 is a stand - alone device separate from the camera configuration 110 and the user terminal 160. In other embodiments, instead, the device 150 may be an integrated part of either the camera configuration 110 or the user terminal 160.
[0033] The user of the terminal 160 may use the first output video stream 152 displayed on the first monitor 161 for situation awareness, while the second output video stream 154 displayed on the second monitor 162 may be used to provide more details about a specific part of the scene. For example, the user may detect a target object or area by looking at the first monitor 161 and then adjust the field of view (FOV) of the second camera 130 so that the second camera 130 captures the target object or area in more detail. For example, the FOV of the second camera 130 can be adjusted by zooming, panning, rotating, and / or tilting the second camera 130, for example, so that the second output video stream 154 displayed on the second monitor 162 provides a zoomed-in view of the specific object or area of interest. To obtain situation awareness, for example, the first output video stream 152 displayed on the first monitor 161 can be used to still track where the zoomed-in object is currently located within the scene.
[0034] The adjustment of the FOV of the second camera 130 may also be automated such that when a target object or area is detected in one or more of the first input video streams 140 and / or the corresponding first output video streams 152 (e.g., by using an object detection and / or object tracking algorithm), the second input video stream 142 (and thus the corresponding second output video stream 154) provides an enlarged view of the detected target object or area when it is displayed on the second monitor 162. Instead of using a physical / optical zoom-in (i.e., by adjusting the lens of the second camera 130), it is also conceivable to use a digital zoom such that, for example, the second input video stream 142 is not changed.
[0035] For privacy reasons, it may be necessary (e.g., by local law) to mask one or more "sensitive" portions of a scene so that identities or other personal characteristics, such as objects or vehicles, that are present in sensitive portions of the scene cannot be inferred from viewing the first monitor 161. Such sensitive portions may correspond to locations where, for example, people are likely to be present in the scene, such as shopping windows, windows of residential buildings, bus or train stops, restaurants, or coffee shops. Thus, for this purpose, the camera system 100 and the device 150 are configured such that one or more privacy masks are applied when generating the first output video stream 152. In the particular example shown in FIG. 1, three such privacy masks 170a, 170b, and 170c are applied. The first privacy mask 170a is in front of the portion of the scene that includes the store, the second privacy mask 170b is in front of the portion of the scene that includes the bus stop, and the third privacy mask 170c is in front of the residential flat located on the ground / first floor of the building. In FIG. 1, these privacy masks are shown as hatches, but of course, instead, they may be completely opaque so that details regarding people, etc. behind the mask cannot be obtained by viewing the first monitor 161. In other embodiments, the privacy masks 170a-c may be created, for example, by pixelating the relevant area, or even further, may be created such that the object to be masked becomes transparent (this is possible, for example, if an image of the scene without the object previously is available). Other ways of masking objects are also envisioned.
[0036] In the specific example shown in FIG. 1, a particular object 180 of interest is present in the first input video stream 140a and constitutes a person who is attempting to enter the store currently. Since the person 180 appears to have broken the store's shopping window, the user of the terminal 160 is interested in tracking the progress of the intrusion so that the location of the person 180 can be followed in real time. For example, by adjusting the FOV of the second camera 130 accordingly, an enlarged view of the person 180 is provided in the second output video stream 154 displayed on the second monitor 162. No privacy mask is applied to the second output video stream 154.
[0037] As described above in this specification, the system 100 is often configured such that the level of permission required to view the first output video stream 152 is lower than the level of permission required to view the second output video stream 154. This is because the overview of the scene provided by the first output video stream 152 contains little detail about the individual objects within the scene, i.e., because the privacy masks 170a - c are applied before the user can view the first output video stream 152 on the first monitor 161. However, in a situation such as that shown in FIG. 1, it is desirable for the user of the terminal 160 to use the first output video stream 152 for situation awareness so that the user can more easily know where the person 180 is currently located within the scene. For example, if the person 180 starts to move, it may be difficult to know where the person 180 is currently located within the scene by just looking at the second terminal 162. Knowing the current location of the person 180 can be important, for example, when contacting law enforcement agencies, etc. However, due to the privacy masks 170a - c, the masks 170a - c hide the person 180 within the first output video stream 152, thus impeding such situation awareness and therefore making it more difficult to utilize the first output video stream 152 displayed on the first monitor 161 for such situation awareness.
[0038] How the methods envisioned in this specification can be useful in improving the above situation will be described in more detail below, also with reference to FIGS. 2A-2C.
[0039] FIG. 2A schematically shows an envisioned embodiment of a method 200 for generating an output video stream, and FIG. 2B schematically shows a flowchart of various steps included in such a method 200. In a first step S201, a first input video stream 140 is received from at least one first camera (such as one or more of cameras 120a-120c), and the first input video stream 140 captures / records a scene. In a second step (S202, which may be executed before step S201 or simultaneously with step S201), a second input video stream 142 is received from a second camera (such as camera 130). The second input video stream 142 captures / records only a portion of the scene. For example, the FOV of the second camera 130 is smaller than the FOV (or combined FOV) of the first camera (or first cameras) 120. Put another way, the second input video stream 142 provides a magnified view of an object or region of the scene. As previously described herein, such a result can be obtained, for example, by mechanically / optically zooming the second camera 130 (i.e., by moving the lens), or, for example, by digitally zooming. In the latter case, the second input video stream 142 generated by the second camera 130 is not necessarily modified to generate a second output video stream 154, but for the purpose of describing the envisioned improved method, the second input video stream 142 can thus be assumed to capture the scene only partially in any case.
[0040] Next, the assumed method 200 proceeds to generate an output video stream 220 from the first input video stream 140 in a third step S203, as described below. To generate the output video stream 220, first, it is determined whether a specific region 212 of the resulting output video stream 220 depicts a portion 210 of the scene that is depicted in the second input video stream 142. If affirmative, the generation of the output video stream 220 aborts masking of the specific region 212 of the output video stream or at least reduces its level.
[0041] In the example shown in FIG. 2A, the second input video stream 142 includes a magnified view of a person 180 and depicts a portion of the scene that is also found in the first input video stream 140, as indicated by the dashed rectangle 210. To find the position and size of the rectangle 210, for example, knowledge of the current FOV of the second camera 130 and knowledge of the current FOV of one or more of the first cameras 120 can be used and compared. In the assumed method 200, it can then be confirmed that a specific region 212 of the corresponding output video stream 220 depicts the same portion of the scene as the second input video stream 142, i.e., the portion as indicated by the dashed rectangle 210. As used herein, having knowledge of the FOV of a camera includes, for example, knowing the position of the camera with respect to the scene (and / or with respect to one or more other cameras), the zoom level of the camera (optical / mechanical and / or digital), and the orientation of the camera (e.g., pan, tilt, and / or roll of the camera). Thus, by comparing the FOVs of two cameras, for example, it can be calculated whether the first camera depicts the same portion of the scene as that depicted by the second camera, and in particular, which region of the video stream generated by the first camera corresponds to this same portion of the scene.
[0042] As already described with reference to FIG. 1, when applying one or more privacy masks 230a-230c to mask, for example, a store, a bus stop, and a residential flat, the envisioned method 200 does not apply masking to a specific region 212 of the output video stream 220. As a result, the situation recognition regarding the scene obtainable from the output video stream 220 is thus enhanced because there is no mask covering the person 180. If the person 180, for example, starts to move, the person 180 is not masked in the output video stream 142 as long as the person 180 remains within the FOV of the second camera 130 that generates the second input video stream 220. For example, it can be envisioned that the second camera 130 tracks the person 180 (either by manual control by the user or by using, for example, object detection and / or object tracking), and the specific region 212 of the output video stream 220 is updated when the FOV of the second camera 130 changes so that the person 180 remains unmasked in the output video stream 220. When using the method envisioned to improve the general camera monitoring system 100 shown in FIG. 1, the device 150 can be configured to, for example, perform the improved method 200 and replace the first output video stream 152 with the output video stream 220 such that the output video stream 220 is displayed on the first monitor 161 instead of the first output video stream 152.
[0043] In the example shown in FIG. 2A, a particular region 212 of the output video stream 220 where masking has been discontinued corresponds to the entire portion of the scene captured / written to the second input video stream 142. Put another way, the portion of the scene written to the second input video stream 142 is defined as all of the scene written to a predetermined region of the second input video stream 142, such as the entire region of the second input video stream 142. In other embodiments, of course, it is also possible for a particular region 212 of the output video stream 220 where masking has been discontinued to correspond only to the portion of the scene captured / written to the second input video stream 142. For example, it can be envisioned that masking is discontinued for only 90%, only 80%, only 75%, etc. of the portion of the scene captured / written to the second input video stream 142. In such a situation, the predetermined region of the second input video stream 142 can be smaller than the entire region of the second input video stream 142. In any embodiment, what is important is that masking is discontinued in a sufficient region of the output video stream 220 to enable the objects and / or regions of the scene of interest to be visible in the output video stream 220.
[0044] Figure 2C schematically shows another embodiment of an envisioned method that shows only the resulting output video stream 220. In the output video stream 220, instead, masking is discontinued in all of the scenes captured by one of the first input cameras 120 (i.e., in the region of the output video stream 220 that may correspond to a larger portion of the scene than that captured by the second camera 130). In this case, the portion of the scene depicted in the second input video stream 142 corresponds only to the region marked by the dashed rectangle 210, but the specific region 212 of the output video stream where masking is discontinued is the region that depicts all of the scenes captured by one of the cameras 120. In other words, in this embodiment, the first input video stream 140 is a composite video stream from a plurality of first cameras 120 (such as cameras 120a - c) each capturing a portion of the scene, and the specific region 212 of the output video stream 220 depicts all of the scenes captured by at least one of the plurality of first cameras 120. For example, in other situations where the dashed rectangle 210 spans the input video streams from the plurality of first cameras 120, it may be envisioned to discontinue masking in all of the portions of the scenes captured by the corresponding first camera 120.
[0045] In contrast to what is shown in FIG. 2C, as already described with reference to FIG. 2A for example, the present disclosure also assumes that a particular region 212 of the output video stream 220 is part of a larger extended region of the output video stream 220 that would otherwise be determined to have masking applied. Said another way, after suspending masking within the particular region 212, there may be other regions of the output video stream 220 where masking is still being performed such that at least a portion of the extended region outside the particular region 212 of the output video stream 220 remains masked. Thus, an advantage of the proposed assumed method 200 is that this method enables improving situation awareness without unmasking regions of the scene that are not related to the object of the currently tracked target, i.e., the privacy of the object can still be respected as long as the position of the object within the scene does not match the tracked object.
[0046] As used herein, the extended region of the output video stream 220 can be, for example, a predetermined region of the first input video stream 140. For example, the extended region can be statically defined by a user manually drawing / marking one or more polygons for which masking should be applied, for example, using a video management system (VMS). This can be useful, for example, when the first camera 120 is stationary and always captures the same part of the scene. In other embodiments, instead, the extended region (excluding a specific region) for which masking should be performed can be dynamic and can change over time, for example, by analyzing various input video streams for detecting / tracking objects classified as sensitive and for which masking would normally be applied, based on an object detection algorithm and / or an object tracking algorithm. Thus, the envisioned method allows such masking decisions to be disabled so that, when the second camera 130 is currently following a target object, the target object remains unmasked in the overview of the scene provided by the output image stream 220 (and, for example, when displayed on the first monitor 162 of the user terminal 160).
[0047] In another embodiment of method 200, instead of completely stopping masking in a particular region 212, it is envisioned that the level of masking is instead changed (i.e., reduced). In such an embodiment, masking is typically done by increasing the level of pixelation of the object. For example, by dividing the input image of the input image stream into blocks of MxM pixels (where M is an integer), each pixel within such a block can be replaced with the average pixel value within the block. A pixel block of size M = 1 corresponds to not performing any masking / pixelation, and when M is equal to the full width and / or full height of the image, the pixel block is rendered such that all of the image has the same color (corresponding to the average color of all the original pixels), so by increasing the size of the block (i.e., by increasing M), the level of detail of the image can be reduced. In the method envisioned herein, it can be assumed that masking is done by performing such MxM block averaging in the region of the output video stream 220 where at least sensitive regions or objects are located. For example, masking can be done with a pixelation block size of 8x8. Thus, in a particular region 212 of the output video stream 220, the level of masking can be obtained by reducing the block size used for pixelation to, for example, 4×4 or 2×2 (or even 1×1, which is equivalent to completely removing the masking). This is particularly useful when the resolution of one or more first cameras 120 is greater than the resolution of the second camera 130, since, for example, the level of masking of the output video stream 220 is adapted to match the resolution level of the second camera 130 (on the first monitor 161). By doing so, the user of the terminal 160 will view the object 180 with the same resolution on both the first monitor 161 and the second monitor 162.As an example, it can be assumed that in the camera monitoring system 100 shown in FIG. 1, each of the four first cameras 120 has a 4K resolution (e.g., 3840×2160), and the second camera 130 has a 1080p resolution (e.g., 1920×1080). The first cameras 120 can be arranged, for example, to each cover a 90-degree horizontal FOV (thus, together generating a 360-degree view of the scene without considering any overlap between the first input video streams 140). In other embodiments, there may be only a single 4K resolution first camera 120 arranged to cover a 90-degree horizontal FOV. The second camera 130 may be configured to cover a 90-degree horizontal FOV when fully zoomed out. As a result, the pixel density of the second camera 130 corresponds to a 2×2 pixelated block size in the output video stream 220. When the second camera 130 is zoomed in, instead, the corresponding pixelated block size in the output video stream 220 becomes 1×1, and for example, pixelation / masking is not performed. Thus, in a situation where the second camera 130 is fully zoomed out and masking is typically performed by averaging over pixel block sizes such as 4×4, 8×8, or 16×16, the masking level in a particular region 212 of the output video stream 220 can be adjusted by reducing the block size, for example, to 2×2 to match the block size of the second camera 130. Thus, the pixel density of the person 180 is the same when viewed on both the first monitor 161 and the second monitor 162 of the terminal 160.
[0048] Herein, it is assumed that the method 200 can be performed by (or within) a camera monitoring system (e.g., the camera monitoring system 100 shown in FIG. 1, although the device 150 has been reconfigured / replaced to perform the method 200), and / or by a device as will be described in more detail below with reference to FIG. 3 as well. Such a camera monitoring system is provided by the present disclosure.
[0049] Figure 3 schematically shows a device 300 for generating an output video stream according to one or more embodiments of the present disclosure. The device 300 includes at least a processor (or "processing circuit") 310 and a memory 312. As used herein, a "processor" or "processing circuit" can be, for example, any suitable central processing unit (CPU), multiprocessor, microcontroller (μC), digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA), graphics processing unit (GPU), etc. that can execute software instructions stored in the memory 312, or any combination of one or more of them. The memory 312 can be external to the processor 310 or internal to the processor 310. As used herein, "memory" can be any combination of random access memory (RAM) and read only memory (ROM), or any other type of memory capable of storing instructions. The memory 312 includes (i.e., stores) instructions that, when executed by the processor 310, cause the device 300 to perform the methods described herein (i.e., method 200 or any embodiment thereof). The device 300 can further include one or more additional items 314 that may be necessary to perform the method in some situations. In some exemplary embodiments, the device 300 can be, for example, a camera configuration as shown in FIG. 1, and the one or more additional items 314 can, in that case, include, for example, one or more first cameras 120 and a second camera 130. The one or more additional items 314 can also include, for example, various other electronic device components necessary to generate the output video stream. Performing the method within the camera assembly can be useful, for example, in that the processing is moved closer to the location where the actual scene is captured compared to when the processing is performed elsewhere (such as in a more centralized processing server, etc.). The device 300 can be connected to a network, for example, so that the generated output video stream can be transmitted to a user (terminal).For this purpose, device 300 may include a network interface 316, which may be, for example, a wireless network interface (such as one that supports Wi-Fi, as defined by either IEEE 802.11 or a subsequent standard), or a wired network interface (such as one that supports Ethernet, as defined by either IEEE 802.3 or a subsequent standard). Network interface 316 can also support any other wireless standard capable of transferring encoded video, such as Bluetooth. The various components 310, 312, 314, and 316 (if present) can be connected via one or more communication buses 318 such that these components can communicate with each other and exchange data as needed.
[0050] The camera configuration as shown in FIG. 1 is advantageous because all cameras are provided as a single unit, and because the relative positions of the cameras are already known from the factory, it may be easier to compare the FOVs of the various cameras. Also, it is envisioned that the camera configuration can include the circuitry required to perform method 200, such as device 300 as described above, and that the camera configuration can output, for example, output video stream 220, and optionally also a second output video stream 152, directly to a user terminal without the need for any intermediate device / processing system. In other embodiments, the device for performing method 200 as contemplated herein may be a separate entity that can be provided between the camera configuration and the user terminal. Such camera configurations that include the functionality to perform the contemplated method 200 are also provided by the present disclosure.
[0051] For example, at least one first camera 120 may be provided as its own entity / unit, and it is also contemplated herein that a second camera 130 may also be provided as a different, its own entity / unit. For example, a fisheye camera can be used to provide an overview of a scene, and for example, the fisheye camera can be mounted on the ceiling of a building or the like, while a PTZ camera can be mounted at another location (e.g., on a wall) and used to provide a more detailed view of the scene. In such a configuration, the devices contemplated herein can be provided, for example, as part of one or more first cameras used to generate an overview of a scene or as part of a second camera used to generate a more detailed view of the scene. In other contemplated embodiments, the device is instead provided as a separate entity or as part of a user terminal and is configured to receive respective input video streams from various cameras and generate one or more output video streams.
[0052] The present disclosure also contemplates providing a computer program and a corresponding computer program product as described above herein.
[0053] Summarizing the various embodiments presented herein, the present disclosure provides an improved way of generating an output video stream including an overview of a scene in a system where there is an additional second camera provided to generate a more detailed view of a portion of the scene. Based on which portion of the scene is currently being captured by the second camera, by stopping masking or at least reducing the level of masking in the output video stream including the overview of the scene, the use of the output video stream to provide an overview of the scene for situation awareness is improved, and thus, using the output video stream, for example, when an object moves within the scene, the position or other whereabouts of the object or region of interest captured by the second camera can be better known.
[0054] The features and elements may be described above in certain combinations, but each feature or element may be used alone, without other features and elements, or in various combinations with or without other features and elements. Further, variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.
[0055] In the claims, the terms "comprising" and "including" do not exclude other elements, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that certain features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be advantageously used.
Explanation of Reference Numerals
[0056] 100 Camera monitoring system 110 Camera configuration 120(a-c) First camera 130 Second camera 140(a-d) First input video stream 142 Second input video stream 150 Device 152 First output video stream 154 Second output video stream 160 User terminal 161 First monitor 162 Second monitor 170a-c (Privacy) mask applied to the first output video stream 180 Object of interest 200 Improved method S201-203 Steps of the improved method 210 Portion of the scene captured by the second camera 212 Specific region of the output video stream 220 Output video stream (Privacy) mask applied to the improved output video stream of 230a-c Improved device 300 Processor 310 Memory 312 One or more optional additional items 314 Optional network interface 316 One or more communication buses 318
Claims
Claim 1 A method (200) for generating an output video stream, comprising: Receiving (S201) a first input video stream (140) from at least one first camera (120a-120c) that captures a scene; Receiving (S202) a second input video stream (142) from a second camera (130) that captures only a portion of the scene; and Generating (S203) an output video stream (220) from the first input video stream, wherein at least one portion of the scene represented in the output video stream is a portion to be masked, and generating the output video stream includes, as part of a process of masking the at least one portion of the scene in the output video stream, in response to determining that a specific region (212) of the at least one portion of the scene corresponds to a portion (210) of the scene represented in the second input video stream, aborting the masking (230a-230c) of the specific region or at least reducing the level of the masking (230a-230c) of the specific region. A method comprising the above. Claim 2 The method according to claim 1, wherein determining that the specific region of the output video stream corresponds to the portion of the scene represented in the second input video stream includes at least comparing the positions, zoom levels, and orientations of the first camera and the second camera with respect to the scene. Claim 3 The method according to claim 1, further comprising generating the output video stream from a plurality of first input video streams (140a-140d) from a plurality of first cameras (120a-120c) each capturing a portion of the scene, wherein all of the portions of the scene captured by at least one of the plurality of first cameras are shown in the specific region of the output video stream. Claim 4 The method according to claim 1, wherein masking includes pixelating the specific region of the output video stream, and reducing the level of the masking includes reducing the block size used for pixelating the specific region of the output video stream. Claim 5 The method according to claim 1, comprising determining at least one part of the scene based on the first input video stream.
6. The method according to claim 1, wherein at least one part of the scene corresponds to a predetermined region of the first input video stream.
7. The method according to claim 1, wherein the part of the scene depicted in the second input video stream is defined as the entire region of the second input video stream.
8. The method according to claim 1, which is executed in a camera configuration including the at least one first camera and the second camera.
9. A device for generating an output video stream, a processor, a memory for storing instructions and, when the instructions are executed by the processor, causing the device to receive a first input video stream from at least one first camera that captures a scene, receive a second input video stream from a second camera that captures only a part of the scene, and generate an output video stream from the first input video stream, wherein at least one part of the scene appearing in the output video stream is a part to be masked, and generating the output video stream includes, as part of a process of masking at least one part of the scene in the output video stream, aborting the masking of a specific region or at least reducing the level of masking of the specific region in response to determining that the specific region of at least one part of the scene corresponds to the part of the scene appearing in the second input video stream. A device that causes the above to be performed.
10. The device according to claim 9, wherein the instructions further cause the device to perform the method according to any one of claims 2 to 8 when executed by the processor.
11. The device according to claim 9, wherein the device is one of the at least one first camera and the second camera, or includes one of the at least one first camera and the second camera.
12. A camera monitoring system comprising at least one first camera, a second camera, and the device according to claim 9.
13. A computer program for generating an output video stream, which, when executed by a processor of a device, causes the device to receive a first input video stream from at least one first camera that captures a scene, receive a second input video stream from a second camera that captures only a portion of the scene, and generate an output video stream from the first input video stream, wherein at least one portion of the scene represented in the output video stream is a portion to be masked, and generating the output video stream includes, as part of a process of masking the at least one portion of the scene in the output video stream, aborting the masking of the specific region or at least reducing the level of masking of the specific region in response to determining that a specific region of the at least one portion of the scene corresponds to a portion of the scene represented in the second input video stream. A computer program configured to cause the above to be performed.
14. The computer program according to claim 13, further configured to cause the device to perform the method according to any one of claims 2 to 8 when executed by the processor of the device.
15. A non-transitory computer-readable storage medium storing the computer program according to claim 13.
Citation Information
Patent Citations
Video monitoring system for multi-target tracking close-up shooting
CN102148965A
Signal dependent video surveillance
EP3640903A1
Polyester crimped fiber and its production
JP1985099027A
Vehicle vibration controlling equipment
JP1986012436A
Crime prevention unit and image provision method
JP2008097379A