Static privacy shield

By estimating the historical performance of object detection and tracking algorithms, regions that are difficult to locate accurately in an image stream are identified and masked, solving the problem of unstable masking of moving objects in existing technologies and achieving reliable masking of objects in an image stream.

CN116777825BActive Publication Date: 2026-01-23AXIS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310249552.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-16
Filing Date
2023-03-10
Publication Date
2026-01-23
Estimated Expiration
2043-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to reliably mask moving objects in an image stream, especially when the object is partially hidden or suddenly appears, leading to masking failure.

Method used

By estimating the historical performance of object detection and tracking algorithms, regions that were historically difficult for the algorithms to locate accurately are identified and temporarily masked. The historical performance of object detection and object tracking algorithms is used to identify problematic regions in the scene and temporarily mask them in the output image stream.

Benefits of technology

It improves the reliability of object masking, avoids masking failure due to sudden loss or uncertainty of object detection and tracking algorithms, and ensures continuous privacy protection of objects in the image stream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116777825B_ABST
    Figure CN116777825B_ABST
Patent Text Reader

Abstract

The invention relates to static privacy shielding. In particular, a method of shielding in an output image stream is provided, comprising receiving an input image stream capturing a scene, processing the input image stream to generate an output image stream, including using a detector to detect objects in the scene and using a tracker to track objects in the scene based on information provided by the detector, and further including generating a particular output image of the output image stream by checking whether a particular region of the scene exists for which at least one condition is met on an estimate of a historical performance of the detector and / or the tracker, and if such a particular region is confirmed to exist, shielding the particular region of the scene in the particular output image. A corresponding apparatus, computer program and computer program product are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to masking of objects in an image stream, such as an image stream captured by a surveillance (video) camera. In particular, the present disclosure relates to such masking in a region of a scene (that is captured in the image stream), where a moving object can be difficult to detect, and thus also to mask. BACKGROUND

[0002] While video surveillance of a particular scene can improve overall security, it can be desirable to prevent a particular object in the scene from being identified (e.g. by a person watching a video recording from a surveillance camera capturing the scene). For example, for privacy reasons, it can be desirable that the identity of a person captured by the camera, or details such as a vehicle license plate, should not be directly available from watching a recorded video footage. Such protection of a particular object can be achieved by masking the object in the image stream before outputting the image stream to, for example, a display or storage device. Such privacy masking can comprise, for example, within each image of the output image stream, covering the object with a solid color, blurring the object, pixelating the object, or even making the object more or less transparent.

[0003] Before a particular object can be masked, the position of the object in the image must first be estimated. This can be achieved by using an object detection algorithm that has been trained / configured to detect a particular class of objects (such as a human face, a person, a license plate, etc.) within, for example, an image. In particular for moving objects, an object tracking algorithm can also be used. The object tracking algorithm can receive regular updates on the position of the object from the object detection algorithm, and can be trained / configured to estimate the position of the object, thereby tracking the movement of the object between receiving such regular updates from the object detection algorithm. Once the position of the object is detected and / or tracked, privacy masking can be applied to the object.

[0004] However, in order to establish trust in such privacy masking, it is important that the object is kept masked in all images of the output image stream, even in situations where it is difficult to detect the position of the object. For example, if the object becomes partially hidden behind another object, or if the object is only partially within the scene described by the image stream captured by the surveillance camera, the object detection algorithm can fail to locate the object correctly. This can result in a failure to mask the object based on information from the object detection algorithm only. Furthermore, this can also prevent the object tracking algorithm from providing regular updates to the object tracking algorithm, and object masking based on information provided by the object tracking algorithm can also fail as a result. SUMMARY

[0005] To at least partly solve the above problem of unreliable object masking in case of difficult object location determination, the present disclosure provides an improved method of masking in an output image stream, a corresponding apparatus, computer program and computer program product, as defined in the independent claims attached hereto. Various optional embodiments of the improved method, apparatus, computer program and computer program product are defined in the dependent claims attached hereto.

[0006] According to a first aspect of the present disclosure, a method of masking in an output image stream is provided. The method comprises receiving an input image stream capturing a scene. The method comprises processing the input image stream to generate an output image stream, including detecting and tracking one or more objects in the scene based on the input image stream. The one or more objects in the scene are tracked using an object detection algorithm and an object tracking algorithm. The object tracking algorithm receives information from the object detection algorithm indicating an object to be tracked. The processing (of the input image stream) further comprises: generating a particular output image of the output image stream, checking whether there is a particular region of the scene that fulfils one or more of the following conditions: a) a historical performance of the object detection algorithm includes that the object detection algorithm has determined that there is an object to be masked, but wherein the object detection algorithm subsequently became uncertain whether there is an object to be masked; b) a historical performance of the object detection algorithm includes that the object detection algorithm is more uncertain whether there is any object to be masked or tracked than whether there is any object to be masked or tracked; c) a historical performance of the object tracking algorithm includes that the object tracking algorithm has stopped receiving information from the object detection algorithm indicating an object to be tracked; d) a historical performance of the object tracking algorithm includes that the object tracking algorithm has started or resumed receiving information from the object detection algorithm indicating an object to be tracked, but cannot yet start or resume tracking the object to be tracked, and / or e) a historical performance of the object tracking algorithm includes that an uncertainty of a position of an object tracked by the object tracking algorithm in the scene has been deemed too large for masking the object. If it is confirmed that there is (i.e. by fulfilling at least one of the above conditions) a particular region of the scene, the method comprises (statically) masking the particular region of the scene in the particular output image.

[0007] The information "indicating an object to be tracked" as used herein may, for example, include a detection of (center) coordinates of an outer shape of the object in the image plus information about the outer shape of the object. Other alternatives may, for example, include (center) coordinates of the object in the image, an orientation of the object, and a dimension of the object along at least one axis, etc. Other alternatives may, for example, include only (center) coordinates of the object, etc. Alternatively or additionally, the indication of the object to be tracked may, for example, include all pixels of the image that are determined to belong to the object, etc.

[0008] As used herein, the "historical performance" of an estimation algorithm means estimating the performance of the algorithm against at least one or more image frames preceding (in time, in the order captured by the camera for example) the particular image frame that will be generated as part of the output image stream. The number of such previous image frames can vary dynamically depending on, for example, one or more conditions of the scene, such as time of day, lighting conditions, etc. The number of previous image frames can also be predefined and correspond to, for example, image frames spanning a certain number of seconds, minutes, hours, days, weeks, months, etc. prior to the particular output image. The input image in the input stream corresponding to the particular output image in the output stream (and its analysis / processing) can of course also be included as part of the historical performance estimation of the algorithm and be considered part of the history in a particular sense for example. Or in other words, estimating the performance of the algorithm can include looking at the performance of the algorithm in one or more previous input images as well as in the current input image corresponding to the particular output image.

[0009] The method proposed and assumed herein improves the current technology in that it uses an estimation of the historical performance of object detection algorithms and / or object tracking algorithms to identify one or more "problematic areas" of the scene where these algorithms historically can have had more difficulty in properly identifying the presence of objects in the scene and determining their location. Such problematic areas can for example correspond to areas where objects are more likely to suddenly disappear from the scene or suddenly appear in the scene. By at least temporarily shielding these areas, the disclosed method allows to avoid being unable to shield objects in these areas simply because the object tracking algorithm and / or the object detection algorithm suddenly lost track of the object, and / or because the object tracking algorithm did not have enough time (i.e. received enough indications from the object detection algorithm) to properly estimate the location of the object.

[0010] In some embodiments of the method, the object detection algorithm can be such that it generates a probability (i.e. a value between 0 and 100%, or a value between 0.0 and 1.0, etc.) that an object is in a particular region of the scene. The higher the probability, the more certain the object detection algorithm is that an object is present in the particular region of the scene (i.e. close to or at 100%). The lower the probability, the more certain the object detection algorithm is that an object is not present in the particular region of the scene (i.e. close to or at 0%). The requirement that the object detection algorithm determines that an object to be tracked is present in a particular region of the scene can thus be rephrased as a requirement that the probability exceeds a first threshold. The requirement that the object detection algorithm determines that an object to be shielded is present in a particular region of the scene can thus be rephrased as a requirement that the probability exceeds a second threshold. The requirement that an object to be tracked is present can be lower (or equal) to the requirement that an object to be shielded is present, i.e. the second threshold can be equal to or greater than the first threshold. For example, the first threshold (for tracking) can be e.g. 60% and the second threshold (for shielding) can be e.g. 80%. As used herein, it is assumed that the object detection algorithm does not detect objects that it knows belong to a class that should not be shielded. The object detection algorithm can of course be trained to also detect such other objects and indicate them as objects to be tracked, but for ease of discussion, these cases will not be discussed herein.

[0011] Instead of a single probability, it is assumed that the object detection algorithm can alternatively output e.g. two probabilities (i.e. a first probability and a second probability). The first probability can indicate how certain the algorithm is that an object is present in a particular region, while the second probability can indicate how certain the algorithm is that an object is not present in a particular region. This can be advantageous, as it can be checked whether both probabilities are high at the same time as well as whether both probabilities are low at the same time, and then such a situation is judged to be unreliable. Similarly, if one probability is high and the other probability is low, such a result can be considered to be more reliable. In other embodiments, such two probabilities can of course be used to construct a single probability as described above. For example, a first probability of 50% and a second probability of also 50% can correspond to a single probability of 50%. A first probability of 100% and a second probability of 0% can correspond to a single probability of 100%. A first probability of 0% and a second probability of 100% can correspond to a single probability of 0%, etc. Of course, other alternatives can also be assumed, as long as the object detection algorithm is at least able to output a particular value from which it can be determined whether it is certain that an object is not present, whether it is certain that an object is present, or anything in between (e.g. more certain or less certain whether any object is present).

[0012] In one or more embodiments of the method, the object detection algorithm being uncertain whether there is any object to be tracked can comprise a probability that is not more than a first threshold but more than a third threshold that is less than the first threshold. Similarly, the object detection algorithm being uncertain whether there is any object to be occluded can comprise a probability that is not more than a second threshold but more than a fourth threshold that is less than the second threshold. Continuing the example provided above, the first threshold can be 60% and the third threshold can be 20%, such that if the probability is between 20% and 60%, the object detection algorithm is uncertain whether there is any object to be tracked. Similarly, the second threshold can be 80% and the fourth threshold can be 20%, such that if the probability is between 20% and 80%, the object detection algorithm is uncertain whether there is any object to be occluded.

[0013] In one or more embodiments of the method, the method can comprise defining the specific region of the scene by requiring that both conditions b) and e) (as defined above) have occurred historically. In other words, defining the specific region of the scene can require that the historical performance of the object detection algorithm includes that the object detection algorithm was more uncertain whether there is any object to be occluded than in determining whether there is any object to be tracked, and further that the historical performance of the object tracking algorithm includes that the object tracking algorithm tracked an object in the scene whose uncertainty of location was deemed too large for the object to be occluded.

[0014] In one or more embodiments of the method, defining the specific region of the scene can comprise / requiring that condition a) defined above occurred, and further that the rate (or speed) at which the object detection algorithm became uncertain whether there is any object to be occluded exceeds a fifth threshold. In other words, defining the specific region of the scene can require that the historical performance of the object detection algorithm includes that the object detection algorithm determined that there is an object to be occluded, but where the object detection algorithm subsequently became more uncertain whether there is an object to be occluded, and the rate at which the object detection algorithm became uncertain is higher than the fifth threshold. The rate at which the object detection algorithm became uncertain can be defined, for example, in terms of how quickly the probability of the presence of the object in the specific region of the scene decreased. The rate and the fifth threshold can be measured, for example, in percentage points per unit of time. For example, if the probability dropped from above 80% to below 80% (e.g. below 60%, or 50%, or even lower) suddenly (e.g. within a few seconds), the rate at which the object detection algorithm became uncertain can be considered to be higher than the fifth threshold. Likewise, if the probability dropped from above 80% to below 80% over a long time (e.g. minutes or hours, etc.), the rate at which the object detection algorithm became uncertain can be considered to be not higher than the fifth threshold. As will be described below, a rapid decrease in certainty of the presence of the object can for example correspond to a case where an occluding object suddenly enters the scene and hides one or more other objects that were previously fully visible and should have been occluded.

[0015] In one or more embodiments of the method, a heat map can be used to estimate the historical performance of the object detection algorithm and / or the object tracking algorithm. In other words, as will be described in more detail below, the heat map can be such that regions of the scene in which at least one of the above-defined conditions a) to e) occurs more frequently are kept at a "warmer temperature" than regions of the scene in which the above-defined conditions a) to e) occur less frequently or not at all. Then, the warm regions (e.g. regions whose "temperature" exceeds a certain predefined value) can be considered to be specific regions of the scene and at least temporarily shielded in the output image stream. The cold regions (e.g. regions whose temperature is below the predefined value) can alternatively be considered not to be specific regions of the scene and any shielding previously applied in that region can be removed in the output image stream.

[0016] In one or more embodiments of the method, estimating the historical performance of the object detection algorithm and / or the object tracking algorithm comprises results of earlier processing of a limited number of input images of the input image stream, which are temporally preceding the specific output image. The limited number of input images of the input image stream, which are temporally preceding the specific output image, is lower than the total number of input images of the input image stream, which are temporally preceding the specific output image. In other words, estimating the historical performance of the object detection algorithm and / or the object detection algorithm can comprise not using all available previous results of these algorithms, but only investigating how the algorithms performed during e.g. the last few seconds, minutes, hours, days, weeks or months. In other embodiments, estimating the historical performance of the algorithms can be the result of all previously analyzed and processed image frames, if possible.

[0017] In one or more embodiments of the method, the method can further comprise checking whether there are any regions of the scene in which none of the above-defined conditions a) to e) historically occurred, and if such regions are confirmed to exist, not shielding such regions in the specific output image of the output image stream. Herein, "historically" can also include looking at only a limited number of previously processed input images, as described above.

[0018] In one or more embodiments of the method, the method can be executed in a surveillance camera. The surveillance camera can be configured to capture the input image stream. In other words, the method can be executed "at the edge" of e.g. a surveillance camera system (rather than e.g. in a central server or the like), in the camera used to capture the input image stream, wherein the method comprises the processing of this input image stream. Executing the method in the surveillance camera itself (i.e. "at the edge") can reduce the need for any subsequent processing of the image stream output from the surveillance camera for purposes of privacy shielding or the like.

[0019] According to a second aspect of the present disclosure, there is provided an apparatus for masking in an output image stream. The apparatus comprises a processor and a memory. The memory stores instructions that, when executed by the processor, cause the apparatus to: receive an input image stream capturing a scene; process the input image stream to generate an output image stream, including detecting and tracking one or more objects in the scene based on the input image stream using an object detection algorithm and an object tracking algorithm, the object tracking algorithm receiving information from the object detection algorithm indicating an object to be tracked, and generating a particular output image of the output image stream to check whether a particular region of the scene exists in which at least one of the following conditions a) to e) has occurred: a) a historical performance of the object detection algorithm includes that the object detection algorithm has determined that an object to be masked exists, but wherein the object detection algorithm subsequently became uncertain whether any object to be masked exists; b) a historical performance of the object detection algorithm includes that the object detection algorithm is more uncertain whether any object to be masked or tracked exists than in determining whether any object to be masked or tracked exists; c) a historical performance of the object tracking algorithm includes that the object tracking algorithm has stopped receiving information from the object detection algorithm indicating an object to be tracked; d) a historical performance of the object tracking algorithm includes that the object tracking algorithm has started or resumed receiving information from the object detection algorithm indicating an object to be tracked, but cannot yet start or resume tracking the object to be tracked, and / or e) a historical performance of the object tracking algorithm includes that an uncertainty of a position of an object tracked by the object tracking algorithm in the scene has been deemed too large for the object to be masked. Further, the instructions are such that, if the particular region of the scene is confirmed to exist, the particular region of the scene is masked in the particular output image.

[0020] In other words, the instructions are such that they cause the apparatus to perform the method according to the first aspect.

[0021] In one or more embodiments of the apparatus, the instructions can be further configured to cause the apparatus to perform any embodiment of the method of the first aspect as disclosed herein.

[0022] In one or more embodiments of the apparatus, the apparatus can be a surveillance camera. The surveillance camera can be configured to capture the input image stream. To this end, the surveillance camera can comprise, for example, one or more lenses, one or more image sensors, and various electronic components required for capturing the input image stream, for example.

[0023] According to a third aspect of the present disclosure, a computer program is provided. The computer program is configured to, when executed by a processor of an apparatus according to the second aspect, cause the apparatus to: receive an input image stream capturing a scene; process the input image stream to generate an output image stream, including detecting and tracking one or more objects in the scene based on the input image stream using an object detection algorithm and an object tracking algorithm, the object tracking algorithm receiving information indicative of an object to be tracked from the object detection algorithm, and generate a particular output image of the output image stream to check whether there is a particular region of the scene in which at least one of the following conditions a) to e) has occurred: a) the historical performance of the object detection algorithm includes that the object detection algorithm has determined that there is an object to be obscured, but wherein the object detection algorithm subsequently became uncertain whether there is any object to be obscured or not; b) the historical performance of the object detection algorithm includes that the object detection algorithm is more uncertain whether there is any object to be obscured or tracked than whether there is any object to be obscured or tracked; c) the historical performance of the object tracking algorithm includes that the object tracking algorithm has stopped receiving information indicative of an object to be tracked from the object detection algorithm; d) the historical performance of the object tracking algorithm includes that the object tracking algorithm has started or resumed receiving information indicative of an object to be tracked from the object detection algorithm, but cannot yet start or resume tracking the object to be tracked, and / or e) the historical performance of the object tracking algorithm includes that the uncertainty of the position of an object tracked by the object tracking algorithm in the scene has been considered too large to obscure the object. Furthermore, the instructions are such that, if it is confirmed that the particular region of the scene exists, the particular region of the scene is obscured in the particular output image.

[0024] In other words, the computer program is such that it causes the apparatus to perform the method according to the first aspect.

[0025] In one or more embodiments of the computer program, the computer program can be further configured to cause the apparatus to perform any embodiment of the method of the first aspect disclosed herein.

[0026] According to a fourth aspect of the present disclosure, a computer program product is provided. The computer program product comprises a computer readable storage medium having stored thereon a computer program according to the third aspect (or any embodiment thereof).

[0027] Other objects and advantages of the present disclosure will become apparent upon reading the following detailed description along with the accompanying drawings. All features and advantages described in connection with the method according to the first aspect are assumed to be relevant, applicable and can be used in connection with any feature and advantage described in connection with the apparatus according to the second aspect, the computer program according to the third aspect and / or the computer program product according to the fourth aspect, and vice versa, within the scope of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0028] Exemplary embodiments will be described below with reference to the accompanying drawings, wherein:

[0029] Figures 1A-1C An example of a situation to which embodiments of the method according to the present disclosure are applicable is schematically illustrated;

[0030] Figures 2A-2C An example of another situation to which embodiments of the method according to the present disclosure are applicable is schematically illustrated;

[0031] Figure 3 A functional block diagram corresponding to embodiments of the method according to the present disclosure is schematically illustrated;

[0032] Figure 4 A flowchart of embodiments of the method according to the present disclosure is schematically illustrated, and

[0033] Figure 5A and Figure 5B Embodiments of the apparatus according to the present disclosure are schematically illustrated.

[0034] In the drawings, like reference numerals will be used to refer to like elements throughout unless otherwise indicated. The drawings are only intended to illustrate examples of the embodiments as required by the specification and are not intended to limit the scope of the embodiments, unless explicitly stated otherwise. As shown in the drawings, the (absolute or relative) dimensions of elements and regions can be exaggerated or de-emphasized for the purpose of illustration, and thus are provided to illustrate the general structure of the embodiments. DETAILED DESCRIPTION

[0035] Reference will now be made to Figures 1A-1C Various example situations will be described in which an object detection algorithm can find it difficult to properly indicate objects, and / or an object tracking algorithm can find it difficult to properly track objects. Reference will also be made to Figure 3 and Figure 4 How the methods assumed herein can be used to overcome the shortcomings of commonly available techniques and solutions for masking in output image streams will be explained. In the following, for the sake of a simplified reading experience, the terms “object detection algorithm” and “detector” will be used interchangeably. The same applies to the terms “object tracking algorithm” and “tracker”, which can also be used interchangeably.

[0036] Herein, the detector is assumed to be implementable using, for example, one or more commonly available algorithms for object detection, as already available in various fields of computer technology, such as, for example, computer vision and / or image processing. Such algorithms can be assumed to include non-neural and neural approaches, for example. However, as a minimum requirement, whatever algorithm (or combination of algorithms) is used, it should be possible, at least under ideal conditions, to determine whether a particular object (such as a human face, body, license plate, etc.) that is deemed to be to be shielded is present in an image, in particular where in the image and / or area the object is located. It does not matter whether the algorithm(s) used are feature-based, template-based, and / or motion-based, as long as the above requirement is met. The detector can be implemented using, for example, one or more neural networks that are specifically trained for this purpose. For the purposes of the present disclosure, it is also assumed that such algorithm(s) used in / for the detector can struggle to properly identify and / or locate objects that are partially hidden within images of a scene (e.g. a person behind a tree, a vehicle, etc. that is partially shielded).

[0037] Similarly, herein, the tracker is assumed to be implementable using, for example, one or more commonly used algorithms for object tracking. Such algorithms can be, for example, bottom-up processes that rely on target representation and localization, and include, for example, kernel-based tracking, contour tracking, etc. Other assumed tracking algorithms can be, for example, top-down processes that include, for example, the use of filtering and data association, and implement, for example, one or more Kalman and / or particle filters. Herein, it is assumed that such tracking algorithms can receive input from the detector, and then use the received input to track objects in a scene over time, also including in the case where no further input / updates from the detector are provided for a limited time. For the purposes of the present disclosure, it is assumed that even if the tracker is able to track / follow objects of at least two images / frames of an image stream after stopping to receive updates from the detector, the quality of such tracking decreases over time as no new input from the detector arrives. After a while, the tracker will not be able to properly track the objects. It is also assumed that the tracker needs a while to lock onto an object and perform successful tracking after receiving its first input / update from the detector. In other words, the tracker needs more than one data point from the detector to draw conclusions about where / into which direction the object will be located next, and can therefore struggle with how to track an object that has just recently appeared in images of a scene (such as, for example, a person entering the scene through a door, via a road / pedestrian walkway, etc.).

[0038] For the examples provided herein, it will be assumed that the detector is configured to only detect objects to be shielded. The detector can of course also be configured to detect other objects for other reasons, but these objects will not be considered in the following. It is further assumed that the detector can generate probabilities which inform how certain the detector is that an object is present at a particular location of an image of the scene. Preferably, the detector can also inform how certain it is that an object is not present at a particular location of an image of the scene. Of course, how exactly the detector conveys this information can vary depending on the specific implementation of the detector. For the examples provided herein, however, it will be assumed that the detector provides probabilities P(x, y) e [0.0, 1.0], where (x, y) is a 2-tuple (pair) corresponding to, e.g., a particular pixel in an image of the scene, or indexing, e.g., a particular region in an image of the scene. It can also be assumed that the detector provides, e.g., probabilities P(x, y) e [0.0, 1.0] for each pixel in the image of the scene. wherein and are arrays of coordinates such that the returned probabilities correspond to the probabilities that an object is present at the pixels of these arrays.

[0039] Higher probability values indicate that the detector is more certain that an object is present at a particular location, and lower probability values indicate that the detector is certain that an object is not present at a particular location. Intermediate values of the probabilities accordingly indicate that the detector is uncertain whether an object is present at a particular location. For example, a value P(x, y) = 1.0 indicates that the detector is most certain that an object is present at (x, y), a value P(x, y) = 0.0 indicates that the detector is most certain that an object is not present at (x, y), and a value, e.g., P(x, y) = 0.5 (or any intermediate value if a linear scale between the two extremes 0.0 and 1.0 is not used) indicates that the detector is most uncertain whether an object is present at (x, y).

[0040] The objects to be tracked and the objects to be shielded at the same time are further distinguished. The requirements for the objects to be tracked can be lower than the requirements for the objects to be shielded. For example, it can be decided that probabilities exceeding a first threshold T1 (e.g., 0.6 or 60%) correspond to objects to be tracked, while probabilities exceeding a higher second threshold T2 (e.g., 0.8 or 80%) correspond to objects to be shielded. Of course, these values are merely provided as examples and can be changed as needed based on, e.g., the particular scene being monitored, the particular objects to be detected / tracked / shielded, etc.

[0041] There can also be respective thresholds below which it is decided that the detector is sufficiently certain that there is no object that should be tracked or shielded. These lower thresholds can be equal or different. For example, if the probability is below a third threshold T3 (e.g. 0.2 or 20%), it can be decided that the detector is certain that there is no object to be tracked, and if the probability is below a fourth threshold T4 (e.g. also 0.2 or 20%), the detector is certain that there is no object to be shielded. If the probability is, for example, between T3 and T1, it can be said that the detector is not certain whether there is an object to be tracked. Likewise, if the probability is, for example, between T4 and T2, it can be said that the detector is not certain whether there is an object to be shielded.

[0042] It is of course possible to assume that the detector can indicate whether it is certain that there is an object (or to be tracked or to be shielded at the same time) in other ways than the ones described above. For example, as mentioned previously herein in the summary of the appended claims, the detector can output two different probabilities, wherein a first probability indicates how certain the detector is that there is an object at a particular region / location / area of the scene, and a second probability indicates how certain the detector is that there is no object at the particular region / location / area of the scene. This can be useful because consistency can be checked, such that any event where the detector indicates that both probabilities are high (or both low) at the same time can be discarded as unreliable information. If the detector is implemented using, for example, a neural network based solution, the value of one output neuron can correspond to the first probability, and the value of another output neuron can correspond to the second probability.

[0043] If the detector considers an object to be an object to be shielded (e.g. when the probability exceeds the second threshold T2), the position of the object can be provided by the detector to a shielding unit configured to apply a suitable shield for the object in the output image. The detector can also inform the tracker that an object that should be tracked is also an object that should be shielded, and the tracker can then provide information about the current position of the object in the scene to the shielding unit (when tracking the object in the scene) so that the shielding unit can apply and / or update the position of the respective shield in the output image.

[0044] Figure 1AA first image 100 from an input image stream capturing a particular scene is shown schematically. The scene e.g. comes from a park and comprises trees 120, a sidewalk 122 extending partly behind the trees 120, a building 124 with a door 126 and at least three persons 130, 131 and 132. When the first image 100 is captured, the persons 130, 131 and 132 are currently moving along the sidewalk 122 according to arrows 140, 141 and 142, respectively. The image stream is captured / generated by e.g. a surveillance camera facing the scene. For privacy reasons, it is desirable that the identities of the persons 130, 131 and 132 are not obtainable from viewing the output stream from the camera, and thus in such an output image stream, the persons 130, 131 and 132 should be shielded.

[0045] In the first image 100, both persons 130 and 131 are fully visible, so the detector is able to note the presence of persons 130 and 131 in the scene. Thus, the detector is deemed to determine that persons 130 and 131 are objects in the scene and that these objects should at least be tracked (and likely also should be shielded). The detector accordingly provides the locations of the detected persons 130 and 131 to the tracker as information indicative of objects to be tracked (and possibly shielded). The tracker manages to track persons 130 and 131, as shown by the dashed boxes 150 and 151 around these persons.

[0046] However, person 132 is not yet fully in the scene, so the detector is not certain whether person 132 is in the scene, and thus is not certain whether person 132 is an object to be tracked. Thus, the tracker has not previously been fed any indication of person 132 as an object to be tracked, and currently there is no available tracking of person 132 (as shown by the lack of a corresponding dashed box around person 132).

[0047] If privacy shielding is applied based only on the locations where objects such as persons 130 and 131 are currently indicated by the detector and / or the tracker, then such shielding would result in shielding being correctly applied to persons 130 and 131 in the first image 100, but not to person 132, because person 132 has not yet been recognized by the detector and the tracker due to not yet being fully in the scene. Based on such a conventional privacy shielding, the images of the output image stream corresponding to the first image 100 would thus fail to at least hide the identities of person 132.

[0048] As previously described herein, the detector "indicating" an object is equivalent to the detector being able to provide information indicative of the object to the tracker. The information may e.g. comprise an estimated center position of the object, a size of the object, a contour of the object, a direction of the object, exact pixels in the image of the scene corresponding to the object, or any combination of the above or other parameters from which the position of the object in the scene and which pixels in the image of the scene should e.g. be altered to shield the object can be derived.

[0049] Figure 1B A second image 101 from the same image stream is schematically shown. The second image 101 is subsequent to the first image 100, i.e. corresponds to the first image 100 and is captured at a later point in time than the first image 100. In the second image 101 the positions of the persons 130, 131 and 132 have changed. A fourth person 133 also appears in the scene, e.g. by entering the scene through the door 126 leaving the house 124. Figure 1A

[0050] Person 130 has partly moved behind the tree 120 and the detector thus becomes uncertain whether person 130 is the object to be tracked. Likewise, person 131 has partly left the scene and the detector has (starting from the processing of the first image 100) become uncertain whether person 131 is the object to be tracked. The detector has thus (starting from the processing of the first image 100) stopped providing further updates / indications to the tracker regarding persons 130 and 131. However, the tracker is still able to track persons 130 and 131 based on the information it previously received (as indicated by the dashed boxes 150 and 151 still present in Figure 1B The tracker correctly estimates that person 130 is partly behind the tree 120 and that person 131 has moved out of half the scene. However, after stopping to receive further indications from the detector regarding persons 130 and 131, the tracker will sooner or later be unable to perform such tracking if the detector does not start providing such indications again. For example, person 130 can stay behind the tree 120 long enough for the tracker to lose track of person 130. If person 130 returns from behind the tree 120 again, there can be one or more moments in time where person 130 is not detected by the detector and not tracked by the tracker.

[0051] On the other hand, person 132 is now fully in the scene and the detector has determined that person 132 is the object to be tracked (and most likely also shielded). The same applies to person 133, which is also fully visible in the scene. The detector has thus started to provide indications of persons 132 and 133 to the tracker as objects to be tracked (and most likely also shielded). However, since these are the first indications of persons 132 and 133 sent to the tracker, the tracker most likely has not yet locked onto these persons, as it has so far only received a single indication of each person from the detector (as persons 132 and 133 both just entered the scene fully). The tracker needs to receive at least one more indication for each of persons 132 and 133 before it is able to properly track persons 132 and 133 in the scene.

[0052] ​If only the second image 101 is used to apply a traditional privacy shield, a proper shielding is possible for persons 132 and 133, as they are likely to both be indicated by the detector as objects that will also be shielded. Based on the output from the tracker, persons 130 and 131 can also be properly shielded, but it is uncertain whether such shielding will also be successful in one or more future images of the image stream, as these persons 130 and 131 are not currently indicated by the detector and there is a risk that they are lost by the tracker. Thus, if a traditional approach is used, future privacy shielding for these persons will fail. Furthermore, if it is assumed that the tracker has not successfully tracked, for example, persons 130 and 131 when going from the first image 100 to the second image 101, the detector or tracker does not provide an indication for these persons 130 and 131, and if an output image (of the output stream) corresponding to the second image 101 is output, shielding for these persons 130 and 131 will thus fail.

[0053] Reference will now again be made to Figure 1A and Figure 1B and to Figure 1C how the assumed method 500 in the present disclosure can reduce the risk of failing to properly apply privacy shielding to all objects, such as persons 130-133, will be explained in more detail.

[0054] When performing the shielding in the output image 102 corresponding to the second image 101, the assumed method 500 does not only consider the single image (second image 101), but relies on an estimate of the historical performance of the tracker and / or detector before a conclusion is drawn on where to apply the privacy shielding. This is obtained in the following way.

[0055] As Figure 1B indicated, the second image 101 will be used to generate a respective particular output image 102 of the output image stream. Instead of performing the shielding in the particular output image 102 using only the information available in the second image 101, the assumed method also takes into account how the detector and / or tracker has historically performed when detecting and / or tracking objects. In the present example, this includes studying how the detector and / or tracker has performed when detecting and / or tracking objects in at least the first image 100. For ease of explanation, the “history” in this example is thus only the first image 100. Of course, it can also be that the history includes one or more images that temporally precede the first image 100.

[0056] When studying the history (i.e. the first image 100), the current image (second image 101), and how the detector and / or tracker has historically performed, the assumed method draws the following conclusions:

[0057] 1) In the area of the first person 130 in the second image 101, the tracker has stopped receiving information from the detector indicating an object (person 130) to be tracked (or masked). This corresponds to condition "c)" described herein.

[0058] 2) In the area of the second person 131 in the second image 101, the tracker has stopped receiving information from the detector indicating an object (person 131) to be tracked (or masked). This also corresponds to condition "c)" described herein.

[0059] 3) In the area of the third person 132 in the second image 101, the tracker has started receiving information from the detector indicating an object (person 132) to be tracked (or masked), but the tracker cannot yet start tracking the object (person 132) to be tracked (or masked). This corresponds to condition "d)" described herein.

[0060] 4) In the area of the fourth person 133 in the second image 101, the tracker has started receiving information from the detector indicating an object (person 133) to be tracked (or masked), but the tracker cannot yet start tracking the object (person 133) to be tracked (or masked). This also corresponds to condition "d)" described herein.

[0061] Based on the above, the assumed method thus concludes that at least conditions "c)" and "d)" have occurred in the areas of persons 130 to 133 in the second image 101, and that these areas thus constitute specific areas in the scene to be masked in the specific output image 102 (of the output image stream) corresponding to the second image 101. The assumed method thus continues to mask these areas in the specific output image 102 of the output image stream, as Figure 1CThe privacy shields 160 to 163 can also be provided in one or more other output images subsequent in time to the particular output image 102, at least until it is decided that no such shield is needed anymore to properly shield the persons 130 to 133. The size of the privacy shields 160 to 163 can for example correspond to the size of the detected persons (objects) 130 to 133, or for example be larger than the respective persons 130 to 133. The shape of the privacy shields 160 to 163 can for example correspond to the shape of the detected persons 130 to 133, or for example form a rectangle, a square, a circle, or have any shape and size that sufficiently covers the respective persons 130 to 133 and possibly also a larger area around each of the persons 130 to 133. In this context, "sufficiently covers" is understood to cover enough objects such that for example the identity of the objects is not obtainable, and / or for example enough license plates such that for example the license plate number written on the license plate is not obtainable, etc. The persons 130 and 131 are detected by the detector as objects that were already in the first image 100 to be shielded. Thus, when the tracker follows the persons 130 and 131 in the second image 101, the tracker knows that the persons 130 and 131 are to be shielded, and can inform the shielding unit accordingly. The persons 132 and 133 are detected by the detector as objects to be shielded during processing of the second image 101, and the detector (or the tracker) can then inform the shielding unit accordingly.

[0062] It should be noted that the (static) privacy shields (e.g. 160 to 163) applied in the assumed approach are associated with regions of the scene (i.e. regions of the particular output image 102), and not with specific objects in the scene. In other words, the applied privacy shields are associated with regions that are considered to be problematic regions of the scene (or output image 102), where objects can be assumed to be more likely to suddenly disappear from the scene or appear in the scene, such that they become difficult or impossible to be detected by the detector (or tracked by the tracker). Thus, the applied privacy shields are static in that they do not move around with the objects. It can be assumed that the regions of the static shields are defined such that if an object moves outside of these regions of the static shields, the detector and / or the tracker will be able to detect and / or track such object without problems, and such that the privacy shields can then be applied using conventional methods.

[0063] The assumed privacy shield herein can for example have a solid / opaque color, be semi-transparent, include applying a motion blur to the object such that the object is no longer easily recognizable, and / or for example force pixilation and / or blur of the object in the output image stream such that the object is no longer easily recognizable, etc. In other assumed embodiments, the privacy shield can include making the object itself at least partially transparent in the output image stream such that the background is visible through the object. This is possible if for example a background image is available (e.g. from an earlier time instant) without the object in front of it.

[0064] By also analyzing the history of the tracker and / or detector, the assumed method manages to identify "problematic" regions of the scene where the risk of failing to properly shield the object is higher than in other regions of the scene. By applying a static shield (i.e. a shield that does not necessarily follow the object, but is fixed relative to the scene itself), the risk of failing to shield the object present in these regions can thereby be reduced. The shield can be applied to at least one subsequent image in the output image stream. When in the future to remove the shield can for example be decided based on an analysis performed on subsequent images of the input image stream, which will be explained in more detail below. The privacy shield can also be left in place in a particular region of the scene for an unforeseeable time, for example if the problems of object detection and / or tracking in that region remain problematic. If reference is made to Figures 1A-1C , such regions can for example be regions where people frequently enter and / or exit the scene (e.g. the regions of shields 161, 162 and 163), and / or regions where objects frequently partially hide or temporarily disappear (e.g. the region of shield 160 behind trees 120 for path 122). It should be noted that in cases where Figures 1A-1CIn the example case shown, the person 131 can not necessarily cause any problems with the shielding, because even if the detector stops indicating the person 131, the tracker can still track the person 131. Thus, the person 131 can be shielded until the person 131 has left the scene completely. However, it can be assumed that the area where the person 131 happens to leave the scene is also an area where other people can enter the scene. Such an area can for example be a sidewalk, road, etc. leading outside the scene, or for example a door, a corner of a building, etc. where people can also enter and leave the scene suddenly. Thus, it can be advantageous to apply a (static) shield, such as the shield 161, in this area. In other words, the assumed method can make use of the fact that in determining an area where a person or other object leaves the scene, it is typically equally likely that the person or other object enters the scene, and vice versa. In one or more other embodiments, it can be assumed that a (static) shield is only executed in an area where it can be determined that a person enters the scene only, and not in an area where it can be determined that a person leaves the scene only. For example, the scene can comprise the outside of a building having an entrance door and an exit door, for example. An area where the tracker historically stopped receiving updates from the detector and thus can have lost track of an object can very likely correspond to the entrance door of the building, and no static shield needs to be applied in such an area. Likewise, an area where the tracker historically started receiving updates from the detector can very likely correspond to the exit door of the building, and a (static) shield can be applied in such an area. If the scene is not inside a building, a shield can instead be applied at the assumed entrance door of the building, and no shield can be applied at the assumed exit door of the building. In other words, the assumed (static) shields can be executed only in areas where they actually work, and for example, not in areas where an object leaves the scene only.

[0065] Reference will now be made to Figures 2A-2C Another example case where the assumed method disclosed herein can be used will now be described in more detail.

[0066] Figure 2A Another case is schematically shown where the assumed method can provide an improved shielding relative to conventional shielding techniques. A first image 200 is taken from an input image stream capturing another scene. This other scene comprises a plurality of people 230-233 waiting at a bus stop, for example. When the first image 200 is captured, all people 230-233 are fully visible in the scene, and thus the detector determines that the people 230-233 are all objects to be tracked (and most likely also shielded), and the detector can inform the tracker accordingly. If it is assumed that the people 230-233 have stayed at their current locations during at least some additional previous images describing this scene, the tracker thus also receives previous indications of the people 230-233 from the detector all the time, and has been tracking the people 230-233, as shown by the dashed boxes 250-253. As the people 230-233 are tracked, the tracker can also determine that the people 230-233 are leaving the scene, and can inform the detector accordingly. If it is assumed that the people 230-233 have stayed at their current locations during at least some additional previous images describing this scene, the detector thus also receives previous indications of the people 230-233 from the tracker all the time, and has been shielding the people 230-233, as shown by the dashed boxes 250-253.Figure 2A The car 228 is about to enter the scene and moves in the direction indicated by arrow 248.

[0067] Figure 2B A second image 201 of the same input image stream capturing the same scene is schematically shown. Figure 2A The second image 201 is temporally subsequent to the first image 200 and the car 228 has now moved far enough into the scene that it now partially screens or blocks the persons 231, 232 and 233. The person 230 is still fully visible in the scene. When the car 228 starts to partially screen or block the persons 231 to 233, the detector will become uncertain whether there are objects to be tracked (and screened) in the area of the persons 231 to 233 that are eventually at least partially hidden behind the car 228. Therefore, the detector stops providing further updates / indications to the tracker about the persons 231 to 233. Since the person 230 is still fully visible in the scene, the detector still determines that the person 230 is an object to be tracked (and screened) and continues to provide update indications to the tracker about the person 230.

[0068] If a conventional privacy screening is applied based on the content of the second image 102 only when creating the respective particular output image 202 of the output image stream, a proper screening can be possible for the person 230, but is less certain for the persons 231 to 233 that are partially hidden behind the car 228. If the tracker is able to lock onto the persons 231 to 233 in advance before the car 228 appears in the scene, the tracker can still be able to correctly guess the positions of the persons 231 to 233 (for a while) and inform the screening unit accordingly. However, if the car 228 remains in front of the persons 231 to 233 for a long time, the tracker will eventually be unable to guess the positions of the persons 231 to 233 (as no new updates from the detector arrive, the tracker becomes more and more uncertain in its estimates) and there is a risk that the screening of the persons 231 to 233 will fail as a result (as neither the detector nor the tracker can provide the required information to the screening unit).

[0069] Figure 2C The above-described problems of conventional privacy screening are schematically helped to be solved by the improved method assumed herein.

[0070] When studying the history (i.e. the first image 200), the current image (the second image 201) and how the detector and / or the tracker performed in the history, the assumed method concludes that 1) in the area of the second person 231, the third person 232 and the fourth person 233, the tracker has stopped receiving information from the detector indicating objects (any of the persons 231 to 233) to be tracked. This corresponds to the condition "c)" described herein.

[0071] Furthermore, the hypothetical method can also lead to the conclusion that: 2) in the regions of the second person 231, the third person 232, and the fourth person 233, the detector has determined (when processing the first image 200) the existence of an object to be masked (any one of persons 231 to 233), but the detector then (when processing the subsequent second image 201) becomes uncertain (in this context, due to the entry of car 228 into the scene) whether an object to be masked exists. This corresponds to condition “a)” described herein.

[0072] In addition to conclusion 2) above, the assumed method also leads to the conclusion that in the region between people 231 and 233, the detector changes from certainty to uncertainty at a rate exceeding the fifth threshold. This is because car 228 suddenly (e.g., within a few seconds) enters the scene, and the detector's certainty about the presence of an object to be masked (or tracked) in the scene decreases much faster than if car 228 were to move slowly into the scene. It should be noted that if car 228 moves slowly into the scene, conclusion 1) above will still hold, because there will still be at least one moment in which the detector becomes sufficiently uncertain about the presence of any object to be masked (or tracked) to stop sending further updates / instructions to the tracker.

[0073] Therefore, as shown above, the assumed method can handle the reference better than traditional shielding techniques. Figures 2A-2B The situation described is such that the sudden entry of car 228 into the scene helps to trigger not just one, but at least two conditions (conditions a) and c) as assumed in this paper. Figure 2C As shown, the assumed method then continues to define the human regions 231 to 233 as specific regions of the scene where (static) masking should be performed, and applies (static) masking 260 to these specific regions in a specific output image 202 (of the output image stream) corresponding to the second image 201. Although in Figure 2C Not shown, but the method could certainly apply masking to person 230, however, followed by, for example, conventional methods, since person 230 remains fully visible in the scene in both the first and second images 200 and 201. Since people 231 to 233 are standing very close to each other, it is assumed that only a single static masking 260 is applied. However, this may not always be the case, and it is certainly possible to assume that separate static masks are applied to people 231 to 233 who are partially hidden behind car 228 or become partially hidden behind car 228.

[0074] The above references Figures 2A-2CThe described scenario can also be used to describe yet another embodiment of the assumed method. Imagine that after the second image 201 is captured, the car 228 stays in front of the persons 231 to 233 long enough so that the tracker eventually loses track of the persons 231 to 233. The car 228 then starts moving again and is present in the scene so that it no longer partially hides the persons 231 to 233. Then, the detector will again be able to detect the persons 231 to 233 and decide that these are objects in the scene to be shielded and resume sending corresponding indications / updates to the tracker. This can be done e.g. when processing a subsequent third image (not shown, but assumed to describe the persons 230 to 233 without any object partially obstructing them). When processing the third image, the tracker just started again to receive indications / updates from the detector and has not yet started to track the corresponding objects (as it waits for at least one more indication / update from the detector). Generating an output image corresponding to the third (input) image, estimating the performance of the detector and the tracker can include investigating how the detector and / or the tracker performed when processing at least the first image 200 and the second image 201, and how the detector and / or the tracker performed during processing of the imagined third image. The assumed method would then conclude that (after the car 228 has left the scene) 3) the historical performance of the tracker includes that the tracker has resumed receiving indications from the detector that an object (any of the persons 231 to 233) is to be tracked, but has not yet resumed tracking the object to be tracked. This also corresponds to condition "d)" described herein. In other cases, it can be assumed that the car 228 leaves the scene before the tracker loses track of the persons 231 to 233. However, the particular area of the persons 231 to 233 can still be shielded with (static) privacy shielding, as it will still be noted that the detector has changed from determining that there is an object to be shielded to being uncertain whether there is an object to be shielded in the scene (condition "a)"), and / or because the tracker has stopped receiving indications from the detector about the persons 231 to 233 in that area of the scene.

[0075] In general, the assumed method suggests defining a "problematic area" of the scene to include a problematic area of one or both of the detector and the tracker. The problematic area of the detector can be defined as a location where the detector historically is more uncertain whether there is any object to be shielded (or tracked) than it is to determine whether there is any object to be shielded (or tracked), corresponding to condition "b)" defined herein. Such a problematic area for the detector can e.g. be an area close to the border of the image describing the scene, where the detector generally is problematic in identifying objects, as objects are generally only partially visible in these areas of the image (e.g. with reference to Fig. 1). Figures 1A-1CThe areas where the sidewalk 122 enters and leaves the scene). These areas can remain in the scene over time, so it can be advantageous to statically mask these areas.

[0076] For the detector, the problematic areas can also be defined as positions where the detector historically changed from being certain to being uncertain due to e.g. a sight blocking object appearing in the scene and blocking one or more objects that the detector previously had identified, such as in reference to the condition "a)" defined herein. Figures 2A-2C The described bus 228 entering the scene. The likelihood of such sight blocking objects appearing in the scene again can be high if e.g. the bus stop, the taxi stop, the train station, etc. is a place where potential sight blocking objects frequently appear in and disappear from the scene, etc., and for this reason it can also be advantageous to statically mask these problematic areas.

[0077] The problematic areas for the tracker can e.g. be defined as areas where the uncertainty of the position of an object tracked by the tracker in the scene is considered too large for the object to be masked (corresponding to the condition "e)" defined herein). Assuming the use of e.g. a threshold, the certainty of the tracker must overcome this threshold for the estimated position of the object to be considered sufficiently certain, and where any estimated position of the object with a certainty below this threshold is considered to have too large uncertainty for the object to be masked. It should be noted that the tracker typically cannot determine on its own whether a particular object is to be masked or not. Instead, the tracker only estimates the current and / or future position of an object it has been told (by the detector) to track, and whether the tracked object is considered to be an object to be masked depends on whether the tracker manages to track the position of the object with a certainty greater than the threshold. For example, for an object o, the tracker can output an estimated position L(o) = (x ′ , y ′ ) of the object "o" and an estimated uncertainty σ(o) = (Δx ′ , Δy ′ ) informing that the position of the object o is at a particular place in the interval (x ′ ± Δx ′ , y ′ ± Δy ′ ). In this case, if the uncertainty (Δx ′ , Δy ′ ) is higher than e.g. a predefined threshold, the object that was indicated by the detector to be tracked and masked can be considered to have too large uncertainty for the object to be masked. In other embodiments, assuming the tracker can instead output e.g. a position L(o) = (x ′ , y ′ ) of the object "o" and a probability density function p(x, y) of the object "o" being at a particular place in the scene, where the probability density function p(x, y) is a function of the position (x, y) in the scene, and where the probability density function p(x, y) is higher in areas where the object is more likely to be than in areas where the object is less likely to be, the object that was indicated by the detector to be tracked and masked can be considered to have too large uncertainty for the object to be masked if the probability density function p(x, y) is below a certain threshold.), and one or more other parameters, such as a confidence level value, a confidence interval, a standard deviation, a mean value, etc., from which a respective uncertainty of the object position estimated by the tracker can be obtained. If the problematic area of the tracker is defined independently from the respective problematic area of the detector, the problematic area of the tracker can be larger than the problematic area of the detector, since the tracker is typically able to continue (at least for some time) tracking an object in the scene even after it has stopped receiving updates / indications about the object from the detector.

[0078] In some embodiments of the assumed approach, it can be advantageous to define the problematic area of the tracker to be, e.g., locations where both conditions “b)” and “e)” have occurred historically, locations where the detector has stopped sending updates / indications of an object to be tracked (or even shielded), and locations where the tracker has subsequently failed to further track the position of the object in its estimation of the position with sufficient certainty. Such an area can, e.g., correspond to a situation where there is a door through which an object leaves the scene and does not return (or only returns after a sufficiently long time such that the tracker loses track of the object), a sufficiently large tree such that a person moving behind the tree remains hidden for a sufficiently long time such that the tracker loses track of the person, etc.

[0079] In some embodiments of the assumed approach, the estimation of the historical performance of the tracker and / or detector can be performed by constructing a so-called “heat map”. Herein, a heat map is, e.g., a two-dimensional map of the scene, wherein each point of the heat map corresponds to a particular area of the scene. Generally herein, an “area” of the scene can be, e.g., a single pixel, a collection of pixels, or any other portion of the scene that does not correspond to the entire scene. For each point of the heat map, a value can then be assigned, and it can be predetermined, e.g., whether a larger value corresponds to a problematic area, and a smaller value corresponds to a non-problematic area, or vice versa. In other words, the value of each point of the heat map can be limited to a particular range of values, and it can be determined whether a value towards or at one end of the range corresponds to a problematic area, and a value towards or at the other end of the range corresponds to a non-problematic area, or vice versa. In the following, it will be assumed, merely for exemplary reasons, that a higher value of a point of the heat map corresponds to a respective area of the scene that is more likely to be a problematic area (for which privacy shielding should be applied), whereas a lower value of the same point of the heat map corresponds to a respective area of the scene that is more likely to be a non-problematic area. It can be checked whether the value has moved sufficiently towards one end of the range, e.g., by comparing the value to a threshold value. Such a heat map can be used as follows.

[0080] As previously explained herein, each time it is determined that at least one of the events "a) to e)" has occurred in a particular region of the scene in the current image (or in a previous image, if the heat map is created based on historical information), the value of the corresponding point of the heat map is increased (in other words, the temperature at this point of the heat map is increased). If it is subsequently determined that the value at a particular point of the heat map exceeds a threshold value, the corresponding region of the scene is decided to be a problematic region and is considered to be the particular region of the scene that is masked in a particular output image of the output image stream. In some embodiments, it can be determined how long historical data should be included in the heat map (and the dynamic reduction of data described below). The heat map can then have to be updated as older data has to be deleted. In other embodiments, all historical data can be kept in the heat map.

[0081] For example, if each image in the input image stream has a resolution of X x Y pixels, a heat map H(x, y) of e.g. the same size can be constructed, each element H(x, y) indicating a value corresponding to the respective pixel (x, y) of the images in the input image stream. If the assumed method determines that one of the events "a) - e)" has occurred in a region of the scene corresponding to a set of pixels S = {(xl, yl), (x2, y2),...}, the heat map can be updated such that the heat map values of all points included in this set S are increased, e.g. by adding 1 to the previous values. These values can also be defined to have an upper limit H max , such that for each point (x, y), H(x, y) < H max . One or more problematic regions of the scene can then be identified as corresponding to all elements H(x, y) that exceed a threshold value T HM1 . By inspecting all elements of the heat map H(x, y), it is thus possible to decide where (or if) (static) masking should be applied in a particular output image of the output image stream. After running the assumed method on a plurality of images in the input image stream, the heat map will thus be "hot" for regions where the detector more frequently starts or stops indicating an object, and "cold" for regions where such events do not typically occur.

[0082] For example, for each indicated object, the detector can return a set of pixels S that is identified as belonging to the object. In other embodiments, the detector can instead return, for example, the (center) coordinates of the object, as well as an estimated size (e.g. height and width) of the object. In other embodiments, the detector can for example return the coordinates of at least two corners of a rectangle that encloses the object, etc. Independently of the exact form / shape of the output from the detector, it is assumed that a respective set of pixels S can be identified and used to update the heat map H(x, y). It should also be noted that each point of the heat map does not necessarily correspond to a pixel, but can correspond to a larger area of the scene. For example, the scene can be divided into a plurality of areas (each area corresponding to a plurality of pixels), while the heat map is instead such that each point of the heat map corresponds to a particular such area of the scene. The detector can for example indicate whether an object is present in each area, but not provide any finer granularity regarding the exact position and / or size and shape of the object.

[0083] As a further development of such embodiments, the method can check whether the detector currently determines (when processing an input image of the input stream that corresponds to a particular output image of the output stream of images) that an object to be shielded is present at / in a particular area of the scene (and / or whether the tracker currently tracks such an object to be tracked in the scene with sufficient certainty). If this is confirmed to be true, the method can then update the heat map H(x, y) by reducing the values of the elements of the heat map that correspond to this particular area of the scene. For example, if the detector indicates that an object is present in a set of pixels of an image of the scene, the heat map can be updated by reducing the values of the corresponding elements at that set of pixels. Thus, positively detecting and / or tracking an object in a particular area of the scene can “cool down” the heat map for the respective area. This can result in one or more values eventually falling below the threshold after previously being above the threshold. In such a case, the method can for example proceed by removing any previously applied shielding for these areas of the scene. This can be advantageous in areas that are temporarily obstructed by an object that is for example occluded, such as the car 228 referred to with reference to Figures 2A-2C The described car 228.

[0084] Instead of having the “hot” areas of the heat map correspond to problematic areas, it is of course also possible that the method has problematic areas that are cooler than non-problematic areas, for example by reducing the values of the heat map that correspond to areas where the conditions “a) to e)” occur, and vice versa.

[0085] The method of shielding in the output stream assumed herein will now be described in more detail with reference to Figure 3 and Figure 4 The figures show various embodiments of the method 400 assumed herein, and have already been described with reference to Figures 1A-1C and Figures 2A-2C .

[0086] Figure 3 A functional block diagram 300 is schematically illustrated, showing various functional blocks for performing the method 400 as assumed herein. Figure 4 A flowchart of the method 400 is schematically illustrated.

[0087] An input image stream 310 is received (in step S401) from an image sensor, e.g. a camera. Currently, an image I n (n, where n is an integer, denotes that this is the n:th image of the input image stream 310) will be analyzed and processed. The image I n is provided to a tracking module 320 and a detection module 330 for processing (in step S402) to generate an output image stream 312 based on the input image stream 310.

[0088] The tracking module 320 is configured to perform object tracking based on the input image stream 310, while the detection module 330 is configured to perform object detection, as previously discussed herein. The tracking module 320 provides tracking data 322 to a masking module 340 regarding one or more objects it is currently tracking in the image I n The tracking module 320 can of course use information from one or more previous images I m<n to perform tracking of one or more objects in the image I n As previously mentioned, the tracking data can for example include an estimated (tracking) position of an object, as well as some measure of how certain this estimate is. In some embodiments, the tracking data 322 can also include an indication as to whether the object being tracked is an object to be masked or not. Such an indication can be provided first from the detection module 330 to the tracking module 320.

[0089] The detection module 330 provides detection data 332 to the masking module 340 regarding one or more objects it believes to be located in the image I n The detection data 332 also includes a probability indicating how certain the detection module 330 is that the object is an object to be tracked (or even to be masked). The detection module 330 also provides similar or identical detection data 333 to the tracking module 320, so that the tracking module 320 can use the detection data 333 to improve its tracking performance. The detection data 333 can also include an indication as to whether an object to be tracked is also an object to be masked.

[0090] The tracking module 320 and the detection module 330 also provide performance data 324 and 334, respectively, to the performance analysis module 350. The performance data 324 can comprise, for example, the tracking data 322, while the performance data 334 can comprise, for example, the detection data 332. In particular, the performance data 324 comprises enough information for the performance analysis module 350 to obtain how the tracking module 320 historically performed, including, for example, the uncertainty of the object position estimated by the tracking module 320. The performance data 324 comprises enough information for the performance analysis module 350 to obtain how the detection module 330 historically performed, including, for example, how certain the detection module 330 was that an object to be tracked was present (or not present) in the scene, and, for example, whether the detection module 330 believed that the object should also be occluded.

[0091] The performance analysis module 350 receives (e.g., stores) the performance data 324 and 334 for each image of the input image stream 310, and can therefore track:

[0092] Whether, when, and how the tracking module 320 stopped, started, or resumed receiving updates / indications 333 from the detection module 330 (e.g., for estimating the conditions “c)” and / or “d)” previously defined herein);

[0093] How certain the tracking module 320 was of its estimates (for estimating the condition “e)” previously defined herein), and / or

[0094] Whether, when, and how the detection module 330 determined or did not determine whether an object to be tracked or to be occluded was present (for estimating the conditions “a)” and / or “b)” previously defined herein).

[0095] The performance analysis module 350 outputs performance estimate data 352 to the occlusion module 340. The performance estimate data 352 can be, or can comprise, for example, a heat map H n (x, y) as previously described, which is used by the occlusion module 340 when occluding the corresponding nthimage O n of the output image stream 312. In other embodiments, the heat map H n (x, y) can be used only internally by the performance estimation module 350, which can then instead provide direct information about the region or regions to be occluded to the occlusion module 340 as part of the performance estimate data 352. The tracking data 324 and detection data 334 sent to the performance estimation module 350 can of course also include data about how the tracker and detection modules 320 and 330 performed when processing / analyzing the most recent input image I n so that this most recent performance can also be considered part of the historical estimates performed by the module 350.

[0096] The shielding module 340 then applies shielding of one or more objects in the scene based on the tracking data 322, the detection data 332, and in particular also based on the performance estimate data 352, and if (in step S403) it confirms that there is one or more such specific regions, it outputs the image O n as part of the output image stream 312. In particular, the shielding module 340 applies static privacy shielding based on the performance estimate data 352. For example, if a heat map H n (x,y) is received, the shielding module 340 can check whether a specific region / pixel of the scene corresponding to a point (x,y) of the heat map H n (x,y) is to be shielded in the image O n by, for example, comparing the value of the point (x,y) of the heat map H n (x,y) with a threshold. In other embodiments, such analysis has been performed in the performance estimate module 350. n

[0097] After the image O n is output as part of the output image stream 312, the method 400 can continue by receiving the next image (e.g. image I n+1 ) of the input image stream 310, and repeating the process to perform shielding in the particular next output image O n+1 , etc. When analyzing the image I n , how many previous images I m<n are considered can be adjusted precisely as needed. For example, since conditions of the scene (e.g. number of moving objects entering and leaving the scene, time of day, etc.) change over time, the number of previous images considered also changes over time. In other embodiments, the number of previous images considered can be static, and for example correspond to a predefined number of seconds, days, hours, weeks, months, etc. of captured and analyzed / processed input images of the scene. As also mentioned herein, the most recent input image I n , and the analysis results of the tracking and detection modules 320 and 330 on it, can also be considered part of the historical performance estimate. n

[0098] In one or more embodiments, the data 322 and 332 sent to the masking module 340 can include various probabilities and certainties of the detection module 330 and the tracking module 320, and these values can be compared by the masking module 340, e.g., to one or more thresholds, to decide whether a particular object is an object to be masked. In other embodiments, such a decision can already be made by the detection module 330 and / or the tracking module 320, such that, e.g., the data 332 and 322 need not contain, e.g., the various probabilities and certainties. For example, the detection module 320 can decide that about an object that the object should be masked with sufficient certainty, and send the position of the object as part of the data 332 to the masking module 340. Conversely, if the detection module 320 is not certain whether an object should be masked, the detection module 320 can choose not to send information, such as the position, of the object to the masking module 340. Likewise, the tracking module 320 can choose to send information about an object to be masked to the masking module 340 only if the detection module 330 has informed the tracking module 320 that the object should be tracked and masked, and has determined that the certainty of the estimated position of the object is sufficiently high. In other words, it does not matter whether the decision about whether to mask an object is made in the detection module 330, the tracking module 320, and / or the masking module 340, as long as the decision is made somewhere. However, specifically, whether one or more regions of the scene that are considered to be problematic regions to be statically masked should be based on the output 352 from the performance estimation module 350.

[0099] In one or more embodiments, the output 352 from the performance estimation module 350 can alternatively (or additionally) be provided to the detection module 330 and / or the tracking module 320, such that they can use, e.g., the data 322 and 332 to communicate where to apply static masking to the masking module 340. In one or more other embodiments, the data 322 and 332 preferably relate to information about objects to be masked and that are not part of a problematic region, while information about additional masking required for a problematic region is instead provided to the masking module 340 via the output 352 from the performance estimation module 350.

[0100] It should be noted that the disclosed method 400 (also illustrated by Figure 3 the functional block diagram) not only uses immediate object detection and / or tracking data (i.e., only from the input image I n obtained) to perform masking, but also relies on historical such data and an estimation of the historical performance of the detection and / or tracking modules. This is because the performance analysis module 350 also estimates the performance of the detection and / or tracking modules for at least some previous images I m<nThe performance of the tracking and detection modules 320 and 330 is tracked. This allows the assumed and disclosed method 400 to apply static (privacy) masking in areas where the tracker and / or detector historically have difficulty to properly track and / or indicate one or more objects, to avoid the risk of failing to mask objects in these areas.

[0101] The present disclosure also assumes an apparatus for masking in an output image stream, as will now be described with reference to Figure 5A and Figure 5B are described in more detail.

[0102] Figure 5A An embodiment of the apparatus 500 is schematically illustrated. The apparatus 500 comprises at least a processor (or “processing circuitry”) 510 and a memory 512. “Processing circuitry” or “processor” as used in the present disclosure can be, for example, any combination of one or more of a suitable central processing unit (CPU), multiprocessor, microcontroller(s) (uC), digital signal processor(s) (DSP), graphics processing unit(s) (GPU), application-specific integrated circuit(s) (ASIC), field-programmable gate array(s) (FPGA), etc., capable of executing software instructions stored in the memory 512. The memory 512 can be external to the processor 410, or internal to the processor 510. Memory as used in the present disclosure can be, for example, any combination of random access memory (RAM) and read only memory (ROM), or any other memory that is suitable for storing instructions for the processor 410. The memory 512 contains (i.e. stores) instructions that, when executed by the processor 510, cause the apparatus 500 to perform, for example, any embodiment of the method 400 previously disclosed herein. The apparatus 500 can further comprise one or more additional items 516, which in some cases are necessary for performing the method. In some embodiments, the apparatus 500 can be, for example, a surveillance camera, and the additional items 516 can then comprise, for example, an image sensor and, for example, one or more lenses for focusing light captured from a scene (in a direction in which the surveillance camera is aimed) on the image sensor to capture (i.e. describe) an input image stream of the scene. The additional items 516 can also comprise, for example, various other electronic components required for capturing the scene, for example, suitably operating the image sensor and / or the lenses. This allows the method to be performed in the surveillance camera itself (i.e. at the “edge”), which can reduce the need for any subsequent processing of the image stream output from the surveillance camera for privacy masking and the like. In other embodiments, for example, in a body-worn camera system (BWC system), the method assumed herein can be performed in a system controller or docking station (if there is sufficient processing capacity) rather than in the camera worn on the body.

[0103] If device 500 is to connect to a network (e.g., if device 400 is a network camera), device 500 may further include a network interface 514. The network interface 514 may be, for example, a wireless interface supporting, for example, Wi-Fi (as defined in, for example, IEEE 802.11 or a later standard), and / or a wired interface supporting, for example, Ethernet (as defined in, for example, IEEE 802.3 or a later standard). For example, a communication bus 518 may be provided to interconnect the various parts 510, 512, 514, and 516, enabling these parts to communicate with each other as needed to achieve desired functionality.

[0104] Figure 5B The reference is shown schematically. Figure 5A The described embodiment of the device 400 is shown, but is illustrated with, for example, references to... Figure 3 The set of functional blocks discussed corresponds to this. Device 500 includes a tracking module 320, a detection module 330, a shielding module 340, and a performance analysis module 350. Modules 320, 330, 340, and 350 are interconnected so that they can communicate with each other as needed (indicated by line 360). One, more, or all of modules 320, 330, 340, and 350 may be implemented, for example, using only software, only hardware, and / or a combination of software and hardware. Such software may be provided, for example, by instructions stored in memory 512. Each module 320, 330, 340, and 350 may be provided as a separate entity, or two or more or all of modules 320, 330, 340, and 350 may be provided as part of the same single entity.

[0105] This paper also provides one or more computer programs based on its assumptions. One such computer program, for example, could be used to perform the masking method 400 discussed and assumed in this paper in the output image stream, for use in reference... Figure 5A and Figure 5BThe described apparatus 500 performs such a method. A computer program may, for example, correspond to the instructions stored in the memory 512 of the apparatus 500, such that when the instructions are executed by the processor (or processing circuitry) 510, the corresponding method is performed by the apparatus 500. In other assumed embodiments, the computer program can be in a form that is not readable by the processor 510 and executable, but is provided as, for example, text specified in a programming language, which needs to be compiled into a format readable by the processor 510, for example by using a suitable compiler. The compiler can of course be performed by the processor 510 itself, or even form part of the processor 510 itself for real-time compilation. The assumptions herein provide one or more computer program products. Each such computer program product comprises a computer-readable storage medium having stored thereon one or more of the above-mentioned computer programs. For example, one computer program product can comprise a computer program for performing the assumed masking method in an output image stream as disclosed and discussed herein. The (computer-readable) storage medium (e.g. the “memory”) can be any combination of, for example, a random access memory (RAM) and a read only memory (ROM). In some embodiments, the computer-readable storage medium can be transitory (e.g. an electrical signal readable by a processor). In other embodiments, the computer-readable storage medium can be non-transitory (e.g. in the form of a non-volatile memory such as a hard disk drive (HDD), a solid-state drive (SSD), a secure digital (SD) card, etc., a USB flash drive, etc., any combination of such as a magnetic memory, an optical memory, a solid-state memory, or even a remotely installed memory). Other types of computer-readable storage media can also be assumed, as long as their functionality allows storing a computer program such that it is readable by a processor and / or an intermediate compiler.

[0106] As an overview of the various embodiments presented herein, the present disclosure provides improved methods of achieving reliable (privacy) masking in an image output stream, particularly in situations where an object detector and / or object tracker can not be able to correctly indicate and / or track one or more objects. By using not only immediate data from the tracker and / or detector, but also an estimate of the historical performance of the tracker and / or detector, the assumed masking approach provides a more reliable process in which the risk of failing to correctly mask an object in one or more images (i.e. image frames) of the output image stream is reduced or eliminated, even in more difficult conditions.

[0107] Although the above can describe features and elements in particular combinations, each feature or element can be used alone without the other features and elements or in various combinations with or without other features and elements. Furthermore, the disclosed embodiments can vary from the described assumptions, as will be appreciated by those skilled in the art, through study of the drawings, the disclosure and the appended claims.

[0108] In the claims, the word "comprising" does not exclude other elements and the indefinite article "a" does not exclude a plurality. Reference to a particular feature does not exclude that the feature is combined with another feature in the same claim.

[0109] List of reference signs

[0110] 100, 200 first image

[0111] 101, 201 second image

[0112] 120 tree

[0113] 122 sidewalk

[0114] 124 house

[0115] 126 door

[0116] 130-132 object (person)

[0117] 140-143 direction of movement

[0118] 150, 151 tracked object

[0119] 160-163 privacy shield

[0120] 228 car

[0121] 230-233 object (person)

[0122] 248 direction of movement

[0123] 250-253 tracked object

[0124] 260 privacy shield

[0125] 310, 312 input image stream, output image stream

[0126] 320, 330 tracking module, detection module

[0127] 322, 332 tracking data, detection data

[0128] 324, 334 tracking performance data, detection performance data

[0129] 333 detection data to tracker

[0130] 340 shielding module

[0131] 350 performance estimation module

[0132] 352 performance estimation data

[0133] 400, S401-S403 method, method steps

[0134] 500 device

[0135] 510 processor

[0136] 512 memory

[0137] 514 network interface

[0138] 516 add-on

[0139] 518 communication bus

[0140] 560 inter-module communication

Claims

1. A method (400) for masking in an output image stream, comprising: Receive (S401) the input image stream of the captured scene (310); Processing (S402) the input image stream to generate an output image stream (312) includes: based on the input image stream, using an object detection algorithm (330) and an object tracking algorithm (320) to detect and track one or more objects in the scene, the object tracking algorithm (320) receiving information (333) from the object detection algorithm indicating objects to be tracked, the processing further includes generating a specific output image of the output image stream, and performing one of the following: a) Determining that in one or more input images preceding the specific output image, the object detection algorithm has generated a probability that an object is located in a specific region of the scene, not exceeding a second threshold and exceeding a fourth threshold less than the second threshold, and the object detection algorithm has also generated a probability that an object is located in the specific region exceeding the second threshold in one or more additional input images preceding the one or more input images, the probability that an object is located in the specific region of the scene, not exceeding the second threshold and exceeding a fourth threshold less than the second threshold, indicates that the object detection algorithm is unsure whether any object exists to be masked, and the probability that an object is located in the specific region exceeding the second threshold indicates that the object detection algorithm determines that an object exists to be masked, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; b) Determine that in the input image prior to the specific output image, the object detection algorithm has: i) generated more frequently, rather than, the probability that an object is located in the specific region of the scene exceeding a second threshold and not exceeding the second threshold, and exceeding a fourth threshold less than the second threshold, in the specific region of the scene. The probability that an object is located in the specific region of the scene exceeding the second threshold indicates that the object detection algorithm determines whether there are any objects to be masked, and the probability that an object is located in the specific region of the scene exceeding the second threshold and exceeding a fourth threshold less than the second threshold indicates that the object detection algorithm is unsure whether there are any objects to be masked, or ii) compared to The probability that an object is located in a specific region exceeding a first threshold is more often generated as the probability that an object is located in the specific region of the scene not exceeding the first threshold and exceeding a third threshold less than the first threshold. The probability that an object is located in a specific region exceeding the first threshold indicates that the object detection algorithm determines whether there is an object to be tracked, and the probability that an object is located in the specific region of the scene not exceeding the first threshold and exceeding a third threshold less than the first threshold indicates that the object detection algorithm is uncertain whether there is any object to be tracked. The specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced. c) Determine that in one or more input images prior to the specific output image, the object tracking algorithm has stopped receiving information from the object detection algorithm indicating an object to be tracked in a specific region of the scene, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; d) Determining that in one or more input images preceding the specific output image, the object tracking algorithm has begun or resumed receiving information from the object detection algorithm indicating an object to be tracked in a specific region of the scene, but cannot yet begin or resume tracking the object to be tracked, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; and e) Determine that in one or more input images preceding the specific output image, the uncertainty in the position of an object tracked by the object tracking algorithm within the scene is considered too large to mask the object in a specific region of the scene, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the position of the object in the scene can be deduced. as well as In the specific output image, the specific area of ​​the scene is masked (S403).

2. The method of claim 1, comprising performing both b) and e).

3. The method of claim 1, comprising performing a), wherein, The determination of a) further includes determining that the rate at which the object detection algorithm becomes uncertain about whether there are any objects to be blocked has exceeded the fifth threshold.

4. The method according to claim 1, wherein, The determination of any one of a) to e) includes the use of a heatmap.

5. The method according to claim 1, wherein, In a) through e), the total number of the one or more input images preceding the specific output image is less than the total number of input images in the input image stream that precede the specific output image in time.

6. The method according to claim 1, wherein, The method is performed in a surveillance camera configured to capture the input image stream.

7. An apparatus (500) for masking in an output image stream, comprising: Processor (510), and A memory (512) stores instructions that, when executed by the processor, cause the device to: Receive (S401) the input image stream of the captured scene (310); Processing (S402) the input image stream to generate an output image stream (312) includes: based on the input image stream, using an object detection algorithm (330) and an object tracking algorithm (320) to detect and track one or more objects in the scene, wherein the object tracking algorithm (320) receives information (333) indicating objects to be tracked from the object detection algorithm, generates a specific output image of the output image stream, and performs one of the following: a) Determine whether, in one or more input images preceding the specific output image, the object detection algorithm has generated a probability that an object is located in a specific region of the scene at a rate not exceeding a second threshold and exceeding a fourth threshold less than the second threshold, and the object detection algorithm has also generated a probability that an object is located in the specific region exceeding the second threshold in one or more additional input images preceding the one or more input images. The probability that an object is located in the specific region of the scene at a rate not exceeding the second threshold and exceeding a fourth threshold less than the second threshold indicates that the object detection algorithm is unsure whether any object exists to be masked, and the probability that an object is located in the specific region exceeding the second threshold indicates that the object detection algorithm determines that an object exists to be masked, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; b) Determine whether, in the input image prior to the specific output image, the object detection algorithm has: i) generated a probability more frequently than the probability that an object is located in the specific region of the scene exceeding a second threshold and not exceeding the second threshold, and exceeding a fourth threshold less than the second threshold, in the specific region of the scene. The probability that an object is located in the specific region of the scene exceeding the second threshold indicates that the object detection algorithm determines whether there are any objects to be masked, and the probability that an object is located in the specific region of the scene exceeding the second threshold and exceeding a fourth threshold less than the second threshold indicates that the object detection algorithm is unsure whether there are any objects to be masked, or ii) compared to The probability of an object being located in a specific region exceeding a first threshold is more often generated as the probability of an object being located in the specific region of the scene not exceeding the first threshold and exceeding a third threshold less than the first threshold. The probability of an object being located in a specific region exceeding the first threshold indicates that the object detection algorithm determines whether there is an object to be tracked, and the probability of an object being located in the specific region of the scene not exceeding the first threshold and exceeding a third threshold less than the first threshold indicates that the object detection algorithm is uncertain whether there is any object to be tracked. The specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced. c) Determine whether, in one or more input images preceding the specific output image, the object tracking algorithm has stopped receiving information from the object detection algorithm indicating an object to be tracked in a specific region of the scene, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; d) Determine whether, in one or more input images preceding the specific output image, the object tracking algorithm has started or resumed receiving information from the object detection algorithm indicating an object to be tracked in a specific region of the scene, but cannot yet start or resume tracking the object to be tracked, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; and e) Determine whether, in one or more input images preceding the specific output image, the uncertainty in the position of an object tracked by the object tracking algorithm within the scene is considered too large to mask the object in a specific region of the scene, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the position of the object in the scene can be deduced, and In response to at least one of the determinations in a) to e) being affirmative, the specific region of the scene is masked (S403) in the specific output image.

8. The apparatus according to claim 7, wherein, The instructions are further configured to cause the device to perform the method (400) according to any one of claims 2 to 6.

9. The apparatus according to claim 7, wherein, The device is a surveillance camera configured to capture the input image stream.

10. A computer program product comprising a computer-readable storage medium storing a computer program for masking in an output image stream, the computer program being configured to, when executed by a processor (510) of a device (500), cause the device to: Receive (S401) the input image stream of the captured scene (310); Processing (S402) the input image stream to generate an output image stream (312) includes: Based on the input image stream, an object detection algorithm (330) and an object tracking algorithm (320) are used to detect and track one or more objects in the scene. The object tracking algorithm (320) receives information (333) indicating the objects to be tracked from the object detection algorithm, generates a specific output image of the output image stream, and performs at least one of the following: a) Determine whether, in one or more input images preceding the specific output image, the object detection algorithm has generated a probability that an object is located in a specific region of the scene at a rate not exceeding a second threshold and exceeding a fourth threshold less than the second threshold, and the object detection algorithm has also generated a probability that an object is located in the specific region exceeding the second threshold in one or more additional input images preceding the one or more input images. The probability that an object is located in the specific region of the scene at a rate not exceeding the second threshold and exceeding a fourth threshold less than the second threshold indicates that the object detection algorithm is unsure whether any object exists to be masked, and the probability that an object is located in the specific region exceeding the second threshold indicates that the object detection algorithm determines that an object exists to be masked, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; b) Determine whether the object detection algorithm has already been used in the input image preceding the specific output image: i) More frequently, a probability is generated that an object is located in the specific region of the scene exceeding a second threshold, rather than the probability that an object is located in the specific region of the scene exceeding a second threshold. The probability that an object is located in the specific region of the scene exceeding a second threshold indicates that the object detection algorithm determines whether there are any objects to be blocked, and the probability that an object is located in the specific region of the scene exceeding a second threshold and exceeding a fourth threshold indicates that the object detection algorithm is uncertain whether there are any objects to be blocked. Or ii) More frequently, a probability is generated that an object is located in the specific region of the scene exceeding a first threshold and exceeding a third threshold less than the first threshold, rather than the probability that an object is located in the specific region of the scene exceeding a first threshold. The probability that an object is located in the specific region of the scene exceeding a first threshold indicates that the object detection algorithm determines whether there are any objects to be tracked, and the probability that an object is located in the specific region of the scene exceeding a first threshold and exceeding a third threshold less than the first threshold indicates that the object detection is uncertain whether there are any objects to be tracked. The specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced. c) Determine whether, in one or more input images preceding the specific output image, the object tracking algorithm has stopped receiving information from the object detection algorithm indicating an object to be tracked in a specific region of the scene, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; d) Determine whether, in one or more input images preceding the specific output image, the object tracking algorithm has started or resumed receiving information from the object detection algorithm indicating an object to be tracked in a specific region of the scene, but cannot yet start or resume tracking the object to be tracked, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the location of the object in the scene can be deduced; and e) Determine whether, in one or more input images preceding the specific output image, the uncertainty in the position of an object tracked by the object tracking algorithm within the scene is considered too large to mask the object in a specific region of the scene, wherein the specific region of the scene is determined by one or more parameters generated by the object detection algorithm from which the position of the object in the scene can be deduced, and In response to at least one of the determinations in a) to e) being affirmative, the specific region of the scene is masked (S403) in the specific output image.

11. The computer program product according to claim 10, wherein, The computer program is further configured to cause the device to perform the method (400) according to any one of claims 2 to 6.

Citation Information

Patent Citations

  • Gradient privacy masks

    CN104660975A

  • Visual tracking method, visual tracking device, UAV(unmanned aerial vehicle) and terminal device

    CN108288281A