Improved masking of objects in image streams

By preserving and utilizing pixel information that would otherwise be discarded when generating the output image stream, the problem of detecting and labeling partially visible or suddenly appearing objects in the image stream is solved, improving the accuracy of object detection and the reliability of tracking, while meeting privacy protection requirements.

CN116778004BActive Publication Date: 2026-01-06AXIS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310232425.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-16
Filing Date
2023-03-10
Publication Date
2026-01-06
Estimated Expiration
2043-03-10

AI Technical Summary

Technical Problem

Existing object detection algorithms cannot correctly identify and locate partially visible or suddenly appearing objects in image streams, leading to reduced performance of object tracking algorithms, especially near the boundaries of image streams, where they cannot effectively mask or label these objects, potentially violating privacy protection rules.

Method used

By preserving and utilizing pixel information about the scene that would otherwise be discarded when generating the output image stream, especially pixel information about low-quality or non-rectangular regions, the object detection and tracking algorithm is assisted, ensuring that objects can be correctly detected and labeled in the output image stream.

Benefits of technology

It improves the accuracy of object detection and the reliability of tracking, ensuring that objects are correctly masked or marked in the image stream, reducing the risk of privacy leakage and meeting privacy protection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778004B_ABST
    Figure CN116778004B_ABST
Patent Text Reader

Abstract

The invention relates to improved masking of objects in image streams. There is provided a method of masking or marking objects in an image stream, comprising: generating one or more output image streams by processing an input image stream capturing a scene, including discarding pixel information about the scene provided by pixels of the input image stream such that the discarded pixel information about the scene is not included in any of the output image streams; and detecting an object in the scene using the discarded pixel information, wherein generating the one or more output image streams comprises masking or marking the detected object in at least one of the output image streams once it is determined that the object is at least partially visible in the at least one output image stream. There are also provided corresponding apparatuses, computer programs and computer program products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to masking or marking objects in an image stream. Specifically, this disclosure relates to how to improve such masking or marking when an object is only partially visible or appears abruptly in an image stream. Background Technology

[0002] Various object detection algorithms are available, which are typically capable of identifying and locating objects in an image. Identification may include, for example, informing the detected object of its category and determining whether it should be masked or labeled, while localization may include, for example, providing the object's coordinates and overall shape within the image. Information from the object detection algorithm can then be fed into an object tracking algorithm, allowing the movement of a particular object to be tracked across multiple images in, for example, an image stream (i.e., a video stream) captured by a camera. Knowing the object's location and / or movement can be useful, for example, when privacy masking needs to be applied to objects so that it is not easily identified by users viewing the image stream.

[0003] However, if an object is only partially visible in the image stream, for example, the object detection algorithm may fail to correctly identify and locate the object. This is especially true if the object is near the boundary of the image stream, such as when the object has just entered or is about to leave the scene captured by the image stream. If the object detection algorithm cannot correctly identify and locate such an object, it will not be able to provide information about the object to the object tracking algorithm. This can lead to a decrease in the performance of the object tracking algorithm for such objects, because the object tracking algorithm typically needs to receive at least several information updates from the object detection algorithm before it can lock onto and begin tracking the object.

[0004] As a result of the above, masking or marking of objects in an image stream becomes less reliable near, or for example, at, the boundaries of the image stream, where objects are more likely to be partially hidden and / or suddenly enter the field of view. This can be particularly problematic if masking is necessary for various privacy reasons, such as, because full compliance with existing privacy rules may no longer be guaranteed. Summary of the Invention

[0005] To at least partially address the aforementioned problem of masking (or, for example, marking) objects that are partially hidden or suddenly appearing in an image stream, this disclosure provides improved methods, improved apparatuses, improved computer programs, and improved computer program products for masking or marking objects in an image stream, as defined in the appended independent claims. Various alternative embodiments of the improved methods, apparatuses, computer programs, and computer program products are defined in the appended dependent claims.

[0006] According to a first aspect of this disclosure, a method for masking or marking objects in an image stream is provided. The method includes generating one or more output image streams by processing a first input image stream that captures at least a portion of a scene. The processing includes discarding pixel information about the scene provided by one or more pixels of the first input image stream, such that the discarded pixel information is not included in any of the one or more output image streams. The method also includes detecting objects in the scene using the discarded pixel information (i.e., pixel information not included in any of the one or more output image streams). Finally, generating one or more output image streams includes at least temporarily masking or marking the detected objects in at least one of the one or more output image streams, wherein the masking or marking is performed after determining that the object is at least partially located within at least one of the one or more output image streams (i.e., in response to this). The method may also optionally include outputting the generated one or more output image streams (e.g., outputting to a server, user terminal, storage device, etc.).

[0007] As will be explained in more detail later in this paper, the envisioned method improves upon currently available techniques, such as conventional processing chains, because it not only utilizes information about the scene included in the output image stream to determine the presence of an object of interest, but also uses information about the scene that might be discarded as part of some processing unrelated to the detected object, without further consideration. This allows object detection (algorithms) to detect objects better before they enter the output image stream. By doing so, object detection can succeed even if, for example, only half of the object is in the output image stream, and the risk of, for example, failing to properly mask such objects is thus reduced or even eliminated. This contrasts with commonly available techniques, in which pixel information about the scene is discarded before object detection (and tracking) is performed, allowing conventional object detection (and tracking) to use only information about the scene that also ends in the output image stream.

[0008] Therefore, the envisioned method is not about discarding pixel information about the scene while processing the input image stream, but rather about detecting such pixel information that has already been discarded (for whatever reason) as part of the processing of the (first) input image stream. Thus, the envisioned method may include an additional step of explicitly detecting the presence of pixel information about the scene such that if it is not used for object detection, this pixel information will be discarded, and then, after this detection, object detection continues based on this pixel information about the scene that has not yet been discarded. In this text, "based on" obviously means that object detection can be performed using pixel information that has not yet been discarded (but is about to be discarded), as well as pixel information that will not be discarded but is part of the output image stream. This can be applied, for example, when detecting an object currently entering the scene (depicted in the output image stream), and in this case, information about the portion of the object outside the output image stream is therefore only included in the pixel information that has not yet been discarded.

[0009] The processing of a first input image stream to generate one or more output image streams as used in this paper may include: processing a specific image (or image frame) of the first input image stream to generate a specific image (or image frame) of one of the one or more output image streams. The specific image of the first input image stream may be referred to as the specific input image, and the specific image of the output image stream in question may be referred to as the specific output image. As will be explained in more detail later in this paper, if, in addition to object detection, object tracking is used, pixel information discarded during the processing of the specific input image can be used when masking / marking objects in a specific output image (which describes a scene at the same time as the specific input image) and in one or more other output images (which describe a scene one or more times later than the specific input image), because the object tracking algorithm can infer the future location of objects based on the received information about objects in earlier images.

[0010] The pixel information used in this article is "scene-related" and should be understood as the pixels of an image (in the input image stream) recorded at a specific moment providing information about the scene at that specific moment. Imagine light rays reflected from, for example, objects in the scene, reaching the image sensor of a camera used to capture the scene, and when the image sensor reads out, providing the values ​​of the corresponding pixels in the image (input image stream), these pixels then provide information about the objects and the scene containing them. Pixel information provided by one or more pixels of the input image stream (the image) can be discarded, for example, if a pixel (as part of the processing of the input image stream) is cropped and not included in the output stream (the image), and / or, for example, if the value / color of a pixel is changed to a predefined value unrelated to its original value. The latter could be, for example, a case where masking is applied as part of the processing of the input image stream, such that one or more pixels, for example, initially describing some details of the scene, are instead forced to have, for example, pure black (or other) colors.

[0011] However, the precise value / color of a pixel as envisioned in this paper may differ from the original value found through readout from the image sensor, due to various functions such as color correction that have already been applied to the pixel. Nevertheless, such a pixel is still considered to provide pixel information about the scene, provided that the pixel is not cropped or masked as part of the processing (as described above), so that the pixel information about the scene it originally carried does not appear in any output image stream.

[0012] The statement that an object is (at least partially) "within the output image stream" should be understood as meaning that, assuming no initial masking of the object occurs, the object is (at least partially) located within at least one particular image of the output image stream.

[0013] As discussed earlier in this paper, “detecting objects in a scene by using discarded pixel information” can include, for example, using discarded pixel information as input to an object detection algorithm. This can correspond to, for example, providing the values ​​of one or more pixels that provide discarded pixel information as input to the object detection algorithm.

[0014] In one or more embodiments of the method, one or more pixels of the first input image stream may belong to a region of the first input image stream that is considered to have lower visual quality than one or more other regions of the first input image stream. Hereinafter, a “region of the first input image stream” is also contemplated as, for example, a region of an image (frame) of the first input image stream, such as a set of pixels of an image of the first input image stream. This region need not be a single, contiguous region, but can be divided into several sub-regions that do not overlap or are not directly adjacent to each other. As will be explained later herein, such regions may, for example, include pixels that provide a lower resolution for the scene and / or are distorted for other reasons. However, it is still assumed that these pixels contain at least some information about the scene, and that this information about the scene can therefore be used to detect objects, even if this information is otherwise discarded as part of the processing of the first input image stream. Conversely, in conventional processing chains, these pixels are typically cropped and / or masked without further consideration, making these lower-quality pixels invisible in one or more output image streams.

[0015] In one or more embodiments of the method, one or more pixels of the first input image stream deemed to have lower visual quality (i.e., one or more pixels of the image of the first input image stream) may originate from the peripheral region of the image sensor used to capture the first input image stream (the image), and / or from light arriving at the image sensor via the peripheral portion of the lens arrangement, and / or from light arriving at the lens arrangement at an angle sufficient to cause a so-called halo. All of these are examples of why some pixels acquire lower visual quality compared to pixels, for example, those originating from the center of the image sensor and / or from light arriving directly at the image sensor, and may result in, for example, lower visual quality in the form of increased blur, reduced sharpness, and / or reduced overall scene resolution in the areas where these pixels are found. Unlike discarding / cropping and / or masking these pixels without further consideration, as is done in conventional processing chains, this disclosure contemplates that these pixels may still contain pixel information about the scene, which could be useful for detecting objects in the scene.

[0016] In one or more embodiments of the method, the description of at least a portion of the scene in the first input image stream may be non-rectangular. One or more pixels of the first input image stream may belong to pixels located outside a rectangular region of the non-rectangular description (i.e., a rectangular region may be defined within the non-rectangular description, and one or more pixels may be pixels located outside that rectangular region). The non-rectangular description (of at least a portion of the scene) may result from at least one of the following: a) attributes of the shot arrangement that capture at least a portion of the scene, and b) transformations applied to the first input image stream (e.g., as part of processing the first input image stream to generate one or more output image streams). Herein, a scene description being non-rectangular means, for example, that some pixels in an image of the scene are, for example, black and do not contain any information about the scene, while the pixels describing the scene and containing information about the scene have a non-rectangular shape around their perimeter. When creating the output image stream, it may often be desirable to crop such an image so that the perimeter of the scene described in one or more output image streams becomes rectangular. In conventional processing chains, pixels located outside the desired rectangular shape are cropped and / or masked without further consideration. However, this disclosure envisions that these pixels may still contain information about the scene, which could be useful when detecting objects in the scene.

[0017] In one or more embodiments of the method, discarding pixel information about the scene may be due to the encoding standard requiring the output image stream to have a specific geometry, which necessitates discarding pixel information about the scene provided by one or more pixels located outside the specific geometry by, for example, cropping or masking these "outer pixels". The specific geometry is typically rectangular, but it is conceivable (or will exist) that the specific geometry is non-rectangular in the encoding standard, such as elliptical, circular, triangular, etc.

[0018] In one or more embodiments of the method, the above transformations may be applied as part of one or more of the following: b-1) a lens distortion correction process; b-2) a bird's-eye view of at least a portion of the scene, and b-3) applied to one or more projections stitching together one or more images, the one or more images capturing different portions of the scene, wherein at least one of these images is included in the first input image stream. The lens distortion correction process may be, for example, a barrel distortion correction (BDC) process, a pincushion distortion correction (PDC) process, or, for example, a tangential distortion correction (TDC) process. The synthesized bird's-eye view may include input image streams from multiple cameras, wherein no camera is actually positioned as having a top-down view of the scene, but various perspective transformations are applied to the respective image streams such that they are combined to simulate a view as if emitted from a camera positioned as having a top-down view of the scene. Stitching multiple images together to create, for example, a panoramic image (stream) may include first applying (mapping) one or more transformations of the form of a projection to each image (wherein, such projection can be, for example, linear, cylindrical, spherical, panoramic, stereoscopic, etc.). Projection can, for example, be used to project all images onto a common surface, such as a spherical or cylindrical surface.

[0019] In all the above-described types of transformations, the perimeter of the scene described in the transformed input image stream may no longer be rectangular, and conventional processing chains may discard / crop and / or mask pixels located outside the desired rectangular region without further consideration. However, this disclosure envisions that pixels outside the desired rectangular region can still provide pixel information about the scene, which can be used to detect objects in the scene.

[0020] In one or more embodiments of the method, generating an output image stream may include concatenating a first input image stream with a second input image stream, the second input image stream capturing scene portions not captured by the first input image stream. One or more pixels of the first input image stream may originate from scene regions captured not only by the first input image stream but also by the second input image stream. When information about a portion of the scene is found in both the first and second input image streams, it is conceivable to simply crop the overlapping pixels of the first input image stream that belong between the first and second input image streams, without performing, for example, blending or fading. Conventional processing chains would discard such “overlapping pixels” of the first input image stream without further consideration, while this disclosure contemplates that these pixels may contain pixel information about the scene, which can be used to detect objects in the scene.

[0021] In one or more embodiments of the method, generating an output image stream may include applying an electronic image stabilization (EIS) process to a first input stream. One or more pixels of the first input image stream may belong to pixels that are cropped or masked by EIS. The reasons why EIS can crop or mask these pixels will be detailed later herein, but it should be noted now that conventional processing chains would discard the information provided by these pixels without further consideration, while this disclosure envisions that these pixels can provide pixel information about the scene, which can be used to detect objects in the scene.

[0022] In one or more embodiments of the method, the method can be executed in a surveillance camera. The surveillance camera may, for example, be part of a surveillance camera system. The surveillance camera can be configured to capture a first input image stream. Executing the method in a surveillance camera can allow so-called "edge computing," where, for example, a server or user terminal receiving the output image stream from the surveillance camera may not require further processing, and where, for example, latency can be reduced because the computation is performed as close as possible to the monitored scene.

[0023] In one or more embodiments of the method, determining that an object is at least partially within the output image stream may include using an object tracking algorithm. For example, when the object is still not within the output image stream, discarded pixel information may be used to detect the object, but when the object moves and enters the output image stream, an object tracking algorithm is used. Therefore, in the method contemplated in this disclosure, using discarded pixel information may be useful even if the object is not yet in the output image stream, and even if the discarded pixel information itself never ends in any output image stream.

[0024] According to a second aspect of this disclosure, an apparatus for masking or marking objects in an image stream is provided. The apparatus includes a processor (or "processing circuitry") and a memory. The memory stores instructions that, when executed by the processor, cause the apparatus to perform the method of the first aspect. In other words, the instructions are such that, (when executed by the processor) they cause the apparatus to: generate one or more output image streams by processing a first input image stream that captures at least a portion of a scene, wherein the processing includes discarding pixel information about the scene provided by one or more pixels of the first input image stream, such that the discarded pixel information is not included in any of the one or more output image streams; detect objects in the scene using the discarded pixel information not included in any of the one or more output image streams; and detect objects in the scene using the discarded pixel information not included in any of the one or more output image streams. Generating one or more output image streams includes: after determining that an object is at least partially located within at least one of the one or more output image streams, at least temporarily masking or marking the detected object in at least one of the one or more output image streams. The apparatus may also optionally be configured to output the output image streams (to, for example, a server, a user terminal, a storage device, etc.).

[0025] In one or more embodiments of the device, the instructions may cause them (when executed by a processor) to cause the device to perform any embodiment of the methods disclosed and contemplated herein in the first aspect.

[0026] In one or more embodiments of the device, the device may be, for example, a surveillance camera. The surveillance camera may be configured to capture a first input image stream. The surveillance camera may, for example, be part of a surveillance camera system.

[0027] According to a third aspect of this disclosure, a computer program for object detection in an image stream is provided. The computer program is configured, when executed by a processor of, for example, a device (wherein the device may be, for example, the device according to the second aspect), to cause the device to perform the method according to the first aspect. In other words, the computer program is configured to, when executed by a processor of the device, cause the device to: generate one or more output image streams by processing a first input image stream capturing at least a portion of a scene, wherein the processing includes discarding pixel information about the scene provided by one or more pixels of the first input image stream, such that the discarded pixel information is not included in any of the one or more output image streams; detect objects in the scene using the discarded pixel information not included in any of the one or more output image streams; and detect objects in the scene using the discarded pixel information not included in any of the one or more output image streams. Generating one or more output image streams includes: after determining that an object is at least partially located within at least one of the one or more output image streams, at least temporarily masking or marking the detected object in at least one of the one or more output image streams.

[0028] In one or more embodiments of the computer program, the instructions also cause them (when executed by a processor) to cause the apparatus to perform any embodiment of the methods disclosed and contemplated herein.

[0029] According to a fourth aspect of this disclosure, a computer program product is provided. The computer program product includes a computer-readable storage medium having the computer program according to the third aspect stored thereon.

[0030] Other objects and advantages of this disclosure will become apparent from the following detailed description, drawings, and claims. Within the scope of this disclosure, it is contemplated that all features and advantages described with reference to, for example, the method of the first aspect, are related to, applicable to, and can be used in combination with any features and advantages described with reference to the apparatus of the second aspect, the computer program of the third aspect, and / or the computer program product of the fourth aspect, and vice versa. Attached Figure Description

[0031] Exemplary embodiments will now be described with reference to the accompanying drawings, in which:

[0032] Figure 1 A functional block diagram illustrating a conventional processing chain / method for masking or marking objects in an image stream is shown;

[0033] Figure 2A Functional block diagrams are shown that schematically illustrate various embodiments of a method for masking or marking objects in an image stream according to the present disclosure;

[0034] Figure 2BA flowchart illustrating one or more embodiments of the method according to this disclosure is shown schematically;

[0035] Figures 3A to 3E Various example situations are illustrated schematically, leading to the discarding of pixel information about the scene, as contemplated in one or more embodiments to which the methods of this disclosure also apply.

[0036] Figure 4A and Figure 4B One or more embodiments of the apparatus according to this disclosure are illustrated schematically.

[0037] In the accompanying drawings, the same reference numerals will be used for the same elements unless otherwise stated. The drawings show only elements necessary to illustrate exemplary embodiments; other elements may be omitted or only suggested for clarity. As shown, for illustrative purposes, the (absolute or relative) dimensions of elements and regions may be exaggerated or underestimated relative to their true values; therefore, they are provided to illustrate the general structure of the embodiments. Detailed Implementation

[0038] In this paper, it is envisioned that object detection can be implemented using one or more commonly available algorithms, which are available in various fields of computer technology, such as computer vision and / or image processing. Such algorithms can be envisioned, for example, including non-neural and neural methods. However, the minimum requirement is that, regardless of the algorithm (or combination of algorithms) used, it is possible to determine the presence of a specific object (e.g., a face, body, license plate, etc.) in an image, specifically its location and / or region within the image. As long as the above requirement is met, it is not important whether the algorithm used is feature-based, template-based, and / or motion-based. For example, object detection can be implemented using one or more neural networks specifically trained for this purpose. For the purposes of this disclosure, it is also assumed that the algorithm used in object detection may have difficulty correctly identifying and / or locating objects partially hidden within an image, where object detection is assumed to be performed, for example when / if a person is partially occluded by trees, vehicles, etc.

[0039] Similarly, in this paper, object tracking is envisioned as being implemented using, for example, one or more commonly available object tracking algorithms. Such algorithms could be, for example, bottom-up processes dependent on target representation and localization, and include, for example, kernel-based tracking, contour tracking, etc. Other envisioned tracking algorithms could be, for example, top-down processes, including, for example, using filtering and data association, and implementing, for example, one or more Kalman and / or particle filters. In this paper, such a tracking algorithm is envisioned to receive input from object detection and, without providing further input / updates from object detection, to use the received input to track objects in the scene over time (i.e., across multiple subsequent images). For the purposes of this disclosure, it is assumed that even if object tracking is able to track / follow an object for at least a few images / frames of the image stream after receiving updates from object detection stops, the quality of such tracking will degrade over time because no new input from detection arrives. After a period of time, tracking will fail to properly track the object. It is also assumed that tracking requires some time after receiving the first input / update from detection before locking onto the object and performing successful tracking. In other words, tracking requires more than a single data point from detection to draw conclusions about where / where the object will be next (because it is difficult, if not impossible, to make a proper inference from a single data point). In the following text, the terms "object detection algorithm," "object detection," "object detection module," and "detector" are used interchangeably. The same applies to the terms "object tracking algorithm," "object tracking," "object tracking module," and "tracker," which are also used interchangeably.

[0040] Now refer to Figure 1 A more detailed description of examples of traditional processing chains / methods for masking or marking objects in an image stream.

[0041] Figure 1 A functional block diagram 100 is shown, schematically illustrating the flow of a method 100 found and used in the prior art. In method 100, an input image stream 110 is received from an image sensor, for example, capturing a scene (using various lens arrangements, for example, as part of a camera aimed at the scene). The input image stream 110 includes an input image I n (where index n indicates a specific input image I) n (Captured at the nth time). The input image stream 110 is received by the image processing module 120, which processes the input image stream 110 and the input image I. n Perform image processing to generate the corresponding output image O at the end of processing chain 100. n As part of the output image stream 114.

[0042] Image processing performed in module 120 produces an intermediate image I′.n For various reasons, as will be detailed later in this article, image processing results in the discarding of images from the input image I. n One or more pixels provide pixel information about the scene. Therefore, in the intermediate image I′ n In this process, one or more pixels 113 are cropped or masked, such that the discarded pixel information about the scene provided by one or more pixels 113 will not be included in any output image O of the output image stream 114. n In the middle. As a result, due to cropping and / or masking of one or more pixels 113, the intermediate image I′ n The region 112 describing the scene may be smaller than the original image I′. n The region in the image. Of course, in other examples, image processing may include zooming / scaling the region 112 after cropping or masking pixel 113, so that the intermediate image I′ n and / or output image O n The size of the part 112 described in the scene is the same as the size of the entire input image I. n The same (or even larger) size. In any case, the intermediate image I′ n Region 112 will not provide the same information as the entire input image I. n The same amount of pixel information about the scene. This is because, for example, even if the area is zoomed / scaled by 112 (using, for example, upsampling), the discarded pixel information about the scene cannot be recovered.

[0043] For the reasons mentioned above, in the traditional method 100, only the intermediate image I′ is used. n The pixel information about the scene provided by the pixels in the remaining region 112 can be used by any subsequent functional module in the flow of method 100. As previously mentioned, as a result of the image processing performed in module 120, the pixel information about the scene initially provided by one or more pixels 113 is therefore discarded. Pixel information discarded by cropping or masking pixels is typically performed early in the processing chain to avoid spending computational resources on further processing or analysis of any pixels that will later be cropped or masked in any way.

[0044] Then, the intermediate image I′ nThe data is passed to object detection module 130, where an attempt is made to identify and locate one or more objects in region 112. The result of object detection performed in module 130 is passed as object detection data 132 to feature addition module 150, where, for example, if the detected object is identified as an object to be masked, a privacy mask is applied to the detected object. Feature addition module 150 may also, or alternatively, add visual indications of the identity and / or location of the detected object (shown as, for example, a frame surrounding the detected object), or any other type of marker, to the output image stream 114. Thus, feature addition module 150, by adding data to the intermediate image I′, [further details about the feature addition module]. n Add one or more such features to modify it, and use the result as the output image O. n Output. Output image O n This forms part of the output image stream 114. The object detection data 132 may include, for example, the location of the identified object, an estimate of how much the object detection module 230 determines that the identified object is an object to be masked (or at least tracked).

[0045] Optionally, the intermediate image I′ n It can also be provided, for example, to an optional object tracking module 140. The object tracking module 140 can also (additionally or alternatively) receive object detection data 133 from the object detection module 130, the object detection data 133 indicating, for example, in the intermediate image I′ n The estimated location of the detected object, and an indication, for example, that the object should be tracked by the object tracking module 140. This can help the object tracking module 140 track objects across several images. For example, the object tracking module 140 can use the intermediate image I′ n The acquired object detection data 133 can also be used as one or more previous intermediate images I′ m<n The acquired object detection data 133 is used to track objects over time (in other words, for intermediate image I′). n The acquired object detection data 133 can be used to track objects, allowing them to be detected in the subsequent output image O. k>n(The image is masked). The result of this tracking can be provided from the object tracking module 140 to the feature addition module 150 as object tracking data 142. Object tracking data 142 may include, for example, the position, shape, etc. of the object estimated by the object tracking module 140, and, for example, the estimation uncertainty of the object position determined by this. The feature addition module 150 may also use such object tracking data 142 when, for example, one or more masks, markers, etc. are applied to one or more objects. In some other examples, it is conceivable that the feature addition module 150 receives data 142 only from the object tracking module 140 and not from the object detection module 130. In such an example, method 100 then relies solely on object tracking to output image O. n Add, for example, masking, marking, etc., to one or more objects in the output image stream 114. As mentioned earlier, when using object tracking, the intermediate image I′ n The data that can be used for tracking can therefore also be output later as one or more subsequent output images O in the output image stream 114. k>n Use when needed.

[0046] However, as discussed earlier in this paper, the object detection module 130 may often struggle with, or even be unable to, properly identify and locate objects, for example, those only partially located in the intermediate image I′. n Objects within region 112. Therefore, object detection may fail to locate specific objects, object tracking module 140 (if used) may become unable to track such objects over time if it stops receiving further object detection data 133 from object detection module 130, and feature addition module 150 may fail to properly mask all objects in the scene, for which, for example, identity should be protected by privacy masking, because data 132 and / or 142 may no longer be available or no longer accurate enough.

[0047] Now refer to Figure 2A and Figure 2B The method envisioned in this disclosure will be explained in more detail how it provides an improvement over conventional method 100.

[0048] Figure 2A A functional block diagram 200 is shown, which schematically illustrates the interaction of various functional modules for performing various embodiments of the method 200 according to the present disclosure. Figure 2B Schematic flowcharts of these embodiments of method 200 are shown.

[0049] In step S201, one or more input image streams are received from one or more image sensors, such as one or more cameras (e.g., surveillance cameras). 210. In this paper, j is the index of the j-th such input image stream, and n is the temporal index such that This represents the input image of the j-th input image stream captured at time n. The "n"-th time can, for example, correspond to time t0 + Δ × n, where t0 is some start time and Δ is the time difference between each captured image (assuming that the time difference Δ is equal for all input image streams and for all input images in each input image stream). At least one input image. Received by image processing module 220 at the end of the processing chain, and via one or more intermediate images Generate and output one or more output image streams 214, of which, The output image of the i-th such output image stream describes at least a portion of the scene at time n. In the following text, for simplicity only, unless otherwise stated, it will be assumed that there is only a single input image stream S. in ={…,I n-1 I n I n+1 , ...}, each input image I n A single intermediate image I′ n and the single generated output image stream S out ={…,O n-1 O n O n+1 , ...}.

[0050] As in the reference Figure 1 In the conventional method 100 described, in order to generate the output image O of the output image stream 214 n Due to various reasons unrelated to, for example, object detection and / or object tracking, the image processing performed by module 220 (in step S202) results in the discarding of pixel information about the scene provided by one or more pixels of the input image of a certain input image stream 210. Since one or more pixels 213 providing the discarded pixel information about the scene have been, for example, cropped or masked, the intermediate image I′ generated by the image processing in module 220... n It also has an occupancy ratio input image I n The remaining description of the scene in small region 212. As previously mentioned, image processing performed in module 220 may also include scaling region 212 so that it again matches (or even exceeds) the input image I. n The entire region (however, it does not reacquire any information about the scene lost during cropping or masking of one or more pixels 213). Therefore, also in the improved method 200, the information is initially obtained from the input image I. nThe pixel information about the scene provided by one or more pixels 213 found in the input image stream 210 is not included in any corresponding output image. The output image stream is 214. It should be noted that in this paper, the output image and the input image have the same time index n, meaning both images represent the scene at the same moment, not necessarily the output image O. n With input image I n Simultaneously generated. For example, due to the completion of input image I. n Due to factors such as the processing time required, the capture of the input image I... n and the final output image O n There may be some delays.

[0051] As envisioned in this paper, when discarded pixel information about the scene is referred to as not included in any output image and output image stream 214, this means for a specific input image of a particular j-th input image stream. The processing is performed, and such an output image does not exist (it does not exist in any output image stream). It includes images from that specific input image. It obtains pixel information about the scene from one or more masked or cropped pixels. Of course, there can be situations where the output image... The pixel information about the scene is still present, but in this case, it has been obtained from the processing of another input image of, for example, another input image stream. For instance, as will be discussed in more detail later in this document, several input image streams can be combined to create a panoramic view of the scene, and if they are captured such that they overlap spatially, some parts of the scene can therefore be seen in multiple images of these multiple input image streams. However, if one or more pixels of a particular input image of a particular such input image stream are cropped or masked during the processing of that particular input image, the pixel information about the scene provided by those pixels is not considered part of the output image stream, even if other pixels in one or more other input images of one or more other input image streams happen to provide the same or similar pixel information about the scene. In other words, the discarded pixel information about the scene provided by one or more pixels of a particular input image of a particular input image stream, as envisioned herein, is not included in any output image of any output image stream because one or more pixels are cropped or masked as a result of the processing of the particular input image and the input image stream. In particular, as envisioned in this paper, pixel information about a scene provided by a specific input image that captures a portion of the scene at a particular moment is not provided in any image of any output image stream that describes a portion of the scene, because it is at that particular moment.

[0052] From input image I n The intermediate image I′ generated by the processing n Provided to object detection module 230, which can use the remaining pixels of region 212 and pixel information about the scene provided by these pixels to detect objects in the scene (in step S203). However, compared with the reference... Figure 1 Compared to the conventional method 100 described, the improved method 200's processing chain does not simply discard the pixel information about the scene provided by one or more (cropped or masked) pixels 113 without further consideration. Instead, and most importantly, in method 200, one or more pixels 213 and the pixel information about the scene they provide are also provided to the object detection module 230, so that the object detection module 230 can also use this pixel information about the scene when detecting objects (i.e., in step S203). As a result, the entire input image I provides the object detection module 230 with information about the scene. n The information, not just in the input image I. n The remaining region 212 after processing. This allows the object detection module 230 to detect objects within region 212 and in the output image O in the output image stream 214. n Only partially visible objects are represented in the input image I, because the objects are not present in the complete input image I. n (That is, the object remains fully visible (or at least sufficiently visible to be properly identified and / or located by object detection) within region 212 plus one or more pixels 213.) This will provide improved, more reliable performance for object detection, making the object more likely to appear in, for example, the output image O in the output image stream 214. n It is appropriately masked or marked.

[0053] If object tracking is used, the entire input image I is used. n The content used to detect objects can also help, for example, in the output image O of the output image stream 214. n One or more subsequent output images O k>n Tracking objects in (not shown). For example, in some embodiments, the object detection module 230 is envisioned analyzing the entire input image I. n At that time, objects visible within one or more pixels 213 but not within region 212 are identified. Therefore, the object is displayed in the output image O. n The middle part will be invisible, and therefore in the output image O n The object is neither masked nor labeled. However, the object detection data 233 is still sent to the object tracking module 240, allowing the module to track the object in, for example, a subsequent image I′. k>nThe object detection module 240 begins tracking objects in the scene before they become visible in region 212 (if it hasn't started yet). When an object becomes visible in region 212, the object detection module 240 will then be already tracking the object, and once the object enters region 212, it can be included in the subsequent output image O. k>m The middle is appropriately blocked. For example... Figure 2A As shown, one or more pixels 213 may of course be optionally provided to the object tracking module 240, not just the object detection module 230, which may be useful if the object tracking module 240 uses pixel information to track objects, rather than, for example, only updating the position of the objects provided by the object detection module 230.

[0054] It is also envisioned that in some embodiments, not all discarded pixels need to be used for object detection and / or tracking. For example, some discarded pixels may belong to a part of a scene that is statistically known not to contain an object of interest (as captured in the input image). It can then be determined not to submit the pixel information about the scene provided by such pixels to, for example, an object detection module, even if the corresponding pixels are masked and / or cropped. This part of the scene may, for example, correspond to the sky or any other area of ​​the scene where the object of interest is unlikely to appear (if the object is, for example, a person, a car, etc.). In other cases, the opposite is of course true, and the area where the object is unlikely to appear may be, for example, a street, a field, etc. (if the object of interest is, for example, an airplane, a drone, or other flying object). By not providing such pixel information to, for example, an object detection algorithm, object detection resources can be avoided in areas that are known (or at least generally known) not to require object detection resources. Whether to use all or only some discarded pixels (and the pixel information they provide about the scene) for detection and / or tracking can be predetermined, for example, by manually indicating one or more regions of the scene as "not of interest" or "of interest," or, for example, by using collected historical data showing regions where the object of interest was previously detected or tracked and regions where it was not. Other types of scene or image analysis can, of course, also be used to determine which regions of the scene are not of interest and which are of interest.

[0055] intermediate image I′ n (wherein, region 212 indicates how the output image O will be...) n The intermediate image I′ is provided to the feature addition module 250 along with object detection data 232 from object detection module 230 and object tracking data 242 from object tracking module 240 (if included). The feature addition module 250 can then (in step S205) add one or more (privacy) masks or tags to the intermediate image I′. n (or on top of it) to output image On One or more detected and / or tracked objects are masked (or marked) before being included in the output image stream 214. (See reference...) Figure 1 In the described conventional method 100, in some embodiments of method 200, the object tracking module 240 may be optional, and if used / available, it may be provided with object detection data 233 from the object detection module 230 to help the object tracking module 240 achieve proper object tracking. Similarly, in some other embodiments, if the object tracking module 240 is included, the feature addition module 250 may rely solely on the object tracking data 242 without relying on the object detection data 232 to know where to add masks or markers (in this case, the object detection data 232 may optionally be sent to the feature addition module 250). For the output image O n The masking or marking of objects in the output image stream 214 can be temporary, for example, indicating that once the object is no longer in the output image O. n In the scenario described in output image stream 214, or for example if an object becomes hidden behind another object in the scene (making it no longer detectable and / or trackable, or making masking no longer necessary), the applied masking or marking can be removed. In other envisioned cases, if an object changes its orientation within the scene, making it no longer possible to infer the object's identity even without masking, the masking of the object can be removed, for example. For instance, if a person's orientation causes their face to face the camera, masking the person might be necessary, while if the person changes their orientation so that their face is away from the camera, masking the person might no longer be necessary. Masking can be removed only if, first determined (in step S204), the object is at least partially present in the output image O. n After the output image stream 214, in the output image O n The object is masked or marked in the output image stream. This determination can be made by, for example, the object detection module 230 and / or the object tracking module 240, and then transmitted to the feature addition module 250. It is also envisioned that such determination can be made by the feature addition module 250 itself, or by having sufficient information to determine whether or at least partially an object is in the output image O. n And any other module within the output image stream 214.

[0056] As previously mentioned, determining whether an object is "in the scene" includes considering whether the object's position in the scene ensures that, without masking, it is present in the output image O. n It is at least partially visible in the output image stream 214.

[0057] Generally speaking, it should be noted that, although in Figure 2AThe diagram illustrates object detection and object tracking performed by different entities / modules, but it's also possible for both functions to be provided by a combined object detection and tracking function, and implemented within, for example, the same combined object detection and tracking module. This can vary depending on the exact algorithm used for object detection and tracking, as some algorithms may track objects based, for example, on their own observations of object locations in some images. In general, it's also possible for the following to be... Figure 2A The two, three, or even all modules 220, 230, 240, and 250 shown are implemented as a single module, or at least implemented as a single module. Figure 2A The number of modules shown is smaller. This is because the various modules do not necessarily represent physical modules, but can also be implemented using only software, or, for example, as a combination of physical hardware and software, as described later in this article.

[0058] As Figure 2A and Figure 2B In addition to the summary of the improved method 200 envisioned herein, it should be noted that by using pixel information about the scene provided by one or more pixels 213 that would otherwise be discarded for other reasons (e.g., by cropping or masking one or more pixels 213), when the object is at least within one or more pixels 213 (i.e., when the object is at least partially within the output image O), n Prior to the scene section described in output image stream 214, object detection (and / or tracking) may have already acquired information about the objects. This means that when (or if) the object later enters the subsequent output image O... k>n When the visible area 212 is reached, the location of the object to be masked (or marked) may already be known. This is in contrast to commonly available methods, where pixel information about the scene provided by one or more pixels 113 is never provided to any object detection and / or tracking function, and where such object detection and / or tracking is therefore unaware of any information about objects included within one or more pixels 113. Therefore, while commonly available methods face a higher risk of failure, such as failing to correctly identify, locate, track, and mask / mark objects that suddenly appear in the output image stream, the method 200 contemplated according to this disclosure can be performed better because it enables the use of discarded pixel information about the scene, even if that pixel information is still not included in any output image and output image stream.

[0059] Now refer to Figures 3A to 3E The image processing module 220 can select various examples of hypothetical situations in which certain pixels can be discarded (i.e., cropped or masked) from the input image and the input image stream.

[0060] Figure 3A This schematically illustrates a situation where electronic image stabilization (EIS) results in the cropping of certain pixels. Input image I n Includes two objects, human 320a and human 320b. Previous input image I n-1 It's the same scene, and from... Figure 3A As can be seen from this, the camera captures image I n-1 and I n The movement is between [the two points]. Assume this movement is due to camera shake. For example... Figure 3A As shown, the detected motion corresponds to the translation vector d. Using image processing 310, camera shake can be corrected by simply applying the inverse translation -d, so that image I... n People 320a and 320b will appear in the previous image I n-1 The middle overlaps with themselves (as shown by the dashed outline). However, in order to keep the size of the scene depicted in the output image constant, this inverse translation will produce the intermediate image I′. n , where the input image I n Multiple pixels 213 have been cropped to keep people 320a and 320b stationary in the scene, and the remaining visible area 212 of the intermediate image is therefore smaller than that of the input image I. n The visible area.

[0061] In commonly available methods for masking or labeling objects in an image stream, the object detection algorithm will be provided solely from the intermediate image I′. n The pixels in region 212 provide pixel information about the scene. Since person 320a remains completely within region 212, the object detection algorithm has no problem locating person 320a and identifying it as an object to be masked or marked, for example. However, since person 320b is now only partially visible in region 212, the same algorithm may have difficulty correctly identifying and locating person 320b. Using the method envisioned in this paper, this will not be a problem, because the object detection algorithm will also be provided with pixel information about the scene from the discarded pixels 213. Since person 320b is fully visible in the joint set of pixels in region 212 and pixels 213, the object detection algorithm will also successfully locate and identify person 320b.

[0062] Figure 3B The following is schematically illustrated, wherein the lens arrangement for capturing the input image stream is such that the input image I... n This is subject to so-called barrel distortion. In reality, the scene being described has a rectangular grid 321a in the background (e.g., a brick wall), and includes objects 321b that should be masked or marked in the foreground. Due to barrel distortion caused by the lens arrangement, the lines of the grid 321a appear as barrel distortion in the input image I. nThe image is no longer straight. Various algorithms exist to correct this barrel distortion, which can be used as input images I. n As part of image processing 311, to create intermediate image I′ n In this case, the lines of grid 321a are again straight. However, a side effect of this algorithm is that while straightening the lines of grid 321a, it also distorts the outer shape of the described scene. For example, as... Figure 3B As shown, input image I n The scene is described as having a rectangular perimeter, but the lines of grid 321a are not straight everywhere. After image processing 311, the lines of grid 321a are straight again, but the perimeter of the described scene is now distorted, making it appear as if the corners of the previously rectangular perimeter of the scene are now pulled radially outward. It is generally undesirable to present this non-rectangular description of the scene to the user, therefore image processing 311 further includes the step of cropping (or masking) all pixels located outside the rectangular region 212 after barrel distortion correction. Figure 3B In this context, these cropped or masked pixels are pixel 213. Therefore, due to barrel distortion correction and subsequent adjustments to the input image I... n Crop / mask one or more pixels and apply to the input image I n The result of the transformation, processing 311, causes the pixel information about the scene provided by pixel 213 to be discarded and not included in the intermediate image I′. n The generated output image O n middle.

[0063] from Figure 3B It can also be seen that object 321b is completely located within the input image I. n Within the scene described, but ultimately only partially located in the intermediate image I′ n The object detection algorithm is located within the rectangular region 212. Therefore, as with conventional methods, an object detection algorithm that relies solely on pixel information about the scene provided by pixels in region 212 may fail to correctly identify and / or locate object 321b, and object 321b is unlikely to be properly masked or labeled in the output image O and the output image stream. However, in the improved method 200 envisioned herein, the object detection algorithm is also provided with pixel information about the scene provided by pixel 213, and therefore has a greater chance of properly identifying and locating object 321b, since object 321b is entirely located within the joint set of pixels in region 212 and pixel 213. As envisioned herein, other types of distortion (such as pincushion distortion, tangential distortion, etc.) may also necessitate the transformation of the input image I. nTo correct this distortion, the algorithm used for correction may result in the periphery of the scene being described after correction being non-rectangular, and one or more pixels 213 thus appearing outside the desired rectangular region 212 in the intermediate image I′. n The content is discarded (i.e., clipped or masked). The hypothetical method described in this article also applies to this alternative scenario.

[0064] Figure 3C This schematically illustrates another scenario where pixel information provided by one or more pixels of the input image stream is ultimately discarded due to transformations applied to the first input image stream. In this paper, it is envisioned that the input image stream is formed from image data recorded by several cameras aimed in different directions, and image I... n Therefore, a panoramic view of the scene, created by stitching together image data from several cameras, is part of the early stage of processing. This early stage of processing the image stream involves applying various transformations to different (mapped) perspective projections, such as stitching these projected images together to form the input image I. n Previously, images from several cameras were projected onto a common surface, such as a cylinder or sphere. As a result of this perspective projection, the input image I... n The perimeter of the scene 322a described in the text is non-rectangular.

[0065] The scene also includes objects 322b and 322c, both of which are in the input image I. n Within the scene 322a described in the image. However, presenting a non-rectangular description of such a scene is undesirable, and another part 312 of the processing of the input image stream therefore includes cropping (or at least masking) the intermediate image I′ produced by such processing. n Input image I outside the rectangular region 212 n All pixels 213. Therefore, object 322b will ultimately be completely outside region 212, while object 322c will only be partially within region 212. As mentioned earlier, conventional methods only provide the object detection algorithm with pixel information about the scene provided by the pixels within region 212, which will fail to identify and locate object 322b, and is also likely to fail or at least have difficulty correctly identifying and locating object 322c. Therefore, for the output image O nProperly masking or marking objects 322c in the output image stream may fail, and conventional processes are not prepared to directly mask or mark objects 322b if (or whenever) they partially enter region 212 in a subsequent output image of the output image stream, because the tracker (if used) has not yet been provided with any prior location of objects 322b. The improved method also provides pixel information about the scene provided by one or more pixels 213 to the object detection algorithm, and thus manages to locate and identify objects 322b even if they are outside region 212, and also helps to better locate and identify objects 322c, even if the object is only partially located in the region that will constitute the output image O. n The content is within region 212. Furthermore, if tracking is used, once an object is detected in pixel 213, the position of object 322b can be provided to the object tracking algorithm, and thus once object 322b becomes at least partially visible in region 212, masking or marking of object 322b can be performed in subsequent output images of the output image stream.

[0066] Figure 3D The following schematically illustrates how the light path passes through the lens arrangement before reaching the image sensor and how light hits the image sensor, potentially leading to the input image I. n One or more pixels are considered to have a higher resolution than the input image I. n The input image stream has lower visual quality than other pixels. This paper uses a fisheye lens to capture the input image I. n And the input image stream, for example, a lens that captures light from multiple directions and provides a wider field of view compared to a more conventional lens with a longer focal length, thus allowing a larger portion of the scene to be captured in a single image. The scene is a parking lot and includes, as in the input image I... n The vehicle 323a described herein moves toward the center of the scene. However, due to the characteristics of this lens arrangement, light emitted from outside the central region 323b of the scene has already passed through the periphery of the lens arrangement, and in particular at a larger angle compared to light emitted from inside the central region 323b of the scene. This may result in lower resolution of the recorded scene in the portion of the image sensor where this (peripheral) light illuminates (most typically in one or more peripheral regions of the image sensor). It may also result in fewer photons hitting these areas of the sensor, which may lead to so-called halos, where, in the input image I... n The farther away from the center of the scene, the darker it appears. To avoid presenting this lower quality of detail to users viewing the output image stream, pixels outside the circular region 323b can be considered less detailed than those belonging to the input image I. nPixels in one or more other regions of the input image stream (such as pixels within circular region 323b) have lower visual quality. As input image I n As part of the processing of the input image stream 313, one or more pixels 213 outside region 212 can therefore be used to form the output image O. n The intermediate image I′ is the result of the output image stream. n The image is cropped or at least masked, and the pixel information about the scene provided by pixel 213 is therefore discarded and not included in any output image or output image stream. Therefore, when car 323a is in the intermediate image I′ n When outside region 212, conventional methods will not be able to detect vehicle 323a, and when (or if) vehicle 323a later enters the subsequent intermediate image I′ k>n When the object tracking algorithm (if used) is in region 212 of the output image, it will not be ready, and the masking of car 323a will therefore fail when (or if) the car first (partially) enters region 212. This will not be a problem for the contemplated improved method 200 described herein, because the object detection algorithm will also be provided with discarded information about the scene provided by one or more pixels 213, and will therefore be able to inform the object tracking algorithm that it is ready to provide the position of car 323a once it becomes at least partially visible within region 212 and within the output image of the output image stream, and car 323a can therefore be appropriately masked or marked in the output image stream.

[0067] Figure 3E The following scenario is illustrated schematically, where discarded pixel information is generated by concatenating one input image stream with another. In this paper, the first input image stream... Input image Capture a portion of the scene including objects 324a, 324b, and 324c. Second input image stream. Another input image Capture another portion of the scene (i.e., the portion of the scene not captured by the first input image stream), and include object 324c and another object 324d. As input images and Part of the processing 314, identifying two input images and The overlap between them (e.g., by identifying which objects in the scene are included in the two input images) and In the input image, for example, object 324c, and how the orientation and size of these common objects are displayed. and (Changes between). Imagine using any suitable, already available technology to stitch multiple images together to form, for example, a panoramic image of a scene, to perform overlap recognition.

[0068] As part of the processing section 314, the obtained intermediate image is determined. Pixels 213 outside region 212, and the resulting intermediate image Pixels 213' outside region 212' are to be cropped and not included in the resulting output image O. n Therefore, at least the first input image The pixel information about the scene provided by one or more pixels 213 will not be included in any output image O of any output image stream. n In other words, the first input image. One or more pixels 213 of the first input image stream (at least partially) originate from the second input image. and the region of the scene captured by the second input image stream, wherein, Figure 3E In the specific case shown, this area is the area of ​​object 324c.

[0069] As a result, in the two adjacent intermediate images and The output image O of the generated output image stream n In this context, traditional methods may struggle to properly mask objects such as 324b and 324c, because object 324b is only partially visible in region 212, and object 324c is in the intermediate image. Only a portion is visible in the first input image, and because of this... The scene-related pixel information provided by pixel 213 is discarded and cannot be used in the object detection algorithm. However, this is not a problem in the improved method envisioned in this paper, because the discarded scene-related pixel information provided by pixel 213 is still provided to the object detection algorithm, and all objects 324a-d can therefore be properly masked or labeled in the output image stream, as described several times previously in this paper.

[0070] therefore, Figures 3A to 3E An example illustrating this situation is provided where, during the processing of the input image stream, pixel information about the scene is discarded to generate one or more output image streams, such that this pixel information (as provided by pixels in a specific input image of a particular input image stream) is not included in any output image or output image stream generated thereby. Of course, besides references... Figures 3A to 3EBeyond the described scenarios, other situations may exist where pixels are cropped or masked as part of the processing of a particular input image for reasons other than object detection and / or tracking, and where the thus discarded pixel information about the scene is used by the object detection algorithms contemplated herein and, for example, those stated and used in the appended claims. Therefore, the improved methods contemplated herein are designed to utilize any pixel information about the scene that is otherwise discarded when processing input images and input image streams to generate one or more output images and output image streams. This is independent of the exact reason why the pixels of the input image stream providing this discarded pixel information are cropped or masked during processing, as long as the discarded pixel information about the scene is not included in any output image of any output image stream generated by processing the input images and input image streams. This allows the method to also utilize this pixel information when detecting and masking one or more objects in the output image stream. In some embodiments of the method, due to... Figures 3A to 3E In one or more (or even all) of the described cases, the method may particularly need to discard pixel information about the scene. In other embodiments of the method, due to, for example Figures 3A to 3E Any combination of any two of the cases described herein, any combination of any three of these cases, any combination of any four of these cases, or any combination of all five cases, the method may specifically require discarding pixel information about the scene. Therefore, the exact reason for discarding pixel information about the scene is conceived as being derived from all possible cases including those described herein (e.g., references...). Figures 3A to 3E The selection is made from a single list, and also includes all possible combinations of two or more such cases.

[0071] It is particularly clear that commonly available methods may include situations where an image stream received from a camera is first processed in some way to generate a processed image stream, and where the processed image stream is then used to generate and output a first output image stream. Specifically, commonly available methods may include using the processed image stream to generate an additional second output image stream in addition to the first output image stream, and where the generation of the second output image stream may include cropping or masking one or more pixels of the processed image stream.

[0072] For example, in this "traditional case," an input image stream from a camera can be received and processed to generate, for example, a first output image stream showing an overview of the scene. The camera could be, for example, a fisheye camera providing a large field of view, or, for example, a bird's-eye view camera with a top-down view of the scene. For an operator viewing the scene overview, if something important appears in a specific area of ​​the scene overview, they might wish to be able to, for example, digitally zoom in on the scene. To perform this digital zoom, a specific area of ​​the scene can first be defined (e.g., by clicking or otherwise marking the region of interest in the first output image stream), then all corresponding pixels outside that area in the first output image stream can be cropped, and the remaining pixels can be upsampled as part of generating an additional second output image stream, which shows a detailed view of the scene thus digitally zoomed in. Therefore, there are two output image streams generated based on a single input image stream, and the processing includes cropping pixels from the first output image stream to generate at least a second output stream.

[0073] However, it is of utmost importance to note that in this "conventional case," the cropping of pixels when generating the second output image stream does not correspond to the discarding of pixel information about the scene as described herein and as used in the appended claims. This is because, in the conventional case, the cropped pixels remain in the scene overview, and therefore the pixel information about the scene provided by one or more pixels of the input image stream is not discarded, such that this "discarded" pixel information is not included in any output image stream.

[0074] If the terminology of this disclosure is used when describing conventional situations, if one or more pixels, for example, recorded by an image sensor, are considered to have low visual quality (e.g., due to the use of a fisheye lens), and the raw image stream from the image sensor is first processed by, for example, cropping or masking pixels considered to have low visual quality, then pixel information about the scene may be discarded, and the result is provided as an input image stream for generating two output image streams. In this disclosure, particularly as used in the appended claims, "input image stream" refers herein to the image stream before the low-quality pixels are masked or cropped, and the discarded pixel information about the scene is the information about the scene provided by these cropped or masked pixels.

[0075] In the conventional case, when either of the two output image streams is generated, this discarded information about the scene is not provided, and therefore it is not used, for example, to detect objects in the overview and / or mask objects in one or both of the overview (first) or detail (second) output image streams, and the conventional case therefore suffers from the disadvantages described herein if the object happens to be at least partially within low-quality pixels. In other words, refer to Figure 1In the traditional case, any detection of an object is based solely on the intermediate image I′ n The pixel information about the scene is provided by the pixels in region 112, rather than by the pixel information about the scene provided by one or more pixels 113.

[0076] The privacy masking envisioned herein could, for example, have a solid / opaque color, be semi-transparent, include applying motion blur to the object to make it less easily identifiable, and / or, for example, forcibly pixelate and / or blur the object in the output image stream to make it less easily identifiable, etc. In other envisioned embodiments, privacy masking could include making the object itself at least partially transparent in the output image stream, allowing the background to be seen through the object. This is possible if, for example, a background image without any preceding object is available (from an input image captured, for example, in an earlier timeframe of the input image stream). The markings envisioned herein could, for example, be any graphical feature added to the output image stream that does not mask the object but instead, for example, provides additional information about the object (such as the object's confirmed identity, the object's identification type, or, for example, visually marking the object's position in the output image stream by adding a rectangle surrounding the object, etc.).

[0077] This disclosure also envisions means for masking or marking objects in an image stream, as will now be referenced. Figure 4A and Figure 4B To provide a more detailed description.

[0078] Figure 4AAn embodiment of apparatus 400 is illustrated schematically. Apparatus 400 includes at least a processor (or “processing circuitry”) 410 and a memory 412. As used herein, “processing circuitry” or “processor” can be, for example, any combination of one or more suitable central processing units (CPUs), multiprocessors, microcontrollers (μCs), digital signal processors (DSPs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc., capable of executing software instructions stored in memory 412. Memory 412 can be external to or internal to processor 410. As used herein, memory can be, for example, any combination of random access memory (RAM) and read-only memory (ROM). Memory 412 contains (i.e., stores) instructions that, when executed by processor 410, cause apparatus 400 to perform any embodiment of method 200, such as those previously disclosed herein. Apparatus 400 may also include one or more additional items 416, which in some cases are necessary for performing the method. In some embodiments, device 400 may be, for example, a surveillance camera (which may form part of, for example, a surveillance camera system and may be configured to capture an input image stream as discussed herein), and additional item 416 may then include, for example, an image sensor and, for example, one or more lenses (or lens arrangements) for focusing light captured from the scene (pointed to by the surveillance camera) onto the image sensor to capture the input image stream of the scene. Additional item 416 may also include, for example, various other electronic components required for capturing the scene, such as suitable operation of the image sensor and / or lenses. This allows the method to be performed within the surveillance camera itself (i.e., at the “edge”), which may reduce the need for any subsequent processing of the image stream output from the surveillance camera for purposes such as privacy shielding. If device 400 is to be connected to a network (e.g., if device 400 is a network camera), device 400 may also include a network interface 414. Network interface 414 may be, for example, a wireless interface supporting, for example, Wi-Fi (as defined in, for example, IEEE 802.11 or a subsequent standard), and / or a wired interface supporting, for example, Ethernet (as defined in, for example, IEEE 802.3 or a subsequent standard). For example, a communication bus 418 can be provided to interconnect the various parts 410, 412, 414, 416 and 418, so that these parts can communicate with each other as needed to obtain the desired function.

[0079] Figure 4B The reference is shown schematically. Figure 4A The described embodiment of the device 400 is shown, but is illustrated with, for example, references to... Figure 2AThe set of corresponding functional blocks / modules discussed herein. Apparatus 400 includes at least an image processing module 220, an object detection module 230, and a feature addition module 250. If object tracking is also used, apparatus 400 may further include an object tracking module 240. Modules 220, 230, 250 (and optionally 240) are interconnected such that they can communicate with each other when necessary (indicated by line 420), so that they can operate to perform the functions contemplated herein and, for example, those described in the references. Figure 2A and Figure 2B The described method. One, more, or all of modules 220, 230, 240, and 250 may be implemented, for example, using only software, using only hardware, and / or a combination of software and hardware. Such software may be provided, for example, by instructions stored in memory 412. Each module 220, 230, 240, and 250 may be provided as a separate entity, or two or more or all of modules 220, 230, 240, and 250 may be provided as part of the same single entity.

[0080] This paper also proposes one or more computer programs. One such computer program, for example, could be used to execute the masking method 400 discussed herein in the output image stream, for use in reference... Figure 4A and Figure 4B Such a method is executed in the described apparatus 400. A computer program may, for example, correspond to instructions stored in the memory 412 of the apparatus 400, such that when the processor (or processing circuitry) 410 executes the instructions, the apparatus 400 performs the corresponding method 200 or any embodiment thereof. In other contemplated embodiments, the computer program may be in a form not readable by the processor 410, but rather provided as text, for example, specified according to a programming language, which needs to be compiled into a format readable by the processor 410, for example, by using a suitable compiler. The compiler may, of course, be executed by the processor 410 itself, or even be formed as part of the processor 410 itself for real-time compilation.

[0081] This document also contemplates one or more computer program products. Each such computer program product includes a computer-readable storage medium on which one or more of the aforementioned computer programs are stored. For example, a computer program product may include a computer program for performing the contemplated masking methods disclosed and discussed herein in an output image stream. The (computer-readable) storage medium (e.g., "memory") may be any combination of, for example, random access memory (RAM) and read-only memory (ROM). In some embodiments, the computer-readable storage medium may be transient (e.g., processor-readable electrical signals). In other embodiments, the computer-readable storage medium may be non-transient (e.g., in the form of non-volatile memory, such as hard disk drive (HDD), solid-state drive (SSD), secure digital card (SD) card, USB flash drive, etc., any combination of magnetic storage, optical storage, solid-state storage, or even remotely mounted memory). Other types of computer-readable storage media are also contemplated, provided that their functionality allows storage of computer programs such that they can be read by a processor and / or intermediate compiler.

[0082] As a generalization of the various embodiments presented herein, this disclosure provides an improved method for masking or marking objects in an image stream, particularly where object detectors and / or object trackers may be unable to properly indicate and / or track one or more objects, for example, because the objects are only partially visible in the output image stream generated from the input image stream. This disclosure is based on the understanding that the input stream may be processed for reasons other than object detection and / or tracking (such as due to image stabilization, image correction, perspective projection, stitching, or any other transformation described herein), and that valuable pixel information about the scene may be lost or discarded when pixels are cropped or masked as part of the input image stream processing. In conjunction with this implementation, by using such pixel information about the scene as input for object detection without further consideration, this disclosure provides the aforementioned advantages over commonly available methods and techniques, thereby reducing the risk of failing to properly mask or mark objects in the output image stream, for example.

[0083] Although features and elements may have been described above in specific combinations, each feature or element may be used alone without other features and elements, or in various combinations with or without other features and elements. Furthermore, those skilled in the art, in practicing the claimed invention, can understand and implement variations of the disclosed embodiments by studying the drawings, the disclosure, and the appended claims.

[0084] In the claims, the word "comprising" does not exclude other elements, and the indefinite article "a" does not exclude multiple elements. The reference to certain features in dissimilar dependent claims does not imply that a combination of those features cannot be used advantageously.

[0085] List of reference numerals

[0086] 100 Traditional masking / marking methods

[0087] 110, 210 input image streams

[0088] 112, 212: Remaining pixel areas after image processing

[0089] Pixels 113 and 213 provide discarded pixel information about the scene.

[0090] 113', 213' are pixels that provide discarded pixel information about the scene.

[0091] 114, 214 output image streams

[0092] 120, 220 image processing modules

[0093] 130, 230 Object Detection Module

[0094] 132,232 object detection data

[0095] 133,233 object detection data

[0096] 140, 240 object tracking modules

[0097] 142,242 object tracking data

[0098] 150, 250 Feature Addition Module

[0099] 200 Improved methods for masking / marking

[0100] 320a, 320b objects

[0101] Image processing 310-314, 314'

[0102] 321a; 321b Objects; Rectangular background grid

[0103] 322a As described in the scenario

[0104] 322b, 322c objects

[0105] 323a object

[0106] 323b refers to the area outside of pixels considered to have lower visual quality.

[0107] 324a-d objects

[0108] 400 device

[0109] 410 processor

[0110] 412 Memory

[0111] 414 Network Interface

[0112] 416 Additional Components

[0113] 418, 420 communication bus

[0114] I n Input image

[0115] I n intermediate image

[0116] O n Output image

Claims

1. A method (200) of masking or marking an object in a stream of images, comprising: identifying, as a result of a processing (S202, 220) of a first input stream of images (210) capturing at least a portion of a scene for generating an output stream of images (214), pixel information that is to be excluded from the output stream of images, wherein the pixel information is provided by one or more pixels (213) of the first input stream of images, and wherein the pixel information relates to a particular peripheral portion of a description of the at least a portion of the scene in the first input stream of images; in response to the identifying, using the excluded pixel information for a detection (S203, 230) and / or tracking (240) of an object in the particular peripheral portion of the description of the at least a portion of the scene in the first input stream of images, and at least temporarily masking or marking (S205, 250) the object in the output stream of images after a decision (S204) that the object is at least partially located within the output stream of images, as part of generating the output stream of images and based on the detection and / or tracking of the object.

2. The method of claim 1, wherein, the one or more pixels of the first input stream of images originate from a peripheral region of an image sensor used for capturing the first input stream of images; and / or from light reaching the image sensor via a peripheral portion of a lens arrangement; and / or from light reaching at the lens arrangement at an angle sufficient to cause a vignetting.

3. The method of claim 1, wherein: the description (322a) of the at least a portion of the scene in the first input stream of images is non-rectangular, the one or more pixels of the first input stream of images belong to pixels that are located outside a rectangular region (212) of the non-rectangular description, and the non-rectangular description is caused by at least one of: a) properties of a lens arrangement through which the at least a portion of the scene is captured, and b) a transformation applied to the first input stream of images.

4. The method of claim 3, wherein, the transformation is applied as part of one or more of: b-1) a lens distortion correction process (311), such as, for example, a barrel distortion correction, a pincushion distortion correction, or a tangential distortion correction; b-2) a generation of an aerial view of the at least a portion of the scene; and b-3) one or more projections applied to stitch together a plurality of images capturing different portions of the scene, and wherein at least one of these plurality of images is included as part of the first input stream of images.

5. The method of claim 1, wherein, the processing of the first input stream of images for generating the output stream of images comprises stitching (314, 314') the first input stream of images with a second input stream of images capturing a portion of the scene that is not captured by the first input stream of images, and wherein the one or more pixels of the first input stream of images originate from a region (324c) of the scene that is also captured by the second input stream of images.

6. The method of claim 1, wherein, The processing of the first input image stream for generating the output image stream comprises applying electronic image stabilization, EIS (310) to generate the first input image stream, and wherein the one or more pixels of the first input image stream belong to pixels cropped or masked by the EIS.

7. The method of claim 1, the method being performed in a surveillance camera configured to capture the first input image stream.

8. The method of claim 1, wherein, Determining that the object is at least partially located within the output image stream comprises using an object tracking algorithm (240).

9. An apparatus (400) for masking or marking an object in an image stream, comprising: a processor (410), and a memory (412) storing instructions which, when executed by the processor, cause the apparatus to: identify, as a result of processing (S202, 210) of a first input image stream (210) of at least a portion of a scene for generating an output image stream (214), pixel information that is to not be included in the output image stream, wherein the pixel information is provided by one or more pixels (213) of the first input image stream, and wherein the pixel information relates to a certain peripheral portion of a description of the at least a portion of the scene in the first input image stream; in response to the identification, use the not included pixel information for detection (S203, 230) and / or tracking (240) of an object in the certain peripheral portion of the description of the at least a portion of the scene in the first input image stream, and at least temporarily mask or mark (S205, 250) the object in the output image stream after determining (S204) that the object is at least partially located within the output image stream as part of generating the output image stream and based on the detection and / or tracking of the object.

10. The apparatus of claim 9, wherein, The instructions, when executed by the processor, cause the apparatus to perform the method of any one of claims 2 to 8.

11. The apparatus of claim 9, wherein, The apparatus is a surveillance camera configured to capture the first input image stream.

12. A computer-readable storage medium storing a computer program for masking or marking an object in an image stream, the computer program being configured to, when executed by a processor (410) of an apparatus (400), cause the apparatus to: As a result of the processing (S202, 210) of the first input image stream (210) of capturing at least a portion of a scene for generating an output image stream (214), pixel information is identified that will not be included in the output image stream, wherein the pixel information is provided by one or more pixels (213) of the first input image stream, and wherein the pixel information relates to a certain peripheral portion of a description of the at least a portion of the scene in the first input image stream; in response to the identification, use the not included pixel information for detection (S203, 230) and / or tracking (240) of an object in the certain peripheral portion of the description of the at least a portion of the scene in the first input image stream, and As part of generating the output image stream and based on the detection and / or tracking of the object, at least temporarily masking or marking (S205, 250) the detected object in the output image stream after deciding (S204) that the object is at least partially located within the output image stream.

13. The computer-readable storage medium of claim 12, wherein, The computer program is further configured to cause the apparatus to perform the method (200) according to any one of claims 2 to 8.

Citation Information

Patent Citations

  • Image processing device, camera device, communication system, image processing method, and program

    CN101510957A

  • Method for processing image, related device and storage medium

    US20210407052A1