A method for controlling an image processing stage for processing image data captured by a surveillance camera.
The method optimizes surveillance camera image processing by detecting ceiling-mounted configurations to selectively apply operations to central scene pixels, enhancing image quality and resource efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- AXIS
- Filing Date
- 2025-10-02
- Publication Date
- 2026-06-02
AI Technical Summary
Surveillance cameras with a field of view greater than 180 degrees, when mounted close to or flush with the ceiling, waste processing resources on uninformative ceiling pixels, degrading image quality and interfering with image analysis.
A method to detect whether the camera is in a ceiling-mounted configuration using pixel analysis, configuring the image processing stage to apply operations only to central scene pixels and exclude peripheral ceiling pixels, thereby optimizing resource use and image quality.
Enhances image quality and resource efficiency by selectively processing only informative central scene pixels, reducing false alarms and improving object detection in surveillance cameras with wide fields of view.
Smart Images

Figure 2026090188000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method, a processing device, and a computer program product for controlling an image processing stage for processing images captured by a surveillance camera. [Background technology]
[0002] Surveillance cameras with a wide field of view (FOV) can be used in applications where it is desirable to monitor a wide area of the scene using a single camera. Some surveillance cameras use fisheye lenses or multi-sensor configurations to provide an FOV of more than 180 degrees. By oriented downwards, for example, by mounting such a camera on a camera pole or suspending it from the ceiling, the camera becomes capable of capturing images of the scene extending both below and above the horizon. The additional FOV beyond 180 degrees can thus provide useful surveillance information about objects or events above the horizon. For example, an attempt to tamper with the camera from above may be detected. [Overview of the Initiative]
[0003] However, as the inventors acknowledge, when a camera with a field of view (FOV) greater than 180 degrees is mounted close to or flush with the ceiling, a relatively large portion of the image is occupied by the ceiling area close to the camera (which would typically be out of focus). Therefore, using a camera in such a configuration may result in portions of the image that do not contribute to useful scene information. As a result, the additional FOV beyond 180 degrees may unnecessarily increase the utilization of processing resources and potentially degrade image quality. Addressing this problem is the objective of the present invention.
[0004] According to a first aspect of the present invention, a method for controlling an image processing stage for processing image data captured by a surveillance camera having a field of view larger than 180 degrees and mounted in a downward-looking configuration for monitoring a scene, Acquiring image data captured by a surveillance camera, wherein the image data includes a first set of pixels that depict the central scene portion located below the horizon line in the scene, and a second set of pixels that depict the peripheral scene portion located above the horizon line. To determine whether the image processing stage should be configured to operate according to the ceiling operating mode, a ceiling detection procedure is performed which involves analyzing the pixels of a second set of pixels, Configuring the image processing stage to operate according to the ceiling operating mode in response to a decision to configure the image processing stage to operate according to the ceiling operating mode while processing subsequently captured image data, wherein the subsequently captured image data comprises a first set of pixels that depict a central scene portion and a second set of pixels that depict a peripheral scene portion, the image processing stage comprises at least one image processing operation, and the ceiling operating mode includes applying at least one image processing operation to the first set of pixels of the captured image data but not to the second set of pixels of the captured image data. A method is provided that includes the following:
[0005] By applying a ceiling detection procedure to the pixels of a second set of pixels, the method enables the use of image analysis to estimate or predict whether the camera is in a ceiling-mounted configuration related to the ceiling operating mode of the image processing stage, and to configure the image processing stage accordingly. As further described herein, a ceiling-mounted configuration can generally correspond to a configuration in which the camera is mounted close to or flush with the ceiling. A ceiling-mounted configuration can generally be such that the optical axis of the camera is directed towards the ground in the scene. Thus, in a ceiling-mounted configuration, the optical axis of the camera can be perpendicular to the ceiling. Furthermore, the horizon in the scene can correspond to an angle of 90 degrees (field of view) with respect to the optical axis.
[0006] By configuring the image processing stage according to the ceiling operating mode, image data subsequently captured by the surveillance camera during surveillance operation can be selectively processed. That is, a first set of pixels from the subsequently captured image data that depicts the central scene portion (which is expected to contain scene information useful from a surveillance perspective) can be supplied and subjected to at least one image processing operation of the image processing stage, while a second set of pixels from the subsequently captured image data that depicts the peripheral scene portion above the horizon (which includes the ceiling and is therefore expected to not contain scene information useful from a surveillance perspective) can be excluded from at least one image processing operation of the image processing stage. This avoids spending valuable processing resources on pixels that do not contribute to scene information useful from a surveillance perspective. For example, the second set of pixels from the subsequently captured image data can be discarded or ignored by the image processing stage.
[0007] It is further assumed that ceilings can generally be rendered as relatively monotonous and / or monochromatic pixel regions with average pixel intensity that deviates from the pixels rendering the central scene, especially in ceiling-mounted configurations where, for example, the peripheral scene is relatively bright and often contains light sources. Therefore, including a second set of pixels in the processed output image data may result in a reduction in perceived image quality and / or be detrimental to the viewer. The second set of pixels may also have adverse effects on image analysis. For example, flickering of light sources may trigger false alarms and / or interfere with object detection and tracking algorithms.
[0008] A pixel analysis-based ceiling detection procedure allows the ceiling operating mode to be automatically configured without relying on configuration information provided by the user or technician installing the surveillance camera. This method therefore facilitates user-friendly deployment.
[0009] On the other hand, this method allows a surveillance camera to utilize its entire field of view, including a second set of pixels, in a non-ceiling-mounted configuration.
[0010] Thus, in some embodiments, the method further comprises configuring an image processing stage to operate according to a non-ceiling operating mode in response to a decision (based on a ceiling detection procedure) not to configure the image processing stage to operate according to a ceiling operating mode, wherein the non-ceiling operating mode includes applying at least one image processing operation to a first set of pixels and a second set of pixels of subsequently captured image data.
[0011] In some embodiments, at least one image processing operation includes at least one of image transformation, encoding, and / or image analysis operations such as object detection and / or object tracking.
[0012] Thus, at least one of the image conversion, encoding operation, and / or image analysis operation can be selectively applied to the first set of pixels, while the second set of pixels can be excluded from that application.
[0013] The image conversion can comprise at least one of an image scaling operation and / or an image warping operation.
[0014] In some embodiments, the method further comprises controlling, based on a first pixel statistic value, a first setting of the surveillance camera that is used by the surveillance camera during the capture of subsequently captured image data, in response to determining to configure the image processing stage to operate according to a ceiling operation mode (i.e., when the image processing stage operates in the ceiling operation mode), wherein the first pixel statistic value is based on pixels that depict a central scene portion but not on pixels that depict a peripheral scene portion.
[0015] Thus, the second set of pixels can be ignored / excluded from consideration for the purpose of controlling (at least) the first setting of the surveillance camera.
[0016] The first setting of the surveillance camera can be a setting of one or more exposure-related control parameters of the surveillance camera (e.g., shutter speed, aperture, ISO value, and / or camera illumination). The first pixel statistic value can indicate an illumination state in the central scene portion.
[0017] In some embodiments, the method further comprises controlling a second setting of the surveillance camera, used by the surveillance camera during capture of subsequently captured image data, based on a second pixel statistic, in response to determining to configure the image processing stage to operate according to a ceiling operation mode (i.e., when the image processing stage is operating in the ceiling operation mode), where the second pixel statistic is based on pixels depicting a peripheral scene portion and not on pixels depicting a central scene portion.
[0018] Processing a second set of pixels in the image processing stage may be wasteful and / or undesirable for the reasons described above, but pixels depicting a peripheral scene portion may still be useful for the purpose of controlling some camera settings of the surveillance camera.
[0019] In some embodiments, the method further comprises discarding a second set of pixels of subsequently captured image data, in response to determining to configure the image processing stage to operate according to a ceiling operation mode. Thus, a second set of pixels depicting a peripheral scene portion, i.e., pixels including the ceiling, may be discarded, thereby avoiding burdening subsequent image processing pipelines.
[0020] In embodiments where a second set of pixels of subsequently captured image data is utilized to determine pixel statistics, such as the second pixel statistic described above, the method may comprise determining the pixel statistics and subsequently discarding the second set of pixels.
[0021] In some embodiments, the image data is an image frame. Performing a ceiling detection procedure Determining a contrast index for each pixel or pixel block of a second set of pixels distributed radially across the image frame, such that the distance from the pixel drawing the horizontal line increases; To analyze the variation in the contrast index in the radial direction in order to identify whether the variation in the contrast index defines a peak, Equipped with, The condition for determining whether to configure the image processing stage to operate according to the ceiling operating mode is that peaks are identified.
[0022] This means that the decision of whether to configure the image processing stage to operate according to the ceiling operating mode can be based on the variation in the radial contrast index. In a ceiling-mounted configuration, at the most extreme field of view covered by the surveillance camera (e.g., corresponding to the edges of the exposed image area of the image frame), pixels tend to be out of focus and therefore produce a low contrast index. The horizon is located at infinity, and therefore, as it is out of focus, the contrast index is also expected to be low for pixels that paint the horizon ("horizon pixels"). Between these field of view angles of "minimum contrast," the contrast index varies, as recognized by the inventors, to define a contrast peak. The contrast peak is generally obtained for pixels in a second set of pixels that paint the portion of the ceiling within the depth of field of the surveillance camera. Therefore, the presence of a contrast peak can be used as an indicator or predictor of a ceiling-mounted configuration, thereby it can be used as a condition for determining whether to configure the image processing stage to operate in ceiling operating mode.
[0023] Contrast-based ceiling detection procedures can be advantageously used in a surveillance camera implementation that comprises a single fisheye lens and a single image sensor positioned behind the fisheye lens, defining the entire field of view of the surveillance camera, where the image data is the image frame captured by the image sensor. Such a single-sensor and fisheye lens-based implementation of a surveillance camera may hereafter be referred to as a "single-sensor implementation" for brevity.
[0024] In some embodiments, the ceiling detection procedure is: For each of at least one further radial directions, determine a contrast index for each pixel or pixel block of a second set of pixels distributed in each radial direction such that the distance from the pixel drawing the horizontal line increases; The analysis of each variation of the contrast index in each radial direction is conducted to identify whether each variation of the contrast index defines a peak. Furthermore, The condition for determining whether to configure the image processing stage to operate according to the ceiling operating mode is that peaks are identified in at least a predetermined minimum number of radial directions.
[0025] The contrast can then be analyzed in two or more radial directions. This can further improve the reliability of the ceiling detection procedure, particularly reducing the risk of false positives (in other words, incorrectly predicting that the camera is in a ceiling-mounted configuration). Requiring that a peak be identified for each radial direction may provide the most effective suppression of false positives.
[0026] In some embodiments, the surveillance camera comprises a lens configuration and at least first and second image sensors, each image sensor positioned behind the lens configuration to capture a portion of the scene from its respective viewpoint. Acquiring image data comprises acquiring a first partial image frame captured by a first image sensor and a second partial image frame captured by a second image sensor, wherein the second set of pixels in the image data comprises a first subset of pixels from the first partial image frame and a second subset of pixels from the second partial image frame, and the first and second subsets of pixels draw the overlapping portion of the surrounding scene. Performing the ceiling detection procedure is Identifying a first set of feature points distributed radially in a first subset of pixels such that their distance from the pixels drawing the horizontal line in the first partial image frame increases, For each first feature point, identify the matching second feature point in the second subset of pixels, For each first feature point, the disparity error with respect to the corresponding second feature point is determined. To determine whether the parallax error increases in the radial direction, we will analyze the variation in the parallax error in the radial direction. Equipped with, The condition for determining that the image processing stage should be configured to operate according to the ceiling operating mode is that the parallax error increases radially.
[0027] Therefore, in the case of such a "multi-sensor implementation configuration" of a camera (in other words, an implementation configuration in which the camera has two or more image sensors), the decision of whether the image processing stage should be configured to operate according to the ceiling operating mode may be based on the variation in radial parallax error.
[0028] In a ceiling-mounted configuration, the first and second partial image frames captured by the first and second image sensors depict overlapping areas of the ceiling, albeit from different viewpoints. When the camera is in a ceiling-mounted configuration, the different viewpoints introduce parallax errors between matching feature points in the first and second subsets of pixels, and this parallax error tends to increase radially, in other words, with increasing field of view. This can be understood by considering that the parallax error between matching feature points approaches zero at the horizon, since the horizon is at infinity. From this point of minimum parallax error, the parallax error increases for matching feature points detected at larger field of view as they are located closer to the surveillance camera. On the other hand, if there is no ceiling, or if the camera is mounted far from the ceiling, the parallax error tends to remain small as the field of view increases above the horizon. The parallax error may be substantially constant as the field of view increases, or it may increase at a relatively low rate in either case. Therefore, an increase in radial parallax error can be used as an indicator or predictor of a ceiling-mounted configuration, and thus can be used as a condition for determining whether the image processing stage should be configured to operate in ceiling operating mode.
[0029] In this disclosure, the term “matching feature points” means pairs of feature points in a first and second subset of pixels that depict the same feature in a surrounding scene portion, albeit from different viewpoints. In this disclosure, the term “(first / second) partial image frame” is used to indicate that each partial image frame covers a portion of the surveillance camera’s entire field of view (FOV). Thus, partial image frames captured by at least the first and second image sensors of the surveillance camera can be stitched together into a composite image frame that covers the entire FOV of the surveillance camera.
[0030] In some embodiments, the method further comprises mapping a first feature point and a corresponding second feature point to a common composite plane, wherein the parallax error for each first feature point is determined by calculating the distance between the first feature point and the corresponding second feature point when mapped to the common composite plane.
[0031] The distance between matching feature points mapped to a common composite plane (in other words, a composite / stitching plane for stitching partial image frames into a composite image frame) provides a convenient and reliable indicator of parallax error.
[0032] In some embodiments, the second set of pixels in the image data further comprises a third subset of pixels from a third partial image frame and a fourth subset of pixels from a fourth partial image frame, wherein the third and fourth subsets of pixels draw the overlapping portion of the surrounding scene. Performing the ceiling detection procedure is Identifying a third set of feature points distributed radially in a third subset of pixels such that the distance from the pixels drawing the horizontal line in the third partial image frame increases, For each third feature point, identify the matching fourth feature point in the fourth subset of pixels, For each third feature point, the disparity error with respect to the corresponding fourth feature point is determined. To determine whether the parallax error increases in the radial direction of the third partial image frame, we will analyze the variation in the parallax error in the radial direction of the third partial image frame. Furthermore, A further condition for determining that the image processing stage is configured to operate according to the ceiling operating mode is that the disparity error determined for the third feature point increases radially in the third partial image frame.
[0033] Parallax errors can then be analyzed for further pairs of overlapping subsets of pixels obtained from the third and fourth partial image frames. Thus, parallax errors can be analyzed in two or more radial directions. This can further improve the reliability of the ceiling detection procedure and reduce the risk of false positives in particular (in other words, incorrectly predicting that the camera is in a ceiling-mounted configuration).
[0034] The third partial image frame may refer, for example, to a third partial image frame captured by a third image sensor (in other words, an image sensor different from the first and second image sensors) positioned behind the lens configuration to render a portion of the scene from each viewpoint. In this case, the fourth partial image frame may refer to either the first or second partial image frame. Alternatively, the fourth partial image frame may refer here to a fourth partial image frame captured by a fourth image sensor (in other words, an image sensor different from the first, second, and third image sensors) positioned behind the lens configuration to render a portion of the scene from each viewpoint. Thus, the terms “third” and “fourth” are used here merely as labels to refer to further pairs of overlapping partial image frames that are different from the first and second partial image frames captured by the first and second image sensors, respectively.
[0035] In some embodiments applicable to multi-sensor implementations of surveillance cameras, subsequently captured image data processed by an image processing stage comprises a sequence of composite image frames, each composite image frame being formed by stitching together partial image frames captured by at least first and second image sensors, such that each composite image frame comprises a first set of pixels for depicting a central scene portion and a second set of pixels for depicting a peripheral scene portion.
[0036] This allows at least one image processing operation of the image processing stage to be applied to the first set of pixels in a composite image frame having first and second sets of pixels.
[0037] In some embodiments applicable to multi-sensor implementations of surveillance cameras, the subsequently captured image data processed by an image processing stage instead comprises partial image frames captured by at least first and second image sensors, wherein the captured partial image frames, when combined, comprise a first set of pixels that depict the central scene portion and a second set of pixels that depict the peripheral scene portion. At least one image processing operation of the image processing stage includes a stitching operation for forming a composite image frame by stitching together partial image frames, wherein the stitching operation is applied to a first set of pixels of the partial image frame but not to a second set of pixels of the partial image frame in ceiling operation mode.
[0038] This allows at least one image processing operation of the image processing stage to be applied to a partial image frame. Furthermore, a stitching operation is implemented by the image processing stage, which, when configured according to the ceiling operation mode, can exclude a second set of pixels from a subsequently captured partial image frame during stitching, such that only a first set of pixels from the subsequently captured partial image frame are stitched to form a composite image frame.
[0039] In some embodiments, the surveillance camera is equipped with an orientation sensor, and the ceiling detection procedure is performed in response to the orientation sensor detecting the downward orientation of the surveillance camera.
[0040] Therefore, an orientation sensor can be used as a first non-image-based technique to detect the possibility that a surveillance camera is installed in a ceiling-mounted configuration. However, since an orientation sensor can only detect the orientation of the surveillance camera, the sensor output from the orientation sensor is not sufficient to distinguish it from configurations where the surveillance camera is suspended from a camera pole or is far from the ceiling. Thus, an image analysis-based technique involving the analysis of a second set of pixels can be employed as a second detection stage to finally detect a ceiling-mounted configuration.
[0041] In some embodiments, the method further comprises obtaining predetermined horizontal line data that points to the pixel coordinates of a horizontal line in image data, and using the predetermined horizontal line data to identify a second set of pixels.
[0042] The pixel coordinates of the horizontal line may be established in advance by calibration measurements or supplied by the manufacturer of the imaging module comprising the image sensor and lens (configuration). This predetermined information may be used as input to the method to determine which parts of the image data correspond to the first and second sets of pixels.
[0043] According to a second aspect, a processing device is provided which is configured to implement a method of the first aspect or an embodiment thereof for controlling an image processing stage for processing image data captured by a surveillance camera. The processing device may be provided in the surveillance camera. The image processing stage may be an image processing stage of the surveillance camera.
[0044] A third aspect provides a computer program product comprising a portion of computer program code configured to carry out the method according to the first aspect or its embodiment.
[0045] In general, the embodiments, features, effects, or benefits discussed in relation to the first embodiment also apply to the monitoring systems and computer program products of the second and third embodiments.
[0046] The above and additional purposes, embodiments, features, and effects of this disclosure can be better understood through the following illustrative and non-limiting detailed description with reference to the accompanying drawings. In the drawings, similar reference numerals are used for similar elements unless otherwise noted. [Brief explanation of the drawing]
[0047] [Figure 1] This is a schematic diagram showing a surveillance camera mounted on the ceiling. [Figure 2] This diagram schematically shows a surveillance camera in a different ceiling-mounted configuration. [Figure 3] This is a schematic diagram showing a surveillance camera mounted on a pole. [Figure 4] This diagram schematically illustrates a system with a single-sensor implementation configuration for a surveillance camera. [Figure 5] This is a block diagram of the image processing stage. [Figure 6a] Figure 6a schematically illustrates image data in the form of an image frame captured by the surveillance camera in Figure 4, and pixels of a second set of pixels in the image frame that render the surrounding scene portion (Figure 6b). [Figure 6b] Figure 6a schematically illustrates image data in the form of an image frame captured by the surveillance camera in Figure 4, and pixels of a second set of pixels in the image frame that render the surrounding scene portion (Figure 6b). [Figure 7a] This is a schematic diagram showing the variation in the contrast index when the surveillance camera is not mounted on the ceiling (Figure 7a) and when the surveillance camera is mounted on the ceiling (Figure 7b). [Figure 7b]This is a schematic diagram showing the variation in the contrast index when the surveillance camera is not mounted on the ceiling (Figure 7a) and when the surveillance camera is mounted on the ceiling (Figure 7b). [Figure 8] This diagram schematically illustrates a system with a multi-sensor implementation configuration for surveillance cameras. [Figure 9] This figure schematically illustrates a composite image frame formed by stitching together partial image frames captured by the surveillance camera shown in Figure 8. [Figure 10a] Figure 10a schematically illustrates the image data from the first and second partial image frames captured by the multi-sensor surveillance camera shown in Figure 8, as well as the determination of the disparity error between matching feature points in the first and second partial image frames (Figure 10b). [Figure 10b] Figure 10a schematically illustrates the image data from the first and second partial image frames captured by the multi-sensor surveillance camera shown in Figure 8, as well as the determination of the disparity error between matching feature points in the first and second partial image frames (Figure 10b). [Figure 11a] This is a schematic diagram of the parallax error when the surveillance camera is not mounted on the ceiling (Figure 11a) and when the surveillance camera is mounted on the ceiling (Figure 11b). [Figure 11b] This is a schematic diagram of the parallax error when the surveillance camera is not mounted on the ceiling (Figure 11a) and when the surveillance camera is mounted on the ceiling (Figure 11b). [Figure 12] This is a flowchart of a method for controlling an image processing stage for processing image frames captured by a surveillance camera. [Figure 13] This is a flowchart illustrating an exemplary method for analyzing pixels in image data captured by a single-sensor surveillance camera. [Figure 14] This is a flowchart illustrating an exemplary method for analyzing pixels in image data captured by a multi-sensor surveillance camera, such as the surveillance camera in the surveillance system shown in Figure 8. [Modes for carrying out the invention]
[0048] Figure 1 shows Scene 10 and the surveillance camera 110. The surveillance camera 110 may be a camera device / image capturing device suitable for image-based surveillance applications such as video surveillance. The surveillance camera 110 may be a networked camera, such as an Internet Protocol (IP) camera. However, non-networked implementations of the surveillance camera 110 are also possible.
[0049] The surveillance camera 110 may be referred to as "camera 110" below for brevity. Camera 110 is positioned in a ceiling-mounted configuration to monitor scene 10. Camera 110 is mounted in the ceiling 12, for example, flush with the ceiling 12, as shown. Camera 100 is mounted in a downward-looking configuration, which means that the optical axis O of the surveillance camera 110 is pointed in the negative vertical direction -Z, for example, toward the floor 14 of scene 10 (or more generally toward the ground), and perpendicular to the ceiling 12 (in other words, perpendicular). Camera 110 is therefore positioned to monitor scene 10 from above.
[0050] Camera 110 has a field of view (FOV) larger than 180 degrees, as shown. Therefore, the scene 10 monitored by camera 110 includes a central scene portion 10a located below the horizon line H in scene 10, and a peripheral scene portion 10b located above the horizon line H. The horizon line H in scene 10 is located at a field of view angle of 90 degrees with respect to the optical axis O, as shown. In the ceiling mounting configuration shown in Figure 1, the peripheral scene portion 10b includes a portion of the ceiling 12, as shown. In other words, the central portion or sub-range of the FOV of surveillance camera 110 covers the central scene portion 10a located below the horizon line H, and the peripheral portion or sub-range of the FOV of surveillance camera 110 covers the peripheral scene portion 10b located above the horizon line H and including the ceiling 12.
[0051] To say that a camera has a “FOV greater than 180 degrees” means that the camera’s FOV is greater than 180 degrees in at least one first direction perpendicular to the camera’s optical axis. Generally, according to the exemplary embodiments shown herein, the FOV is greater than 180 degrees in each of the first and second directions, the first and second directions are perpendicular to each other, and both are perpendicular to the camera’s optical axis. An FOV greater than 180 degrees can thus cover more than a hemisphere of the scene monitored by the camera. For the purposes of this specification, the optical axis is assumed to be located at the center of the FOV.
[0052] Figure 2 shows camera 110 in another ceiling-mounted and downward-looking configuration, where camera 110 is suspended from ceiling 12 by a suspended camera mount, such as a pole, instead of being mounted flush with ceiling 12. This positions camera 110 further away from ceiling 12 than in the flush-mounted configuration of Figure 1.
[0053] Figure 3 shows a camera 110 in a pole-mounted configuration, in other words, suspended from a pole. In this configuration, the camera 110 is facing downwards, i.e., towards the ground 14. In contrast, in the scenario shown in Figure 3, there is no ceiling above the camera 110. Figure 3 may correspond to a use case where the camera 110 is used in an outdoor environment, for example. The surrounding scene portion 10b may be formed, for example, by an open sky.
[0054] As can be understood, in the ceiling-mounted configuration of Figure 1, the central scene portion 10a is generally the area of interest of scene 10 from a surveillance perspective. On the other hand, the peripheral scene portion 10b mainly comprises the ceiling 12 and may be of little interest from a surveillance perspective. Furthermore, it is thought that objects of interest moving along the ceiling 12 may also be present in the central scene portion 10a, in other words, the area of interest. Moreover, the distance to the ceiling 12 along the maximum field of view within the camera 110's FOV is smaller than the focal length of the camera 110, and therefore, the area surrounding the ceiling 12 near the camera 110 is assumed to be out of focus, in other words, outside the depth of field (DOF) of the camera 110. As a result, in Figure 1, the additional FOV exceeding 180 degrees may not provide useful surveillance information regarding objects or events above the horizon line H. On the other hand, in Figure 2, due to the increased distance to the ceiling 12, the ceiling 12 covers a smaller sub-range of the camera 110's FOV than in the case of Figure 1. Additionally, the ceiling 12 tends to be less out of focus, and a wider area is expected to be within the camera 110's DOF. Therefore, in this case, the additional FOV exceeding 180 degrees can provide useful surveillance information about objects or events above the horizon H, provided that the distance to the ceiling 12 is relatively large. Similarly, in the scenario in Figure 3, where there is no ceiling 12, the additional FOV exceeding 180 degrees can, in this scenario as well, provide useful surveillance information about objects or events above the horizon H.
[0055] As can be understood, Figure 1 shows the camera 110 mounted flush with the ceiling 12, but the issues discussed with reference to Figure 1 may also apply to configurations in which the camera 110 is suspended from the ceiling 12, as in Figure 2, however, this applies to relatively short distances such that the peripheral scene portion 10b mainly consists of the ceiling 12, thereby also including the object of interest moving along the ceiling 12 in the central scene portion 10a. In this specification, the term “ceiling-mounted configuration” may therefore be used to refer to either the flush-mounted scenario corresponding to Figure 1, or the scenario corresponding to Figure 2, where the distance between the ceiling 12 and the camera 110 is small (for example, smaller than the lower boundary of the camera 110’s DOF).
[0056] The set of pixels in the image frame captured by camera 110 and rendering the central scene portion 10a will hereafter be referred to as the "first set of pixels" or interchangeably as the "central pixels." The set of pixels in the image frame rendering the peripheral scene portion 10b will hereafter be referred to as the "second set of pixels" or interchangeably as the "peripheral pixels."
[0057] As can be understood from the above, when the camera 110 is mounted close to or flush with the ceiling 12, the second set of pixels / peripheral pixels may be of little value from a surveillance perspective. In addition, peripheral pixels can result in several problems, as described above, including unnecessary use of processing resources, degradation of the image quality of the first set of pixels / central pixels, and obstacles to image analysis. This disclosure provides techniques to address or mitigate one or more of these problems. Furthermore, it should be noted that in Figures 1 and 2, the camera 110 is mounted so that the optical axis O is perpendicular to the ceiling 12, but similar problems may occur if the angle between the ceiling 12 and the optical axis O is 90 degrees within a certain tolerance (e.g., ±5 or ±10 degrees). As can be understood, the techniques described herein are also applicable in this case, because even when the angle between the optical axis O and the ceiling is approximately 90 degrees, the peripheral pixels can still primarily depict the ceiling.
[0058] It should be noted that the explanations and meanings of the terms “FOV,” “optical axis,” and “horizon,” provided above in relation to Figures 1 and 2, also apply to the embodiments described below with reference to Figures 4 and 8.
[0059] Figure 4 schematically shows a block diagram of a system 100 comprising a surveillance camera 110, an image processing stage 140, and a processing device 150. The dotted line schematically indicates possible positions on the ceiling 12 when the surveillance camera 110 is in a ceiling-mounted configuration. The image processing stage 140 is configured to process image frames captured by the camera 110. The processing device 150 is configured to implement a method for controlling the image processing stage 140, as described below. In the illustrated example, the camera 110 is in a single-sensor configuration, comprising a fisheye lens 120 and a single image sensor 130 positioned behind the fisheye lens 120. This allows the image sensor 130 to capture an image frame 300 depicting the scene 10, which is imaged onto the image sensor 130 by the fisheye lens 120, during operation. As further shown, the camera 110 may optionally include an orientation sensor 160 configured to detect the physical orientation of the camera 110. The orientation sensor 160 may, for example, include one or more accelerometers and / or gyroscopes.
[0060] Figure 4 shows the image processing stage 140 and processing device 150 as blocks outside the camera 110 for clarity of illustration, but note that both collocated and distributed implementations of the system 100 being depicted are possible. For example, in a typical configuration, both the image processing stage 140 and processing device 150 may be located on the camera 110. In another configuration, the processing device 150 may be located on the camera 110, and the image processing stage 140 may be located in an external device outside the camera 110. For example, the image processing stage 140 may be an image processing stage of an external camera controller, or a remote or non-edge device (such as a server-side image processing stage). The camera 110 and the external device may be connected via a network (wired or wireless) or via a non-networked communication interface (such as a USB interface) to receive and process image frames captured by the camera 110. In another configuration, the image processing stage 140 may be provided on the camera 110, and the processing device 150 may be located outside the camera 110 in an external device, for example, in one of the external devices of the types described above. In yet another configuration, both the image processing stage 140 and the processing device 150 may be located outside the camera 110 in one of the external devices of the types described above. A distributed configuration may be useful when the computing resources of the camera 110 are limited, and it is desirable to offload some image processing operations and / or processing associated with the methods performed by the processing device 150 from the camera 110.
[0061] Figure 5 shows the image processing stage 140 in more detail in block diagram form. The image processing stage 140 is configured to process image data, such as image frames 300 captured by camera 110, and to output processed image data, such as image frames 301. Each image frame 300 may be, for example, a sequence of image frames 300 captured by camera 110 (at a fixed or variable frame rate), and the image processing stage 140 may output the processed image frames 301 in the form of a video stream. As will be further described below, the image frames processed by the image processing stage 140 may more generally be image frames captured by a single-sensor implementation camera, such as camera 110, or composite or partial image frames captured by a multi-sensor implementation camera, such as camera 210 in Figure 8.
[0062] The image processing stage 140 comprises several image processing operations 141, 142, ..., 14n, as shown. The number of image processing operations in the image processing stage 140 may vary depending on the application. In any case, the image processing stage 140 comprises at least one image processing operation. The image processing stage 140 may comprise, for example, an image transformation operation, an encoding operation, and / or an image analysis operation.
[0063] The image transformation may comprise at least one of an image scaling operation and / or an image warping operation. The scaling operation may comprise scaling of the received image frame 300, such as resizing by upsampling or downsampling. The warping operation may comprise optical distortion correction. This allows the image transformation in the form of mapping to be applied to pixels of the received image frame 300 to reduce the effects of optical aberrations or distortions introduced into the image frame 300 by the lens or lens configuration. As a non-limiting example, the warping operation may comprise mapping a non-rectangular image field with typical distortions caused by a fisheye lens to a rectangular image area.
[0064] The encoding operation may include encoding the received image frame 300 into an encoded image frame. The image frame 300 may be encoded into a format suitable for, for example, transmission over an IP network, storage, and / or viewing on a monitor. If each image frame 300 forms part of a sequence of image frames 300 in a video stream, the encoding operation may encode the sequence of image frames 300 into an encoded video stream.
[0065] Image analysis operations may include image and / or video analysis, such as object detection and / or object tracking operations. The received image frame 300 may be processed, for example, to detect and / or track objects in the image frame 300. Any conventional type of object detection and tracking algorithm known in the art may be used.
[0066] If the image processing stage 140 comprises two or more image processing operations, the image processing operations may be applied sequentially to the image frame 300 received by the image processing stage 140. For example, the image processing stage 140 may comprise an image transformation operation (e.g., scaling and / or warping), an encoding operation, and an image analysis operation. In this case, the image processing stage 140 may process the captured image frame by first applying the image transformation operation to the captured image, then applying the encoding operation to the output from the image transformation operation, and then applying the image analysis operation to the encoded output from the encoding operation. However, parallel processing is not ruled out. For example, the image analysis operation may be applied to the captured image frame in parallel with applying the image transformation operation and / or encoding operation to the captured image frame. This allows the image analysis operation to be applied to the pixel data of the captured image frame that has not been modified by the image transformation and / or encoding operation. However, this is just one example, and other combinations of sequential and / or parallel image processing operations are possible.
[0067] The operation of the image processing stage 140 can be implemented in both hardware and software. In the software implementation, the operation may be performed by one or more processors, such as one or more central processing units, which work with computer program code instructions stored in a (non-temporary) computer-readable medium, such as non-volatile memory, to perform the image processing operations 141, 142, etc., of the image processing stage 140. Examples of non-volatile memory include read-only memory, flash memory, ferroelectric RAM, magnetic computer storage devices, and optical discs. In the hardware implementation, the image processing stage 140 may instead be realized by a dedicated circuit configured to implement the image processing operations 141, 142, etc. The circuit may take the form of one or more integrated circuits, such as one or more application-specific integrated circuits (ASICs) or one or more field-programmable gate arrays (FPGAs). It is also possible to have a combination of hardware and software implementations, which means that some operations may be implemented by dedicated circuits and others by software.
[0068] In an example where the image processing stage 140 is provided in the camera 110, the one or more processors and computer-readable media and / or dedicated circuitry implementing the image processing stage 140 may be provided in the camera 110, for example, as part of the camera 110's (overall) image processing pipeline.
[0069] In examples where the image processing stage 140 is provided on an external device, the one or more processors and computer-readable media described above, or dedicated circuitry implementing the image processing stage 140, may be provided on the external device, for example, as a processing block for a server-side image processing stage. In this case, it should be noted that the camera 110 may still have an image processing pipeline and at least some basic image processing stages to facilitate further transmission, storage, processing, and other handling of the captured image frames. For example, the camera 110 may have a raw image transformation stage with raw image mosaic removal and / or noise reduction. To reduce the bandwidth requirements for transmitting image data (e.g., captured image frames 300) from the camera 110 to an external device, the camera 110 may additionally include an encoding block for encoding the image data prior to transmission to the remote device. The external device may decode the encoded image data to provide the decoded image data for further processing by the image processing stage 140.
[0070] As described above, the processing device 150 is configured to implement a method for controlling the image processing stage 140. More specifically, the processing device 150 is configured to implement an image processing-based technique for predicting whether the camera 110 is positioned in a ceiling-mounted configuration. The processing device 150 is configured to cause the image processing stage 140 to operate according to a first operating mode, referred herein as the “ceiling operating mode,” in response to such a detection, and otherwise according to a second operating mode, referred herein as the “non-ceiling operating mode.” This will be described in more detail below with reference to Figure 4, along with Figures 6a-6b, 7a-7b, and the flowcharts in Figures 12 and 13.
[0071] Figure 6a is a schematic diagram of image data in the form of an image frame 300 captured by camera 110 in scene 10.
[0072] Camera 110 is assumed in the illustrated example to produce a circular image of exposed pixels 310, in other words, a circular image area 310 within a rectangular frame. The image or image area 310 is bounded by an edge or perimeter pixel indicated by a solid line E. Edge pixels are pixels that capture the scene at the maximum field of view within the camera 110's FOV. The area of the image frame 300 outside edge E is formed by unexposed pixels. The image area 310 of the image frame 300 further comprises a first set of pixels or central pixels 312 that depict the central scene portion 10a, and a second set of pixels or peripheral pixels 314 that depict the peripheral scene portion 10b. The dashed line H indicates the location of the horizontal line in the image frame 300, in other words, the “horizontal line pixel” corresponding to the boundary between the central pixel 312 and the peripheral pixels 314. Point O corresponds to the optical center of the image frame 300 and / or image area 310 and coincides with the optical axis O of the camera 110, thereby allowing them to share the same reference numeral.
[0073] The location of the horizontal line H in the image frame 300 (in other words, the coordinates of the horizontal line pixels) can be known in advance, for example, by determining which pixels in the image frame 300 correspond to a 90-degree field of view with respect to the optical axis O of the camera 110. Thus, the processing device 150, and optionally the image processing stage 140, may have access to predetermined horizontal line data (e.g., predetermined horizontal line coordinate data) that indicates which pixels in the image frame 300 belong to / constitute a first set of pixels 312 and a second set of pixels, respectively. The horizontal line data may indicate, for example, the coordinates of the horizontal line pixels, where pixels inside the horizontal line pixels may be associated with the first set of pixels 312, and pixels outside the horizontal line pixels may be associated with the second set of pixels 314. While horizontal line pixels generally belong to the second set of pixels 314, it is also possible to consider horizontal line pixels as belonging to the first set of pixels 312.
[0074] For simplicity, the illustrated example shows the image frame 300 having a circular image area 310, but it should be noted that other types of fisheye lenses are also possible, such as cropped circular fisheye lenses or diagonal fisheye lenses, as long as the (cropped) image area still depicts the horizon line H and the peripheral scene portion 10b. The pixels of the image area 310 may be mapped to cover the entire area of the image frame 300 (for example, as a preprocessing step for camera 110), and the edges E of the image area 310 become the edges of the image frame 300. In general, this disclosure is applicable to cameras 110 having fisheye lenses or other lens configurations with an FOV greater than 180 degrees, thereby the image data 300 captured by the surveillance camera 110 when mounted in a downward-looking configuration includes a first set of pixels 312 depicting the central scene portion 10a located below the horizon line H and a second set of pixels 314 depicting the peripheral scene portion 10b located above the horizon line H.
[0075] Figure 12 is a flowchart of method 400 for controlling the image processing stage 140. Some steps of method 400 (for example, steps S1 to S5) may be performed as part of an initialization procedure for camera 110, following, for example, the installation or deployment of camera 110. The initialization procedure may be triggered by an operator inputting an initialization signal, for example, via a dedicated button, switch or other actuator on camera 110, or when an initialization signal is received from a remote control device (for example, a server) via a communication network. Camera 110 may be configured to automatically start the initialization procedure when powered on.
[0076] As described in relation to Figure 4, the camera 110 may be equipped with an orientation sensor 160. In this case, the method 400 may optionally include, as an initial step S1, detection of whether the camera 110 is in a downward-looking configuration. This can be detected based on an orientation signal output by the orientation sensor 160. For example, a downward-looking configuration can be detected in response to an orientation signal indicating that the camera 110 is oriented at an angle of 90 degrees with respect to the horizontal plane (within some predetermined tolerance, such as ±5 or ±10 degrees). The detection may be performed by the orientation sensor 160, and the detection signal indicating that the downward-looking orientation of the camera 110 has been detected can be supplied to the processing device 150 as a trigger to proceed with the method 400. It is also possible for the orientation sensor 160 to output an orientation signal to the processing device 150, which can then perform the detection.
[0077] In response to detecting the downward orientation of the camera 110 in step S1, the method proceeds so that the processing device 150 acquires image data from the surveillance camera 110 (step S2). In this example, where the camera 110 comprises a fisheye lens 120 and a single image sensor 130, the image data is acquired in the form of image frames 300 of the scene 10 captured by the image sensor 130. The processing device 150 may acquire the image frames 300 by outputting a control signal that causes the camera 110 to capture the image frames 300. The image frames 300 may then be provided to the processing device 150 for analysis. If the processing device 150 is provided with the camera 110, the image frames 300 may be stored in the camera 110's memory or buffer, and the processing device 150 may simply read the image frames 300 from there. If the processing device 150 is located on an external device (for example, an external camera controller or server), the camera 110 may transmit image frames 300 to the processing device 150 via a communication interface (for example, a network).
[0078] In step S3, the processing device 150 performs a ceiling detection procedure which comprises analyzing a subset of pixels of at least a second set of pixels 314 in order to determine whether the image processing stage 140 should be configured to operate according to the ceiling operating mode. More specifically, as will be further described below, this analysis allows the processing device 150 to reasonably estimate or predict whether the camera 110 is in a ceiling-mounted configuration, for example, as shown in Figure 1, or whether it is not in a ceiling-mounted configuration but is located at a relatively close distance from the ceiling 12, as shown in Figure 2. This prediction can be used as a basis for determining the configuration of the image processing stage 140.
[0079] Next, a method for analyzing pixels in step S3 of method 400, which is applicable to image data captured by a single-sensor surveillance camera such as camera 110, is disclosed in detail with further reference to the flowcharts in Figures 7a-7b and Figure 13.
[0080] Figures 7a and 7b are schematic diagrams of the contrast index C determined for the pixels of a second set 314 of pixels in the image frame 300 of Figure 6a under two different mounting scenarios, which will be discussed further below. More specifically, referring further to Figure 6b, the contrast index C is determined for each of the pixels or sets of pixel blocks 316a of the subset 316 of pixels in the second set 314.
[0081] A set of pixels or pixel blocks 316a are distributed in the radial direction R of the image frame 300 (for example, from the center O toward the edge E) as their distance from the horizontal line pixel H increases. The set of pixels or pixel blocks 316a are therefore distributed between the horizontal line pixel H and the edge E of the image area 310. This allows for the determination of a sequence of contrast indices C, where each contrast indices in the sequence is determined for each pixel or pixel block 316a in the corresponding sequence of pixels or pixel blocks 316a distributed in the radial direction R. Thus, Figures 7a and 7b illustrate how the contrast indices C of pixels or pixel blocks 316a vary depending on their location X along the radial direction R. Location X can be expressed, for example, in terms of distance from the horizontal line pixel H (pixel distance), or it can represent the position (index) of the contrast indices C in the sequence of contrast indices (for example, a greater distance from the horizontal line pixel H corresponds to a later position in the sequence). Similarly, this means that each pixel or pixel block in the set of pixels or pixel blocks 316a, and its associated contrast index C, corresponds to the respective field of view within the subrange of the field of view of the camera 110's FOV covering the peripheral scene portion 10b.
[0082] Whether the contrast index C should be determined for a set of individual pixels or for a set of pixel blocks may depend on factors such as available computing resources and the number of contrast samples expected to enable reliable analysis. For pixel block-based contrast index C, similar considerations apply to the dimensions and number of pixel blocks 316a. For example, a pixel block 316a may have dimensions of 4x4 pixels, 16x16 pixels, or more, to give some non-limiting examples.
[0083] The contrast index C for each pixel or pixel block 316a may be calculated as the difference between the pixel value and the average pixel value of each pixel or pixel block divided by the average pixel value. The pixel value may be pixel intensity (e.g., luminance). The average pixel value may be the average pixel value of a second set of pixels 314, or the average pixel value of a subset of pixels 316 comprising a set of pixels or pixel blocks 316a. If the contrast index C is determined for each pixel, the pixel value may be the pixel value of each pixel. If the contrast index C is determined for a pixel block, the pixel value of the pixel block may be a representative pixel value of the pixel block, such as the average pixel value of the pixel block, or a single sampled pixel value of the pixel block. In the case of a pixel block-based contrast index C, the contrast index C for each pixel block 316a may be calculated as a local contrast index C for each pixel block 316a, in other words, as the difference between the representative pixel value and the average pixel value of each pixel block divided by the average pixel value of each pixel block. The representative pixel value for each pixel block can, in this case, be a single sampled pixel value for the pixel block (for example, the pixel value of the center pixel of the pixel block).
[0084] If the surrounding scene portion 10b includes a ceiling, the contrast may be defined, for example, by structural features (e.g., beams, light sources, ceiling tiles, etc.) and / or by local variations in the ceiling texture. In general, even ceilings with a relatively uniform visual appearance can produce different contrasts at the pixel level.
[0085] Figure 7a shows the contrast index C when the peripheral scene portion 10b does not include the ceiling (calculated, for example, according to one of the methods described above). Figure 7b shows the contrast index C when the peripheral scene portion 10b includes the ceiling 12. That is, Figure 7a may correspond to the scenario shown in Figure 3. Figure 7b may correspond to the ceiling mounting configuration shown in Figure 1, for example.
[0086] In each figure, the horizontal line H coincides with the C-axis. In all cases, the contrast index C shown is expected to be low for pixels that draw the horizontal line H, as the horizontal line H is located at infinity. At larger field of view angles (for example, further from the horizontal line H towards edge E), the contrast index C is expected to gradually increase as the distance to the object being imaged decreases and therefore gradually approaches the DOF of camera 110. If the peripheral scene portion 10b does not include the ceiling, the contrast index C may continue to increase and reach a maximum value at the maximum field of view, or it may level off. Leveling off may occur, for example, if camera 110 is mounted on an outdoor pole (for example, as in Figure 3) and the peripheral scene portion 10b drawn in the image frame 300 includes a cloudless or overcast sky (which may have low contrast). On the other hand, the maximum contrast index C at the widest field of view may occur, for example, when the camera 110 is suspended below the ceiling 12 at a certain distance (as in Figure 2) and the ceiling 12 is close to or within the camera 110's DOF. However, as shown in Figure 7b, when the camera 110 is mounted flush with or close to the ceiling 12 (as in Figure 1), the pixels rendering the scene at the most extreme field of view (in other words, edge pixels E) tend to be out of focus and therefore produce a low contrast index C.
[0087] Therefore, the scenarios discussed above tend to result in different variations in the contrast index C. Thus, by analyzing the variation in the contrast index C in the radial direction R, it can be determined whether the camera 110 is in a ceiling-mounted configuration. In particular, mounting the camera 110 flush with or close to the ceiling 12, as shown in Figure 7b, results in a peak P of the contrast index C. C This tends to result in the following: Therefore, the variation in the contrast index C in the radial direction R is at its peak P C By detecting whether a second set of pixels 314 is defined, it can be determined whether the ceiling 12 is drawn.
[0088] Therefore, as shown in Figure 13, the ceiling detection procedure performed in step S3 may comprise several substeps.
[0089] In step S31, the processing device 150 determines a contrast index C for each pixel of the second set of pixels 314 or for each set of pixel blocks 316a using one of the methods described above.
[0090] In step S32, the processing device 150 performs the peak P C To identify the presence of [something], the variation in contrast index C in the radial direction R is analyzed. More specifically, the analysis shows that the sequence of contrast index C peaks P C It may include the ability to identify whether or not it has the necessary features.
[0091] For example, in response to the processing device 150 identifying at least one pixel or pixel block 316a where the contrast index C exceeds the contrast index C by at least a threshold amount for one or more adjacent pixels in the pixel block 316a on both sides (for example, both the side closer to the horizontal line pixel and the side further away from the horizontal line pixel, or correspondingly both the front and back sides in the sequence of contrast index C), the processing device 150 generates a peak PC The presence of can be determined. To enhance robustness, such peaks are considered candidate peaks and the identified peak P C may be subject to additional conditions to ultimately be considered. For example, a candidate peak is one or more of the following conditions: the candidate peak defines the global maximum of the sequence of contrast metric C, the candidate peak has a contrast metric C that exceeds a global threshold or is different from the global minimum of the sequence of contrast metric C by at least a threshold amount, the rate of increase of the sequence of contrast metric C on both sides of the candidate peak exceeds a rate threshold. Only when these are satisfied, the identified peak P C can ultimately be considered.
[0092] In a further example, peak detection can be achieved, for example, by identifying a pixel or pixel block 316a where the derivative of the contrast metric C (dc / dX) as a function of location X has a zero crossing. Each zero crossing can be considered a candidate peak. One or more filtering steps can be applied to reduce the risk of minor and / or slow fluctuations that cause false positives. For example, the second derivative of the contrast metric C (d 2 C / dX 2 ) can also be calculated, and zero crossings with a second derivative having an absolute value smaller than a threshold can be excluded. Further, the local height of the candidate peak at each zero crossing can be calculated, and candidate peaks having a local height smaller than a threshold can be excluded. The local height of a candidate peak can be calculated by subtracting the local surrounding contrast metric C of the candidate peak from the maximum value of the candidate peak (in other words, the contrast metric C at the location X corresponding to the zero crossing).
[0093] The peak identification algorithm discussed above is only an example, and any other algorithm for identifying the presence of peaks in a one-dimensional dataset can be used.
[0094] In Figure 6b, pixels or pixel blocks 316a are shown as continuous, but the set of pixels or pixel blocks 316a for which the contrast index C is determined may be more sparsely distributed, thereby separating subsequent pixels or pixel blocks by only a few pixels in the radial direction R. Furthermore, it should be noted that, generally, it is not necessary to determine the sequence of contrast index C over the entire distance from the horizontal pixel H to the edge E in the radial direction R. Rather, the contrast index C may be determined for pixels of pixel blocks 316a distributed over only a portion of the distance, such as the major portion of the distance. However, the contrast index C may be advantageously determined for pixels of pixel blocks 316a distributed both in front of and behind the focal length of the camera 110, thereby determining the peak P in the sequence of contrast index C. C The presence of [something] may be detected.
[0095] Processing device 150, Peak P C In response to identifying (for example, at least one peak remaining after filtering), method 400 proceeds to step S33 according to the "yes" branch. Processing device 150 identifies peak P in contrast index C. C In response to not identifying the answer, method 400 proceeds to step S34 according to the "no" branch.
[0096] Referring again to Figure 12, the method steps for the "ceiling mode" branch and the "non-ceiling mode" branch are shown, respectively. Method 400 reaches peak P in step S3 / S33. C If identified, proceed according to the ceiling mode branch. At the ceiling mode branch S4, the processing device 150 configures the image processing stage 140 to operate according to the ceiling operating mode. Method 400 performs peak P in step S3 / S34. CIf none is identified, the process proceeds according to the non-ceiling mode branch. At S5 of the non-ceiling mode branch, the processing device 150 configures the image processing stage 140 to operate according to the non-ceiling operating mode.
[0097] After configuring the image processing stage 140 in S4 or S5, the initialization procedure may be terminated. The camera 110 may then enter a monitoring operation, which may proceed to monitor scene 10 by capturing image frames 300 of scene 10. The captured image frames 300 are provided to the image processing stage 140 and subjected to at least one of its image processing operations (e.g., 141, 142, ..., 14n shown in Figure 5), which may provide processed image frames 301, for example, in the form of a video stream as discussed above. That is, the image processing stage 140 processes each image frame 300 captured by the camera 110 during the monitoring operation after the configuration of the image processing stage 140. Each of such image frames 300 will hereafter be referred to as a subsequently captured image frame, or interchangeably, a “captured image frame”. Each captured image frame 300 has content corresponding to the image frame 300 shown in Figure 6a, and it should be noted that this includes a first set of pixels 312 for rendering the central scene portion 10a and a second set of pixels 314 for rendering the peripheral scene portion 10b. Therefore, Figure 6a and the reference numerals therein are also used with respect to subsequently captured image frames 300.
[0098] In accordance with the non-ceiling mode branch, in step S5, the processing device 150 configures the image processing stage 140 to operate according to the non-ceiling operating mode. The non-ceiling operating mode means that the image processing stage 140 is configured to apply at least one of each of the image processing operations 141, 142, 14n to both a first set of pixels 312 and a second set of pixels 314 of each subsequently captured image frame 300 (step S9). As a result, each processed image frame 301 output by the image processing stage 140 includes processed first and second sets of pixels corresponding to the first and second sets of pixels 312, 314. If the image processing stage 140 includes an image analysis operation, such as object detection and / or object tracking, the image analysis operation may involve analysis of both the first and second sets of pixels 312, 314. As a result, for example, an object may be detected and / or tracked even within the surrounding scene portion 10b. In a sense, this means that in non-ceiling operation mode, the image processing stage 140 operates essentially as in conventional implementations, in other words, it processes each captured image frame 300 entirely.
[0099] In contrast, according to the ceiling mode branch, in step S4, the processing device 150 configures the image processing stage 140 to operate according to the ceiling operating mode. The ceiling operating mode means that the image processing stage 140 is configured to apply each of at least one image processing operation 141, 142, 14n to a first set of pixels 312 of each captured image frame 300, but not to a second set of pixels 314 of each captured image frame 300 (step S8). The second set of pixels 314 is therefore excluded from processing by at least one image processing operation 141, 142, 14n so that at least one image processing operation 141, 142, 14n is applied selectively to the first set of pixels 312 / only to the first set of pixels 312. Thus, processing of the second set of pixels 314 (in this case, including ceiling pixels) can be avoided. Therefore, each processed image frame 301 output by the image processing stage 140 contains processed pixels corresponding only to the first set of pixels 312 in this case.
[0100] If the image processing stage 140 includes an image transformation operation and / or encoding operation, the image transformation operation and / or encoding operation may be applied only to the first set of pixels 312. Therefore, the output of such an operation will include only the processed counterpart of the first set of pixels 312, and may not include data derived from or corresponding to the second set of pixels 314.
[0101] If the image processing stage 140 includes image analysis operations, such as object detection and / or object tracking, the image analysis operations may involve analysis of only a first set of pixels 312. This allows, for example, an object to be detected and / or tracked only within the central scene portion 10a.
[0102] In any case, the ceiling mode branch of this method may optionally include a step S7 for discarding a second set 314 of pixels in each captured image frame 300. The discard step can be carried out, for example, by a pixel discard or cropping block located at the input of the image processing stage 140, for example, upstream of the first image processing operation 141 of the image processing stage 140 (in other words, before the first image processing operation 141). The discard step can also be implemented by configuring at least one of the image processing blocks 141, 142, 14n to ignore or skip the second set 314 of pixels (when in ceiling operation mode) and output image data consisting only of the processed counterparts of the first set 312 of pixels. The processing device 150 can also be configured as a relay for the captured image frames 300 between the image sensor 130 and the image processing stage 140. In this case, the discard step S7 may instead be performed by the processing device 150, which may forward the cropped image frame 300, which thus contains only the first set of pixels 312 and omits the second set of pixels 314, to the image processing stage 140, thereby configuring the image processing stage 140 to operate in ceiling operation mode by providing the cropped image frame 300 to the image processing stage 140.
[0103] It should be noted that at least one image processing operation 141, 142, and 14n discussed above refers to an image processing operation of the image processing stage 140, each of which responds to the ceiling and non-ceiling mode configurations of the image processing stage 140. However, it is not ruled out that the system may include one or more further image processing operations that are configured "statically," in other words, whose operations are independent of the operating modes of the image processing stage 140. One non-limiting example of such an image processing operation may be one or more operations of raw image transformation. As one non-limiting example, subjecting all captured image frames 300 to raw transformation may facilitate image processing for image-based ceiling detection performed by the processing device 150, as well as subsequent operations of the image processing stage 140 and the calculation of pixel statistics (discussed below).
[0104] In addition to configuring the image processing stage 140, it is further possible to control one or more settings of the camera 110 in response to the ceiling detection procedure. Thus, the method 400 may include controlling a first setting of the camera 110 based on a first pixel statistic in an optional step S6 of the ceiling mode branch (which is thus performed in response to configuring the image processing stage 140 according to the ceiling operating mode), where the first pixel statistic is based on pixels that draw the central scene portion 10a but not on pixels that draw the peripheral scene portion 10b. Thus, pixels that draw the peripheral scene portion 10b may also be ignored / excluded from consideration for the purpose of controlling the first setting of the camera 110. The first setting here refers to the setting used by the camera 110 in step S8, in other words, during the capture of image frames 300 during the monitoring operation.
[0105] The first setting of camera 110 may be the setting of one or more exposure-related control parameters of camera 110 (e.g., shutter speed, aperture, ISO value, and / or camera illumination). The first pixel statistic may indicate the illumination state in the central scene portion 10a. The first pixel statistic may be based on the intensity (e.g., luminance) of the pixels that render the central scene portion 10a. As stated above, in a ceiling-mounted configuration, the peripheral scene portion 10b may be relatively bright and often contain light sources. Therefore, controlling the exposure-related control parameters based on the pixels that render the peripheral scene portion 10b (and thus the ceiling 12) may result in the exposure setting of camera 110 being unsuitable or at least not optimal for the illumination state in the central scene portion 10a, for example, thereby causing the first set of pixels 312 in the captured image frame 300 to be underexposed on average. However, better exposure of the first set of pixels can be obtained by excluding the second set of pixels 314 that render the peripheral scene portion 10b from the derivation of the first pixel statistic.
[0106] An example of a further (first) setting of the camera 110, which may be set based on a first pixel statistic determined while excluding pixels that render the peripheral scene portion 10b, is the white balance. Thus, the white balance can be set while avoiding undesirable bias introduced by ceiling pixels.
[0107] Each of the first pixel statistics mentioned above may be determined based on a first set 312 of one or more pixels in the subsequently captured image frame 300, in other words, the image frame 300 that is processed by the image processing stage 140 to become the processed image frame 301. However, the first pixel statistics may also be determined based on a first set of pixels in one or more dedicated measurement image frames of the scene 10 captured by the camera 110. The image frame 300 shown in Figure 6a is also a representative example of such a measurement image frame. The measurement image frame may be captured and interleaved with the image frame 300 for the purpose of collecting one or more pixel statistics to facilitate control of the camera 110.
[0108] As an addition or alternative, method 400 may include, in step S6, controlling a second setting of camera 110 based on a second pixel statistic, the second pixel statistic being based on pixels that draw the peripheral scene portion 10b but not on pixels that draw the central scene portion 10a. Thus, pixels that draw the peripheral scene portion 10b may be used for the purpose of controlling some camera settings of camera 110. The second setting here refers to the setting used by camera 110 in step S8, in other words, during the capture of image frames 300 during surveillance operation.
[0109] A second setting for camera 110 may be the frame rate setting for capturing image frames 300. A second pixel statistic may indicate the frequency of time variations in the lighting conditions in the peripheral scene portion 10b. The second pixel statistic may be based on the intensity (e.g., luminance) of the pixels that render the peripheral scene portion 10b. If the ceiling 12 contains a light source, flickering of the light source may have an adverse effect on image quality. Therefore, by controlling the frame rate for capturing image frames 300 based on the pixels that render the peripheral scene portion 10b, the capturing process can be controlled to reduce the effect of flickering light in the peripheral scene portion 10b. By excluding the pixels that render the central scene portion 10a from the derivation of the second pixel statistic, the amount of pixel data that needs to be processed to detect flickering light can be reduced.
[0110] Similar to the above description of the first pixel statistics, the second pixel statistics may be determined based on a second set 314 of pixels in a sequence of image frames 300 that are subsequently captured, in other words, processed by the image processing stage 140 to become a processed image frame 301, or a sequence of measured image frames.
[0111] While the image processing stage 140 operates in ceiling operation mode, the method may further be equipped with determining additional pixel statistics, such as a third pixel statistic, based on pixels that render the peripheral scene portion 10b, but not on pixels that render the central scene portion 10a. The pixels here may be either a second set 314 of pixels in the captured image frame 300, or a second set 314 of pixels in the measurement image frame. The third pixel statistic does not need to be used to control the camera 110, as in the example above. Instead, the third pixel statistic may be output as diagnostic data. In one example, the third pixel statistic may show the trend of the time variation of the pixel intensity of one or more pixels in the second set of pixels in each image frame of a sequence of image frames captured by the camera 110. For example, dirt may accumulate on the lens and / or image sensor of the camera 110 over time, leading to a degradation of overall image quality. Such degradation can be detected and indicated in diagnostic data by monitoring the trend of the intensity of one or more pixels in a second set of pixels in an image frame captured over a period of time (such as several days, several weeks, or several months).
[0112] The calculation of the different pixel statistics discussed above, and the associated control of camera settings (where applicable), may be performed by the processing device 150. However, the pixel statistics may also be calculated by a separate pixel statistics block provided in the camera 110. This may be useful, for example, when the processing device 150 is located outside the camera 110.
[0113] To further improve the reliability of the ceiling detection procedure, and in particular to reduce the risk of false positives (in other words, incorrectly predicting that camera 110 is in a ceiling-mounted configuration), the contrast-based analysis discussed above may be applied in two or more directions. For example, step S3 of method 400 in Figure 12 may comprise performing steps S31-S32 of the flowchart in Figure 13 for each set of pixels or pixel blocks in one or more further radial directions. For example, steps S31-S32 may be performed for two or more sets of pixels or pixel blocks located along two or more radial directions, such as along two or more of the radial directions R, R', and R'' indicated in Figure 6a. The processing device 150 then processes the peak P C However, if detection occurs in at least a majority of the radial directions, or in all radial directions, or in at least a predetermined minimum number of radial directions, it may be determined to configure the image processing stage 140 to operate according to the ceiling operating mode.
[0114] The operation of the processing device 150 disclosed above and below can be implemented both in hardware (e.g., in one or more integrated circuits such as an ASIC or FPGA) and in software (e.g., as computer program code instructions stored in a non-temporary computer-readable medium, executed by one or more processors of the processing device 150), similar to the image processing device 140.
[0115] In the above description of the contrast index-based method, for simplicity, the references mainly refer to the mounting configurations in Figures 1 and 3, respectively. However, as previously indicated, the camera 110 is suspended from the ceiling 12, as in Figure 2, but at a relatively short distance, so that the ceiling operating mode may also be useful when the peripheral scene portion 10b mainly consists of the ceiling 12. As can be understood from the above description of Figures 7a and 7b, when the camera 110 is mounted as in Figure 2, the greater the distance to the ceiling 12, the more the variation in the contrast index will resemble that of Figure 7a. Conversely, the smaller the distance to the ceiling 12, the more the variation in the contrast index will resemble that of Figure 7b. The determinant for the method (in other words, what peak P in the sequence of contrast index C) C Parameters for peak identification that can be used to adjust (how to consider) include, for example, the various thresholds discussed in relation to peak identification algorithms. For example, if the peak identification algorithm becomes more inclusive, this will allow for the identification of more slowly fluctuating contrast index C (in other words, wider peaks) as well as peak P C By adjusting one or more parameters (e.g., thresholds) so that it can be considered as such, the ceiling operating mode can also be applied to mounting configurations where the camera 110 is mounted at a greater distance from the ceiling 12.
[0116] Figure 8 shows a block diagram of system 200, which corresponds to system 100 in Figure 4, but differs in that it has a multi-sensor (surveillance) camera 210 instead of a single-sensor camera 110. The camera 210 thus comprises a lens configuration 220 and at least first and second image sensors 231, 232, 233, each image sensor 231, 232, 233 positioned behind the respective lenses 221, 222, 223 of the lens configuration 220 to capture a portion of the scene 10 from its respective viewpoint. The lens configuration 220 and the image sensors 231, 232, 233 may be positioned behind a transparent cover 212 (which may be dome-shaped, for example). The illustrated example depicts three lenses 221, 222, and 223 and image sensors 231, 232, and 233, but this is just one example, and other configurations are possible, such as only two lenses and two image sensors, or four lenses and four image sensors, or more lenses and image sensors. In any case, each lens and associated image sensor ("lens-image pair") may have its respective partial FOV corresponding to a portion of the camera 210's total FOV. Lens-image pairs may be arranged such that the respective partial FOVs of adjacent lens-image pairs partially overlap, and so that the lens-image pairs collectively define the total FOV beyond 180 degrees. The reference numeral O here refers to the common optical axis of the combined optical system of the lens-image pairs of the camera 210.
[0117] During operation, image sensors 231, 232, and 233 can each capture their respective partial image frames 321, 322, and 323, which depict their respective parts of scene 10, such that the partial image frames 321, 322, and 323 collectively cover the entire FOV of camera 210. Thus, the image data 320 captured by camera 210 in each capture occasion comprises each of the partial image frames 321, 322, and 323. During monitoring operation, the partial image frames can thus be merged into a composite image frame using image stitching, as is known in the art. Image stitching can be performed by a stitching block 260. As indicated in Figure 8, the stitching block 260 may be located upstream of the image processing stage 140, or optionally, as an image processing block 260 of the image processing stage 140. These implementation options will be discussed further below.
[0118] Figure 9 schematically shows an example of such a composite image frame 330 formed by stitching together partial image frames captured by each image sensor of camera 210. In the illustrated example, the composite image frame 330 is formed from the first, second, and third partial image frames 321, 322, and 323 captured by the first, second, and third image sensors 231, 232, and 233, respectively, and also from a fourth partial image frame 324 captured by each of the fourth image sensors of camera 210, which is not shown in Figure 8 for clarity of illustration. Each of the first, second, and third image sensors 221, 222, and 223, as well as the fourth image sensor, has a respective partial FOV that substantially covers each quadrant of scene 10 (when viewed in the horizontal plane) and overlaps to some extent with its adjacent image sensor.
[0119] In the illustrated example, the rendering of scene 10 in the composite image frame 330 is shown as covering the entire rectangular area of the image frame 330. However, it is also possible to form a composite image frame 330 that renders the scene within a circular image area, for example, as in Figure 6a, which covers only a portion of the image frame 330. That is, the shape and dimensions of the image area on which scene 10 is rendered may vary depending on the type of stitching algorithm used to form the composite image frame 330.
[0120] Similar to the image frame 300 in Figure 6a, the composite image frame 330 comprises a first set of pixels or central pixel 331 that renders the central scene portion 10a, and a second set of pixels or peripheral pixels 332 that renders the peripheral scene portion 10b. The dashed line H points to the location of the horizontal line in the composite image frame 330, in other words, the "horizontal line pixel" corresponding to the boundary between the central pixel 331 and the peripheral pixels 332. Furthermore, point O corresponds to the optical center of the composite image frame 330 and coincides with the common optical axis O of the camera 210, thereby allowing them to share the same reference numeral.
[0121] Reference numerals 321, 322, 323, and 324 generally refer to the respective portions of the composite image frame 330 to which the pixels of the respective partial image frames 321, 322, 323, and 324 contribute. More specifically, the portion (sector) of the composite frame 330 indicated by reference numeral 321 and defined by a dotted line contains pixels from the first partial image frame 321. The portion indicated by reference numeral 322 and defined by a dashed line contains pixels from the second partial image frame 322. The portion indicated by reference numeral 323 and defined by a dashed line contains pixels from the third partial image frame 323. The portion indicated by reference numeral 324 and defined by a dashed line contains pixels from the fourth partial image frame 324.
[0122] The shaded areas indicate portions of the composite image frame 330 based on pixels from overlapping partial image frames. For example, the shaded area 334 indicates a portion of pixels based on pixels from the overlapping portion of the first partial image 321 and the second partial image frame 322. For example, in area 334, the pixels of the composite image frame 330 may be formed by mixing or combining pixels from the overlapping portions of the first and second partial image frames 321, 322 in some other way. The mixing or combining may be performed with the aim of producing a seamless transition between the partial image frames 321, 322, ideally without stitching errors. This is also true for any further shaded areas and each of the associated partial image frames. In Figure 9, the mixed / combined regions extend from the center O toward the approximate midpoint of each side of the rectangular image frame 330, but this is just an example, and the extension and orientation of the regions generally depend on the mapping used when stitching the partial image frames 321, 322, 323, and 324, and how the stitched image data is cropped.
[0123] The method 400 of Figure 12, described above with reference to camera 110 of Figure 4, can be applied in a corresponding manner to control the image processing stage 140 of camera 210 of Figure 8. Thus, steps S2 to S5 (and optional step S1) of method 400 can be performed as part of the initialization procedure for camera 210. This allows the processing device 150 to acquire image data from camera 210 in S2 of method 400 (for example, in response to detecting the downward viewing orientation of camera 210 using orientation sensor 160). Furthermore, in S3, a ceiling detection procedure can be performed, comprising analyzing the pixels of the image data according to the contrast-based technique described with reference to Figure 13. The acquired image data may in this case consist of one or more of the partial image frames 231, 232, 233, 234, or a composite image frame 330 generated by the stitching block 260, as shown in Figure 9. Therefore, in the former case, the second set of pixels analyzed in S3 and substeps S31-S34 may be the second set of pixels from any one of the partial image frames 231, 232, 233, and 234 that depict each portion of the surrounding scene portion 10b. That is, the contrast index C may be determined in S31 for each pixel or set of pixel blocks of the second set of pixels of the partial image frame (for example, any one of the partial image frames 231, 232, 233, and 234) that are distributed in the radial direction R of each partial image frame such that the distance from the horizontal line pixels of each partial image frame increases. Alternatively, in the latter case, the contrast index may be determined in S31 for each pixel or set of pixel blocks of the second set of pixels 332 of the composite image frame 330 that are distributed in the radial direction R of the composite image frame 330 such that the distance from the horizontal line pixels H of the composite image frame 330 increases. In either case, the variation of the determined contrast index is peak P CIn response to the identification of defining peak P, the image processing stage 140 may be configured to operate in S4 according to the ceiling operating mode. C In response to failing to identify the ceiling mode, the image processing stage 140 may instead be configured to operate according to the non-ceiling mode in S5. Method 400 may then proceed according to the "ceiling mode" branch or the "non-ceiling mode" branch described above.
[0124] The different viewpoints of the image sensors of camera 210 result in parallax between their respective partial views. As acknowledged by the inventors, this allows for an alternative to the contrast-based method described above for analyzing the image data in step S3 of method 400, namely a parallax-based method. This method will be described below with reference to Figure 8, and further with reference to Figures 10a-10b, 11a-11b, and the flowcharts in Figures 12 and 14. The method will be described with respect to first and second partial image frames 321, 322 captured by first and second image sensors 231, 232, respectively. However, the method is applicable to overlapping pairs of image frames captured by pairs of image sensors of camera 210.
[0125] In S2, the processing device 150 acquires image data 320 comprising first and second partial image frames 321 and 322 captured by first and second image sensors 231 and 232.
[0126] Figure 10a shows, in schematic form, the first and second partial image frames 321 and 322 of the image data 320 acquired by the processing device 150. The dashed line H indicates the location of the horizontal line pixels in the first and second partial image frames 321 and 322. Reference numerals E1 and E2 indicate the respective edges of the first and second partial image frames 321 and 322. Reference numerals R1 and R2 indicate the respective radial directions of the first and second partial image frames 321 and 322. The radial directions R1 and R2 may extend from their respective optical centers to their respective edges E1 and E2 of the first and second partial image frames 321 and 322.
[0127] The first and second partial image frames 321 and 322 each comprise a first set of pixels 3211 and 3221 for rendering each portion of the central scene portion 10a, and a second set of pixels 3221 and 3222 for rendering each portion of the surrounding scene portion 10b. As a result, the image data 320 comprises a first set of pixels 325 for rendering the central scene portion 10a, each comprising the first set of pixels 3211 and 3221 of the first and second partial image frames 321 and 322, and a second set of pixels 326 for rendering the surrounding scene portion 10b, each comprising the second set of pixels 3212 and 3222 of the first and second partial image frames 321 and 322.
[0128] In Figure 10a, the first and second partial image frames 321 and 322 are shown aligned. More specifically, the first and second partial image frames 321 and 322 are mapped to a common composite plane 360 so that the respective partial views drawn in the first and second partial image frames 321 and 322 are aligned. The first and second partial image frames 321 and 322 are thus mapped to a common coordinate system of the composite plane 360, schematically represented by the axis (u,v) in Figure 10a.
[0129] The shaded region 341 points to the pixels of the first and second partial image frames 321 and 322 that depict the overlapping portion of scene 10. The further shaded region 342, a sub-region of region 341, points to the respective subsets of pixels from the first and second partial image frames 321 and 322 that depict the overlapping portion of the surrounding scene portion 10b. That is, each subset of pixels within region 342 is formed by the pixels of the first and second partial image frames 321 and 322 that depict a portion of the surrounding scene portion 10b and are mapped to the same set of coordinates on the composite plane 360. These respective subsets of pixels from the first and second partial image frames 321 and 322 are hereafter referred to as the first and second subsets of pixels 3212a and 3222a, respectively (shown in Figure 10b). From the above, the first and second subsets of pixels 3212a and 3222a will form part of the second set of pixels 326 of the image data 320.
[0130] As schematically indicated by the displacement between edges E1 and E2 of each partial image frame 321 and 322, the alignment of each of those partial views does not necessarily result in the alignment of their respective edges E1 and E2.
[0131] The composite surface 360 may be defined, for example, as a spherical or cylindrical surface, and may generally be defined to be located at an infinite distance from the camera 210. The composite surface 360 may generally be the same composite surface used by the stitching block 260 during the monitoring operation to stitch together the partial image frames 321, 322, 323, and 324 captured by the camera 210 to form the stitched image frame 330 shown in Figure 9.
[0132] The mapping of pixels from the first and second partial image frames 321, 322 to the composite plane 360 may be based on the spatial relationship between the respective partial FOVs of the first and second image sensors 231, 232. If the arrangement of the first and second image sensors 231, 232 is fixed, the mapping may be predetermined by the manufacturer during the assembly of the camera 210, for example. In this case, the mapping may be implemented as a lookup table that defines the mapping between the coordinates of each pixel in each partial image frame 321, 322 and the coordinate system (u,v) of the composite plane 360. The lookup table may be stored as part of the configuration information that the processing device 150 can acquire (for example, by retrieving it from the memory area of the camera 210 or the stitching block 260). If the image sensors 231, 232 are equipped with motors so that they are movable with respect to the scene 10 (for example, by being rotated about the optical axis O of the camera 210), the mapping may be based on pose data that points to the current spatial configuration of each image sensor 231, 232. The spatial configuration can be represented with respect to the coordinate system (u,v), or to an external spatial reference frame having a predefined relationship with the coordinate system (u,v). Attitude data may be provided, for example, by the motor controller of camera 210, and may indicate the current attitude of each image sensor 231, 232.
[0133] Similar to the description of image frame 300 in Figure 6a, the processing device 150 may have access to horizontal line data that points to the location of each of the first and second partial image frames 321 and 322. The processing device 150 can then determine which pixels in the first and second partial image frames 321 and 322 belong to their respective second sets of pixels 3212 and 3222. The horizontal line data may, for example, point to the coordinates of the horizontal line pixels H in the first and second partial image frames 321 and 322, and pixels radially outside each horizontal line pixel H may be associated with their respective second sets of pixels 3212 and 3222.
[0134] The processing device 150 can thus determine or identify the first and second subsets 3212a, 3222a of pixels by mapping only the first and second partial image frames 321, 322, or the respective second sets of pixels 3212, 3222, to the composite plane 360, and then determining the pixels of the second sets of pixels 3212, 3222 of pixels from the first and second partial image frames 321, 322 that overlap with each other when mapped to the composite plane 360 (in other words, mapped to region 342), as the first and second subsets 3212a, 3222a of pixels.
[0135] After acquiring the image data 320, the processing device 350 proceeds to perform a ceiling detection procedure in step S3 by analyzing the first and second subsets 3212a, 3222a of pixels in a second set of pixels 326 in order to determine whether the image processing stage 140 should be configured to operate according to the ceiling operating mode. The analysis in step S3 comprises several substeps, which will be further described with reference to the flowcharts in Figures 10b and 14.
[0136] In step S31' of Figure 14, as shown in Figure 10b, a first set of feature points 351 is identified in a first subset of pixels 3212a. The first set of feature points 351 is distributed in the radial direction R1 such that the distance from the horizontal line pixels H in the first partial image frame 321 increases.
[0137] In step S32', as shown in Figure 10b, the matching second feature point 352 is then identified in the second subset 3222a of pixels for each first feature point 351. This allows steps S31' and S32' to determine sequences of matching pairs of first and second feature points. The second feature point 352, like the first feature point 351, may be distributed in the radial direction R2 of the second partial image frame 322, with increasing distances from the horizontal line pixels H of the second partial image frame 322.
[0138] A greater distance between the first feature point 251 and the horizontal pixel H of the first partial image frame 321 means that the first feature point 251 is located in a larger field of view within both the partial FOV of the first image sensor 231 and the full FOV of the camera 210. Similarly, a greater distance between the second feature point 252 and the horizontal pixel H of the second partial image frame 322 means that the second feature point 252 is located in a larger field of view within both the partial FOV of the second image sensor 232 and the full FOV of the camera 210. Thus, the first feature point 251 is distributed across a subrange of the field of view of the partial FOV of the first image sensor 231, the second feature point 252 is distributed across a subrange of the field of view of the partial FOV of the second image sensor 232, and the first and second feature points are distributed across a subrange of the field of view of the FOV of the camera 210. More specifically, each subrange here refers to a subrange of the field of view above the horizontal line H.
[0139] The first feature point 351 can be identified, for example, by analyzing a set of pixel blocks (indicated by the dotted outline in Figure 10b) of a first subset 3212a of pixels in the first partial image frame 321, which are distributed radially in R1 such that the distance from the horizontal line pixels H of the first partial image frame 321 increases. Within each pixel block, one feature point or group of feature points can be identified. For example, each pixel block can be analyzed using an edge detection algorithm, an angle detection algorithm, or a scale-invariant feature transformation (SIFT) to identify the feature points within it. If the surrounding scene portion 10b includes a ceiling, the feature point 351 can be defined, for example, by structural features (e.g., beams, light sources, ceiling tiles, etc.) and / or by local variations in the ceiling texture. The pixel blocks being analyzed in the first subset 3212a of pixels may be contiguous or more sparsely distributed such that subsequent pixel blocks are separated by only a few pixels radially in R1. The dimensions of the pixel blocks being analyzed in the first subset of pixels, 3212a, may depend on factors such as available computing resources and the number of feature points expected to be required to enable a reliable analysis. For example, a pixel block may have dimensions of 4x4 pixels, 16x16 pixels, 32x32 pixels, or larger, to give a few non-limiting examples.
[0140] The matching second feature point 352 can then be determined by searching for matching features in a second subset 3222a of pixels in the second partial image frame 322. The search for matching features can be performed on a pixel-block basis, as indicated in Figure 10b. Any conventional preferred feature matching algorithm can be used.
[0141] After determining the sets of matching pairs of first and second feature points 351 and 352, the analysis proceeds to step S33', where, for each first feature point 351, the disparity error of each first feature point 351 with respect to its matching second feature point 352 is determined.
[0142] The parallax error can be determined by calculating the distance between each first feature point 351 and the corresponding second feature point 352 when mapped to the common composite plane 360, in other words, within the common coordinate system (u,v). In Figure 10b, the distances between the corresponding first and second feature points 351, 352 are D1, D2, ... D n-1 , D n This is indicated by . When a first group of feature points is identified in a pixel block, it may generally suffice to calculate a single representative distance between the first group of feature points and a corresponding second group of feature points. For example, the representative distance can be calculated as the distance between each centroid of the group of feature points. This then gives the disparity errors D1, D2, ... D corresponding to the sequence of corresponding pairs of first and second feature points. n-1 , D n The sequence can be determined.
[0143] Figures 11a and 11b are schematic diagrams illustrating how the parallax error (calculated, for example, according to the method described above) tends to vary when the peripheral scene portion 10b does not include the ceiling (Figure 11a) and when the peripheral scene portion 10b includes the ceiling 12 (Figure 11b). Specifically, Figure 11a may correspond to the scenario shown in Figure 3. Figure 11b may correspond to the ceiling mounting configuration shown in Figure 1, for example. The figures show the parallax error D as a function of location X along the radial direction R1. Location X may be expressed, for example, in terms of the distance from the horizontal line pixel H in the first partial image frame 321, or it may represent the position (index) of the parallax error D in the sequence of parallax errors (for example, a greater distance from the horizontal line pixel H corresponds to a later position in the sequence). In each figure, the horizontal line H coincides with the D axis.
[0144] In both mounting scenarios in Figures 1 and 3, the parallax error D shown in Figures 11a and 11b is expected to be low among matching feature points close to the horizon line H, as the horizon line H is located at infinity. At larger field of view angles, the parallax error D is expected to gradually increase as the distance to the imaged object decreases and therefore gradually approaches the DOF of camera 210. As a result, as the field of view angle increases, the different viewpoints of the first and second image sensors 231 and 232 cause a gradually increasing parallax error between matching feature points in the first and second partial image frames. When the peripheral scene portion 10b includes the ceiling 12, the mismatch tends to increase more strongly than when there is no ceiling, as shown in Figures 10a and 10b. In fact, when the peripheral scene portion 10b in Figure 3 is formed by an open, clear sky, there may be no noticeable increase in parallax error at all.
[0145] Therefore, the scenarios discussed above tend to result in different variations in the parallax error D. Thus, the variation in the parallax error D in the radial direction R1 provides a useful basis for determining whether or not the image processing stage 140 should be configured according to the ceiling operating mode.
[0146] Therefore, the method proceeds, and in step S34', the variation of the parallax error D in the radial R1 is analyzed to determine whether the parallax error increases in the radial R1. For example, the processing device 150 analyzes the parallax errors D1, D2, ... D determined in step S33'. n-1 , D n This sequence can determine whether it defines a sequence that increases the disparity error.
[0147] To increase robustness, the analysis determines that the parallax error D for at least one of the first 351 feature points is equal to the magnitude threshold T. M Whether it exceeds and / or the rate of increase in parallax error in the radial R1 is the rate threshold T RIt may be possible to determine whether it exceeds a certain threshold. These thresholds are schematically indicated in Figures 11a and 11b. As can be seen, both the magnitude and rate of increase of the parallax error (e.g., dD / dX) are smaller in Figure 11a (where the peripheral scene portion 10b does not include the ceiling) than in Figure 11b (where the peripheral scene portion 10b includes the ceiling 12). Therefore, the condition for determining whether to configure the image processing stage 140 to operate according to the ceiling operating mode is the magnitude threshold T M However, this must be exceeded for at least a subset of matching feature points 251, 252, and / or rate threshold T R However, this may be exceeded over at least a portion of the sequence of parallax errors D.
[0148] Size threshold T M and rate threshold T R The value of can be set according to the desired sensitivity of this method. That is, the magnitude threshold T M and rate threshold T R A smaller value of may more frequently predict a ceiling mounting configuration, which may result in applying the ceiling operating mode in a wider range of scenarios (for example, even when the distance to ceiling 12 is greater, as in the scenario in Figure 2). Conversely, the magnitude threshold T M and rate threshold T R A larger value may result in less frequent prediction of ceiling mounting configurations, which could lead to a more selective application of ceiling operating modes.
[0149] Furthermore, it should be noted that, generally speaking, it is not necessary to determine and analyze the parallax error over the entire radial distance / field of view range from the horizontal line pixel H to edge E1 in the radial direction R1. Rather, it may be sufficient to limit the analysis to a portion of the distance / range. For example, the parallax error D may only be determined for the pixel block distributed between the horizontal line pixel H and a point located approximately 50-60% of the radial distance to edge E1. If the parallax error D over this radial distance increases at a rate exceeding a rate threshold and / or exceeds a magnitude threshold, it can be concluded with a relatively high degree of certainty that the camera 110 is in a ceiling-mounted configuration, and therefore the image processing stage 140 is configured to operate according to the ceiling operating mode.
[0150] To further improve the reliability of the ceiling detection procedure, and in particular to reduce the risk of false positives (in other words, incorrectly predicting that camera 210 is in a ceiling-mounted configuration), the disparity-based analysis discussed above may be applied in two or more radial directions, or more specifically, to two or more pairs of partial image frames. Thus, steps S31' to S34' may be applied to overlapping subsets of pixels from two or more pairs of partial image frames captured by each of two or more pairs of image sensors. For example, referring to Figure 9, the analysis may be applied to overlapping subsets of pixels (corresponding to subsets 3212a and 3222a in Figure 10b) from (overlapping) pairs of partial image frames, including one or more of the following: second partial image frame 322 and third partial image frame 323, third partial image frame 323 and fourth partial image frame 324, and first partial image frame 321 and fourth partial image frame 324. Therefore, a further condition for configuring the image processing stage 140 so that the processing device 150 operates according to the ceiling operating mode may be that the variation in the parallax error determined for each pair of partial image frames (in other words, the sequence of parallax errors) increases for at least a minimum number of pairs of partial image frames, such as for at least a majority of pairs of partial image frames or for all pairs of partial image frames.
[0151] In S34', in response to the determination that the disparity error increases in the radial direction R1 (or, if two or more pairs of partial image frames are analyzed, for at least a minimum number of pairs of partial image frames), the method proceeds to step S4 in S35', thereby configuring the image processing stage 140 to operate according to the ceiling operating mode (ceiling mode branch). On the other hand, in response to the determination in S34' that the disparity error does not increase in the radial direction R1 (or, if two or more pairs of partial image frames are analyzed, does not increase for at least a minimum number of pairs of partial image frames), the method proceeds to step S5 in S36', thereby configuring the image processing stage 140 to operate according to the non-ceiling operating mode (non-ceiling mode branch). The method may then proceed as discussed above with respect to contrast-based methods. Therefore, camera 210 can proceed to a monitoring operation, thereby capturing partial image frames (substantially simultaneously) by each of its image sensors (for example, partial image frames 321, 322, 323 captured by image sensors 231, 232, 233), and stitching the partial image frames by the stitching block 260 to form a composite image frame 330. This process can be repeated (for example, at the frame rate of camera 210) to generate a video sequence of the composite image frame 330.
[0152] As described above, the stitching block 260 may be positioned upstream of the image processing stage 140, or optionally, as an image processing block 260 of the image processing stage 140. In the former case, the stitching block 260 can thereby stitch together the simultaneously captured partial image frames and provide the stitched composite image frame 330 to the image processing stage 140. In the latter case, the stitching block 260 may be configured to exclude a second set of pixels of the partial image frame during stitching (when the image processing stage 140 is configured according to the ceiling configuration mode), thereby stitching together only the first set of pixels of the partial image frame to form the composite image frame 330. The composite image frame 330 can then be provided to one or more further downstream image processing operations of the image processing stage 140, such as image processing operations 141, 142, 14n, etc., as discussed above with reference to Figure 5.
[0153] The image processing-based analysis described above with reference to camera 110 and the flowcharts in Figures 12 and 13 is considered particularly useful and effective when the camera 110's FOV is at least 190 degrees (e.g., at least 200 degrees), and when the camera 110 is mounted on the ceiling such that the portion of the ceiling 12 closest to the camera 110, within the camera 110's FOV, is in front of (and therefore outside) the camera 110's DOF. In a typical scenario, the lower limit of the camera 110's DOF can be at least 2 meters, generally at least 3 meters, from the camera 110 (more specifically, from the image sensor 130). The upper limit of the DOF can vary, but for example, it can be at least 5 meters from the camera 110. Meanwhile, the distance between the image sensor 130 and the ceiling 12 (e.g., the distance measured along the optical path from the image sensor 130 to the closest drawing portion of the ceiling 12) can be up to 1 m, up to 0.5 m, or up to 0.1 m. This explanation also applies to further image processing-based techniques described with reference to the multi-sensor camera 210 in Figure 8 and the flowcharts in Figures 12 and 14.
[0154] Those skilled in the art will recognize that the present invention is by no means limited to the examples described above. Rather, many modifications and variations are possible within the scope of the appended claims.
Claims
1. A method for controlling an image processing stage (140) for processing image data (300) captured by a surveillance camera (110) having a field of view (FOV) larger than 180 degrees and mounted in a downward-looking configuration for monitoring a scene (10), The acquisition of image data (300) captured by the surveillance camera (110), wherein the image data includes a first set of pixels (312) that renders the central scene portion (10a) located below the horizontal line (H) in the scene (10), and a second set of pixels (314) that renders the peripheral scene portion (10b) located above the horizontal line (H). To determine whether the image processing stage (140) should be configured to operate according to the ceiling operating mode, a ceiling detection procedure is performed which involves analyzing the pixels (316a) of the second set (314) of pixels, In response to a decision to configure the image processing stage (140) to operate according to the ceiling operating mode, the image processing stage (140) is configured to operate according to the ceiling operating mode while processing subsequently captured image data (300), wherein the subsequently captured image data (300) comprises a first set (312) of pixels that draw the central scene portion (10a) and a second set (314) of pixels that draw the peripheral scene portion (10b), the image processing stage (140) comprises at least one image processing operation (141, 142), and the ceiling operating mode includes applying the at least one image processing operation (141, 142) to the first set (312) of pixels in the captured image data (300) but not to the second set (314) of pixels in the captured image data (300). A method that includes [a certain feature].
2. The method according to claim 1, wherein the optical axis (O) of the surveillance camera (110) is directed toward the ground (14) in the scene (10), and the horizontal line (H) in the scene (10) corresponds to an angle of 90 degrees with respect to the optical axis (O).
3. The method according to claim 2, further comprising: obtaining predetermined horizontal line data that indicates the pixel coordinates of the horizontal line (H) in the image data; and identifying the second set (314) of pixels using the predetermined horizontal line data.
4. The method according to any one of claims 1 to 3, wherein the at least one image processing operation (141, 142) includes at least one of an image conversion operation, an encoding operation, and / or an image analysis operation.
5. In response to the decision to configure the image processing stage (140) to operate according to the ceiling operating mode, To control a first setting of the surveillance camera (110) used by the surveillance camera (110) during the capturing of the subsequently captured image data (300), based on a first pixel statistic, wherein the first pixel statistic is based on pixels (312) that draw the central scene portion (10a) but not on pixels (314) that draw the peripheral scene portion (10b), and / or To control a second setting of the surveillance camera (110) used by the surveillance camera (110) during the capturing of the subsequently captured image data (300), based on a second pixel statistic, wherein the second pixel statistic is based on pixels (312) that render the peripheral scene portion (10b), but not on pixels that render the central scene portion (10a). The method according to any one of claims 1 to 4, further comprising:
6. The method according to any one of claims 1 to 5, further comprising discarding the second set (314) of pixels of the subsequently captured image data (300) in response to a decision to configure the image processing stage (140) to operate in accordance with the ceiling operating mode.
7. The method according to any one of claims 1 to 6, further comprising configuring the image processing stage (140) to operate according to a non-ceiling operating mode in response to a decision not to configure the image processing stage (140) to operate according to the ceiling operating mode, wherein the non-ceiling operating mode includes applying the at least one image processing operation (141, 142) to the first set (312) and the second set (314) of pixels of the subsequently captured image data (300).
8. The aforementioned image data is an image frame (300), Implementing the aforementioned ceiling detection procedure means Determining a contrast index for each of the pixels or pixel blocks (316a) of the second set of pixels (314) distributed in the radial direction (R) such that the distance from the pixel drawing the horizontal line (H) increases, The variation of the contrast index in the radial direction (R) is defined as the peak (P) of the variation of the contrast index. C To analyze whether to define ) and Equipped with, The conditions for determining that the image processing stage (140) is configured to operate according to the ceiling operating mode are the peak (P C The method according to any one of claims 1 to 7, wherein the ) is identified.
9. The ceiling detection procedure is, For each of at least one further radial direction (R', R''), a contrast index is determined for each of the pixels or pixel blocks (316a) of the second set of pixels (314) distributed in each of the radial directions (R', R'') such that the distance from the pixel drawing the horizontal line (H) increases; The respective fluctuations of the contrast index in each of the radial directions (R', R'') are defined as the respective peaks (P C To analyze whether to define ) and Furthermore, The conditions for determining that the image processing stage (140) is configured to operate according to the ceiling operating mode are the peak (P C The method according to claim 8, wherein the ) is arbitrarily identified for each of the radial directions (R, R', R'') with respect to at least a minimum number of the radial directions (R, R', R'').
10. The subsequently captured image data (300) comprises a sequence of image frames (300), each image frame comprising a first set of pixels (312) for rendering the central scene portion (10a) and a second set of pixels (314) for rendering the surrounding scene portion (10b), and the method further comprises processing the sequence of image frames (300) by the image processing stage (140). The method according to claim 8 or 9.
11. The surveillance camera (210) comprises a lens configuration (220) and at least first and second image sensors (231, 232, 233), each image sensor (231, 232, 233) positioned behind the lens configuration (220) such that each image sensor (231, 232, 233) captures a portion of the scene (10) from its respective viewpoint. Acquiring the image data (320) comprises acquiring a first partial image frame (321) captured by the first image sensor (231) and a second partial image frame (322) captured by the second image sensor (232), wherein the second set (326) of pixels in the image data (320) comprises a first subset (3212a) of pixels from the first partial image frame (321) and a second subset (3222a) of pixels from the second partial image frame (322), and the first and second subsets (3212a, 3222a) of pixels draw the overlapping portion (334) of the surrounding scene portion (10b). Implementing the aforementioned ceiling detection procedure means Identifying a set of first feature points (351) distributed radially (R1) such that the distance from the pixels that draw the horizontal line (H) in the first partial image frame (321) increases in the first subset of pixels (3212a), For each first feature point (351), identify the corresponding second feature point (352) in the second subset (3222a) of the pixels, For each first feature point (351), the disparity error with respect to the corresponding second feature point (352) is determined. In order to determine whether the parallax error increases in the radial direction (R1), the variation in the parallax error in the radial direction (R1) is analyzed. Equipped with, The method according to any one of claims 1 to 7, wherein the condition for determining that the image processing stage (140) is configured to operate according to the ceiling operating mode is that the parallax error increases in the radial direction (R1).
12. Mapping the first feature point (351) and the matching second feature point (352) onto the common composite surface (360). Furthermore, The method according to claim 10, wherein the parallax error for each first feature point (351) is determined by calculating the distance between the first feature point (351) and the second feature point (352) that coincides with the first feature point (351) when mapped to the common composite plane (360).
13. The second set (326) of pixels in the image data (320) further comprises a third subset of pixels from a third partial image frame and a fourth subset of pixels from a fourth partial image frame, wherein the third and fourth subsets of pixels draw the overlapping portion of the surrounding scene portion (10b). Implementing the aforementioned ceiling detection procedure means Identifying a third set of feature points distributed radially in the third subset of pixels such that the distance from the pixels drawing the horizontal line in the third partial image frame increases; For each third feature point, identify the matching fourth feature point in the fourth subset of pixels, For each third feature point, the disparity error with respect to the corresponding fourth feature point is determined. In order to determine whether the parallax error increases in the radial direction of the third partial image frame, the variation of the parallax error in the radial direction of the third partial image frame is analyzed. Furthermore, The method according to claim 11 or 12, wherein a further condition for determining that the image processing stage (140) is configured to operate according to the ceiling operating mode is that the parallax error determined for the third feature point increases in the radial direction of the third partial image frame.
14. The subsequently captured image data processed by the image processing stage (140) comprises a sequence of composite image frames (330), each composite image frame (330) being formed by stitching together partial image frames (321, 322, 323) captured by at least the first and second image sensors (231, 232, 233), such that each composite image frame (330) comprises a first set of pixels (331) for rendering the central scene portion (10a) and a second set of pixels (332) for rendering the peripheral scene portion (10b), or The subsequently captured image data processed by the image processing stage (140) comprises partial image frames (321, 322, 323) captured by at least the first and second image sensors (231, 232, 233), wherein the captured partial image frames (321, 322, 323), when combined, comprise a first set of pixels (325) for rendering the central scene portion (10a) and a second set of pixels (326) for rendering the peripheral scene portion (10b). The image processing stage (140) comprises at least one image processing operation (260) for forming a composite image frame (330) by stitching together the partial image frames (321, 322, 323), wherein the stitching operation (260) is applied to the first set (325) of pixels in the partial image frames (321, 322, 323) but not to the second set (326) of pixels in the partial image frames (321, 322, 323) in the ceiling operation mode. The method according to any one of claims 11 to 13.
15. The method according to any one of claims 1 to 14, wherein the surveillance cameras (110, 210) are equipped with an orientation sensor (160), and the ceiling detection procedure is performed in response to the orientation sensor detecting a downward orientation of the surveillance cameras (110, 210).
16. A processing device (150) configured to carry out the method according to any one of claims 1 to 15 for controlling an image processing stage (140) for processing image data (300) captured by a surveillance camera (110).
17. A non-temporary computer-readable storage medium comprising a portion of computer program code configured to perform the method described in any one of claims 1 to 15 when executed on a device having processing capabilities.