Camera noise reduction

By selecting reference images in wearable and gimbal cameras and performing time noise filtering for viewing angle transformation, combining motion and light sensor optimization processing flow, the problem of noise and resource consumption of wearable cameras in low-light environments is solved, achieving efficient noise reduction and resource conservation.

CN113542533BActive Publication Date: 2025-07-18AXIS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110334743.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2021-03-29
Publication Date
2025-07-18
Estimated Expiration
2041-03-29

AI Technical Summary

Technical Problem

Images captured by wearable cameras and gimbal cameras in low-light environments are prone to noise problems, and existing time noise filtering techniques can cause blurred moving objects and excessive consumption of computing resources and power.

Method used

By selecting reference images and multiple images for viewing angle transformation, performing time noise filtering, reusing viewing angle transformation to reduce the calculation amount and storage needs, combining motion sensors and light sensors to determine whether to perform filtering, and optimizing the image processing flow.

Benefits of technology

Effectively reduce image noise, reduce calculation workload and energy consumption, reduce image processing delay, improve image quality and optimize resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113542533B_ABST
    Figure CN113542533B_ABST
Patent Text Reader

Abstract

The present disclosure relates to camera noise reduction. The present disclosure relates to cameras and, in particular, to methods for reducing noise in images captured by a camera, where the same perspective transformation can be reused to perform temporal noise filtering on multiple images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to cameras, and more particularly to methods for reducing noise in images captured by a camera. Background Art

[0002] Wearable cameras are becoming increasingly common and are used in a wide variety of applications ranging from artistic or entertainment uses to those related to security and documentation. Like many other cameras, a common problem with wearable cameras is image noise problems that occur somewhere in the camera, e.g., in the optics, at the image sensor, or in the circuitry. Noise typically appears as random variations in brightness and color in the image. Noise is particularly prevalent in images captured in low-light environments, i.e., when fewer photons are recorded on the image sensor. Low-light environments typically correspond to images with a low signal-to-noise ratio (SNR). One technique for mitigating the effects of image noise is referred to as temporal noise filtering (TNF). TNF relies on averaging multiple images captured at different points in time. However, such averaging can result in blurring caused by the motion of objects that move in other static scenes, and if the entire scene is dynamic, e.g., if the camera is moving, the entire scene may ultimately be somewhat blurred to some extent. Wearable cameras worn by a wearer are always subject to some motion due to the wearer's movement, whether that movement is based on conscious movement (e.g., if the wearer is walking), or subconscious movement (e.g., through the wearer's breathing). Various techniques for adapting motion compensation based on the cause of the motion are known. However, motion compensation techniques may require excessive computational resources and are therefore also more power-consuming. For wearable cameras that rely on some form of limited electrical storage, e.g., a battery, this can be a significant problem. Mechanical motion stabilization can also be considered, but it is difficult to fully address the blurring caused by all types of motion. Such systems can make the wearable camera more complex and bulky. Therefore, improvements are needed in the technical field. Summary of the Invention

[0003] An object of the present invention is to at least mitigate some of the above problems and provide improved noise reduction for a wearable camera.

[0004] According to a first aspect of the present invention, there is provided a method for reducing noise in an image captured by a wearable camera, the method comprising:

[0005] A. providing a sequence of temporally consecutive images,

[0006] B. selecting a first reference image among the temporally consecutive images,

[0007] C. Select a first plurality of images from among the images that are temporally continuous for temporal noise filtering of a first reference image.

[0008] D. Form a first plurality of transformed images by transforming each of the first plurality of images to have the same perspective as the first reference image.

[0009] E. Perform temporal noise filtering TNF on the first reference image using the first plurality of transformed images.

[0010] F. Select a different second reference image from among the images that are temporally continuous.

[0011] G. Select a second plurality of images from among the images that are temporally continuous for temporal noise filtering of the second reference image, wherein at least one of the images in the second plurality of images is also included in the first plurality of images, and wherein the second plurality of images includes the first reference image.

[0012] H. Determine whether TNF of the second reference image should be performed, wherein when it is determined that TNF should be performed, the method further comprises:

[0013] I. Form a second plurality of transformed images by:

[0014] I1. For each image in the second plurality of images that is not included in the first plurality of images, transform the image to have the same perspective as the second reference image;

[0015] I2. For each image in the second plurality of images that is also included in the first plurality of images, transform the corresponding transformed image among the first plurality of transformed images to have the same perspective as the second reference image;

[0016] J. Perform temporal noise filtering TNF on the second reference image using the second plurality of transformed images.

[0017] The term "wearable camera" can be understood as a camera configured to be worn by a wearer during use. A camera worn on the body, a camera mounted on glasses, and a camera mounted on a helmet should be considered non-limiting examples of the term wearable camera. The wearer can be, for example, a human or an animal.

[0018] The term "temporally continuous images" can be understood as images or image frames captured continuously at different time points. In other words, the first image is captured earlier in time compared to the capture time of the second image, and the second image is captured earlier in time compared to the capture time of the third image, and so on. The images or image frames can together form a sequence of temporally continuous images or a sequence of image frames of a video stream.

[0019] The term "transform an image to have the same perspective" can refer to transforming an image or creating a projection of the image as if it were captured by a camera positioned and oriented in the same way as when another image was captured. This wording should be understood to refer to substantially the same perspective. Various ways to achieve the perspective transformation can include using homography and image projection. Readings from sensors such as, for example, an accelerometer and a gyroscope can be used to perform the perspective transformation. The perspective transformation can include calculating a homography based on corresponding candidate point pairs in two different images.

[0020] Note that steps A to J may not need to be performed in the order presented in this disclosure. Both step I1 and step I2 should be understood as part of step I.

[0021] The above method provides a way to optimize the TNF of the images captured by a wearable camera. In particular, the method can reduce the computational workload and energy usage. This is achieved by creating a processing flow in which, when forming the second plurality of transformed images, the second transformation step can benefit from being able to reuse the same perspective transformation for at least two images. This should be understood to be the case where, since at least one of the first plurality of transformed images and the first reference image are present in the second plurality of images, the same perspective transformation can be used when transforming these images to have the same perspective as the second reference image. This means that when generating the second plurality of transformed images, fewer new perspective transformations may need to be calculated / determined. Additionally, the reuse of the instructions for performing the perspective transformation can reduce the storage amount or cache storage amount required for the transformed images.

[0022] The present invention should be understood in such a way that, for example, for a video stream including a large number of temporally consecutive images, more reuse of the perspective transformation is advantageously achieved.

[0023] When the processing is expanded by successive iterations, the benefit of reduced workload will become more significant, as is the case when performing the method on a continuous image stream (i.e., a video).

[0024] According to some embodiments of the first aspect, the method may further include: after transforming one of the temporally consecutive images to have the same perspective as the first reference image or the second reference image, deleting the one of the temporally consecutive images from the memory of the wearable camera.

[0025] An advantageous effect of such an embodiment may be the effect of reducing computer memory requirements. This is possible because, according to the method, once the original perspective images have been transformed into features of another image perspective, they may be redundant, since the transformed images are actually used if TNF is to be performed on other images among temporally consecutive images (see, for example, step I2).

[0026] According to some embodiments of the first aspect, each image of the first plurality of images may be temporally before a first reference image, and each image of the second plurality of images may be temporally before a second reference image.

[0027] By such an embodiment, latency can be reduced. This is because only temporally prior image information is used for TNF of the reference image, which means a reduction in the delay between providing an image and performing TNF on that image. Reducing latency can be particularly advantageous when the method is applied to a real-time video stream.

[0028] According to some embodiments of the first aspect, the first plurality of images may include 4 to 8 images, and the second plurality of images may include 4 to 8 images.

[0029] Generally, the more images used for TNF can improve the result of the noise reduction method. More images can also enable more reuse of the perspective transformation. However, when using more images, the computational complexity of TNF generally increases. The range of images according to this embodiment represents a good consideration between the quality of TNF and the computational complexity.

[0030] According to some embodiments of the first aspect, forming the first plurality of transformed images and forming the second plurality of transformed images may include: transforming each image to have the same perspective as the associated reference image based on comparing motion data associated with each image in a sequence of temporally consecutive images.

[0031] The term "motion data" may refer to any data or information related to the physical motion of a camera relative to the scene it monitors.

[0032] Such an embodiment provides flexibility in performing the method because there can be various types of motion data and various means for determining motion data can be applied.

[0033] According to some embodiments of the first aspect, the motion data associated with each image in a sequence of temporally consecutive images is determined by at least one of a motion sensor, an accelerometer, and a gyroscope, or the motion data is determined based on image analysis of each image.

[0034] When installed on or connected to a camera, an accelerometer can provide accurate motion data related to the direction, speed, and acceleration of the camera's movement. When installed on or connected to a camera, a gyroscope can provide accurate motion data related to the orientation of the camera. Motion sensors can provide similar and / or further motion data. Image analysis can provide motion data based on the analysis of images captured by the camera (e.g., by comparing consecutively captured images). The above methods for determining motion data can be combined or performed in combination with each other.

[0035] According to some embodiments of the first aspect, step H may further include: determining a perspective difference between at least two images of the first plurality of images, wherein, when it is determined that the perspective difference is less than or equal to a predetermined perspective difference threshold, TNF of the second reference image is performed, and wherein, when it is determined that the perspective difference is greater than the predetermined perspective difference threshold, TNF of the second reference image is not performed.

[0036] If the motion is within a certain compliance range, it may be preferable to perform only TNF. TNF depends on images that are consecutive over time. Performing TNF on consecutive images that include too much motion may cause distortion of image details. In this case, the effect of not performing the second round of TNF at all may be beneficial. In addition, as long as the motion is too large, by essentially bypassing the remaining noise reduction steps, this can also provide a beneficial effect in reducing the required computational workload.

[0037] According to some embodiments of the first aspect, the perspective difference may be based on at least one of the following:

[0038] Motion data associated with each image over time, wherein the motion data is determined by a motion sensor, an accelerometer, or a gyroscope; and

[0039] Image data related to how many pixels have changed between subsequent images of the first plurality of images.

[0040] By using another sensor other than the camera (such as a motion sensor, an accelerometer, or a gyroscope), the determination can be made more reliable and is not dependent on the operation of the camera. By using the image data of the camera, the system can be made less complex and less dependent on other sensors. Determining how many pixels have changed between subsequent images can easily provide the applicability and suitability for determining whether TNF is performed.

[0041] According to some embodiments of the first aspect, the method may further include: before step A, determining capture conditions for a sequence of images that are consecutive over time, wherein, when it is determined that the capture conditions meet the requirements of a predetermined capture condition, only steps A to J are performed.

[0042] If the capture conditions are within a certain compliance range, it may be preferable to perform only TNF. It is understood that due to the expected distortion of image details, some image capture conditions are not suitable for TNF. Therefore, it may be advantageous to check whether the capture conditions are good or at least acceptable before performing each step of the method. Similar to the above embodiments, as long as the capture conditions do not meet the predetermined capture condition requirements, by substantially not performing the remaining method steps, this can also provide an advantageous effect in reducing the required computational workload.

[0043] According to some embodiments of the first aspect, the capture conditions may be determined by at least one of the following:

[0044] The level of motion determined by a motion sensor, accelerometer, gyroscope, or positioning device; and

[0045] The light level determined by a light sensor or by image analysis.

[0046] If the level of motion is too high, it is preferably not to perform method steps A to step J. This can be understood as similar to the embodiments discussed above, when the motion is too large and is expected to distort the image details, the second round of TNF is not performed. Too high a level of motion may reduce the feasibility of successfully performing a perspective transformation, i.e., resulting in a distorted transformed image. Too high a level of motion may also reduce the feasibility of successful TNF by reducing the number of common pixels between temporally consecutive images.

[0047] A positioning device (e.g., a Global Navigation Satellite System (GNSS), tracker / receiver) can be used to determine whether the level of motion is too high. Such a device can be advantageously used in combination with a camera to determine the speed of the wearer's motion. The speed can be used to determine whether the wearer is, for example, running or walking. For example, if the wearer is walking, steps A to step J can be performed, but if the wearer is running, due to the expected higher level of motion in the latter activity, steps A to step J may not be performed.

[0048] It may also preferably not perform method steps A to method step J based on the light level. When applied to low-light images, TNF may cause distortion of image details, which is why it may be good to completely avoid method steps A to method step J if the light level is not satisfactory. On the other hand, if the light level is too high, TNF may not be needed or not necessary. Therefore, TNF can be advantageously avoided.

[0049] If the light level exceeds a predetermined threshold, it can be further determined that there is no need to record the level of motion. Therefore, sensors and devices for recording motion, motion data, and / or the level of motion can be turned off to save, for example, the battery energy of a wearable camera.

[0050] The light level and the motion level can be combined (e.g., as a quality factor) or evaluated in a combined manner to determine the capture condition.

[0051] According to some embodiments of the first aspect, the capture condition is determined by the light level, which is determined by a light sensor or by image analysis, wherein the predetermined capture condition requires a light level below a predetermined level.

[0052] Under low light conditions, TNF may generally be more necessary. This may be because cameras or image sensors often exhibit a lower signal-to-noise ratio under low light conditions. Thus, an upper threshold for performing method steps A to method steps J may be preferred.

[0053] According to some embodiments of the first aspect, the method may further include: storing the first reference image on the memory of the wearable camera after performing TNF on the first reference image, and storing the second reference image on the memory of the wearable camera after performing TNF on the second reference image.

[0054] Thus, the reference image on which TNF has been performed can form a final video stream that can be stored on the memory of the wearable camera.

[0055] According to some embodiments of the first aspect, the method may further include: sending the first reference image from the wearable camera to a remote device after performing TNF on the first reference image, and sending the second reference image from the wearable camera to the remote device after performing TNF on the second reference image.

[0056] Thus, the reference image on which TNF has been performed can form a final video stream that is sent to the remote device for display or storage. Advantageously, the reference image may thus not need to be stored on the memory of the wearable camera for a long time.

[0057] According to a second aspect of the present invention, there is provided a wearable camera including an image capture unit and a computing unit, the wearable camera being configured to:

[0058] A. Capture a sequence of temporally consecutive images by the image capture unit,

[0059] B. Select a first reference image from the temporally consecutive images by the computing unit,

[0060] C. Select a first plurality of images from the temporally consecutive images by the computing unit for temporal noise filtering of the first reference image;

[0061] D. Transform each of the first plurality of images into a first plurality of transformed images having the same perspective as the first reference image by the computing unit;

[0062] E. The computing unit performs temporal noise filtering TNF on the first reference image using the first plurality of transformed images.

[0063] F. The computing unit selects a different second reference image from among the temporally consecutive images.

[0064] G. The computing unit selects a second plurality of images from among the temporally consecutive images for temporal noise filtering of the second reference image, wherein at least one of the images in the second plurality of images is also included in the first plurality of images, and wherein the second plurality of images includes the first reference image.

[0065] H. The computing unit determines whether TNF of the second reference image should be performed, and wherein when it is determined that TNF should be performed, the wearable camera is further configured to:

[0066] I. The computing unit forms the second plurality of transformed images by:

[0067] I1. For each image in the second plurality of images that is not included in the first plurality of images, transforming each image to have the same perspective as the second reference image.

[0068] I2. For each image in the second plurality of images that is also included in the first plurality of images, transforming the corresponding transformed image among the first plurality of transformed images to have the same perspective as the second reference image.

[0069] J. The computing unit performs temporal noise filtering on the second reference image using the second plurality of transformed images.

[0070] The wearable camera described in the second aspect provides similar advantages as the method described in the first aspect due to their corresponding features. The wearable camera can be considered as a device configured to implement the method of the first aspect.

[0071] According to a third aspect of the present invention, there is provided a non-transitory computer-readable storage medium having instructions stored on the storage medium that, when executed by a device having processing capabilities, implement the method of the first aspect.

[0072] The non-transitory computer-readable storage medium described in the third aspect provides advantages similar to those of the method described in the first aspect.

[0073] The further scope of application of the present invention will become apparent from the specific embodiments given below. However, it should be understood that when indicating the preferred embodiments of the present invention, only the specific embodiments and specific examples are given by way of example, since various changes and modifications within the scope of the present invention will be apparent to those skilled in the art from such specific embodiments.

[0074] It should be noted that, as used in the specification and the appended claims, the articles "a", "an", "the", and "said" are intended to mean that there is one or more of the elements, unless the text clearly indicates otherwise. Thus, for example, reference to "a unit" or "the unit" may include several devices, etc. In addition, the words "comprising", "including", "containing", and similar words do not exclude other elements or steps. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Hereinafter, the above and other aspects of the present invention will be described in more detail with reference to the drawings. These drawings should not be considered restrictive; rather, they should be considered for explanatory and understanding purposes.

[0076] As shown, the dimensions of the layers and regions may be enlarged for illustrative purposes and are thus provided to illustrate the general structure. Throughout the text, the same reference numerals denote the same elements.

[0077] Figure 1 Exemplarily shown is a method for performing noise reduction on temporally continuous images captured by a wearable camera or a pan / tilt camera, wherein a reference image to be subjected to temporal noise filtering is temporally after a plurality of images on which the noise filtering is based.

[0078] Figure 2 Exemplarily shown is a method for performing noise reduction on temporally continuous images captured by a wearable camera or a pan / tilt camera, wherein a reference image to be subjected to temporal noise filtering is not temporally after a plurality of images on which the noise filtering is based.

[0079] Figure 3 A flowchart of a method for performing noise reduction on images captured by a wearable camera or a pan / tilt camera is shown.

[0080] Figure 4 A wearable camera that optionally communicates with a remote device is schematically shown.

[0081] Figure 5 A pan / tilt camera that optionally communicates with a remote device is schematically shown. DETAILED DESCRIPTION

[0082] The present invention will now be described more fully hereinafter with reference to the drawings, in which current preferred embodiments of the invention are shown. However, the present invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0083] Figure 1 andFigure 2 Disclosed is a noise reduction method in an image captured by a wearable camera, such as a body-worn camera (BWC). Figure 3 A flowchart of the same method for noise reduction in an image captured by a wearable camera is shown. The method includes steps A to step J. As Figure 4 shown, the method can be implemented in the wearable camera 200.

[0084] Now, the method will be explained in conjunction with Figure 4 the wearable camera 200. Figure 1 , Figure 2 and Figure 3 . Reference numerals starting from Figure 1 (e.g., 113) refer to Figure 1 the images shown in Figure 2 . Reference numerals starting from Figure 4 (e.g., 202) refer to Figure 4 the features shown in

[0085] Figure 1 and Figure 2 The boxes shown in Figure 1 and Figure 2 represent images or image frames. Figure 1 and Figure 2 contain a horizontal time component related to when the original images (i.e., temporally continuous images 101, 102, 103, 104) are provided or captured. Boxes located at different horizontal positions in the drawing indicate that images are captured or provided at different time points. Time t is shown as advancing from left to right in the drawing, i.e., box 101 (one of the temporally continuous images) is provided or captured before the rest of the temporally continuous images 102, 103, 104. The temporally continuous images 101, 102, 103, 104 can be provided as raw format images from an image sensor device, i.e., there is no prior image processing of the images before being provided. Alternatively, this method can be performed before the image processing of the temporally continuous images 101, 102, 103, 104 before being provided as input to the method. Non-limiting examples of the aforementioned image processing include image data adjustment or correction, such as defective pixel removal and column fixed pattern noise filtering. In other words, the temporally continuous images 101, 102, 103, 104 can be provided as raw image data or processed image data. However, it should be noted that the method according to the present invention is performed before the video coding processing of the images (i.e., the temporally continuous images 101, 102, 103, 104 are not encoded / are not video encoded).

[0086] Figure 1 and Figure 2Further includes a vertical / column component. An image below other images in the same column indicates a selected, processed, transformed, or filtered version of the image above it in the same column. These lower images need not correspond to any chronological order of processing / transformation, etc., and can be processed / transformed, etc. at any point in time. These images should be understood to be mainly based on the temporally consecutive images 101, 102, 103, 104 above them.

[0087] In Figure 1 and Figure 2 the temporally consecutive images 101, 102, 103, 104 are provided in the top row. This should be understood to correspond to step A of the method in Figure 3 .

[0088] It should be noted that step A in the second aspect of the present invention (i.e., the aspect of providing the wearable camera 200) specifies that when the wearable camera 200 is configured to provide a sequence of temporally consecutive images 101, 102, 103, 104 in the first aspect (i.e., the aspect shown in Figure 3 ), a sequence of temporally consecutive images 101, 102, 103, 104 is captured by the image capture unit 202. These two aspects can still be considered related to each other. When the method in the first aspect does not require the camera itself to capture the temporally consecutive images 101, 102, 103, 104, the wearable camera 200 in the second aspect uses its own image capture unit 202 to capture a sequence of temporally consecutive images 101, 102, 103, 104. The method can be remotely executed on the wearable camera 200 that captures the images, so only the temporally consecutive images 101, 102, 103, 104 need to be provided.

[0089] The method can be executed by or in the wearable camera 200. Step A can be executed by the image capture unit 202 of the wearable camera 200. Steps B to J can be executed by the capture unit 204 of the wearable camera 200.

[0090] Select a first reference image 113 among the temporally consecutive images 101, 102, 103, 104. This should be understood to correspond to step B. The first reference image 113 should be understood to be an image that should be filtered using TNF. In Figure 1 image 103 is selected as the first reference image 113 from the temporally consecutive images 101, 102, 103, 104. In Figure 2 image 101 is selected as the first reference image 113 from the temporally consecutive images 101, 102, 101, 104.

[0091] To assist TNF, a first plurality of images 121, 122 among temporally consecutive images 101, 102, 103, 104 are selected for temporal noise filtering of the first reference image 113. This should be understood as corresponding to step C. In Figure 1 , the first plurality of images 121, 122 are shown as being temporally before the selected first reference image 113, corresponding to images 101 and 102 among the temporally consecutive images 101, 102, 103, 104. In Figure 2 , the first plurality of images 121, 122 are shown as being temporally after the selected first reference image 113, corresponding to images 102 and 103 among the temporally consecutive images 102, 103, 103, 104. There are other embodiments where the first plurality of images 121, 122 are selected from images before and after the first reference image 113. It may be advantageous that the first plurality of images 121, 122 can include 4 to 8 images. For the sake of helping the understanding of the general concept, Figure 1 and Figure 2 show a first plurality of images 121, 122 including two images. Those skilled in TNF have the knowledge to apply the concepts disclosed herein to any (reasonable) number of images.

[0092] After the first plurality of images 121, 122 have been selected, a first plurality of transformed images 131, 132 are formed by transforming each of the first plurality of images 121, 122 to have the same perspective as the first reference image 113. This should be understood as corresponding to step D.

[0093] The transformation step D can be performed using homography. The transformation step D can include calculating the homography based on corresponding candidate point pairs in two different images. For homography, for the purpose of generating the homography matrix, images captured at different time points can be considered as images captured from different cameras in a stereo camera arrangement.

[0094] The transformation step D can be based on comparing the motion data associated with each image in the sequence of temporally consecutive images 101, 102, 103, 104. The motion data associated with each image in the sequence of temporally consecutive images 101, 102, 103, 104 can be determined by the motion sensor 210, the accelerometer 211, and / or the gyroscope 212. The motion data can alternatively or additionally be determined based on the image analysis of the temporally consecutive images 101, 102, 103, 104.

[0095] Motion data can relate to the motion of the wearable camera 200 or the motion of the wearer of the wearable camera 200. The motion data can include the acceleration, velocity, and / or direction of the motion. The motion data can include the orientation and / or position of the wearable camera 200 or the orientation and / or position of the wearer of the wearable camera 200. The motion data can include data related to the rotational motion of the wearable camera 200. The motion data can include data related to the fall / spin, move / yaw, tilt / pitch, and / or roll of the wearable camera 200. The motion data can include data related to the translational motion of the wearable camera 200. The motion data can include data related to the track of the camera, the dolly, the pedestal / jib / cantilever, and / or the truck / crane.

[0096] In Figure 1 and Figure 2 , the box labeled PT indicates that a perspective transformation is being performed. This works by transforming the base image to have the same perspective as an image from a different vertical column indicated by the dashed arrow pointing to the PT box. The base image is shown being transformed through the PT box with a solid arrow. Sometimes, the same PT box can be used to transform several images, as is the case when the images of the second plurality 152, 153 are transformed to form the second plurality of transformed images 162, 163 (see further below). In this case, image 152 can be understood as the base image for transformed image 162. This same logic can be applied to image 153 and its corresponding transformed image 163.

[0097] After forming the first plurality of transformed images 131, 132, the first plurality of transformed images 131, 132 can be used to perform the TNF of the first reference image 113. This should be understood as corresponding to Figure 3 step E in

[0098] In Figure 1 and Figure 2 , the box labeled TNF indicates that a temporal noise filtering step is being performed. In these cases, the base image for which noise filtering is to be performed is indicated by a dashed arrow. The dashed arrow pointing to the TNF box indicates the images from different vertical columns used in the temporal noise filtering process. The TNF process can include, for example, averaging the image content in images captured at different time points. Generally, these figures exemplarily depict transforming the images to have the same perspective before performing the temporal noise filtering.

[0099] In the drawings, the reference numerals may be the same before and after performing the TNF step. This may be because the first reference image 113 and the second reference image 144 are original / unfiltered before / above the noise filtering step and are substantially the same image after / below the noise filtering step, although ideally with less noise. For subsequent steps of the method, starting from step F, the original / unfiltered or temporally noise-filtered reference images 113, 144 may be used.

[0100] After performing TNF on the first reference image 113, a different second reference image 144 may be selected among the temporally consecutive images. This should be understood as corresponding to Figure 3 step F in. The second reference image 144 is different from the first reference image 113. Similar to the first reference image 113, the second reference image 144 should be understood as an image that should be filtered using TNF.

[0101] In Figure 1 the embodiment of, the image 104 is selected as the second reference image 144 from the temporally consecutive images 101, 102, 104, 104. In Figure 2 the embodiment of, the image 102 is selected as the second reference image 144 from the temporally consecutive images 101, 102, 102, 104.

[0102] To assist TNF again, a second plurality of images 152, 153 in the temporally consecutive images 101, 102, 103, 104 are selected for the temporal noise filtering of the second reference image 144. However, this time, at least one image in the image 152 of the second plurality of images 152, 153 is also included in the first plurality of images 121, 122. The second plurality of images 152, 153 also includes the first reference image 113. This should be understood as corresponding to Figure 3 step G in. In Figure 1 the embodiment of, the second plurality of images 152, 153 are shown to be temporally before the selected second reference image 144, corresponding to the images 102 and 103 in the temporally consecutive images 102, 103, 103, 104. In Figure 2 the embodiment of, the images in the second plurality of images 152, 153 are shown to be temporally before and after the second reference image 144, corresponding to the images 101 and 103 in the temporally consecutive images 101, 103, 103, 104. There are other embodiments where the images in the second plurality of images 152, 153 are all after the second reference image 144. The second plurality of images 152, 153 may include 4 to 8 images. Figure 1 and Figure 2An embodiment is shown in which the second plurality of images 152, 153 includes two images.

[0103] As disclosed, at least one of the images 152 in the second plurality of images 152, 153 is also included in the first plurality of images 121, 122. This can be understood as referring to any corresponding images based on the same selected images in the temporally continuous images 101, 102, 103, 104. In Figure 1 this case, this should be understood as any images based on images 101 and 102. For example, one or more of the first plurality of transformed images 131, 132 can also be selected for the second plurality of images 152, 153, as Figure 1 shown, where the transformed image 132 is selected as the image 152 in the second plurality of images.

[0104] The method includes the step of determining whether the TNF of the second reference image should be performed. This should be understood as corresponding to Figure 3 step H in. Step H can be performed after the second plurality of images 152, 153 are selected.

[0105] Step H can include determining the perspective difference between at least two of the first plurality of images 121, 122. Alternatively, the perspective difference can be determined based on any of the temporally continuous images 101, 102, 103, 104. When it is determined that the perspective difference is less than or equal to a predetermined perspective difference threshold, the TNF of the second reference image 144 can be performed. When it is determined that the perspective difference is greater than the predetermined perspective difference threshold, the TNF of the second reference image 144 should not be performed. In this case, step H can be performed before steps F and G. If it is determined that the TNF of the second reference image 144 is not performed, steps F and G can be omitted from the method together.

[0106] Step H can include determining the perspective difference between at least two of the second plurality of images 121, 122. When it is determined that the perspective difference is less than or equal to a predetermined perspective difference threshold, the TNF of the second reference image 144 can be performed. When it is determined that the perspective difference is greater than the predetermined perspective difference threshold, the TNF of the second reference image 144 should not be performed.

[0107] The perspective difference can be based on motion data associated with each image in time. The motion data can be determined by the motion sensor 210, the accelerometer 211, and / or the gyroscope 212. The perspective difference can be based on image data related to how many pixels have been directly or otherwise changed between subsequent images of the first plurality of images 121, 122, between subsequent images of the second plurality of images 152, 153, or between subsequent images of the temporally continuous images 101, 102, 103, 104.

[0108] A predetermined perspective difference threshold may be related to the following distinction: whether the motion data or the image data is expected to be the result of an action such as running, walking, or breathing performed by the wearer of the camera 200. In one embodiment, TNF is performed when determining certain types of low-activity actions (such as walking and / or breathing), and TNF is not performed when determining certain types of high-activity actions (such as running).

[0109] Motion data may be determined periodically and evaluated to determine whether TNF should be performed. Motion data may be determined at a rate that matches the value of frames per second FPS, at which rate the wearable camera 200 acquires temporally consecutive images 101, 102, 103, 104. The value of FPS may preferably be in the range of 1 - 60, and more preferably in the range of 20 - 40.

[0110] When it is determined during step H that TNF for the second reference image 144 should not be performed, the method may terminate after step H. This would mean that steps I through J are not performed. In this case, the method may restart at step A.

[0111] When it is determined during step H that TNF for the second reference image 144 should be performed, the method may proceed as shown in steps I through J. Figure 3 as shown in steps I through J.

[0112] The method proceeds by forming a second plurality of transformed images 162, 163. This should be understood as corresponding to step I. Step I includes two sub-steps, depending on how the images selected for the second plurality of images 152, 153 were previously used in the method.

[0113] Each image 153 of the second plurality of images 152, 153 that is not included in the first plurality of images 121, 122 is transformed to have the same perspective as the second reference image 144, thereby forming a transformed image 163. This should be understood as corresponding to sub-step I1.

[0114] For each image 152 of the second plurality of images 152, 152 that is also included in the first plurality of images 121, 122, the corresponding transformed image 132 (transformed in step D of Figure 3 is transformed to have the same perspective as the second reference image 144, thereby forming a transformed image 162. This should be understood as corresponding to sub-step I2. The transformed image 132 should be understood to temporally correspond to image 122 and image 152.

[0115] In Figure 1In this case, this correspondence relates to how they are all derived from image 102 of temporally consecutive images 101, 102, 103, 104. In Figure 2 In this case, images 122, 132, and 152 are all derived from image 103 of temporally consecutive images 101, 102, 103, 104. The formation details of the second plurality of transformed images 162, 163 can be similar to the formation details of the first plurality of transformed images 131, 132 (step D) discussed above, for example, regarding the use of homography calculation and motion data.

[0116] According to this method, when forming the second plurality of transformed images 162, 163, the same perspective transformation can be advantageously used for at least two of the second plurality of images 152, 153.

[0117] To summarize Figure 1 and Figure 2 the differences between, these embodiments have slightly different timings or orders of the first plurality of images and the second plurality of images 121, 122, 152, 153 and their corresponding reference images 113, 144 in Figure 1 and Figure 2 .

[0118] Figure 1 shows each of the first plurality of images 121, 122 that are temporally before the first reference image 113, and each of the second plurality of images 152, 153 that are temporally before the second reference image 144.

[0119] Figure 2 indicates an alternative embodiment, where the first reference image 113 is temporally before the first plurality of images 121, 122. Although not shown in Figure 2 the second reference image 144 can similarly be temporally before the second plurality of images 152, 153.

[0120] Figure 2 indicates that the second reference image 144 can be temporally sandwiched between the images of the second plurality of image frames 152, 153. Similarly, the second reference image 113 can be temporally sandwiched between the images of the first plurality of image frames 121, 122.

[0121] After forming the second plurality of transformed images 162, 163, TNF is performed on the second reference image 144 using the second plurality of transformed images 162, 163. The details related to performing TNF on the second reference image 144 can be similar to the details discussed above regarding performing TNF on the first reference image 113 (step E).

[0122] This method can be used to form a video stream with temporal noise filtering. In this case, the reference images 113, 144 will form the video stream. Each reference image 113, 144 can be associated with different time points in the video stream.

[0123] As Figure 4 shown, this method can be implemented in the wearable camera 200. However, as described above, this method can also be executed in a device external to the wearable camera 200.

[0124] This method can include: before step A, determining the capture conditions for a sequence of temporally consecutive images 101, 102, 103, 104. In this case, steps A to J are only executed when it is determined that the capture conditions meet the predetermined capture condition requirements. Therefore, it can be considered that method steps A to method step J are delayed until the requirements are met.

[0125] The capture conditions can be determined by the motion level. The motion level can be determined by the motion sensors 210, accelerometers 211, gyroscopes 212, and / or positioning devices 213. The capture conditions can be determined by the light level. The light level is determined by the light sensor 220 or by image analysis.

[0126] The predetermined capture condition requirements can be a requirement for a light level below a predetermined level. Such a predetermined level can be in the range of 50 - 200 lux. The predetermined level can be more preferably in the range of 75 - 125 lux, for example 100 lux. Higher light level values can be associated with less noise, thus reducing the need for TNF for these higher light level values.

[0127] The predetermined capture condition requirements can be a requirement for a light level above a predetermined level. The predetermined capture condition requirements can further include a minimum acceptable light level and a maximum acceptable light level. The predetermined capture condition requirements can include an intermediate light level exclusion range, where light levels outside this range are acceptable.

[0128] This method can further include: after performing TNF on the first reference image 113, storing the first reference image 113 on the memory 206 of the wearable camera 200. This method can include: after performing TNF on the second reference image 144, storing the second reference image 144 on the memory 206 of the wearable camera 200.

[0129] This method can further include: after transforming one of the temporally consecutive images 101, 102, 103, 104 to have the same perspective as the first reference image 113 or the second reference image 144, deleting that one of the temporally consecutive images 101, 102, 103, 104 from the memory 206 of the wearable camera 200.

[0130] The method may further include: after performing TNF on the first reference image 113, sending the first reference image 113 from the wearable camera 200 to the remote device 230. The method may include: after performing TNF on the second reference image 144, sending the second reference image 144 from the wearable camera 200 to the remote device 230.

[0131] The method may be implemented by a computer, a decoder, or another device having processing capabilities. A non-transitory computer-readable storage medium may be provided with instructions stored thereon that, when executed by a device having processing capabilities, implement the method.

[0132] Figure 4 A wearable camera 200 including an image capture unit 202 and a computing unit 204 is shown. The wearable camera 200 may be configured to perform the above-described methods and method steps.

[0133] The image capture unit 202 may be understood as any device capable of capturing images. The image capture unit 202 may include a charge-coupled device (CCD), an image sensor, or a complementary metal-oxide-semiconductor (CMOS)-based active pixel image sensor.

[0134] The computing unit 204 may include any device capable of performing the processing and calculations according to the method. The computing unit itself may include several sub-units for performing different actions or steps of the method.

[0135] The wearable camera 200 may be worn by the wearer of the wearable camera. The wearer of the wearable camera 200 may be a human. The wearer of the wearable camera 200 may be a law enforcement professional. Further examples of the wearer of the wearable camera 200 include security personnel and individuals performing activities in a hazardous environment (e.g., a road construction site). The wearer of the wearable camera 200 may be a professional or amateur photographer / camera operator recording for aesthetic, documentary, sports, or entertainment purposes. For some uses (e.g., used by law enforcement professionals), a longer battery life and better detail capture quality of the camera may be more desirable. For other uses (e.g., entertainment / aesthetic purposes), better color capture and cognitive viewing convenience may be more desirable.

[0136] Alternatively, the wearer can be an animal, such as for example a dog, a cat or a horse. The wearer can be a service animal, such as for example a law enforcement animal. Law enforcement animals can include for example police dogs trained to detect illegal substances, or police horses deployed for crowd control tasks. The wearer can be a hunting dog. The wearer can be a wild animal provided with the wearable camera 200 for surveillance or scientific purposes. The wearer can be a pet. The wearer can be provided with the wearable camera 200 to prevent the wearer from running away, getting lost or being injured.

[0137] The wearable camera 200 can be mounted on the wearer's belt or harness. Alternatively, the wearable camera 200 can be fixedly mounted on a piece of clothing, or on a protective device such as a helmet or a vest.

[0138] As Figure 4 shown, the wearable camera 200 can include a motion sensor 210 configured to determine any type of motion of the wearable camera 200. As Figure 4 shown, the wearable camera 200 can include an accelerometer 211 configured to determine the acceleration, speed and / or direction of motion of the wearable camera 200. As Figure 4 shown, the wearable camera 200 can include a gyroscope 212 configured to determine the orientation of the wearable camera 200. As Figure 4 shown, the wearable camera 200 can include a positioning device 213 configured to determine the position, speed and / or direction of motion of the wearable camera 200. The positioning device 213 can include a GNSS sensor or receiver. The positioning device 213 can include an inertial navigation system. The wearable camera 200 can include a compass configured to determine the orientation of the camera. As Figure 4 shown, the wearable camera 200 can include a light sensor 220 (photo-detector) configured to determine the light conditions or light level of the wearable camera 200. As Figure 4 shown, the wearable camera 200 can be configured to communicate with a remote device 230. The communication can be wireless or wired.

[0139] The present disclosure further relates to a gimbal camera and particularly to a method for reducing noise in images captured by a gimbal camera.

[0140] Similar problems as those disclosed for wearable cameras can occur with gimbal cameras.

[0141] The aim of the present disclosure is to at least mitigate some of the above problems and provide improved noise reduction for gimbal cameras.

[0142] According to a fourth aspect of the present disclosure, a method for reducing noise in an image captured by a pan-tilt camera is provided, the method comprising:

[0143] A. Providing a sequence of temporally consecutive images,

[0144] B. Selecting a first reference image among the temporally consecutive images,

[0145] C. Selecting a first plurality of images among the temporally consecutive images for temporal noise filtering of the first reference image,

[0146] D. Forming a first plurality of transformed images by transforming each of the first plurality of images to have the same perspective as the first reference image,

[0147] E. Performing temporal noise filtering TNF on the first reference image using the first plurality of transformed images,

[0148] F. Selecting a different second reference image among the temporally consecutive images,

[0149] G. Selecting a second plurality of images among the temporally consecutive images for temporal noise filtering of the second reference image, wherein at least one image among the images in the second plurality of images is further included in the first plurality of images, and wherein the second plurality of images includes the first reference image,

[0150] H. Determining whether TNF of the second reference image should be performed, wherein when it is determined that TNF should be performed, the method further comprises:

[0151] I. Forming a second plurality of transformed images by:

[0152] I1. For each image in the second plurality of images that is not included in the first plurality of images, transforming the image to have the same perspective as the second reference image;

[0153] I2. For each image in the second plurality of images that is further included in the first plurality of images, transforming the corresponding transformed image in the first plurality of transformed images to have the same perspective as the second reference image;

[0154] J. Performing temporal noise filtering TNF on the second reference image using the second plurality of transformed images.

[0155] The term "pan-tilt camera" can be understood as a camera that is configured to be fixedly mounted and can pan and tilt during use, so as to capture different fields of view. A fixedly mounted pan-tilt camera for surveillance or monitoring (which can be manually or automatically controlled to pan and / or tilt to capture different fields of view) should be considered a non-limiting example of the term pan-tilt camera. The pan-tilt camera can also include a zoom function and can then be referred to as a pan-tilt-zoom camera or a PTZ camera.

[0156] The term "temporally continuous images" can be understood as images or image frames captured continuously at different time points. In other words, compared with the capture time of the second image, the first image is captured earlier in time, and compared with the capture time of the third image, the second image is captured earlier in time in turn, and so on. The images or image frames can together form a sequence of temporally continuous images or a sequence of image frames of a video stream.

[0157] The term "transform an image to have the same perspective" can refer to transforming an image or creating a projection of the image as if it were captured by a camera positioned and oriented in the same way as when capturing another image. This wording should be understood to refer to substantially the same perspective. Various ways of achieving the perspective transformation can include using homography and image projection. Readings from sensors such as, for example, accelerometers and gyroscopes can be used to perform the perspective transformation. The perspective transformation can include calculating a homography based on corresponding candidate point pairs in two different images.

[0158] Note that steps A to J may not need to be performed in the order presented in this disclosure. Both step I1 and step I2 should be understood as part of step I.

[0159] The above method provides a way to optimize the TNF of images captured by a pan-tilt camera. In particular, the method can reduce the computational workload and energy usage. This is achieved by creating a processing flow in which, when forming a second plurality of transformed images, the second transformation step can benefit from being able to reuse the same perspective transformation for at least two images. This should be understood as the case where, since at least one of the first plurality of transformed images and the first reference image are present in the second plurality of images, the same perspective transformation can be used when transforming these images to have the same perspective as the second reference image. This means that when generating the second plurality of transformed images, fewer new perspective transformations may need to be calculated / determined. In addition, the reuse of instructions for performing the perspective transformation can reduce the storage amount or cache storage amount required for the transformed images.

[0160] This disclosure should be understood in such a way that, for example, for a video stream including a large number of temporally continuous images, it is advantageously possible to achieve more reuse of perspective transformations.

[0161] When expanding the processing through successive iterations, the benefit of workload reduction will become more important, as is the case when performing the method on a continuous image stream (i.e., video).

[0162] According to some embodiments of the fourth aspect, the method may further include: after transforming one of the temporally consecutive images into an image having the same perspective as the first reference image or the second reference image, deleting the one image from the memory of the pan-tilt camera among the temporally consecutive images.

[0163] An advantageous effect of such an embodiment may be the effect of reducing the computer memory requirements. This is possible because according to the method, once the original perspective images have been transformed into features of another image perspective, they may be redundant, since the transformed images are actually used if TNF needs to be performed on other images among the temporally consecutive images (see, for example, step I2).

[0164] According to some embodiments of the fourth aspect, each image in the first plurality of images may be temporally before the first reference image, and each image in the second plurality of images may be temporally before the second reference image.

[0165] By such an embodiment, the time delay can be reduced. This is because only the temporally prior image information is used for the TNF of the reference image, which means that the delay between providing the image and performing TNF on the image is reduced. When the method is applied to a real-time video stream, reducing the time delay can be particularly advantageous.

[0166] According to some embodiments of the fourth aspect, the first plurality of images may include 4 to 8 images, and the second plurality of images may include 4 to 8 images.

[0167] Generally, the more images used for TNF can improve the effect of the noise reduction method. More images can also enable more reuse of the perspective transformation. However, when using more images, the computational complexity of TNF usually increases. The range of images according to this embodiment represents a good consideration between the quality of TNF and the computational complexity.

[0168] According to some embodiments of the fourth aspect, forming the first plurality of transformed images and forming the second plurality of transformed images may include: based on comparing the motion data associated with each image in the sequence of temporally consecutive images, transforming each image into an image having the same perspective as the associated reference image.

[0169] The term "motion data" may refer to any data or information related to the physical motion of the camera relative to the scene it monitors.

[0170] Such an embodiment provides flexibility for implementing the method because there can be various types of motion data and various apparatuses for determining motion data can be applied.

[0171] According to some embodiments of the fourth aspect, the motion data associated with each image in a sequence of temporally consecutive images can be pan / tilt data, including values of translation and tilt of a pan / tilt camera associated with each image in the sequence of temporally consecutive images. The values of translation and tilt associated with an image can be absolute values (i.e., values with respect to a coordinate system fixed to a reference pan / tilt camera) or relative values (i.e., relative values with respect to the translation and tilt of a pan / tilt camera associated with a different image).

[0172] According to some embodiments of the fourth aspect, step H can further include: determining a viewing angle difference between at least two images among a first plurality of images, wherein when it is determined that the viewing angle difference is less than or equal to a predetermined viewing angle difference threshold, the TNF of the second reference image is performed, and wherein when it is determined that the viewing angle difference is greater than the predetermined viewing angle difference threshold, the TNF of the second reference image 144 is not performed.

[0173] If the motion is within a certain compliance range, it may be preferable to perform only the TNF. The TNF depends on temporally averaged consecutive images. Performing the TNF on consecutive images including excessive motion may cause distortion of image details. In this case, the effect of not performing the second round of TNF at all may be beneficial. In addition, as long as the motion is too large, this can also provide a beneficial effect in reducing the required computational workload by substantially bypassing the remaining noise reduction steps.

[0174] According to some embodiments of the fourth aspect, the viewing angle difference is based on pan / tilt data associated with each image in time, the pan / tilt data including values of translation and tilt of a pan / tilt camera associated with each image in the sequence of temporally consecutive images.

[0175] The values of translation and tilt of the pan / tilt camera associated with an image can be with respect to a coordinate system fixed to a reference pan / tilt camera, or can be with respect to the translation and tilt of a pan / tilt camera associated with a different image.

[0176] By using pan / tilt data including values of translation and tilt of a pan / tilt camera, the viewing angle difference can be determined very precisely.

[0177] According to some embodiments of the fourth aspect, the method can further include: before step A, determining capture conditions for a sequence of temporally consecutive images, wherein when it is determined that the capture conditions meet predetermined capture condition requirements, only steps A to J are performed.

[0178] If the capture conditions are within a certain compliance range, it may be preferable to perform only TNF. It is understood that due to the expected image detail distortion, some image capture conditions are not suitable for TNF. Therefore, it may be advantageous to check whether the capture conditions are good or at least acceptable before performing each step of the method. Similar to the above embodiments, as long as the capture conditions do not meet the predetermined capture condition requirements, by substantially not performing the remaining method steps, this can also provide an advantageous effect in reducing the required computational workload.

[0179] According to some embodiments of the fourth aspect, the capture conditions are determined by at least one of the following:

[0180] The motion level determined from the gimbal data including values of the pan and tilt of the gimbal camera associated with different time points; and

[0181] The light level determined by a light sensor or by image analysis.

[0182] If the motion level is too high, it is preferably not to perform method steps A to step J. This can be understood as similar to the embodiments discussed above. When the motion is too large and is expected to distort the image details, the second round of TNF is not performed. Too high a motion level may reduce the feasibility of successfully performing a perspective transformation, that is, resulting in a distorted transformed image. Too high a motion level may also reduce the feasibility of successful TNF by reducing the number of common pixels between temporally consecutive images.

[0183] It may also preferably not perform method steps A to method steps J based on the light level. When applied to low-light images, TNF may cause image detail distortion, which is why it may be good to completely avoid method steps A to method steps J if the light level is not satisfactory. On the other hand, if the light level is too high, TNF may not be needed or not necessary. Therefore, TNF can be advantageously avoided.

[0184] If the light level exceeds a predetermined threshold, it can be further confirmed that there is no need to record the motion level. Therefore, the sensors and devices for recording motion, motion data, and / or motion level can be turned off to reduce energy consumption.

[0185] The light level and the motion level can be combined (e.g., as a quality factor) or evaluated in a combined manner to determine the capture conditions.

[0186] According to some embodiments of the fourth aspect, the capture conditions are determined by the light level, which is determined by a light sensor or by image analysis, wherein the predetermined capture condition requirement is a light level below a predetermined level.

[0187] Under low light conditions, TNF may generally be more necessary. This may be because cameras or image sensors often exhibit a lower signal-to-noise ratio under low light conditions. Thus, an upper threshold for performing method steps A to method steps J may be preferred.

[0188] According to some embodiments of the fourth aspect, the method may further include: after performing TNF on the first reference image, storing the first reference image in the memory of the pan-tilt camera, and after performing TNF on the second reference image, storing the second reference image in the memory of the pan-tilt camera.

[0189] Thus, the reference images on which TNF has been performed can form a final video stream that can be stored in the memory of the pan-tilt camera.

[0190] According to some embodiments of the fourth aspect, the method may further include: after performing TNF on the first reference image, sending the first reference image from the pan-tilt camera to a remote device, and after performing TNF on the second reference image, sending the second reference image from the pan-tilt camera to the remote device.

[0191] Thus, the reference images on which TNF has been performed can form a final video stream that is sent to the remote device for display or storage. Advantageously, the reference images may thus not need to be stored in the memory of the pan-tilt camera for a long time.

[0192] According to a fifth aspect of the present disclosure, there is provided a zoom pan-tilt camera including an image capture unit and a computing unit, the pan-tilt camera being configured to:

[0193] a. capture a sequence of temporally consecutive images by the image capture unit,

[0194] b. select a first reference image from the temporally consecutive images by the computing unit,

[0195] c. select a first plurality of images from the temporally consecutive images by the computing unit for temporal noise filtering of the first reference image;

[0196] d. form a first plurality of transformed images by the computing unit by transforming each of the first plurality of images to have the same perspective as the first reference image;

[0197] e. perform temporal noise filtering TNF on the first reference image by the computing unit using the first plurality of transformed images,

[0198] f. select a different second reference image from the temporally consecutive images by the computing unit,

[0199] g. The computing unit selects a second plurality of images from temporally consecutive images for temporal noise filtering of the second reference image, wherein at least one of the images in the second plurality of images is also included in the first plurality of images, and wherein the second plurality of images includes the first reference image.

[0200] h. The computing unit determines whether TNF of the second reference image should be performed. When it is determined that TNF should be performed, the pan-tilt camera is further configured for:

[0201] i. The computing unit forms a second plurality of transformed images by:

[0202] i1. For each image in the second plurality of images that is not included in the first plurality of images, transforming each image to have the same viewing angle as the second reference image;

[0203] i2. For each image in the second plurality of images that is also included in the first plurality of images, transforming the corresponding transformed image among the first plurality of transformed images to have the same viewing angle as the second reference image;

[0204] j. The computing unit performs temporal noise filtering on the second reference image using the second plurality of transformed images.

[0205] The pan-tilt camera described in the fifth aspect and the method described in the fourth aspect provide similar advantages due to their corresponding features. The pan-tilt camera can be considered as a device configured to implement the method of the fourth aspect.

[0206] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having instructions stored thereon, which when executed by a device having processing capabilities, implement the method of the fourth aspect.

[0207] The non-transitory computer-readable storage medium described in the sixth aspect provides advantages similar to those of the method described in the fourth aspect.

[0208] Now, the fourth, fifth, and sixth aspects of the present disclosure will be described more fully hereinafter with reference to Figure 1 、 Figure 2 、 Figure 3 and Figure 5 wherein the presently preferred embodiments are related to the fourth, fifth, and sixth aspects of the present disclosure. However, these aspects can be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided to be thorough and complete and to fully convey the scope of these aspects to those skilled in the art.

[0209] Figure 1 and Figure 2A method for reducing noise in an image captured by a pan-tilt camera is shown. Figure 3 A flowchart of the same method for reducing noise in an image captured by a pan-tilt camera is shown. The method includes steps A to step J. The method can be implemented in a pan-tilt camera 300 as Figure 5 shown.

[0210] Now, the Figure 5 pan-tilt camera 300 will be used to explain Figure 1 , Figure 2 and Figure 3 of the method. Reference numerals (such as 113) starting from Figure 1 refer to the images shown in Figure 1 and Figure 2 . Reference numerals (such as 302) starting from Figure 5 refer to the features shown in Figure 5 .

[0211] Figure 1 and Figure 2 The boxes shown in represent an image or an image frame. Figure 1 and Figure 2 contain a horizontal time component related to when the original image (i.e., temporally consecutive images 101, 102, 103, 104) is provided or captured. Boxes located at different horizontal positions in the drawing indicate that images are captured or provided at different time points. Time t is shown as advancing from left to right in the drawing, i.e., box 101 (one of the temporally consecutive images) is provided or captured before the remainder of the temporally consecutive images 102, 103, 104. The temporally consecutive images 101, 102, 103, 104 can be provided as raw format images from an image sensor device, i.e., there is no prior image processing of the image before it is provided. Alternatively, the method can be performed before the image processing of the temporally consecutive images 101, 102, 103, 104 before they are provided as input to the method. Non-limiting examples of the aforementioned image processing include image data adjustment or correction, such as defective pixel removal and column fixed pattern noise filtering. In other words, the temporally consecutive images 101, 102, 103, 104 can be provided as raw image data or processed image data. However, it should be noted that the method according to the present invention is performed before the video coding processing of the image (i.e., the temporally consecutive images 101, 102, 103, 104 are not encoded / are not video encoded).

[0212] Figure 1 and Figure 2Further includes a vertical / column component. An image below other images in the same column indicates a selected, processed, transformed, or filtered version of the image above in the same column. These lower images do not have to correspond to any chronological order of processing / transformation, etc., and can be processed / transformed, etc. at any point in time. These images should be understood to be mainly based on the temporally consecutive images 101, 102, 103, 104 above them.

[0213] In Figure 1 and Figure 2 the temporally consecutive images 101, 102, 103, 104 are provided in the top row. This should be understood to correspond to step A of the method in Figure 3 .

[0214] It should be noted that with respect to step A of the fifth aspect of the present disclosure (i.e., the aspect of providing the gimbal camera 300), it is specified that when the gimbal camera 300 is configured to provide a sequence of temporally consecutive images 101, 102, 103, 104 in step A of the fourth aspect (i.e., the aspect shown in Figure 3 ), a sequence of temporally consecutive images 101, 102, 103, 104 is captured by the image capture unit 302. These two aspects can still be considered related to each other. When the method in the fourth aspect does not require the camera itself to capture the temporally consecutive images 101, 102, 103, 104, the gimbal camera 300 of the fifth aspect uses its own image capture unit 302 to capture a sequence of temporally consecutive images 101, 102, 103, 104. The method can be remotely executed by the gimbal camera 300 that captures images, so only a sequence of temporally consecutive images 101, 102, 103, 104 needs to be provided.

[0215] The method can be executed by the gimbal camera 300 or executed in the gimbal camera 300. Step A can be executed by the image capture unit 302 of the gimbal camera 300. Steps B to J can be executed by the capture unit 304 of the gimbal camera 300.

[0216] Select a first reference image 113 among the temporally consecutive images 101, 102, 103, 104. This should be understood to correspond to step B. The first reference image 113 should be understood as an image that should be filtered using TNF. In Figure 1 image 103 is selected as the first reference image 113 from the temporally consecutive images 101, 102, 103, 104. In Figure 2 image 101 is selected as the first reference image 113 from the temporally consecutive images 101, 102, 101, 104.

[0217] To assist TNF, a first plurality of images 121, 122 among temporally consecutive images 101, 102, 103, 104 are selected for temporal noise filtering of the first reference image 113. This should be understood as corresponding to step C. In Figure 1 it, the first plurality of images 121, 122 are shown as being temporally before the selected first reference image 113, corresponding to images 101 and 102 among the temporally consecutive images 101, 102, 103, 104. In Figure 2 it, the first plurality of images 121, 122 are shown as being temporally after the selected first reference image 113, corresponding to images 102 and 103 among the temporally consecutive images 102, 103, 103, 104. There are other embodiments where the first plurality of images 121, 122 are selected from images before and after the first reference image 113. It may be advantageous that the first plurality of images 121, 122 can include 4 to 8 images. For the sake of helping the understanding of the general concept, Figure 1 and Figure 2 show a first plurality of images 121, 122 including two images. Those skilled in TNF have the knowledge to apply the concepts disclosed herein to any (reasonable) number of images.

[0218] After the first plurality of images 121, 122 have been selected, a first plurality of transformed images 131, 132 are formed by transforming each of the first plurality of images 121, 122 to have the same perspective as the first reference image 113. This should be understood as corresponding to step D.

[0219] The transformation step D can be performed using homography. The transformation step D can include calculating the homography based on corresponding candidate point pairs in two different images. For homography, for the purpose of generating the homography matrix, images captured at different time points can be considered as images captured from different cameras in a stereo camera arrangement.

[0220] The transformation step D can be based on comparing the motion data associated with each image in the sequence of temporally consecutive images 101, 102, 103, 104. The motion data associated with each image in the sequence of temporally consecutive images 101, 102, 103, 104 can be pan-tilt data including values of the pan and tilt of the pan-tilt camera 300 associated with each image in the sequence of temporally consecutive images.

[0221] By comparing the values of the pan and tilt of the pan-tilt camera 300 when capturing the first reference image 113 with the values of the pan and tilt of the pan-tilt camera 300 when capturing each of the first plurality of images 121, 122, a transformation can be determined to transform each of the first plurality of images 121, 122 into having the same viewing angle as the first reference image 113, thereby forming the first plurality of transformed images 131, 132.

[0222] Pan-tilt data can be obtained from a pan-tilt controller (such as the pan-tilt control function in the computing unit 304) to control the pan-tilt camera for the desired pan and tilt. The pan-tilt data including the values of the pan and tilt of the pan-tilt camera associated with the image can be related to the values of the pan and tilt of the pan-tilt camera 300 according to the expected pan and tilt of the pan-tilt camera 300 when capturing the image. The pan-tilt data can be further corrected by the deviation from the desired pan and tilt recognized by the pan-tilt sensor when capturing the image, as indicated in the feedback from the pan-tilt sensor.

[0223] The values of the pan and tilt associated with the image can be absolute values (i.e., values relative to a coordinate system fixed to the reference pan-tilt camera) or relative values (i.e., relative values of the pan and tilt of the pan-tilt camera associated with different images).

[0224] In Figure 1 and Figure 2 the boxes marked PT indicate that a viewing angle transformation is being performed. This works by transforming the base image into having the same viewing angle as the image from a different vertical column indicated by the dashed arrow pointing to the PT box. The base image is shown being transformed through the PT box by the solid arrow. Sometimes, the same PT box can be used to transform several images, as is the case when the images of the second plurality of images 152, 153 are transformed to form the second plurality of transformed images 162, 163 (see further below). In this case, the image 152 can be understood as the base image of the transformed image 162. This same logic can be applied to the image 153 and its corresponding transformed image 163.

[0225] After forming the first plurality of transformed images 131, 132, the first plurality of transformed images 131, 132 can be used to perform the TNF of the first reference image 113. This should be understood as corresponding to Figure 3 step E in

[0226] In Figure 1 and Figure 2In it, the box labeled TNF indicates that the time noise filtering step is being performed. In these cases, the base image for which the noise filtering is to be performed is indicated by the dashed arrow. The dashed arrow pointing to the TNF box indicates the images from different vertical columns used in the time noise filtering process. The TNF process may include, for example, averaging the image content in the images captured at different time points. Generally, these figures exemplarily depict transforming the images to have the same perspective before performing the time noise filtering.

[0227] In the figures, the reference numerals may be the same before and after performing the TNF step. This may be because the first reference image 113 and the second reference image 144 are original / unfiltered before / above the noise filtering step and are substantially the same images after / below the noise filtering step, although ideally with less noise. For the subsequent steps of the method, starting from step F, the original / unfiltered or time noise filtered reference images 113, 144 may be used.

[0228] After performing TNF on the first reference image 113, a different second reference image 144 may be selected among the temporally consecutive images. This should be understood as corresponding to Figure 3 step F in. The second reference image 144 is different from the first reference image 113. Similar to the first reference image 113, the second reference image 144 should be understood as an image that should be filtered using TNF.

[0229] In Figure 1 's embodiment, the image 104 is selected as the second reference image 144 from the temporally consecutive images 101, 102, 104, 104. In Figure 2 's embodiment, the image 102 is selected as the second reference image 144 from the temporally consecutive images 101, 102, 102, 104.

[0230] To assist TNF again, a second plurality of images 152, 153 in the temporally consecutive images 101, 102, 103, 104 are selected for the time noise filtering of the second reference image 144. However, this time, at least one image in the image 152 of the second plurality of images 152, 153 is also included in the first plurality of images 121, 122. The second plurality of images 152, 153 also includes the first reference image 113. This should be understood as corresponding to Figure 3 step G in. In Figure 1 's embodiment, the second plurality of images 152, 153 are shown to be temporally before the selected second reference image 144, corresponding to the images 102 and 103 in the temporally consecutive images 102, 103, 103, 104. In Figure 2In the embodiments, the images in the second plurality of images 152, 153 are shown as being before and after the second reference image 144 in time, corresponding to images 101 and 103 in the temporally consecutive images 101, 103, 103, 104. There are other embodiments in which the images in the second plurality of images 152, 153 are all after the second reference image 144. The second plurality of images 152, 153 may include 4 to 8 images. Figure 1 and Figure 2 An embodiment is shown in which the second plurality of images 152, 153 includes 2 images.

[0231] As disclosed, at least one of the images 152 in the second plurality of images 152, 153 is also included in the first plurality of images 121, 122. This can be understood as referring to any corresponding images based on the same selected images in the temporally consecutive images 101, 102, 103, 104. In Figure 1 this case, this should be understood as any images based on images 101 and 102. For example, one or more of the first plurality of transformed images 131, 132 can also be selected for the second plurality of images 152, 153, as Figure 1 shown, where the transformed image 132 is selected as the image 152 in the second plurality of images.

[0232] The method includes the step of determining whether the TNF of the second reference image should be performed. This should be understood as corresponding to Figure 3 step H in. Step H can be performed after the second plurality of images 152, 153 are selected.

[0233] Step H may include determining the perspective difference between at least two of the first plurality of images 121, 122. Alternatively, the perspective difference can be determined based on any of the temporally consecutive images 101, 102, 103, 104. When it is determined that the perspective difference is less than or equal to a predetermined perspective difference threshold, the TNF of the second reference image 144 can be performed. When it is determined that the perspective difference is greater than the predetermined perspective difference threshold, the TNF of the second reference image 144 should not be performed. In this case, step H can be performed before steps F and G. If it is determined that the TNF of the second reference image 144 is not performed, steps F and G can be omitted from the method together.

[0234] Step H may include determining the perspective difference between at least two of the second plurality of images 121, 122. When it is determined that the perspective difference is less than or equal to a predetermined perspective difference threshold, the TNF of the second reference image 144 can be performed. When it is determined that the perspective difference is greater than the predetermined perspective difference threshold, the TNF of the second reference image 144 should not be performed.

[0235] The perspective difference can be based on motion data associated with each image in time. The motion data can be pan-tilt data including values of the pan and tilt of the pan-tilt camera 300 associated with each of at least two of the second plurality of images 121, 122.

[0236] The pan-tilt data can be obtained from a pan-tilt controller (such as the pan-tilt control function in the computing unit 304) to control the pan-tilt camera to perform the expected pan and tilt. The pan-tilt data including the values of the pan and tilt of the pan-tilt camera associated with the image can be related to the values of the pan and tilt of the pan-tilt camera 300 according to the expected pan and tilt of the pan-tilt camera 300 when capturing the image. The pan-tilt data can be further corrected by the deviation from the expected pan and tilt recognized by the pan-tilt sensor when capturing the image as indicated in the feedback from the pan-tilt sensor.

[0237] The values of the pan and tilt associated with the image can be absolute values (i.e., values relative to a coordinate system fixed to the reference pan-tilt camera) or relative values (i.e., relative values of the pan and tilt relative to the pan and tilt of the pan-tilt camera associated with different images).

[0238] The motion data can be determined regularly and evaluated to determine whether TNF should be executed. The motion data can be determined at a rate matching the value of the frames per second FPS, at which rate the pan-tilt camera 300 acquires temporally consecutive images 101, 102, 103, 104. The value of FPS can preferably be in the range of 1 - 60, and more preferably in the range of 20 - 40.

[0239] When it is determined during step H that TNF for the second reference image 144 should not be executed, the method can terminate after step H. This would mean that steps I to J are not executed. In this case, the method can restart at step A.

[0240] When it is determined during step H that TNF for the second reference image 144 should be executed, the method can proceed as shown in steps I to J. Figure 3 shown in steps I to J.

[0241] The method is performed by forming a second plurality of transformed images 162, 163. This should be understood as corresponding to step I. Step I includes two sub-steps, depending on how the images selected for the second plurality of images 152, 153 were used previously in the method.

[0242] Each image 153 in the second plurality of images 152, 153 that is not included in the first plurality of images 121, 122 is transformed to have the same perspective as the second reference image 144, thereby forming the transformed image 163. This should be understood as corresponding to sub-step I1.

[0243] For each of the second plurality of images 152, 152 that are also included in the first plurality of images 121, 122, the corresponding transformed image 132 (transformed in step D of Figure 3 ) is transformed to have the same perspective as the second reference image 144, thereby forming a transformed image 162. This should be understood as corresponding to the local step I2. The transformed image 132 should be understood as corresponding in time to the image 122 and the image 152.

[0244] In Figure 1 , this correspondence is related to how they are all derived from the image 102 of the temporally continuous images 101, 102, 103, 104. In Figure 2 , the images 122, 132, and 152 are all derived from the image 103 of the temporally continuous images 101, 102, 103, 104. The formation details of the second plurality of transformed images 162, 163 can be similar to the formation details of the first plurality of transformed images 131, 132 (step D) discussed above, for example, regarding the use of homography calculation and motion data.

[0245] According to this method, when forming the second plurality of transformed images 162, 163, the same perspective transformation can be advantageously used for at least two of the second plurality of images 152, 153.

[0246] To summarize the differences between Figure 1 and Figure 2 , these embodiments have slightly different timings or orders of the first plurality of images and the second plurality of images 121, 122, 152, 153 and their corresponding reference images 113, 144 in Figure 1 and Figure 2 .

[0247] Figure 1 shows each of the first plurality of images 121, 122 that are before the first reference image 113 in time, and each of the second plurality of images 152, 153 that are before the second reference image 144 in time.

[0248] Figure 2 indicates an alternative embodiment, where the first reference image 113 is before the first plurality of images 121, 122 in time. Although not shown in Figure 2 , the second reference image 144 can similarly be before the second plurality of images 152, 153 in time.

[0249] Figure 2Indicate that the second reference image 144 can be temporally sandwiched between the images of the second plurality of image frames 152, 153. Similarly, the second reference image 113 can be temporally sandwiched between the images of the first plurality of image frames 121, 122.

[0250] After forming the second plurality of transformed images 162, 163, perform TNF on the second reference image 144 using the second plurality of transformed images 162, 163. Details related to performing TNF on the second reference image 144 can be similar to the details related to performing TNF on the first reference image 113 (step E) discussed above.

[0251] The method can be used to form a temporally noise-filtered video stream. In this case, the reference images 113, 144 will form the video stream. Each reference image 113, 144 can be associated with a different time point in the video stream.

[0252] As Figure 5 shown, the method can be implemented in the pan-tilt camera 300. However, as described above, the method can also be executed in a device external to the pan-tilt camera 300.

[0253] The method can include: before step A, determining the capture conditions for a sequence of temporally consecutive images 101, 102, 103, 104. In this case, steps A to J are only executed when it is determined that the capture conditions meet the predetermined capture condition requirements. Thus, it can be considered that method steps A to method step J are delayed until the requirements are met.

[0254] The capture conditions can be determined by the motion level. The motion level can be determined from pan-tilt data including values of the pan and tilt of the pan-tilt camera 300 associated with different time points (such as the time point of capturing an image or any time point). The capture conditions can be determined by the light level. The light level is determined by the light sensor 220 or through image analysis.

[0255] The pan-tilt data can be obtained from a pan-tilt controller (such as the pan-tilt control function in the computing unit 304) to control the pan-tilt camera for the expected pan and tilt. The pan-tilt data including values of the pan and tilt of the pan-tilt camera associated with a time point can be related to the values of the pan and tilt of the pan-tilt camera 300 according to the expected pan and tilt of the pan-tilt camera 300 at that time point. The pan-tilt data can be further corrected by the deviation from the expected pan and tilt identified by the pan-tilt sensor at the time point as indicated in the feedback from the pan-tilt sensor.

[0256] The values of the pan and tilt associated with a time point can be absolute values (i.e., values relative to a coordinate system fixed to the reference pan-tilt camera) or relative values (i.e., relative values of the pan and tilt of the pan-tilt camera associated with different time points).

[0257] The predetermined capture condition requirement can be a requirement for a light level below a predetermined level. Such a predetermined level can be in the range of 50 - 200 lux. The predetermined level can be more preferably in the range of 75 - 125 lux, for example 100 lux. Higher light level values can be associated with less noise, thus reducing the need for TNF for these higher light level values.

[0258] The predetermined capture condition requirement can be a requirement for a light level above a predetermined level. The predetermined capture condition requirement can further include a minimum acceptable light level and a maximum acceptable light level. The predetermined capture condition requirement can include an intermediate light level exclusion range, where light levels outside this range are acceptable.

[0259] The method can further include: after performing TNF on the first reference image 113, storing the first reference image 113 on the memory 306 of the pan-tilt camera 300. The method can include: after performing TNF on the second reference image 144, storing the second reference image 144 on the memory 306 of the pan-tilt camera 300.

[0260] The method can further include: after transforming one of the temporally consecutive images 101, 102, 103, 104 to have the same perspective as the first reference image 113 or the second reference image 144, deleting that one of the temporally consecutive images 101, 102, 103, 104 from the memory 306 of the pan-tilt camera 300.

[0261] The method can further include: after performing TNF on the first reference image 113, sending the first reference image 113 from the pan-tilt camera 300 to the remote device 230. The method can include: after performing TNF on the second reference image 144, sending the second reference image 144 from the pan-tilt camera 300 to the remote device 230.

[0262] The method can be implemented by a computer, a decoder, or another device with processing capabilities. A non-transitory computer-readable storage medium can be provided with instructions stored on the storage medium that, when executed by a device with processing capabilities, implement the method.

[0263] Figure 5 A pan-tilt camera 300 including an image capture unit 202 and a computing unit 204 is shown. As combined with Figure 5As explained for the pan-tilt camera 300, the pan-tilt camera 300 may be configured to perform Figure 1 , Figure 2 and Figure 3 the above methods.

[0264] The image capture unit 302 may be understood as any device capable of capturing images. The image capture unit 302 may include a charge-coupled device (CCD), an image sensor, or a complementary metal-oxide-semiconductor (CMOS)-based active pixel image sensor.

[0265] The computing unit 304 may include any device capable of performing the processing and calculations according to the method. The computing unit itself may include several sub-units for performing different actions or steps of the method.

[0266] The pan-tilt camera 300 is a camera that, in use, is configured to be fixedly mounted and can pan and tilt to capture different fields of view. A fixedly mounted pan-tilt camera for surveillance or monitoring (which can be manually or automatically controlled to pan and / or tilt to capture different fields of view) should be considered a non-limiting example of the term pan-tilt camera. In addition to the pan-tilt function, the pan-tilt camera may also include a zoom function and may be referred to as a pan-tilt-zoom camera or a PTZ camera.

[0267] As Figure 5 shown, the pan-tilt camera 300 may include a pan-tilt motor 310 configured to pan and tilt the pan-tilt camera 300 to a desired pan and tilt. The pan and tilt may be controlled by a pan-tilt controller (such as the pan-tilt control function in the computing unit 304). The pan-tilt motor may include a pan-tilt sensor (not shown) for providing feedback to the pan-tilt function in the computing unit 304 for feedback control to achieve the desired pan and tilt of the pan-tilt camera. Pan-tilt data including the values of the pan and tilt of the pan-tilt camera associated with an image may be related to the values of the pan and tilt of the pan-tilt camera 300 according to the desired pan and tilt when capturing the image. The pan-tilt data may be further corrected by the deviation from the desired pan and tilt recognized by the pan-tilt sensor during image capture as indicated in the feedback from the pan-tilt sensor. As Figure 5 shown, the pan-tilt camera 300 may include a light sensor 320 (photoelectric detector) configured to determine the light conditions or light level of the pan-tilt camera 300. As Figure 5 shown, the pan-tilt camera 300 may be configured to communicate with a remote device 330. The communication may be wireless or wired.

[0268] List of embodiments:

[0269] 1. A method for reducing noise in an image captured by a pan-tilt camera, the method comprising:

[0270] a. Provide a sequence of temporally continuous images,

[0271] b. Select a first reference image among the temporally continuous images,

[0272] c. Select a first plurality of images among the temporally continuous images for temporal noise filtering of the first reference image,

[0273] d. Form a first plurality of transformed images by transforming each of the first plurality of images to have the same perspective as the first reference image,

[0274] e. Perform temporal noise filtering TNF on the first reference image using the first plurality of transformed images,

[0275] f. Select a different second reference image among the temporally continuous images,

[0276] g. Select a second plurality of images among the temporally continuous images for temporal noise filtering of the second reference image, wherein at least one image among the images in the second plurality of images is also included in the first plurality of images, and wherein the second plurality of images includes the first reference image,

[0277] h. Determine whether TNF of the second reference image should be performed, and wherein when it is determined that TNF should be performed, the method further comprises:

[0278] i. Form a second plurality of transformed images by:

[0279] i1. For each image in the second plurality of images that is not included in the first plurality of images, transform the image to have the same perspective as the second reference image, wherein the perspective transformation is determined and used to transform the first reference image to have the same perspective as the second reference image;

[0280] i2. For each image in the second plurality of images that is also included in the first plurality of images, transform the corresponding transformed image in the first plurality of transformed images to have the same perspective as the second reference image using the perspective transformation used to transform the first reference image;

[0281] j. Perform temporal noise filtering on the second reference image using the second plurality of transformed images.

[0282] 2. The method according to Embodiment 1 further includes: after transforming one of the temporally consecutive images into an image having the same perspective as the first reference image or the second reference image, deleting the one image from the memory of the pan-tilt camera among the temporally consecutive images.

[0283] 3. The method according to any one of Embodiment 1 and Embodiment 2, wherein each image in the first plurality of images is before the first reference image in time, and each image in the second plurality of images is before the second reference image in time.

[0284] 4. The method according to any one of Embodiments 1 to 3, wherein the first plurality of images includes 4 to 8 images, and wherein the second plurality of images includes 4 to 8 images.

[0285] 5. The method according to any one of Embodiments 1 to 4, wherein forming the first plurality of transformed images and forming the second plurality of transformed images includes: transforming each image into an image having the same perspective as the associated reference image based on comparing the motion data associated with each image in the sequence of the temporally consecutive images.

[0286] 6. The method according to Embodiment 5, wherein the motion data associated with each image in the sequence of the temporally consecutive images is pan-tilt data, and the pan-tilt data includes the values of the pan and tilt of the pan-tilt camera associated with each image in the sequence of the temporally consecutive images.

[0287] 7. The method according to any one of Embodiments 1 to 6, wherein step h further includes: determining the perspective difference between at least two images in the first plurality of images, wherein when it is determined that the perspective difference is less than or equal to a predetermined perspective difference threshold, the TNF of the second reference image is performed, and wherein when it is determined that the perspective difference is greater than the predetermined perspective difference threshold, the TNF of the second reference image is not performed.

[0288] 8. The method according to Embodiment 7, wherein the perspective difference is based on the pan-tilt data associated with each image in time, and the pan-tilt data includes the values of the pan and tilt of the pan-tilt camera associated with each image in the sequence of the temporally consecutive images.

[0289] 9. The method according to any one of Embodiments 1 to 8 further includes: before step a, determining a capture condition for the sequence of the temporally consecutive images, wherein when it is determined that the capture condition meets the requirements of a predetermined capture condition, only steps a to j are performed.

[0290] 10. The method according to embodiment 9, wherein the capture condition is determined by at least one of the following:

[0291] The level of motion determined from gimbal data associated with each image over time, the gimbal data including values of translation and tilt of the gimbal camera associated with different time points; and

[0292] The light level determined by a light sensor or by image analysis.

[0293] 11. The method according to embodiment 10, wherein the capture condition is determined by the light level, the light level being determined by a light sensor or by image analysis, wherein the predetermined capture condition requires a light level below a predetermined level.

[0294] 12. The method according to any one of embodiments 1 to 11, further comprising: storing the first reference image on a memory of the gimbal camera after performing TNF on the first reference image, and storing the second reference image on the memory of the gimbal camera after performing TNF on the second reference image.

[0295] 13. The method according to any one of embodiments 1 to 12, further comprising: sending the first reference image from the gimbal camera to a remote device after performing TNF on the first reference image; and sending the second reference image from the gimbal camera to the remote device after performing TNF on the second reference image.

[0296] 14. A gimbal camera including an image capture unit and a computing unit, the gimbal camera being configured to:

[0297] a. Capture a sequence of temporally consecutive images by the image capture unit,

[0298] b. Select a first reference image among the temporally consecutive images by the computing unit,

[0299] c. Select a first plurality of images among the temporally consecutive images by the computing unit for temporal noise filtering of the first reference image;

[0300] d. Form a first plurality of transformed images by the computing unit by transforming each of the first plurality of images to have the same perspective as the first reference image;

[0301] e. Perform temporal noise filtering TNF on the first reference image by the computing unit using the first plurality of transformed images,

[0302] r. The computing unit selects different second reference images from the images consecutive in time;

[0303] g. The computing unit selects a second plurality of images from the images consecutive in time for temporal noise filtering of the second reference image, wherein at least one image among the images in the second plurality of images is also included in the first plurality of images, and wherein the second plurality of images includes the first reference image;

[0304] h. The computing unit determines whether TNF of the second reference image should be performed, and when it is determined that TNF should be performed, the pan-tilt camera is further configured for:

[0305] i. The computing unit forms a second plurality of transformed images by:

[0306] i1. For each image in the second plurality of images that is not included in the first plurality of images, transforming the image to have the same viewing angle as the second reference image, wherein the viewing angle transformation is determined and used to transform the first reference image to have the same viewing angle as the second reference image;

[0307] i2. For each image in the second plurality of images that is also included in the first plurality of images, using the viewing angle transformation for transforming the first reference image to transform the corresponding transformed image in the first plurality of transformed images to have the same viewing angle as the second reference image;

[0308] j. The computing unit performs temporal noise filtering on the second reference image using the second plurality of transformed images.

[0309] 15. A non - transitory computer - readable storage medium having instructions stored thereon, the instructions, when executed by a device having processing capabilities, implement the method according to any one of Embodiments 1 to 13.

[0310] In addition, by studying the drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments when practicing the claimed invention.

Claims

1. A method for reducing noise in images captured by a wearable camera, the method comprising: A. Providing a sequence of temporally continuous images, B. Selecting a first reference image among the temporally continuous images, C. Selecting a first plurality of images among the temporally continuous images for temporal noise filtering of the first reference image, D. Forming a first plurality of transformed images by transforming each of the first plurality of images to have the same perspective as the first reference image, E. Performing temporal noise filtering TNF on the first reference image using the first plurality of transformed images, F. Selecting a different second reference image among the temporally continuous images, G. Selecting a second plurality of images among the temporally continuous images for temporal noise filtering of the second reference image, wherein at least one of the images in the second plurality of images is also included in the first plurality of images, and wherein the second plurality of images includes the first reference image, H. Determining whether TNF of the second reference image should be performed, wherein when it is determined that TNF should be performed, the method further comprises: I. Forming a second plurality of transformed images by: I1. For each image in the second plurality of images that is not included in the first plurality of images, transforming the image to have the same perspective as the second reference image, wherein the perspective transformation is determined and used to transform the first reference image to have the same perspective as the second reference image; I2. For each image in the second plurality of images that is also included in the first plurality of images, transforming the corresponding transformed image in the first plurality of transformed images to have the same perspective as the second reference image using the perspective transformation used to transform the first reference image; J. Performing temporal noise filtering on the second reference image using the second plurality of transformed images.

2. The method according to claim 1, further comprising: After transforming an image in the temporally continuous images to have the same perspective as the first reference image or the second reference image, deleting the image in the temporally continuous images from the memory of the wearable camera.

3. The method according to claim 1, wherein Each image in the first plurality of images is temporally before the first reference image, and each image in the second plurality of images is temporally before the second reference image.

4. The method according to claim 1, wherein, The first plurality of images includes 4 to 8 images, and wherein the second plurality of images includes 4 to 8 images.

5. The method according to claim 1, wherein Forming the first plurality of transformed images and forming the second plurality of transformed images includes: transforming each image to have the same perspective as the associated reference image based on comparing motion data associated with each image in the sequence of temporally continuous images.

6. The method according to claim 5, wherein, The motion data associated with each image in the sequence of temporally continuous images is determined by at least one of a motion sensor, an accelerometer, and a gyroscope, or the motion data is determined based on image analysis of the respective images.

7. The method according to claim 1, wherein Step H further includes: determining a perspective difference between at least two images among the first plurality of images, wherein when it is determined that the perspective difference is less than or equal to a predetermined perspective difference threshold, the TNF of the second reference image is performed, and wherein when it is determined that the perspective difference is greater than the predetermined perspective difference threshold, the TNF of the second reference image is not performed.

8. The method according to claim 7, wherein The perspective difference is based on at least one of the following: Motion data associated with each image in time, wherein the motion data is determined by a motion sensor, an accelerometer, or a gyroscope; and Image data related to how many pixels have changed between subsequent images of the first plurality of images.

9. The method according to claim 1, further comprising: Before step A, capture conditions are determined for a sequence of temporally consecutive images, and wherein when it is determined that the capture conditions meet the predetermined capture condition requirements, only steps A to J are performed.

10. The method according to claim 9, wherein, The capture conditions are determined by at least one of the following: The level of motion determined by a motion sensor, an accelerometer, a gyroscope, or a positioning device; and The level of light determined by a light sensor or by image analysis.

11. The method according to claim 9, wherein, The capture conditions are determined by the level of light, which is determined by a light sensor or by image analysis, wherein the predetermined capture condition requirement is a level of light below a predetermined level.

12. The method according to claim 1, further comprising: After performing TNF on the first reference image, store the first reference image in the memory of the wearable camera; and after performing TNF on the second reference image, store the second reference image in the memory of the wearable camera.

13. The method according to claim 1, further comprising: After performing TNF on the first reference image, send the first reference image from the wearable camera to a remote device; and after performing TNF on the second reference image, send the second reference image from the wearable camera to the remote device.

14. A wearable camera including an image capture unit and a computing unit, the wearable camera being configured to: A. Capture a sequence of temporally consecutive images by the image capture unit, B. Select a first reference image among the temporally consecutive images by the computing unit, C. Select a first plurality of images among the temporally consecutive images by the computing unit for temporal noise filtering of the first reference image; D. Transform each of the first plurality of images by the computing unit to have the same perspective as the first reference image to form a first plurality of transformed images; E. Perform temporal noise filtering TNF on the first reference image by the computing unit using the first plurality of transformed images, F. Select a different second reference image among the temporally consecutive images by the computing unit, G. Select a second plurality of images among the temporally consecutive images by the computing unit for temporal noise filtering of the second reference image, wherein at least one image among the images in the second plurality of images is also included in the first plurality of images, and wherein the second plurality of images includes the first reference image, H. Determine, by the computing unit, whether TNF of the second reference image should be performed, wherein, when it is determined that TNF should be performed, the wearable camera is further configured to: I. Form, by the computing unit, a second plurality of transformed images by: I1. For each image in the second plurality of images that is not included in the first plurality of images, transform the image to have the same perspective as the second reference image, wherein the perspective transformation is determined and used to transform the first reference image to have the same perspective as the second reference image; I2. For each image in the second plurality of images that is also included in the first plurality of images, transform the corresponding transformed image in the first plurality of transformed images to have the same perspective as the second reference image using the perspective transformation used to transform the first reference image; J. Perform, by the computing unit, temporal noise filtering on the second reference image using the second plurality of transformed images.

15. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by a device having processing capabilities, implement the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • System and method for multi-frame temporal de-noising using image alignment

    US20150262341A1