Camera Noise Reduction

The method for noise reduction in wearable cameras optimizes temporal noise filtering by transforming images to a common viewpoint, addressing motion-induced blur and power consumption, thus enhancing image quality and reducing computational demands.

JP7718832B2Active Publication Date: 2025-08-05AXIS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021048021
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2021-03-23
Publication Date
2025-08-05
Estimated Expiration
2041-03-23

AI Technical Summary

Technical Problem

Wearable cameras face challenges in noise reduction, particularly in low-light conditions, due to motion-induced blur and high computational demands of existing noise filtering techniques, which consume excessive power and increase complexity.

Method used

A method for noise reduction in wearable cameras that involves transforming consecutive images to have the same viewpoint, allowing reuse of transformations and optimizing temporal noise filtering (TNF) to reduce computation and energy usage, while minimizing motion-induced blur.

Benefits of technology

The method effectively reduces noise in wearable camera images by optimizing TNF, conserving computational resources and power, and reducing the complexity and size of wearable cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007718832000001
    Figure 0007718832000001
  • Figure 0007718832000002
    Figure 0007718832000002
  • Figure 0007718832000003
    Figure 0007718832000003
Patent Text Reader

Abstract

To provide a method of noise reduction in images captured by a wearable camera.SOLUTION: A method includes: selecting a first reference image 113 among a sequence of temporally successive images 101-104; selecting a first plurality of images 121, 122 among the temporally successive images to be used for temporal noise filtering of the first reference image; and, after the selection of the first plurality of images, forming a first plurality of transformed images 131, 132 by transforming each of the first plurality of images to have a same perspective as the first reference image. The method also includes performing TNF on the first reference image using the first plurality of transformed images. The method further includes: selecting a second plurality of images 152, 153 among the temporally successive images to be used for temporal noise filtering of a second reference image 144, in order to promote TNF again; and, upon determining that TNF of the second reference image should be performed, forming a second plurality of transformed images 162, 163.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to cameras and, in particular, to methods for noise reduction in images captured by cameras. [Background technology]

[0002] Wearable cameras are becoming increasingly common and are used in a wide variety of applications, ranging from artistic or recreational uses to security and documentation. A common problem with wearable cameras, as with other cameras, is that of image noise, which can be generated somewhere within the camera, such as in the image sensor, optics, or circuitry. Noise is typically seen as random variations in brightness and color in an image. Noise is particularly prevalent in images captured in low-light environments, i.e., when few photons are recorded by the image sensor. Low-light environments generally correspond to images with a low signal-to-noise ratio (SNR). One technique for mitigating the effects of image noise is called temporal noise filtering (TNF). TNF relies on averaging over multiple images captured at different times. However, such averaging can cause motion-induced blur for moving objects in an otherwise static scene, and even when the entire scene is dynamic, such as when the camera is moving, causing the entire scene to feature some degree of blur. A wearable camera worn by a wearer is expected to be constantly exposed to some motion due to the wearer's motion, whether that motion is conscious, such as when the wearer is walking, or subconscious, such as when the wearer is breathing. Various techniques are known for adapting motion compensation depending on the source of the motion. However, motion compensation techniques can require excessive computing resources and therefore consume more power. This can pose a significant challenge for wearable cameras with limited power storage, such as batteries. Mechanical motion stabilization is also considered to adequately address all types of motion-induced blur. However, such systems can make wearable cameras more complex and larger. Therefore, there is a need for improvement in this field. Summary of the Invention

[0003] It is an object of the present invention to alleviate at least some of the above problems and to provide improved noise reduction for wearable cameras.

[0004] According to a first aspect of the present invention, there is provided a method for noise reduction in images captured by a wearable camera, the method comprising: A. Providing a sequence of time-sequential images; B. Selecting a first reference image from the temporal sequence of images; C. Selecting a first plurality of images from the temporally consecutive images to be used for temporal noise filtering of the first reference image; D. forming a first plurality of transformed images by transforming each of the first plurality of images to have the same viewpoint as the first reference image; E. Performing temporal noise filtering (TNF) on a first reference image using the first plurality of transformed images; F. Selecting a second, different, reference image from the temporally consecutive images; G. Selecting a second plurality of images from a temporally consecutive image sequence to be used for temporal noise filtering of a second reference image, wherein at least one of the images of the second plurality of images is also included in the first plurality of images, and the second plurality of images includes the first reference image; H. Determining whether TNF of a second reference image should be performed; Including, Once it is determined that TNF should be administered, I. forming a second plurality of transformed images, I1. For each image of a second plurality of images that is not included in the first plurality of images, transforming the image to have the same viewpoint as a second reference image; I2. For each image of a second plurality of images that is also included in the first plurality of images, transforming a corresponding transformed image of the first plurality of transformed images to have the same viewpoint as the second reference image; forming a second plurality of transformed images by J. performing temporal noise filtering on a second reference image using the second plurality of transformed images; Further includes:

[0005] The term "wearable camera" may be understood as a camera that is configured to be worn by a wearer when in use. Body-worn cameras, eyewear-worn cameras, and helmet-worn cameras should be considered as non-limiting examples of the term wearable camera. The wearer may be, for example, a person or an animal.

[0006] The term "temporally successive images" may be understood as images or image frames taken successively at different points in time. In other words, for example, a first image is taken earlier in time compared to the time of taking for a second image, which is subsequently taken earlier in time compared to the time of taking for a third image. The images or image frames together may form a sequence of temporally successive images or image frames of a video stream.

[0007] The term "transforming images to have a same perspective" may refer to transforming an image or creating a projection of an image as if it were captured using a camera positioned and oriented in the same way as when capturing another image. This phrase should be understood to refer to substantially the same perspective. Various methods for achieving perspective transformation may include using homographies and image projections. For example, readouts from sensors such as accelerometers and gyroscopes may be employed to perform the perspective transformation. The perspective transformation may include calculating a homography based on each corresponding pair of candidate points in two different images.

[0008] It should be noted that steps A through J do not necessarily have to be performed temporally in the order in which they are presented in this disclosure. Steps I1 and I2 should both be understood as being parts of step I, respectively.

[0009] The above method provides a way to optimize the TNF of images captured by a wearable camera. In particular, the method may reduce the amount of computation and energy used. This is achieved by creating a process flow. Here, the second transformation step may have the advantage that the same viewpoint transformation can be reused for at least two images when forming the second plurality of transformed images. This should be understood as the case where at least one of the first plurality of transformed images and the first reference image are present in the second plurality of images, allowing the same viewpoint transformation to be used when transforming these images to the same viewpoint transformation as the second reference image. This means that fewer new viewpoint transformations may need to be calculated / determined when generating the second plurality of transformed images. Furthermore, reusing instructions for performing viewpoint transformations may reduce the amount of storage or cache storage required to transform images.

[0010] It should be understood that the present invention advantageously allows for the repeated iteration of the provided method for video streams containing a large number of temporally consecutive images, allowing for multiple reuse of the viewpoint transformation.

[0011] The larger the scale of the process, which is the case when performing the method on a continuous stream of images, i.e., video, the more significant the gains from the reduction efforts through successive iterations.

[0012] According to some embodiments of the first aspect, the method may further include deleting one of the temporally consecutive images from the memory of the wearable camera after transforming the one of the temporally consecutive images to have the same perspective as the first or second reference image.

[0013] A favorable effect of such an embodiment may be that of reducing computer memory requirements, which is possible according to the present method, where the original viewpoint images may become unnecessary once they have been transformed and another image viewpoint is characterized, since the transformed images are actually used when needed to perform TNF on other images in the temporal sequence (see, for example, step I2).

[0014] According to some embodiments of the first aspect, each image of the first plurality of images may precede in time the first reference image, and each image of the second plurality of images may precede in time the second reference image.

[0015] Such an embodiment may reduce latency, which is the case when only temporally preceding image information is used for the TNF of the reference image, meaning that the delay between providing an image and performing TNF on that image is reduced. Reducing latency may be particularly advantageous when the method is applied to live video streams.

[0016] According to some embodiments of the first aspect, the first plurality of images may include four to eight images, and the second plurality of images may include four to eight images.

[0017] Increasing the number of images used in the TNF can generally improve the results of the noise reduction method. Increasing the number of images allows for more reuse of viewpoint transformations. However, using more images generally increases the computational complexity of the TNF. The range of images in this embodiment represents a good compromise between the quality of the TNF and the computational complexity.

[0018] According to some embodiments of the first aspect, forming the first plurality of transformed images and forming the second plurality of transformed images may include transforming the images to have the same viewpoint as an associated reference image based on comparing motion data associated with each image in a sequence of temporally consecutive images.

[0019] The term "motion data" may refer to any data or information relating to the physical movement of a camera relative to the scene it monitors.

[0020] Such an embodiment provides flexibility for performing the method, as various types of motion data may be present and a wide range of means for determining the motion data may be applicable.

[0021] According to some embodiments of the first aspect, the motion data associated with each image in the sequence of temporally consecutive images may be determined by at least one of a motion sensor, an accelerometer, and a gyroscope, or the motion data is determined based on image analysis of the images.

[0022] An accelerometer, when mounted on or associated with a camera, can provide accurate motion data regarding the direction, speed, and acceleration of camera movement. A gyroscope, when mounted on or associated with a camera, can provide accurate motion data regarding the orientation of the camera. A motion sensor can provide similar and / or additional motion data. Image analysis can provide motion data based on analysis of images captured by the camera, for example, by comparing successively captured images. The above methods of determining motion data can be combined with or performed in combination with each other.

[0023] According to some embodiments of the first aspect, step H may further include determining a viewpoint difference between at least two images of the first plurality of images, wherein the TNF of the second reference image is performed if it is determined that the viewpoint difference is less than or equal to a predetermined viewpoint difference threshold, and the TNF of the second reference image is not performed if it is determined that the viewpoint difference is greater than the predetermined viewpoint difference threshold.

[0024] It may be preferable to perform TNF only if the motion is within certain compliance limits. TNF relies on averaging temporally consecutive images. Performing TNF on consecutive images that contain too much motion may result in distorted image details. In such cases, it may be preferable to not expend the effort of performing TNF a second time at all. Furthermore, this may also provide a favorable effect in reducing the amount of computation required by essentially not performing the remainder of the noise reduction step as long as the motion is too large.

[0025] According to some embodiments of the first aspect, the viewpoint difference is motion data temporally associated with each image, the motion data being determined by a motion sensor, an accelerometer, or a gyroscope; image data regarding how many pixels have changed between subsequent images of the first plurality of images; and It may be based on at least one of the following:

[0026] By using sensors other than the camera, such as a motion sensor, accelerometer, or gyroscope, the determination may be more stable and not dependent on the camera's operation. By using camera image data, the system may be less complex and not dependent on other sensors. Determining how many pixels have changed between subsequent images may provide easy applicability and adaptability of the determination of whether TNF should be performed.

[0027] According to some embodiments of the first aspect, the method may further include, before step A, determining imaging conditions for the sequence of temporally consecutive images, and only steps A to J are performed if it is determined that the imaging conditions satisfy predetermined imaging condition requirements.

[0028] It may be preferable to perform TNF only if the imaging conditions are within certain compliance limits. Certain image imaging conditions may be recognized as unsuitable for TNF due to expected distortion of image details. It may therefore be preferable to check whether the imaging conditions are beneficial, or at least acceptable, before performing the method steps. As with the above embodiment, this may also provide the advantageous effect of reducing the amount of computation required by essentially not performing the remaining steps of the method unless the imaging conditions meet certain imaging condition requirements.

[0029] According to some embodiments of the first aspect, the imaging conditions include: a level of motion as determined by a motion sensor, accelerometer, gyroscope, or positioning device; and a light level determined by a light sensor or by image analysis; It is determined by at least one of the following.

[0030] If the level of motion is too high, it may be preferable not to perform steps A to J of the method. This may be understood as not performing a second TNF if the motion is too large and is expected to distort image details, similar to the embodiment described above. A level of motion that is too high may reduce the feasibility of a good viewpoint transformation, i.e., may distort the transformed image. A level of motion that is too high may also reduce the feasibility of a good TNF by reducing the number of common pixels between temporally consecutive images.

[0031] To determine whether the level of motion is too high, for example, a positioning device such as a global navigation satellite system (GNSS) and a tracker / receiver may be used. Such a device may be suitably used in conjunction with a camera to determine the speed at which the wearer is moving. This speed may be used to determine whether the wearer is, for example, running or walking. Steps A through J may be performed, for example, if the wearer is walking, but should not be performed if the wearer is running, because the latter activity involves a higher expected level of motion.

[0032] It may also be preferable to not perform steps A through J of the method based on light levels. TNF, when applied to low-light images, can lead to increased distortion of image detail. If the light levels are not satisfactory, that may be beneficial to avoid steps A through J of the method altogether. On the other hand, if the light levels are too high, TNF may not be required or unnecessary. TNF is therefore preferably avoided.

[0033] It may further be derived that when the light level exceeds a predetermined threshold, the motion level does not need to be recorded, and therefore the sensors and means for recording the motion, motion data, and / or motion level may be powered down, for example to conserve battery energy of the wearable camera.

[0034] Light levels and motion levels may be combined, for example, as a figure of merit, or evaluated in combination to determine imaging conditions.

[0035] According to some embodiments of the first aspect, the imaging condition is determined by a light level determined by a light sensor or by image analysis, and the predetermined imaging condition requirement is a light level below a predetermined level.

[0036] TNF may generally be more needed in low light conditions because cameras or image sensors often exhibit lower signal-to-noise ratios in low light conditions, so an upper threshold for performing steps A through J of the method may be preferable.

[0037] According to some embodiments of the first aspect, the method may further include storing the first reference image on a memory of the wearable camera after performing TNF on the first reference image, and storing the second reference image on a memory of the wearable camera after performing TNF on the second reference image.

[0038] The reference images, once TNF is performed on them, may therefore form the final video stream, which may be stored on the wearable camera's memory.

[0039] According to some embodiments of the first aspect, the method may further include transmitting a first reference image from the wearable camera to a remote device after performing TNF on the first reference image, and transmitting a second reference image from the wearable camera to the remote device after performing TNF on the second reference image.

[0040] Once the reference images have been TNF'd, they may therefore form the final video stream that is sent to a remote device for display or storage. Preferably, the reference images therefore do not need to be stored for extended periods on the wearable camera's memory.

[0041] According to a second aspect of the present invention, there is provided a wearable camera including an imaging unit and a computing unit. A. capturing a sequence of time-sequential images with an imaging unit; B. selecting, by a computing unit, a first reference image from the temporally consecutive images; C. selecting, by a computing unit, a first plurality of images from the temporally consecutive images to be used for temporal noise filtering of the first reference image; D. forming a first plurality of transformed images by transforming, with a computing unit, each of the first plurality of images to have the same viewpoint as the first reference image; E. performing, by a computing unit, temporal noise filtering (TNF) on a first reference image using the first plurality of transformed images; F. selecting, by a computing unit, a second, different, reference image from the temporally consecutive images; G. selecting, by a computing unit, a second plurality of images from a temporally consecutive image sequence to be used for temporal noise filtering of the second reference image, wherein at least one of the images of the second plurality of images is also included in the first plurality of images, and the second plurality of images includes the first reference image; H. determining by a calculation unit whether TNF of the second reference image should be performed; Once it is determined that TNF should be administered, I. forming, by a computing unit, a second plurality of transformed images; I1. For each image of a second plurality of images that is not included in the first plurality of images, transforming the image to have the same viewpoint as a second reference image; I2. For each image of a second plurality of images that is also included in the first plurality of images, transforming a corresponding transformed image of the first plurality of transformed images to have the same viewpoint as the second reference image; forming a second plurality of transformed images by J. performing, by a computing unit, temporal noise filtering on a second reference image using the second plurality of transformed images; It is further composed of:

[0042] The wearable camera described in the second aspect provides similar advantages to those of the method described in the first aspect due to their corresponding features, and may be considered as a device configured to implement the method of the first aspect.

[0043] According to a third aspect of the present invention, there is provided a non-transitory computer-readable storage medium having stored thereon instructions which, when executed by a device having processing capability, perform the method of the first aspect.

[0044] The non-transitory computer-readable storage medium described in the third aspect provides advantages similar to those of the method described in the first aspect.

[0045] Further scope of applicability of the present invention will become apparent from the following detailed description. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the present invention, are given for purposes of illustration only, since various changes and modifications within the scope of the invention will become apparent to those skilled in the art from this detailed description.

[0046] It should be noted that, as used in this specification and the appended claims, the articles "a," "an," "the," and "said" are intended to mean that there are one or more elements, unless the context clearly dictates otherwise. Thus, for example, reference to "a unit" or "the unit" may include several devices, etc. Furthermore, the words "comprising," "including," "containing," etc. do not exclude other elements or steps.

[0047] These and other aspects of the present invention will now be described in more detail with reference to the accompanying drawings, which should not be considered limiting, but instead should be considered for purposes of illustration and understanding.

[0048] As shown in these drawings, the size of each layer and region may be exaggerated for illustrative purposes and, therefore, is provided to illustrate the general structure. Like reference symbols denote like elements throughout the drawings. [Brief explanation of the drawings]

[0049] [Figure 1]FIG. 1 illustrates an exemplary method for noise reduction on temporally consecutive images captured by a wearable camera or a pan / tilt camera, where a temporally noise-filtered reference image temporally follows the images on which the noise filtering is based. [Figure 2] FIG. 2 illustrates an exemplary method for performing noise reduction on temporally consecutive images captured by a wearable camera or a pan / tilt camera, where the reference image that is temporally noise filtered does not temporally follow the images on which the noise filtering is based. [Figure 3] FIG. 3 shows a flowchart of a method for performing noise reduction on images captured by a wearable camera or a pan / tilt camera. [Figure 4] FIG. 4 shows a schematic of a wearable camera optionally in communication with a remote device. [Figure 5] FIG. 5 shows a schematic diagram of a pan-tilt camera optionally in communication with a remote device. DETAILED DESCRIPTION OF THE INVENTION

[0050] The present invention will now be described in more detail with reference to the accompanying drawings, which illustrate presently preferred embodiments of the invention. The invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided for completeness and completeness, and to fully convey the scope of the invention to those skilled in the art.

[0051] 1 and 2 show a method for noise reduction in images captured by a wearable camera, such as a body worn camera (BWC). FIG. 3 shows a flowchart for the method for noise reduction in images captured by a wearable camera. The method includes steps A to J. The method may be implemented in a wearable camera 200, as illustrated in FIG. 4.

[0052] The methods of Figures 1, 2, and 3 are described below in conjunction with the wearable camera 200 of Figure 4. Beginning in Figure 1, reference numerals such as 113 refer to images shown in Figures 1 and 2. Beginning in Figure 2, reference numerals such as 202 refer to features shown in Figure 4.

[0053] The blocks shown in Figures 1 and 2 represent images or image frames. Figures 1 and 2 include a horizontal time component relative to when the original images, i.e., the time-sequential images 101, 102, 103, and 104, were provided or captured. Blocks located at different horizontal positions within the figures indicate that the images were captured or provided at different times. Time t is shown progressing from left to right in the figures, i.e., block 101, which is one of the time-sequential images, is provided or captured before the remaining time-sequential images 102, 103, and 104. The time-sequential images 101, 102, 103, and 104 may be provided as raw format images from an image sensor device, i.e., there is no prior image processing of the images before providing them as input to the method. Alternatively, the method may be preceded by image processing of the time-sequential images 101, 102, 103, and 104 before providing them as input to the method. Non-limiting examples of preceding image processing include adjustments or modifications of image data, such as removal of bad pixels and fixed pattern noise filtering of columns. In other words, the temporally consecutive images 101, 102, 103, 104 may be provided as raw image data or as processed image data. However, it should be noted that the method according to the present invention is performed before the image video encoding process, i.e., the temporally consecutive images 101, 102, 103, 104 are not encoded / video encoded.

[0054] 1 and 2 also include a vertical / column component. Images placed below one another in the same column indicate that they are selected, processed, transformed, or filtered versions of the image above them in that same column. The images below them do not necessarily correspond to any temporal order of processing / transformation, etc., and may be processed / transformed, etc., at any time. Rather, they should be understood as being primarily based on the temporally consecutive images 101, 102, 103, 104 above them.

[0055] 1 and 2, the temporally successive images 101, 102, 103, 104 are provided in the top row, which should be understood as corresponding to step A of the method in FIG.

[0056] It should be noted that step A in the second aspect of the present invention, i.e., providing a wearable camera 200, specifies that the wearable camera 200 is configured to capture a sequence of temporally consecutive images 101, 102, 103, 104 by means of an imaging unit 202, while step A in the first aspect, i.e., providing the present method (shown in FIG. 3), specifies providing a sequence of temporally consecutive images 101, 102, 103, 104. Both of these aspects may still be considered to be interrelated. While the wearable camera 200 of the second aspect uses its own imaging unit 202 to capture a sequence of temporally consecutive images 101, 102, 103, 104, the present method of the first aspect does not require the camera itself to capture the temporally consecutive images 101, 102, 103, 104. The method may be performed remotely from the wearable camera 200 capturing the images, and therefore only requires the time-sequential images 101, 102, 103, 104 to be provided.

[0057] The method may be performed by or within wearable camera 200. Step A may be performed by imaging unit 202 of wearable camera 200. Steps B through J may be performed by computing unit 204 of wearable camera 200.

[0058] A first reference image 113 is selected from the temporally consecutive images 101, 102, 103, 104. This should be understood as corresponding to step B. The first reference image 113 should be understood as the image to be filtered using TNF. In Figure 1, image 103 is selected as the first reference image 113 from the temporally consecutive images 101, 102, 103, 104. In Figure 2, image 101 is selected as the first reference image 113 from the temporally consecutive images 101, 102, 103, 104.

[0059] To promote TNF, a first plurality of images 121, 122 are selected from the temporally consecutive images 101, 102, 103, 104 and used for temporal noise filtering of the first reference image 113. This should be understood as corresponding to step C. In FIG. 1, the first plurality of images 121, 122 are shown temporally preceding the selected first reference image 113 and correspond to images 101 and 102 of the temporally consecutive images 101, 102, 103, 104. In FIG. 2, the first plurality of images 121, 122 are shown temporally following the selected first reference image 113 and correspond to images 102 and 103 of the temporally consecutive images 101, 102, 103, 104. In other embodiments, the first plurality of images 121, 122 are selected from both images preceding and following the first reference image 113. It may be useful to note that the first plurality of images 121, 122 may include between four and eight images. To facilitate understanding of this general concept, Figures 1 and 2 show the first plurality of images 121, 122 including two images. Those skilled in the art and familiar with TNF will have the knowledge to apply the concepts disclosed below to any (reasonable) number of images.

[0060] After the first plurality of images 121, 122 are selected, a first plurality of transformed images 131, 132 are formed by transforming each of the first plurality of images 121, 122 to have the same viewpoint as the first reference image 113. This should be understood as corresponding to step D.

[0061] The transformation step D may be performed using a homography, which may include calculating a homography based on each corresponding pair of candidate points in two different images, for which images taken at different times may be considered as images taken from different cameras in a stereo camera arrangement for the purposes of generating a homography matrix.

[0062] The conversion step D may be based on comparing motion data associated with each image in the sequence of temporally consecutive images 101, 102, 103, 104. The motion data associated with each image in the sequence of temporally consecutive images 101, 102, 103, 104 may be determined by the motion sensor 210, the accelerometer 211 and / or the gyroscope 212. The motion data may alternatively or additionally be determined based on image analysis of the temporally consecutive images 101, 102, 103, 104.

[0063] The motion data may relate to the motion of the wearable camera 200 or the wearer of the wearable camera 200. The motion data may include the acceleration, speed, and / or direction of the motion. The motion data may include the orientation and / or position of the wearable camera 200 or the wearer of the wearable camera 200. The motion data may include data regarding the rotational motion of the wearable camera 200. The motion data may include data regarding the tumble / orbit, pan / yaw, tilt / pitch, and / or roll of the wearable camera 200. The motion data may include data regarding the translational motion of the wearable camera 200. The motion data may include data regarding the camera trajectory, dolly, platform / boom / jib, and / or track / crab.

[0064] In Figures 1 and 2, the block labeled PT indicates that a perspective transform is taking place. This works by having a base image transformed to have the same perspective as an image from a different vertical column, indicated by a dashed arrow pointing to the PT block. The base image is shown by a solid arrow as being transformed by the PT block. Sometimes, the same PT block may be used to transform multiple images, where each image in the second plurality of images 152, 153 is transformed to form a second plurality of transformed images 162, 163 (see further below). In this case, image 152 may be understood as a base image for transformed image 162. This same logic may be applied to image 153 and its transformed equivalent image 163.

[0065] After the first plurality of transformed images 131, 132 are formed, TNF of the first reference image 113 may be performed using the first plurality of transformed images 131, 132. This should be understood as corresponding to step E in FIG.

[0066] In Figures 1 and 2, the blocks labeled TNF indicate that a temporal noise filtering step is being performed. In these cases, the base image on which noise filtering is performed is indicated by a dot-dash arrow. The dashed arrows pointing to the TNF blocks indicate images from different vertical columns that are being used in the temporal noise filtering process. The TNF process may include, for example, averaging image content from images captured at different times. In general, these figures illustratively illustrate transforming an image to have the same perspective as before temporal noise filtering.

[0067] In these figures, the reference numbers may be the same before and after the TNF step. This may be motivated by the raw / unfiltered first and second reference images 113, 144 before / above the noise filtering step, which may be essentially the same as the images after / below the noise filtering step, but ideally with less noise. For later steps of the method, i.e., from step F onwards, both the raw / unfiltered or temporally noise-filtered reference images 113, 144 may be used.

[0068] After performing TNF on the first reference image 113, a second, different, reference image 144 may be selected from the temporally consecutive images. This should be understood as corresponding to step F in Figure 3. The second reference image 144 is different from the first reference image 113. The second reference image 144 should be understood as an image that, like the first reference image 113, should be filtered using TNF.

[0069] In the embodiment of Figure 1, image 104 is selected from the temporally consecutive images 101, 102, 103, 104 as the second reference image 144. In the embodiment of Figure 2, image 102 is selected from the temporally consecutive images 101, 102, 103, 104 as the second reference image 144.

[0070] To again promote TNF, a second plurality of images 152, 153 is selected from the temporally consecutive images 101, 102, 103, 104 and used for temporal noise filtering of the second reference image 144. This time, however, at least one of the images 152 of the second plurality of images 152, 153 is also included in the first plurality of images 121, 122. The second plurality of images 152, 153 also includes the first reference image 113. This should be understood as corresponding to step G in Figure 3. In the embodiment of Figure 1, the second plurality of images 152, 153 is shown temporally preceding the selected second reference image 144 and corresponds to images 102 and 103 of the temporally consecutive images 101, 102, 103, 104. In the embodiment of Figure 2, each of the images in the second plurality of images 152, 153 is shown both temporally preceding and temporally following the second reference image 144, and corresponds to images 101 and 103 in the temporally consecutive images 101, 102, 103, 104. In other embodiments, both images in the second plurality of images 152, 153 follow the second reference image 144. The second plurality of images 152, 153 may include four to eight images. Figures 1 and 2 show an embodiment of the second plurality of images 152, 153 including two images.

[0071] As disclosed, at least one of the images 152 of the second plurality of images 152, 153 is also included in the first plurality of images 121, 122. This may be understood to refer to any corresponding images based on the same, selected image of the temporally consecutive images 101, 102, 103, 104. In the case of Figure 1, this should be understood as any image based on images 101 and 102. For example, one or more images in the first plurality of transformed images 131, 132 may also be selected for the second plurality of images 152, 153, as shown in Figure 1, where transformed image 132 is selected as image 152 in the second plurality of images.

[0072] The method includes determining whether TNF of a second reference image should be performed, which should be understood as corresponding to step H in Figure 3. Step H may be performed after selecting the second plurality of images 152, 153.

[0073] Step H may include determining a viewpoint difference between at least two images of the first plurality of images 121, 122. The viewpoint difference may alternatively be determined based on any of the temporally consecutive images 101, 102, 103, 104. TNF of the second reference image 144 may be performed if it is determined that the viewpoint difference is less than or equal to a predetermined viewpoint difference threshold. TNF of the second reference image 144 should not be performed if it is determined that the viewpoint difference is greater than the predetermined viewpoint difference threshold. Step H may, in such a case, be performed before steps F and G. If it is determined that TNF of the second reference image 144 will not be performed, steps F and G may be omitted entirely from the method.

[0074] Step H may include determining a viewpoint difference between at least two images of the second plurality of images 121, 122. The TNF of the second reference image 144 may be performed if it is determined that the viewpoint difference is less than or equal to a predetermined viewpoint difference threshold. The TNF of the second reference image 144 should not be performed if it is determined that the viewpoint difference is greater than the predetermined viewpoint difference threshold.

[0075] The perspective difference may be based on motion data temporally associated with each image. The motion data may be determined by the motion sensor 210, the accelerometer 211, and / or the gyroscope 212. The perspective difference may be based on image data relating to how many pixels have changed in the first plurality of images 121, 122, in the second plurality of images 152, 153, or between temporally consecutive images 101, 102, 103, 104, directly or otherwise between subsequent images.

[0076] The predetermined viewpoint difference threshold may relate to the difference between whether the motion data or image data is expected to result from the wearer of camera 200 performing an action, such as running, walking, breathing, etc. In one embodiment, TNF is performed upon determining a particular type of low activity action, such as walking and / or breathing, and is not performed upon determining a particular type of high activity action, such as running.

[0077] The motion data may be periodically determined and evaluated to determine whether TNF should be performed. The motion data may be determined at a rate that matches the number of frames-per-second (FPS) at which wearable camera 200 captures time-sequential images 101, 102, 103, 104. The value for FPS may preferably be in the range of 1 to 60, more preferably in the range of 20 to 40.

[0078] If during step H it is determined that TNF of the second reference image 144 should not be performed, the method may end after step H, which means that steps I to J are not performed. The method may restart at step A in such a case.

[0079] If during step H it is determined that TNF of the second reference image 144 should be performed, the method proceeds to steps I to J as shown in FIG.

[0080] The method proceeds by forming a second plurality of transformed images 162, 163. This should be understood as corresponding to step I. Step I includes two sub-steps depending on how the images selected for the second plurality of images 152, 153 have been used previously in the method.

[0081] Each image 153 of the second plurality of images 152, 153 that is not included in the first plurality of images 121, 122 has been transformed to have the same viewpoint as the second reference image 144, thus forming a transformed image 163. This should be understood as corresponding to partial step I1.

[0082] For each image 152 of the second plurality of images 152, 153 that is also included in the first plurality of images 121, 122, the corresponding transformed image 132 (transformed in step D of FIG. 3 ) has been transformed to have the same viewpoint as the second reference image 144, thus forming a transformed image 162. This should be understood as corresponding to partial step I2. The transformed image 132 should be understood as corresponding in time to the images 122 and 152.

[0083] In Figure 1, this correspondence concerns how they all arise from image 102 of temporally consecutive images 101, 102, 103, 104. In Figure 2, images 122, 132, and 152 all arise from image 103 of temporally consecutive images 101, 102, 103, 104. The details of forming the second plurality of transformed images 162, 163 may be similar to those of forming the first plurality of transformed images 131, 132 (step D), for example, as described above with respect to using homography calculations and motion data.

[0084] According to this method, the same viewpoint transformation may be preferably used for at least two images of the second plurality of images 152,153 in forming the second plurality of transformed images 162,163.

[0085] To summarize the differences between FIGS. 1 and 2, the embodiments differ slightly with respect to the time sequence or order of the first and second plurality of images 121, 122, 152, 153 and their corresponding reference images 113, 144 in FIGS.

[0086] FIG. 1 shows each image of a first plurality of images 121, 122 that precedes in time a first reference image 113, and each image of a second plurality of images 152, 153 that precedes in time a second reference image 144.

[0087] 2 shows an alternative embodiment in which the first reference image 113 temporally precedes the first plurality of images 121, 122. The second reference image 144 may similarly temporally precede the second plurality of images 152, 153. However, this is not shown in FIG.

[0088] 2 shows that the second reference image 144 may be temporally interleaved between each image of the second plurality of image frames 152, 153. Similarly, the first reference image 113 may be temporally interleaved between each image of the first plurality of image frames 121, 122.

[0089] After the second plurality of transformed images 162, 163 are formed, TNF is performed on the second reference image 144 using the second plurality of transformed images 162, 163. Details regarding performing TNF on the second reference image 144 may be similar to those described above regarding performing TNF (step E) on the first reference image 113.

[0090] The method may be used to form a temporally noise-filtered video stream. In such a case, it is the reference images 113, 144 that form the video stream. All of the reference images 113, 144 may be associated with different points in time in the video stream.

[0091] The method may be implemented in wearable camera 200, as illustrated in Figure 4. However, as noted above, the method may also be performed in a device external to wearable camera 200.

[0092] The method may include, prior to step A, determining imaging conditions for the sequence of temporally consecutive images 101, 102, 103, 104. Steps A to J may be performed only if it is determined that the imaging conditions meet predetermined imaging condition requirements. Steps A to J of the method may therefore be considered to be delayed until those requirements are met.

[0093] The imaging conditions may be determined by the level of motion, which may be determined by the motion sensor 210, the accelerometer 211, the gyroscope 212, and / or the positioning device 213. The imaging conditions may be determined by the light level, which may be determined by the light sensor 220 or by image analysis.

[0094] The predetermined imaging condition requirement may be a requirement for a light level below a predetermined level. Such a predetermined level may be in the range of 50 to 200 lux. More preferably, the predetermined level may be in the range of 75 to 125 lux, such as 100 lux. Higher light level values may be associated with lower noise, thus mitigating the need for TNF for those higher light level values.

[0095] The predetermined imaging condition requirement may be a requirement for a light level above a predetermined level. The predetermined imaging condition requirement may further include both a lowest allowable light level and a highest allowable light level. The predetermined imaging condition requirement may include an intermediate light level exclusion range, outside which light levels are allowable.

[0096] The method may further include storing first reference image 113 on memory 206 of wearable camera 200 after performing TNF on first reference image 113. The method may include storing second reference image 144 on memory 206 of wearable camera 200 after performing TNF on second reference image 144.

[0097] The method may further include deleting one of the temporally consecutive images 101, 102, 103, 104 from the memory 206 of the wearable camera 200 after transforming the one of the temporally consecutive images 101, 102, 103, 104 to have the same perspective as the first or second reference image 113, 144.

[0098] The method may further include transmitting first reference image 113 from wearable camera 200 to remote device 230 after performing TNF on first reference image 113. The method may include transmitting second reference image 144 from wearable camera 200 to remote device 230 after performing TNF on second reference image 144.

[0099] The method may be performed by a computer, a decoder, or another device having processing capability. A non-transitory computer-readable storage medium may be provided with instructions stored thereon that, when executed by a device having processing capability, perform the method.

[0100] 4 shows a wearable camera 200 including an imaging unit 202 and a computing unit 204. The wearable camera 200 may be configured to perform the methods and method steps described above.

[0101] The imaging unit 202 may be understood as any device capable of capturing an image, and may include a charged coupled device (CCD) image sensor or a complementary metal-oxide-semiconductor (CMOS) based active pixel image sensor.

[0102] The computing unit 204 may include any device capable of performing the processing and calculations according to the present methods. The computing unit may itself include multiple sub-units for performing different actions or steps of the present methods.

[0103] The wearable camera 200 may be worn by a wearer. The wearer of the wearable camera 200 may be a person. The wearer of the wearable camera 200 may be a law enforcement professional. Further examples of wearers of the wearable camera 200 include security guards and workers operating in hazardous environments, such as road construction sites. The wearer of the wearable camera 200 may be a professional or amateur photographer / video camera operator recording for aesthetic, documentary, athletic, or recreational purposes. For some uses, such as by police officers, a long camera battery life and high quality detail capture may be more desirable. For other uses, such as for recreational / aesthetic purposes, color capture and easy visual recognition may be more desirable.

[0104] The wearer may alternatively be an animal, such as a dog, cat, or horse. The wearer may be a service animal, such as a law enforcement animal. Law enforcement animals may include, for example, police dogs trained to detect illegal substances or police horses deployed for crowd control duties. The wearer may be a hunting dog. The wearer may be a wild animal provided with wearable camera 200 for monitoring or scientific purposes. The wearer may be a pet animal. The wearer may be provided with wearable camera 200 to prevent the wearer from escaping, losing sight of the wearer, or injuring the wearer.

[0105] Wearable camera 200 may be mounted on a strap or harness on the wearer, or alternatively, may be fixedly mounted to a piece of clothing or protective equipment such as a helmet or vest.

[0106] Wearable camera 200 may include motion sensor 210 configured to determine any type of motion of wearable camera 200, as shown in FIG. 4. Wearable camera 200 may include accelerometer 211 configured to determine acceleration, speed, and / or direction of movement of wearable camera 200, as shown in FIG. 4. Wearable camera 200 may include gyroscope 212 configured to determine orientation of wearable camera 200, as shown in FIG. 4. Wearable camera 200 may include positioning device 213 configured to determine position, speed, and / or direction of movement of wearable camera 200, as shown in FIG. 4. Positioning device 213 may include a global navigation satellite system (GNSS) sensor or receiver. Positioning device 213 may include an inertial navigation system. Wearable camera 200 may include a compass configured to determine camera orientation. Wearable camera 200 may include a light sensor 220 (photodetector) configured to determine light conditions or light levels for wearable camera 200, as shown in Figure 4. Wearable camera 200 may be configured to communicate with a remote device 230, as shown in Figure 4. The communication may be wireless or wired.

[0107] The present disclosure further relates to pan / tilt cameras and, in particular, to methods for noise reduction in images captured by pan / tilt cameras.

[0108] Similar challenges to those disclosed for wearable cameras may arise for pan / tilt cameras.

[0109] It is an object of the present disclosure to alleviate at least some of the above problems and provide improved noise reduction for pan / tilt cameras.

[0110] According to a fourth aspect of the present disclosure, there is provided a method for noise reduction in images captured by a pan / tilt camera, the method comprising: A. Providing a sequence of time-sequential images; B. Selecting a first reference image from the temporal sequence of images; C. Selecting a first plurality of images from the temporally consecutive images to be used for temporal noise filtering of the first reference image; D. forming a first plurality of transformed images by transforming each of the first plurality of images to have the same viewpoint as the first reference image; E. Performing temporal noise filtering (TNF) on a first reference image using the first plurality of transformed images; F. Selecting a second, different, reference image from the temporally consecutive images; G. Selecting a second plurality of images from a temporally consecutive image sequence to be used for temporal noise filtering of a second reference image, wherein at least one of the images of the second plurality of images is also included in the first plurality of images, and the second plurality of images includes the first reference image; H. Determining whether TNF of a second reference image should be performed; Including, Once it is determined that TNF should be administered, I. forming a second plurality of transformed images, I1. For each image of a second plurality of images that is not included in the first plurality of images, transforming the image to have the same viewpoint as a second reference image; I2. For each image of the second plurality of images that is also included in the first plurality of images, transforming the corresponding transformed image in the first plurality of transformed images to have the same viewpoint as the second reference image; forming a second plurality of transformed images by J. performing temporal noise filtering on a second reference image using the second plurality of transformed images; Further includes:

[0111] The term "pan / tilt camera," when used, may be understood as a camera that is fixedly mounted and capable of panning and tilting, and is therefore configured to capture different fields of view. A pan / tilt camera used for monitoring or surveillance that is fixedly mounted, manually or automatically controlled, and can be panned and / or tilted to capture different fields of view should be considered a non-limiting example of the term pan / tilt camera. Pan / tilt cameras also include zoom functionality and therefore may also be referred to as pan-tilt-zoom (PTZ) cameras.

[0112] The term "temporally successive images" may be understood as images or image frames taken successively at different points in time. In other words, for example, a first image is taken earlier in time compared to the time of taking for a second image, which is subsequently taken earlier in time compared to the time of taking for a third image. The images or image frames together may form a sequence of temporally successive images or image frames of a video stream.

[0113] The term "transforming images to have a same perspective" may refer to transforming an image or creating a projection of an image as if it were captured using a camera positioned and oriented in the same way as when capturing another image. This phrase should be understood to refer to substantially the same perspective. Various methods for achieving perspective transformation may include using homographies and image projections. For example, readouts from sensors such as accelerometers and gyroscopes may be employed to perform the perspective transformation. The perspective transformation may include calculating a homography based on each corresponding pair of candidate points in two different images.

[0114] It should be noted that steps A through J do not necessarily have to be performed temporally in the order in which they are presented in this disclosure. Steps I1 and I2 should both be understood as being parts of step I, respectively.

[0115] The above method provides a way to optimize the TNF of images captured by a pan-tilt camera. In particular, the method may reduce the amount of computation and energy used. This is achieved by creating a process flow. Here, the second transformation step may have the advantage that the same viewpoint transformation can be reused for at least two images when forming the second plurality of transformed images. This should be understood as the case where at least one of the first plurality of transformed images and the first reference image are present in the second plurality of images, allowing the same viewpoint transformation to be used when transforming these images to the same viewpoint transformation as the second reference image. This means that fewer new viewpoint transformations may need to be calculated / determined when generating the second plurality of transformed images. Furthermore, reusing instructions for performing viewpoint transformations may reduce the amount of storage or cache storage required to transform images.

[0116] The present disclosure should be understood such that, for example, for video streams containing a large number of temporally consecutive images, repeated iteration of the provided method advantageously allows for multiple reuse of viewpoint transformations.

[0117] The larger the scale of the process, which is the case when performing the method on a continuous stream of images, i.e., video, the more significant the gains from the reduction efforts through successive iterations.

[0118] According to some embodiments of the fourth aspect, the method may further include deleting one of the temporally consecutive images from the memory of the pan / tilt camera after transforming the one of the temporally consecutive images to have the same perspective as the first or second reference image.

[0119] A favorable effect of such an embodiment may be that of reducing computer memory requirements, which is possible according to the present method, where the original viewpoint images may become unnecessary once they have been transformed and another image viewpoint is characterized, since the transformed images are actually used when needed to perform TNF on other images in the temporal sequence (see, for example, step I2).

[0120] According to some embodiments of the fourth aspect, each image of the first plurality of images may precede the first reference image in time, and each image of the second plurality of images may precede the second reference image in time.

[0121] Such an embodiment may reduce latency, which is the case when only temporally preceding image information is used for the TNF of the reference image, meaning that the delay between providing an image and performing TNF on that image is reduced. Reducing latency may be particularly advantageous when the method is applied to live video streams.

[0122] According to some embodiments of the fourth aspect, the first plurality of images may include between four and eight images, and the second plurality of images may include between four and eight images.

[0123] Increasing the number of images used in the TNF can generally improve the results of the noise reduction method. Increasing the number of images allows for more reuse of viewpoint transformations. However, using more images generally increases the computational complexity of the TNF. The range of images in this embodiment represents a good compromise between the quality of the TNF and the computational complexity.

[0124] According to some embodiments of the fourth aspect, forming the first plurality of transformed images and forming the second plurality of transformed images may include transforming the images to have the same viewpoint as an associated reference image based on comparing motion data associated with each image in a sequence of temporally consecutive images.

[0125] The term "motion data" may refer to any data or information relating to the physical movement of a camera relative to the scene it monitors.

[0126] Such an embodiment provides flexibility for performing the method, as various types of motion data may be present and a wide range of means for determining the motion data may be applicable.

[0127] According to some embodiments of the fourth aspect, the motion data associated with each image in the sequence of temporally consecutive images may be pan / tilt data including pan and tilt values of a pan / tilt camera associated with each image in the sequence of temporally consecutive images. The pan and tilt values associated with an image may be absolute values, i.e., values relative to a coordinate system fixed with respect to the pan / tilt camera, or relative values, i.e., values for the pan and tilt of pan / tilt cameras associated with different images.

[0128] According to some embodiments of the fourth aspect, step H may further include determining a viewpoint difference between at least two images of the first plurality of images, wherein the TNF of the second reference image is performed if it is determined that the viewpoint difference is less than or equal to a predetermined viewpoint difference threshold, and the TNF of the second reference image is not performed if it is determined that the viewpoint difference is greater than the predetermined viewpoint difference threshold.

[0129] It may be preferable to perform TNF only if the motion is within certain compliance limits. TNF relies on averaging temporally consecutive images. Performing TNF on consecutive images that contain too much motion may result in distorted image details. In such cases, it may be preferable to not expend the effort of performing TNF a second time at all. Furthermore, this may also provide a favorable effect in reducing the amount of computation required by essentially not performing the remainder of the noise reduction step as long as the motion is too large.

[0130] According to some embodiments of the fourth aspect, the viewpoint difference is based on pan / tilt data temporally associated with each image, the pan / tilt data including pan and tilt values of a pan / tilt camera associated with each image in a sequence of temporally consecutive images.

[0131] The pan and tilt values of a pan / tilt camera associated with an image may be relative to a coordinate system fixed with respect to the pan / tilt camera, or it may be relative to the pan and tilt of a pan / tilt camera associated with a different image.

[0132] By using pan / tilt data, which includes the pan and tilt values of a pan / tilt camera, the determination of viewpoint difference can be made very accurate.

[0133] According to some embodiments of the fourth aspect, the method may further include, before step A, determining imaging conditions for the sequence of temporally consecutive images, and only steps A to J are performed if it is determined that the imaging conditions satisfy predetermined imaging condition requirements.

[0134] It may be preferable to perform TNF only if the imaging conditions are within certain compliance limits. Certain image imaging conditions may be recognized as unsuitable for TNF due to expected distortion of image details. It may therefore be preferable to check whether the imaging conditions are beneficial, or at least acceptable, before performing the method steps. As with the above embodiment, this may also provide the advantageous effect of reducing the amount of computation required by essentially not performing the remaining steps of the method unless the imaging conditions meet certain imaging condition requirements.

[0135] According to some embodiments of the fourth aspect, the imaging conditions include: a level of motion as determined by pan / tilt data including pan and tilt values of a pan / tilt camera associated with different points in time; a light level determined by a light sensor or by image analysis; It is determined by at least one of the following.

[0136] If the level of motion is too high, it may be preferable not to perform steps A to J of the method. This may be understood as not performing a second TNF if the motion is too large and is expected to distort image details, similar to the embodiment described above. A level of motion that is too high may reduce the feasibility of a good viewpoint transformation, i.e., may distort the transformed image. A level of motion that is too high may also reduce the feasibility of a good TNF by reducing the number of common pixels between temporally consecutive images.

[0137] It may also be preferable to not perform steps A through J of the method based on light levels. TNF, when applied to low-light images, can lead to increased distortion of image detail. If the light levels are not satisfactory, that may be beneficial to avoid steps A through J of the method altogether. On the other hand, if the light levels are too high, TNF may not be required or unnecessary. TNF is therefore preferably avoided.

[0138] It may further be derived that once the light level exceeds a predetermined threshold, the motion level does not need to be recorded, and therefore the sensors and means for recording motion, motion data and / or motion levels may be powered down to reduce energy consumption.

[0139] Light levels and motion levels may be combined, for example, as a figure of merit, or evaluated in combination to determine imaging conditions.

[0140] According to some embodiments of the fourth aspect, the imaging condition is determined by a light level determined by a light sensor or by image analysis, and the predetermined imaging condition requirement is a light level below a predetermined level.

[0141] TNF may generally be more needed in low light conditions because cameras or image sensors often exhibit lower signal-to-noise ratios in low light conditions, so an upper threshold for performing steps A through J of the method may be preferable.

[0142] According to some embodiments of the fourth aspect, the method may further include storing the first reference image on a memory of the pan / tilt camera after performing TNF on the first reference image, and storing the second reference image on a memory of the pan / tilt camera after performing TNF on the second reference image.

[0143] The reference images, once TNF is performed on them, may therefore form the final video stream, which may be stored on the memory of the pan / tilt camera.

[0144] According to some embodiments of the fourth aspect, the method may further include transmitting a first reference image from the pan / tilt camera to a remote device after performing TNF on the first reference image, and transmitting a second reference image from the pan / tilt camera to the remote device after performing TNF on the second reference image.

[0145] Once the reference images have been TNF'd, they may therefore form the final video stream that is sent to a remote device for display or storage. Advantageously, the reference images may therefore not need to be stored for extended periods on the pan / tilt camera's memory.

[0146] According to a fifth aspect of the present disclosure, there is provided a pan-tilt-zoom pan / tilt camera, including an imaging unit and a computing unit. a. capturing a sequence of time-sequential images with an imaging unit; b. selecting, by a computing unit, a first reference image from the temporally consecutive images; c. selecting, by a computing unit, a first plurality of images from the temporally consecutive images to be used for temporal noise filtering of the first reference image; d. forming, with a computing unit, a first plurality of transformed images by transforming each of the first plurality of images to have the same viewpoint as the first reference image; e. performing, by a computing unit, temporal noise filtering (TNF) on a first reference image using the first plurality of transformed images; f. selecting, by the computing unit, a second, different, reference image from the temporally consecutive images; g. selecting, by a computing unit, a second plurality of images from a temporally consecutive image sequence to be used for temporal noise filtering of the second reference image, wherein at least one of the images of the second plurality of images is also included in the first plurality of images, and the second plurality of images includes the first reference image; h. determining, by a calculation unit, whether TNF of the second reference image should be performed; Once it is determined that TNF should be administered, i. forming, by a computing unit, a second plurality of transformed images; i1. for each image of a second plurality of images that is not included in the first plurality of images, transforming the image to have the same viewpoint as a second reference image; i2. For each image of a second plurality of images that is also included in the first plurality of images, transforming a corresponding transformed image of the first plurality of transformed images to have the same viewpoint as the second reference image; forming a second plurality of transformed images by j. performing, by a computing unit, temporal noise filtering on a second reference image using the second plurality of transformed images; It is further composed of:

[0147] The pan / tilt camera described in the fifth aspect provides similar advantages to those of the method described in the fourth aspect due to their corresponding features, and the pan / tilt camera may be considered as a device configured to implement the method of the fourth aspect.

[0148] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by a device having processing capability, perform the method of the fourth aspect.

[0149] The non-transitory computer-readable storage medium described in the sixth aspect provides advantages similar to those of the method described in the fourth aspect.

[0150] The fourth, fifth, and sixth aspects of the present disclosure are described more fully below with reference to Figures 1, 2, 3, and 5, which are presently preferred embodiments of the fourth, fifth, and sixth aspects of the present disclosure. These aspects may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided for completeness and completeness, and to fully convey the scope of these aspects to those skilled in the art.

[0151] 1 and 2 show a method for noise reduction in images captured by a pan / tilt camera. FIG. 3 shows a flowchart for the same method for noise reduction in images captured by a pan / tilt camera. The method includes steps A through J. The method may be implemented in a pan / tilt camera 300, as illustrated in FIG. 5.

[0152] The method of Figures 1, 2, and 3 is described below in conjunction with the pan / tilt camera 300 of Figure 5. Beginning in Figure 1, reference numerals such as 113 refer to images shown in Figures 1 and 2. Beginning in Figure 3, reference numerals such as 302 refer to features shown in Figure 5.

[0153] The blocks shown in Figures 1 and 2 represent images or image frames. Figures 1 and 2 include a horizontal time component relative to when the original images, i.e., the time-sequential images 101, 102, 103, and 104, were provided or captured. Blocks located at different horizontal positions within the figures indicate that the images were captured or provided at different times. Time t is shown progressing from left to right in the figures, i.e., block 101, which is one of the time-sequential images, is provided or captured before the remaining time-sequential images 102, 103, and 104. The time-sequential images 101, 102, 103, and 104 may be provided as raw format images from an image sensor device, i.e., there is no prior image processing of the images before providing them as input to the method. Alternatively, the method may be preceded by image processing of the time-sequential images 101, 102, 103, and 104 before providing them as input to the method. Non-limiting examples of prior image processing include adjustments or modifications of image data, such as removal of bad pixels and fixed pattern noise filtering of columns. In other words, the temporally consecutive images 101, 102, 103, 104 may be provided as raw image data or as processed image data. However, it should be noted that the method according to the present disclosure is performed before the image video encoding process, i.e., the temporally consecutive images 101, 102, 103, 104 are not encoded / video encoded.

[0154] 1 and 2 also include a vertical / column component. Images placed below one another in the same column indicate that they are selected, processed, transformed, or filtered versions of the image above them in that same column. The images below them do not necessarily correspond to any temporal order of processing / transformation, etc., and may be processed / transformed, etc., at any time. Rather, they should be understood as being primarily based on the temporally consecutive images 101, 102, 103, 104 above them.

[0155] 1 and 2, the temporally successive images 101, 102, 103, 104 are provided in the top row, which should be understood as corresponding to step A of the method in FIG.

[0156] It should be noted that step A of the fifth aspect of the present disclosure, i.e., providing a pan / tilt camera 300, specifies that the pan / tilt camera 300 is configured to capture time-sequential images 101, 102, 103, 104 with an imaging unit 302, while step A of the fourth aspect, i.e., providing a method (shown in FIG. 3), specifies providing a sequence of time-sequential images 101, 102, 103, 104. Both of these aspects may still be considered interrelated. While the pan / tilt camera 300 of the fifth aspect uses its own imaging unit 302 to capture the sequence of time-sequential images 101, 102, 103, 104, the method of the fourth aspect does not require the camera itself to capture the time-sequential images 101, 102, 103, 104. The method may be performed remotely from the pan / tilt camera 300 capturing the images, and therefore only requires the time-sequential images 101, 102, 103, 104 to be provided.

[0157] The method may be performed by or within the pan / tilt camera 300. Step A may be performed by the imaging unit 302 of the pan / tilt camera 300. Steps B through J may be performed by the computing unit 304 of the pan / tilt camera 300.

[0158] A first reference image 113 is selected from the temporally consecutive images 101, 102, 103, 104. This should be understood as corresponding to step B. The first reference image 113 should be understood as the image to be filtered using TNF. In Figure 1, image 103 is selected as the first reference image 113 from the temporally consecutive images 101, 102, 103, 104. In Figure 2, image 101 is selected as the first reference image 113 from the temporally consecutive images 101, 102, 103, 104.

[0159] To promote TNF, a first plurality of images 121, 122 are selected from the temporally consecutive images 101, 102, 103, 104 and used for temporal noise filtering of the first reference image 113. This should be understood as corresponding to step C. In FIG. 1, the first plurality of images 121, 122 are shown temporally preceding the selected first reference image 113 and correspond to images 101 and 102 of the temporally consecutive images 101, 102, 103, 104. In FIG. 2, the first plurality of images 121, 122 are shown temporally following the selected first reference image 113 and correspond to images 102 and 103 of the temporally consecutive images 101, 102, 103, 104. In other embodiments, the first plurality of images 121, 122 are selected from both images preceding and following the first reference image 113. It may be useful to note that the first plurality of images 121, 122 may include between four and eight images. To facilitate understanding of this general concept, Figures 1 and 2 show the first plurality of images 121, 122 including two images. Those skilled in the art and familiar with TNF will have the knowledge to apply the concepts disclosed below to any (reasonable) number of images.

[0160] After the first plurality of images 121, 122 are selected, a first plurality of transformed images 131, 132 are formed by transforming each of the first plurality of images 121, 122 to have the same viewpoint as the first reference image 113. This should be understood as corresponding to step D.

[0161] The transformation step D may be performed using a homography, which may include calculating a homography based on each corresponding pair of candidate points in two different images, for which images taken at different times may be considered as images taken from different cameras in a stereo camera arrangement for the purposes of generating a homography matrix.

[0162] The conversion step D may be based on comparing motion data associated with each image in the sequence of temporally consecutive images 101, 102, 103, 104. The motion data associated with each image in the sequence of temporally consecutive images 101, 102, 103, 104 may be pan / tilt data comprising pan and tilt values of the pan / tilt camera 300 associated with each image in the sequence of temporally consecutive images.

[0163] By comparing the panning and tilting values of the pan / tilt camera 300 when capturing the first reference image 113 with the panning and tilting values of the pan / tilt camera 300 when capturing each of the first plurality of images 121, 122, a transformation may be determined for transforming each of the first plurality of images 121, 122 to have the same perspective as the first reference image 113, thus forming the first plurality of transformed images 131, 132.

[0164] The pan / tilt data may be obtained from a pan / tilt controller, such as a pan / tilt control function in the computing unit 304, that controls the pan / tilt camera to pan and tilt as desired. The pan / tilt data, including pan and tilt values of the pan / tilt camera associated with an image, may relate to the pan and tilt values of the pan / tilt camera 300 according to the desired pan and tilt of the pan / tilt camera 300 when capturing the image. The pan / tilt data may further be modified using deviations from the desired pan and tilt as determined by the pan / tilt sensor when capturing the image, as indicated by feedback from the pan / tilt sensor.

[0165] The pan and tilt values associated with an image may be absolute, i.e., values relative to a coordinate system fixed with respect to the pan / tilt camera, or relative, i.e., values relative to the pan and tilt of the pan / tilt camera associated with different images.

[0166] In Figures 1 and 2, the block labeled PT indicates that a perspective transform is taking place. This works by having a base image transformed to have the same perspective as an image from a different vertical column, indicated by a dashed arrow pointing to the PT block. The base image is shown by a solid arrow as being transformed by the PT block. Sometimes, the same PT block may be used to transform multiple images, where each image in the second plurality of images 152, 153 is transformed to form a second plurality of transformed images 162, 163 (see further below). In this case, image 152 may be understood as a base image for transformed image 162. This same logic may be applied to image 153 and its transformed equivalent image 163.

[0167] After the first plurality of transformed images 131, 132 are formed, TNF of the first reference image 113 may be performed using the first plurality of transformed images 131, 132. This should be understood as corresponding to step E in FIG.

[0168] In Figures 1 and 2, the blocks labeled TNF indicate that a temporal noise filtering step is being performed. In these cases, the base image on which noise filtering is performed is indicated by a dot-dash arrow. The dashed arrows pointing to the TNF blocks indicate images from different vertical columns that are being used in the temporal noise filtering process. The TNF process may include, for example, averaging image content from images captured at different times. In general, these figures illustratively illustrate transforming an image to have the same perspective as before temporal noise filtering.

[0169] In these figures, the reference numbers may be the same before and after the TNF step. This may be motivated by the raw / unfiltered first and second reference images 113, 144 before / above the noise filtering step, which may be essentially the same as the images after / below the noise filtering step, but ideally with less noise. For later steps of the method, i.e., from step F onwards, both the raw / unfiltered or temporally noise-filtered reference images 113, 144 may be used.

[0170] After performing TNF on the first reference image 113, a second, different, reference image 144 may be selected from the temporally consecutive images. This should be understood as corresponding to step F in Figure 3. The second reference image 144 is different from the first reference image 113. The second reference image 144 should be understood as an image that, like the first reference image 113, should be filtered using TNF.

[0171] In the embodiment of Figure 1, image 104 is selected from the temporally consecutive images 101, 102, 103, 104 as the second reference image 144. In the embodiment of Figure 2, image 102 is selected from the temporally consecutive images 101, 102, 103, 104 as the second reference image 144.

[0172] To again promote TNF, a second plurality of images 152, 153 is selected from the temporally consecutive images 101, 102, 103, 104 and used for temporal noise filtering of the second reference image 144. This time, however, at least one of the images 152 of the second plurality of images 152, 153 is also included in the first plurality of images 121, 122. The second plurality of images 152, 153 also includes the first reference image 113. This should be understood as corresponding to step G in Figure 3. In the embodiment of Figure 1, the second plurality of images 152, 153 is shown temporally preceding the selected second reference image 144 and corresponds to images 102 and 103 of the temporally consecutive images 101, 102, 103, 104. In the embodiment of Figure 2, each of the images in the second plurality of images 152, 153 is shown both temporally preceding and temporally following the second reference image 144, and corresponds to images 101 and 103 in the temporally consecutive images 101, 102, 103, 104. In other embodiments, both images in the second plurality of images 152, 153 follow the second reference image 144. The second plurality of images 152, 153 may include four to eight images. Figures 1 and 2 show an embodiment of the second plurality of images 152, 153 including two images.

[0173] As disclosed, at least one of the images 152 of the second plurality of images 152, 153 is also included in the first plurality of images 121, 122. This may be understood to refer to any corresponding images based on the same, selected image of the temporally consecutive images 101, 102, 103, 104. In the case of Figure 1, this should be understood as any image based on images 101 and 102. For example, one or more images in the first plurality of transformed images 131, 132 may also be selected for the second plurality of images 152, 153, as shown in Figure 1, where transformed image 132 is selected as image 152 in the second plurality of images.

[0174] The method includes determining whether TNF of a second reference image should be performed, which should be understood as corresponding to step H in Figure 3. Step H may be performed after selecting the second plurality of images 152, 153.

[0175] Step H may include determining a viewpoint difference between at least two images of the first plurality of images 121, 122. The viewpoint difference may alternatively be determined based on any of the temporally consecutive images 101, 102, 103, 104. TNF of the second reference image 144 may be performed if it is determined that the viewpoint difference is less than or equal to a predetermined viewpoint difference threshold. TNF of the second reference image 144 should not be performed if it is determined that the viewpoint difference is greater than the predetermined viewpoint difference threshold. Step H may, in such a case, be performed before steps F and G. If it is determined that TNF of the second reference image 144 will not be performed, steps F and G may be omitted entirely from the method.

[0176] Step H may include determining a viewpoint difference between at least two images of the second plurality of images 121, 122. The TNF of the second reference image 144 may be performed if it is determined that the viewpoint difference is less than or equal to a predetermined viewpoint difference threshold. The TNF of the second reference image 144 should not be performed if it is determined that the viewpoint difference is greater than the predetermined viewpoint difference threshold.

[0177] The perspective difference may be based on motion data temporally associated with each image, which may be pan / tilt data including pan and tilt values of the pan / tilt camera 300 associated with each of at least two images of the second plurality of images 121, 122.

[0178] The pan / tilt data may be obtained from a pan / tilt controller, such as a pan / tilt control function in the computing unit 304, that controls the pan / tilt camera to pan and tilt as desired. The pan / tilt data, including pan and tilt values of the pan / tilt camera associated with an image, may relate to the pan and tilt values of the pan / tilt camera 300 according to the desired pan and tilt of the pan / tilt camera 300 when capturing the image. The pan / tilt data may further be modified using deviations from the desired pan and tilt as determined by the pan / tilt sensor when capturing the image, as indicated by feedback from the pan / tilt sensor.

[0179] The pan and tilt values associated with an image may be absolute, i.e., values relative to a coordinate system fixed with respect to the pan / tilt camera, or relative, i.e., values relative to the pan and tilt of the pan / tilt camera associated with different images.

[0180] The motion data may be periodically determined and evaluated to determine whether TNF should be performed. The motion data may be determined at a rate that matches the number of frames-per-second (FPS) at which the pan / tilt camera 300 acquires time-sequential images 101, 102, 103, 104. The value for FPS may preferably be in the range of 1 to 60, more preferably in the range of 20 to 40.

[0181] If during step H it is determined that TNF of the second reference image 144 should not be performed, the method may end after step H, which means that steps I to J are not performed. The method may restart at step A in such a case.

[0182] If during step H it is determined that TNF of the second reference image 144 should be performed, the method proceeds to steps I to J as shown in FIG.

[0183] The method proceeds by forming a second plurality of transformed images 162, 163. This should be understood as corresponding to step I. Step I includes two sub-steps depending on how the images selected for the second plurality of images 152, 153 have been used previously in the method.

[0184] Each image 153 of the second plurality of images 152, 153 that is not included in the first plurality of images 121, 122 has been transformed to have the same viewpoint as the second reference image 144, thus forming a transformed image 163. This should be understood as corresponding to partial step I1.

[0185] For each image 152 of the second plurality of images 152, 153 that is also included in the first plurality of images 121, 122, the corresponding transformed image 132 (transformed in step D of FIG. 3 ) has been transformed to have the same viewpoint as the second reference image 144, thus forming a transformed image 162. This should be understood as corresponding to partial step I2. The transformed image 132 should be understood as corresponding in time to the images 122 and 152.

[0186] In Figure 1, this correspondence concerns how they all arise from image 102 of temporally consecutive images 101, 102, 103, 104. In Figure 2, images 122, 132, and 152 all arise from image 103 of temporally consecutive images 101, 102, 103, 104. The details of forming the second plurality of transformed images 162, 163 may be similar to those of forming the first plurality of transformed images 131, 132 (step D), for example, as described above with respect to using homography calculations and motion data.

[0187] According to this method, the same viewpoint transformation may be preferably used for at least two images of the second plurality of images 152,153 in forming the second plurality of transformed images 162,163.

[0188] To summarize the differences between FIGS. 1 and 2, the embodiments differ slightly with respect to the time sequence or order of the first and second plurality of images 121, 122, 152, 153 and their corresponding reference images 113, 144 in FIGS.

[0189] FIG. 1 shows each image of a first plurality of images 121, 122 that precedes in time a first reference image 113, and each image of a second plurality of images 152, 153 that precedes in time a second reference image 144.

[0190] 2 shows an alternative embodiment in which the first reference image 113 temporally precedes the first plurality of images 121, 122. The second reference image 144 may similarly temporally precede the second plurality of images 152, 153. However, this is not shown in FIG.

[0191] 2 shows that the second reference image 144 may be temporally interleaved between each image of the second plurality of image frames 152, 153. Similarly, the first reference image 113 may be temporally interleaved between each image of the first plurality of image frames 121, 122.

[0192] After the second plurality of transformed images 162, 163 are formed, TNF is performed on the second reference image 144 using the second plurality of transformed images 162, 163. Details regarding performing TNF on the second reference image 144 may be similar to those described above regarding performing TNF (step E) on the first reference image 113.

[0193] The method may be used to form a temporally noise-filtered video stream. In such a case, it is the reference images 113, 144 that form the video stream. All of the reference images 113, 144 may be associated with different points in time in the video stream.

[0194] The method may be implemented in a pan / tilt camera 300, as illustrated in Figure 5. However, as noted above, the method may also be implemented in a device external to the pan / tilt camera 300.

[0195] The method may include, prior to step A, determining imaging conditions for the sequence of temporally consecutive images 101, 102, 103, 104. Steps A to J may be performed only if it is determined that the imaging conditions meet predetermined imaging condition requirements. Steps A to J of the method may therefore be considered to be delayed until those requirements are met.

[0196] The imaging conditions may be determined by a level of motion, which may be determined from pan / tilt data including pan and tilt values of the pan / tilt camera 300 associated with different points in time, such as the time the image was captured or any time. The imaging conditions may be determined by a light level, which may be determined by the light sensor 220 or by image analysis.

[0197] The pan / tilt data may be obtained from a pan / tilt controller, such as a pan / tilt control function in the computing unit 304, that controls the pan / tilt camera to pan and tilt as desired. The pan / tilt data, including pan and tilt values of the pan / tilt camera associated with a point in time, may relate to the pan and tilt values of the pan / tilt camera 300 according to the desired pan and tilt of the pan / tilt camera 300 at that point in time. The pan / tilt data may further be modified with deviations from the desired pan and tilt as determined by the pan / tilt sensors at that point in time, as indicated by feedback from the pan / tilt sensors.

[0198] The pan and tilt values associated with a point in time may be absolute values, i.e., values relative to a fixed coordinate system relative to the pan / tilt camera, or relative values, i.e., values relative to the pan and tilt of the pan / tilt camera associated with different points in time.

[0199] The predetermined imaging condition requirement may be a requirement for a light level below a predetermined level. Such a predetermined level may be in the range of 50 to 200 lux. More preferably, the predetermined level may be in the range of 75 to 125 lux, such as 100 lux. Higher light level values may be associated with lower noise, thus mitigating the need for TNF for those higher light level values.

[0200] The predetermined imaging condition requirement may be a requirement for a light level above a predetermined level. The predetermined imaging condition requirement may further include both a lowest allowable light level and a highest allowable light level. The predetermined imaging condition requirement may include an intermediate light level exclusion range, outside which light levels are allowable.

[0201] The method may further include storing the first reference image 113 on the memory 306 of the pan / tilt camera 300 after performing TNF on the first reference image 113. The method may include storing the second reference image 144 on the memory 306 of the pan / tilt camera 300 after performing TNF on the second reference image 144.

[0202] The method may further include deleting one of the time-sequential images 101, 102, 103, 104 from the memory 306 of the pan / tilt camera 300 after transforming the one of the time-sequential images 101, 102, 103, 104 to have the same perspective as the first or second reference image 113, 144.

[0203] The method may further include transmitting the first reference image 113 from the pan-tilt camera 300 to the remote device 230 after performing TNF on the first reference image 113. The method may include transmitting the second reference image 144 from the pan-tilt camera 300 to the remote device 230 after performing TNF on the second reference image 144.

[0204] The method may be performed by a computer, a decoder, or another device having processing capability. A non-transitory computer-readable storage medium may be provided with instructions stored thereon that, when executed by a device having processing capability, perform the method.

[0205] 5 shows a pan / tilt camera 300 including an imaging unit 202 and a computing unit 204. The pan / tilt camera 300 may be configured to perform the above methods of FIGS. 1, 2, and 3, as described in conjunction with the pan / tilt camera 300 of FIG.

[0206] The imaging unit 302 may be understood as any device capable of capturing an image, and may include a charged coupled device (CCD) image sensor or a complementary metal-oxide-semiconductor (CMOS) based active pixel image sensor.

[0207] The computing unit 304 may include any device capable of performing the processing and calculations according to the present methods. The computing unit may itself include multiple sub-units for performing different actions or steps of the present methods.

[0208] Pan / tilt camera 300 is a camera that is fixedly mounted and capable of panning and tilting in use, thus configured to capture different fields of view. A pan / tilt camera used for monitoring or surveillance that is fixedly mounted, manually or automatically controlled, and can be panned and / or tilted to capture different fields of view should be considered a non-limiting example of the term pan / tilt camera. In addition to pan / tilt functionality, pan / tilt cameras also include zoom functionality and therefore may also be referred to as pan-tilt-zoom cameras (or PTZ cameras).

[0209] The pan / tilt camera 300 may include a pan / tilt motor 310 configured to pan and tilt the pan / tilt camera 300 to achieve a desired pan and tilt, as shown in FIG. 5 . The panning and tilting may be controlled by a pan / tilt controller, such as a pan / tilt control function in the computing unit 304. The pan / tilt motor may include a pan / tilt sensor (not shown) to provide feedback to the pan / tilt function in the computing unit 304 for feedback control to achieve the desired pan and tilt of the pan / tilt camera. Pan / tilt data, including pan and tilt values of the pan / tilt camera associated with an image, may relate the pan and tilt values of the pan / tilt camera 300 according to the desired pan and tilt when capturing the image. The pan / tilt data may be further modified using deviations from the desired panning and tilting as determined by the pan / tilt sensor when capturing an image, as indicated by feedback from the pan / tilt sensor. The pan / tilt camera 300 may include a light sensor 320 (photodetector) configured to determine the light conditions or light levels of the pan / tilt camera 300, as shown in FIG. 5. The pan / tilt camera 300 may be configured to communicate with a remote device 330, as shown in FIG. 5. The communication may be wireless or wired. List of embodiments:

[0210] 1. A method of noise reduction in an image captured by a pan / tilt camera, comprising: a. Providing a sequence of time-sequential images; b. selecting a first reference image from the temporal sequence of images; c. selecting a first plurality of images from the temporally consecutive images to be used for temporal noise filtering of the first reference image; d. forming a first plurality of transformed images by transforming each of the first plurality of images to have the same viewpoint as the first reference image; e. performing temporal noise filtering (TNF) on a first reference image using the first plurality of transformed images; f. selecting a second, different, reference image from the temporally consecutive images; g. selecting a second plurality of images from a temporally consecutive image sequence to be used for temporal noise filtering of the second reference image, wherein at least one of the images of the second plurality of images is also included in the first plurality of images, and the second plurality of images includes the first reference image; h. determining whether TNF of a second reference image should be performed; Including, Once it is determined that TNF should be administered, i. forming a second plurality of transformed images; i1. for each image of a second plurality of images not included in the first plurality of images, transforming the image to have the same viewpoint as a second reference image, wherein the viewpoint transformation is determined and used to transform the first reference image to have the same viewpoint as the second reference image; i2. for each image of a second plurality of images that is also included in the first plurality of images, transforming a corresponding transformed image in the first plurality of transformed images to have the same viewpoint as the second reference image using the viewpoint transformation used to transform the first reference image; forming a second plurality of transformed images by j. performing temporal noise filtering on a second reference image using the second plurality of transformed images; The method further comprises:

[0211] 2. The method of embodiment 1, further comprising deleting one of the time-sequential images from the memory of the pan / tilt camera after transforming the one of the time-sequential images to have the same viewpoint as the first or second reference image.

[0212] 3. Each image of the first plurality of images temporally precedes the first reference image; each image of the second plurality of images temporally precedes the second reference image; The method according to any one of embodiments 1 and 2.

[0213] 4. The first plurality of images includes four to eight images; the second plurality of images includes four to eight images; 4. The method of any one of embodiments 1 to 3.

[0214] 5. A method according to any one of embodiments 1 to 4, wherein forming a first plurality of transformed images and forming a second plurality of transformed images includes transforming the images to have the same viewpoint as the associated reference image based on comparing motion data associated with each image in a sequence of temporally consecutive images.

[0215] 6. The method of embodiment 5, wherein the motion data associated with each image in the sequence of time-sequential images is pan / tilt data including pan and tilt values of a pan / tilt camera associated with each image in the sequence of time-sequential images.

[0216] 7. Step h further includes determining a viewpoint difference between at least two images of the first plurality of images; The TNF of the second reference image is performed when it is determined that the viewpoint difference is equal to or less than a predetermined viewpoint difference threshold value; The TNF of the second reference image is not performed if it is determined that the viewpoint difference is greater than a predetermined viewpoint difference threshold. 7. The method of any one of embodiments 1 to 6.

[0217] 8. The viewpoint difference is based on the pan / tilt data temporally associated with each image. 8. The method of embodiment 7, wherein the pan / tilt data includes pan and tilt values of a pan / tilt camera associated with each image in a sequence of temporally consecutive images.

[0218] 9. Before step A, further comprising determining imaging conditions for the sequence of temporally consecutive images; 9. The method according to any one of embodiments 1 to 8, wherein only steps A to J are performed if it is determined that the imaging conditions satisfy the predetermined imaging condition requirements.

[0219] 10. Imaging conditions: a level of motion determined from pan / tilt data temporally associated with each image, the pan / tilt data including pan and tilt values of the pan / tilt camera associated with different points in time; a light level determined by a light sensor or by image analysis; 10. The method of embodiment 9, wherein the method is determined by at least one of the following:

[0220] 11. The imaging conditions are determined by a light sensor or by light levels determined by image analysis; the predetermined imaging condition requirement is a light level below a predetermined level; 11. The method of embodiment 10.

[0221] 12. Storing the first reference image in the memory of the pan / tilt camera after performing TNF on the first reference image; storing the second reference image on the memory of the pan / tilt camera after performing TNF on the second reference image; 12. The method of any one of embodiments 1 to 11, further comprising:

[0222] 13. Transmitting a first reference image from the pan / tilt camera to a remote device after performing TNF on the first reference image; transmitting a second reference image from the pan / tilt camera to a remote device after performing TNF on the second reference image; 13. The method of any one of embodiments 1 to 12, further comprising:

[0223] 14. A pan / tilt camera including an imaging unit and a computing unit, a. capturing a sequence of time-sequential images with an imaging unit; b. selecting, by a computing unit, a first reference image from the temporally consecutive images; c. selecting, by a computing unit, a first plurality of images from the temporally consecutive images to be used for temporal noise filtering of the first reference image; d. forming, with a computing unit, a first plurality of transformed images by transforming each of the first plurality of images to have the same viewpoint as the first reference image; e. performing, by a computing unit, temporal noise filtering (TNF) on a first reference image using the first plurality of transformed images; r. selecting, by a computing unit, a second, different, reference image from the temporally consecutive images; g. selecting, by a computing unit, a second plurality of images from a temporally consecutive image sequence to be used for temporal noise filtering of the second reference image, wherein at least one of the images of the second plurality of images is also included in the first plurality of images, and the second plurality of images includes the first reference image; h. determining, by a calculation unit, whether TNF of the second reference image should be performed; Once it is determined that TNF should be administered, i. forming, by a computing unit, a second plurality of transformed images; i1. for each image of a second plurality of images not included in the first plurality of images, transforming the image to have the same viewpoint as a second reference image, wherein the viewpoint transformation is determined and used to transform the first reference image to have the same viewpoint as the second reference image; i2. for each image of a second plurality of images that is also included in the first plurality of images, transforming a corresponding transformed image in the first plurality of transformed images to have the same viewpoint as the second reference image using the viewpoint transformation used to transform the first reference image; forming a second plurality of transformed images by j. performing, by a computing unit, temporal noise filtering on a second reference image using the second plurality of transformed images; forming a second plurality of transformed images by A pan / tilt camera configured for

[0224] 15. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by a device having processing capabilities, perform a method according to any one of embodiments 1 to 13.

[0225] Additionally, variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.

Claims

1. A method for noise reduction in an image captured by a wearable camera, comprising: A. Providing a sequence of time-sequential images; B. selecting a first reference image from the temporal sequence of images; C. selecting a first plurality of images from the temporally consecutive images to be used for temporal noise filtering of the first reference image; D. forming a first plurality of transformed images by transforming each of the first plurality of images to have the same viewpoint as the first reference image; E. using the first plurality of transformed images to perform temporal noise filtering (TNF) on the first reference image by averaging multiple images taken at different times; F. selecting a second, different, reference image from the temporal sequence of images; G. Selecting a second plurality of images from the temporally consecutive images to be used for temporal noise filtering of the second reference image, wherein at least one of the images of the second plurality of images is also included in the first plurality of images, and the second plurality of images includes the first reference image; H. Determining whether TNF of the second reference image should be performed; Including, If it is determined that TNF should be administered, I. forming a second plurality of transformed images, I1. For each image of the second plurality of images that is not included in the first plurality of images, transforming the image to have the same viewpoint as the second reference image, wherein a viewpoint transformation is determined and used to transform the first reference image to have the same viewpoint as the second reference image; I2. For each image in the second plurality of images that is also included in the first plurality of images, transforming the corresponding transformed image in the first plurality of transformed images to have the same viewpoint as the second reference image using the viewpoint transformation used to transform the first reference image; forming a second plurality of transformed images by J. performing temporal noise filtering on the second reference image using the second plurality of transformed images; The method further comprises:

2. 2. The method of claim 1, further comprising: deleting one of the temporally consecutive images from the wearable camera memory after transforming the one of the temporally consecutive images to have the same viewpoint as the first or second reference image.

3. each image of the first plurality of images temporally precedes the first reference image; each image of the second plurality of images temporally precedes the second reference image; 3. The method according to any one of claims 1 to 2.

4. the first plurality of images includes four to eight images; the second plurality of images includes four to eight images; 4. The method according to any one of claims 1 to 3.

5. 5. The method of claim 1, wherein forming the first plurality of transformed images and forming the second plurality of transformed images comprises transforming images to have the same viewpoint as the associated reference image based on comparing motion data associated with each image in the sequence of temporally consecutive images.

6. 6. The method of claim 5, wherein the motion data associated with each image in the sequence of temporally consecutive images is determined by at least one of a motion sensor, an accelerometer, and a gyroscope, or the motion data is determined based on image analysis of the images.

7. Step H further includes determining a viewpoint difference between at least two images of the first plurality of images; The TNF of the second reference image is performed when it is determined that the viewpoint difference is equal to or less than a predetermined viewpoint difference threshold value; The TNF of the second reference image is not performed if it is determined that the viewpoint difference is greater than a predetermined viewpoint difference threshold.

7. The method according to any one of claims 1 to 6.

8. The viewpoint difference is motion data temporally associated with each image, the motion data being determined by a motion sensor, an accelerometer, or a gyroscope; image data regarding how many pixels have changed between subsequent images of the first plurality of images; and The method of claim 7, based on at least one of the following:

9. before step A, further comprising determining an imaging condition for the sequence of temporally consecutive images; If it is determined that the imaging conditions satisfy the predetermined imaging condition requirements, only steps A to J are performed.

9. The method according to any one of claims 1 to 8.

10. The imaging conditions are: a level of motion as determined by a motion sensor, accelerometer, gyroscope, or positioning device; and a light level determined by a light sensor or by image analysis; The method of claim 9 , wherein the determination is made by at least one of the following:

11. the imaging conditions are determined by a light sensor or by light levels determined by image analysis; the predetermined imaging condition requirement is a light level below a predetermined level; 10. The method of claim 9.

12. storing the first reference image on a memory of the wearable camera after performing TNF on the first reference image; storing the second reference image on the memory of the wearable camera after performing TNF on the second reference image; 12. The method of any one of claims 1 to 11, further comprising:

13. transmitting the first reference image from the wearable camera to a remote device after performing TNF on the first reference image; transmitting the second reference image from the wearable camera to the remote device after performing TNF on the second reference image; 13. The method of any one of claims 1 to 12, further comprising:

14. A wearable camera including an imaging unit and a computing unit, A. capturing a sequence of time-sequential images with the imaging unit; B. selecting, by the computing unit, a first reference image from the temporally consecutive images; C. selecting, by the computing unit, a first plurality of images from the temporally consecutive images to be used for temporal noise filtering of the first reference image; D. forming, by the computing unit, a first plurality of transformed images by transforming each of the first plurality of images to have the same viewpoint as the first reference image; E. performing temporal noise filtering (TNF) on the first reference image using the first plurality of transformed images by the computing unit, the TNF averaging multiple images taken at different times; F. selecting, by said computing unit, a second, different, reference image from among said temporally consecutive images; G. selecting, by the computing unit, a second plurality of images from the temporally consecutive images to be used for temporal noise filtering of the second reference image, wherein at least one of the images of the second plurality of images is also included in the first plurality of images, and the second plurality of images includes the first reference image; H. determining, by the computing unit, whether TNF of the second reference image should be performed; If it is determined that TNF should be administered, I. forming, by said computing unit, a second plurality of transformed images, I1. For each image of the second plurality of images that is not included in the first plurality of images, transforming the image to have the same viewpoint as the second reference image, wherein a viewpoint transformation is determined and used to transform the first reference image to have the same viewpoint as the second reference image; I2. For each image in the second plurality of images that is also included in the first plurality of images, transforming the corresponding transformed image in the first plurality of transformed images to have the same viewpoint as the second reference image using the viewpoint transformation used to transform the first reference image; forming a second plurality of transformed images by J. performing, by the computing unit, temporal noise filtering on the second reference image using the second plurality of transformed images; Further configured for wearable cameras.

15. A non-transitory computer readable storage medium having stored thereon instructions that, when executed by a device having processing capabilities, perform the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Image processor and image processing method

    JP2006246309A

  • Image processor

    JP2009071679A

  • Image processor, image processing method, and program

    JP2014039169A