Method and system for removing banding artefacts while capturing a video

WO2026177258A1PCT designated stage Publication Date: 2026-08-27SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/006671
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-18
Filing Date
2025-05-16
Publication Date
2026-08-27

Smart Images

  • Figure KR2025006671_27082026_PF_FP_ABST
    Figure KR2025006671_27082026_PF_FP_ABST
Patent Text Reader

Abstract

Methods and system for removing a banding artefact appearing in a video. The system comprises a band-mask generating module configured for generating, for a plurality of buffer frames of the video, band-masks for indicating band-locations of one or more bands in the buffer frames. The system further comprises a frame selecting module configured for selecting, from the plurality of buffer frames, a first set of frames having first band-location different from a current band-location in a current frame. The band-location is determined based upon the indications in the corresponding generated band-masks. The system comprises a frame regeneration module configured for fusing, by applying machine learning models, current frame with at least one frame from the first set of frames to remove the banding artefacts from the video.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR REMOVING BANDING ARTEFACTS WHILE CAPTURING A VIDEO

[0001] The present disclosure generally relates to the field of videography and more particularly, the present disclosure relates to a method and a system for removing banding artefacts from a video.

[0002] With the advent of technology, videography has become increasingly prevalent, encompassing a wide range of users including photographers, astronomers, content creators, and even average social media users. The trend of videography is driven by a continuous improvement in the quality and availability of consumer cameras.

[0003] While capturing images or videos under artificial lighting, dark bands may appear in the captured frames in the form of dark regions. Especially under the presence of a flickering light source based on alternating current, vertical dark bands may appear to move across the frames being captured due to the variation in the light source. These moving bands appearing in the captured frames may be referred to as banding artefacts, band artefacts or simply, bands. Appearance of bands may be attributed to a mismatch of frequency between the flickering light source and a shutter speed of a capturing device. However, even in the presence of a non-flickering light source, bands may appear because of the reflection of the light from randomly moving objects in the scene being captured.

[0004] Appearance of bands remains a persistent challenge for image / frames capturing devices such as camera devices (e.g. smartphone cameras) in every video mode, including a portrait video mode, a hyperlapse video mode, a super slow-motion video mode, a slow-motion video mode, or a standard 30 frame per second (fps) video mode. The issue is more pronounced when the images are captured in an indoor environment under artificial lights with high shutter speeds.

[0005] Detection and removal of bands regions are essential for capturing good-quality frames for videos and images. There are methods known in the art for addressing the banding artefacts. However, such methods are dependent upon controlling shooting / capturing parameters such as angle of shooting, shutter speed, rate of capturing, mode of the capturing device, and the like. Besides, such methods may also require inputs regarding the ambient conditions and capturing parameters from the user capturing the video. Therefore, such methods become dependent on the skill of the user to manage and control the various capturing parameters. Moreover, additional skills are also required in calculating and inputting the correct parameters. Further, the ambient conditions may be dynamic, and a object and other associated entities in a scene may be moving randomly. Therefore, the skills of the user become even more significant in such scenarios and such methods are prone to human errors.

[0006] Another hardware-based solution is a phase-based capture technique, which utilizes sensors to detect a peak time of light oscillation and starts capturing the video at that time for different parts of the video. On the other hand, a gamma application is one of the software-based solutions, which uses adaptive gamma control to remove dark bands. The technique of gamma application involves edge map analysis and histogram analysis of an illumination map to apply different gamma curves in different parts of images, thereby increasing the illumination in band regions. Additionally, conventional solutions for real-time anti-banding fail to provide good quality output. Such solutions often require additional light detection sensors, which is not always feasible.

[0007] Some methods aim to improve the illumination in the dark regions but are unable to recover lost details. The effect of lost details becomes problematic when a single light source is present. Further, some methods suggest using post-processing tasks such as video HDR and video denoising. However, such solutions suffer from inconsistencies in illumination and details in temporal frames. Moreover, conventional solutions fail to work for all video recording modes (e.g., super slow, slow, hyperlapse, and normal videos) and become ineffective when the frame rate (frame-per-second (fps)) exceeds the ambient light frequency. Further, such solutions are inadequate for high frame rate and high shutter speed settings, where the banding effect is particularly severe. Finally, many conventional solutions rely on additional hardware sensors for monitoring light flickering or detecting light phase, adding complexity and cost to the system.

[0008] Therefore, there is a need for a system and method to overcome the above-mentioned problems and other associated problems in overcoming the banding artefacts.

[0009] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the disclosure. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended to determine the scope of the disclosure.

[0010] According to an embodiment of the present disclosure, a method for removing banding artefacts while capturing a video is disclosed. The method includes generating, for a plurality of buffer frames of the video being captured, band-masks for indicating band-locations of the bands appearing in the plurality of buffer frames. The method also includes determining a first set of frames from the plurality of buffer frames having band-locations different from a band-location in a current frame. The band-locations are determined based upon the indicated band-locations in the corresponding generated band-masks. The method further includes fusing, by applying machine learning models, the current frame with at least one frame from the first set of frames to remove the banding artefacts from the video.

[0011] According to an embodiment of the present disclosure, a system for removing banding artefact while capturing a video is disclosed. The system may include a band-mask generating module configured for generating, for a plurality of buffer frames of the video being captured, band-masks for indicating band-locations of the bands appearing in the plurality of buffer frames. The system may also include a frame selecting module configured for selecting, from the plurality of buffer frames, a first set of frames having band-locations different from a band-location in a current frame of the video. The band-locations are determined based upon the indicated band-locations in the corresponding generated band-masks. The system further includes a frame regeneration module configured for fusing, by applying machine learning models, the current frame with at least one frame from the first set of frames to remove the banding artefacts from the video.

[0012] To further clarify the advantages and features of the present disclosure, a more particular description of the disclosure will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered limiting its scope. The disclosure will be described and explained with additional specificity and detail with the accompanying drawings.

[0013] These and other features, aspects, and advantages of the present disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:

[0014] Figure 1A illustrates an environment comprising a system for removing banding artefacts while capturing a video under a light source using a capturing device, in accordance with an embodiment of the present disclosure;

[0015] Figure 1B illustrates exemplary implementations of the system for a video capturing device such as the capturing device for removing banding artefacts, in accordance with an embodiment of the present disclosure.

[0016] Figure 1C illustrates exemplary implementations of the system for a video capturing device such as the capturing device for removing banding artefacts, in accordance with an embodiment of the present disclosure.

[0017] Figure 2 illustrates the system for removing banding artefacts while capturing the video using the device, in accordance with an embodiment of the present disclosure;

[0018] Figure 3 illustrates a process flow of a band mask generating module of the system, in accordance with an embodiment of the present disclosure;

[0019] Figure 4A illustrates a process flow of a frame selecting module of the system, in accordance with an embodiment of the present disclosure;

[0020] Figure 4B illustrates a process flow of the frame selecting module of the system, in accordance with an embodiment of the present disclosure;

[0021] Figure 5A illustrate the working of a band removing module of the system, in accordance with an embodiment of the present disclosure;

[0022] Figure 5B illustrates the working of a spatial sync map block to determine common regions between a pair of current frame and another frame, in accordance with an embodiment of the present disclosure;

[0023] Figure 5C illustrates the working of a synergetic fusion block configured to create a regenerated current frame having the band removed from the current frame, in accordance with an embodiment of the present disclosure; and

[0024] Figure 6 is a flowchart illustrating a method for removing banding artefacts while capturing the video using the device, in accordance with an embodiment of the present disclosure.

[0025] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help to improve understanding of aspects of the present disclosure. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present disclosure so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.

[0026] For the purpose of promoting an understanding of the principles of the disclosure, reference will now be made to the various embodiments, and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the disclosure as illustrated therein being contemplated as would normally occur to one skilled in the art to which the disclosure relates.

[0027] The term "some" or "one or more" as used herein is defined as "one", "more than one", or "all." Accordingly, the terms "more than one," "one or more" or "all" would all fall under the definition of "some" or "one or more". The term "an embodiment", "another embodiment", "some embodiments", or "in one or more embodiments" may refer to one embodiment or several embodiments, or all embodiments. Accordingly, the term "some embodiments" is defined as meaning "one embodiment, or more than one embodiment, or all embodiments."

[0028] The terminology and structure employed herein are for describing, teaching, and illuminating some embodiments and their specific features and elements and do not limit, restrict, or reduce the spirit and scope of the claims or their equivalents. The phrase "exemplary" may refer to an example.

[0029] More specifically, any terms used herein such as but not limited to "includes," "comprises," "has," "consists," "have" and grammatical variants thereof do not specify an exact limitation or restriction and certainly do not exclude the possible addition of one or more features or elements, unless otherwise stated, and must not be taken to exclude the possible removal of one or more of the listed features and elements, unless otherwise stated with the limiting language "must comprise" or "needs to include".

[0030] Whether or not a certain feature or element was limited to being used only once, either way, it may still be referred to as "one or more features", "one or more elements", "at least one feature", or "at least one element." Furthermore, the use of the terms "one or more" or "at least one" feature or element does not preclude there being none of that feature or element unless otherwise specified by limiting language such as "there needs to be one or more " or "one or more element is required."

[0031] Unless otherwise defined, all terms, and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.

[0032] Figure 1A illustrates an environment of a space 100 and a system 110 for removing banding artefacts while capturing a video 190 in accordance with an embodiment of the present disclosure. In the space 100, a capturing device 150 may capture the video 190 for a real-world scene 100S under a light source 100L. The light source 100L may be an artificial light source and may be attributed for appearance of a banding artefact (e.g. band 190B) in the video 190. The system 110 may be communicably coupled with the capturing device 150 (or the device 150) for removing banding artefacts which may appear in the video 190 to create a regenerated video 190G.

[0033] The real-world scene 100S includes entities such as one or more objects 102 including a first object 102A, a second object 102B and a third object 102C. The objects 102 may include any object in the real-world scene 100S of the space 100. For example, the objects 102 may include person, window, ball, umbrella, etc. The objects 102 may be dynamically and randomly moving at a variable speed within the scene 100S.

[0034] In an embodiment, the device 150 may include, but not limited to, a smartphone, a camera, or any other electronic device having one or more cameras compatible with capturing images or recording video, etc. of the scene 100S without departing from the scope of the present disclosure.

[0035] In an embodiment, the device 150 may include multiple layers, for example, an application layer, a file system layer, etc. The application layer may include a video player application, a gallery application, or a camera application, without departing from the scope of the present disclosure. Further, the file system layer may include a file reader, a coder-decoder, a frame data, and a file writer. The file reader may be configured to read a video recorded by the application layer.

[0036] Figure 1B illustrates exemplary implementations 100B of the system 110 for a video capturing device such as the device 150 for removing banding artefacts, in accordance with an embodiment of the present disclosure. Specifically, the exemplary implementation 100B shows the system 110 for YUV input. YUV is a color model, where Y represents the brightness or 'luma' value, and UV represents the color or 'chroma' value. YUV input may include brightness(luma) value and color(chroma) value for each pixel. The frames being processed by the system 110 may require preprocessing 198-1 such as bad pixel correction, lens shading correction, demosaicing, white balance correction and the like, but is not limited to. Further, the system 110 is located before other modules such as denoising modules, High Dynamic Range (HDR) modules and the like, but is not limited to.

[0037] Figure 1C illustrates exemplary implementations 100C of the system 110 for a video capturing device such as the device 150 for removing banding artefacts, in accordance with an embodiment of the present disclosure. Specifically, the exemplary implementation 100C shows the system 110 for Bayer / Tetra / Nona input. A Bayer filter is a color filter array for arranging RGB (Red / Green / Blue) color filters on grids of photosensors. A Tetra filter and a Nona filter are advanced color filter array (CFA) patterns that arrange four and nine sub-pixels, respectively, within a single pixel unit to improve light sensitivity, color accuracy, and noise performance in image sensors. Bayer / Tetra / Nona input may include luminance value for each pixel, corresponding to a filter applied to the pixel. The frames being processed by the system 110 may require preprocessing 198-2 such as bad pixel correction, lens shading correction and the like, but is not limited to. Further, the system 110 is located before other modules such as denoising modules, High Dynamic Range (HDR) modules and the like, but is not limited to.

[0038] Figure 2 illustrates the system 110 for removing banding artefacts while capturing the video 190 using the device 150, in accordance with an embodiment of the present disclosure.

[0039] In an embodiment, the system 110 includes a plurality of modules 200 including a band mask generating module 210, a frame selecting module 220, and a frame regeneration module 230. The band mask generating module 210 is configured for generating band-masks for a plurality of buffer frames of the video 190 being captured. The band-masks are generated for indicating band-locations of the bands appearing in the plurality of buffer frames. The frame selecting module 220 is configured for selecting a first set of frames having band-location different from a band-location in a current frame from the plurality of buffer frames. The band may be appearing in the current frame of the video 190. For example, the band appeared in the current frame may be called a current band, but is not limited to. The band-location in the buffer frames of the video 190 is determined based upon the indicated band-locations in the corresponding generated band-masks. The frame regeneration module 230 is configured for fusing, by applying machine learning models 230AI, the current frame with at least one frame from the first set of frames to remove the banding artefacts from the video 190.

[0040] At least one of the plurality of modules 200 may be implemented through an Artificial Intelligence (AI) or a machine learning model. For example, in an embodiment, the frame regeneration module 230 is configured to apply the AI or machine learning models 230AI for removing the bands. A function associated with the AI model 230AI may be performed through a memory 208 (including a non-volatile memory and / or a volatile memory) and a processor 204. The memory 208 may be configured to store a program or at least one instruction for performing operations according to one or more embodiments of the disclosure. The processor 204 may be configured to execute the program or the at least one instruction.

[0041] In an embodiment, the system 110 includes a processor 204, a memory 208, a transceiver 226, and an I / O interface 228. The processor 204 may be disposed in communication with a communication network via a network interface. The processor 204 may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or the AI model stored in the non-volatile memory and the volatile memory. The predefined operating rule or the AI model is provided through training or learning. Here, being provided through learning means that, by applying a learning technique to a plurality of learning data, a predefined operating rule or AI model of a desired characteristic is made. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system.

[0042] The AI model may consist of a plurality of neural network layers. Each layer has a plurality of weight values and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, a Convolutional Neural Network (CNN), a Deep Neural Network (DNN), a Recurrent Neural Network (RNN), a Restricted Boltzmann Machine (RBM), a Deep Belief Network (DBN), a Bidirectional Recurrent Deep Neural Network (BRDNN), a Generative Adversarial Networks (GAN), and deep Q-networks.

[0043] The learning technique is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0044] According to the disclosure, a method for generating a plurality of instructions may use an AI model to recommend / execute the plurality of instructions by using sensor data. The processor may perform a pre-processing operation on the data to convert into a form appropriate for use as an input for the AI model. The AI model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or AI model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training technique. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.

[0045] Reasoning prediction is a technique of logical reasoning and predicting by determining information and includes, e.g., knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.

[0046] In an embodiment, the network interface may be the I / O interface 228. In an embodiment, the network interface may connect to the network to enable the connection of the system 110 with the device 150. The network interface may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 702.11a / b / g / n / x, etc. The communication network may include, without limitation, a direct interconnection, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, etc. Using the network interface and the communication network, the system 110 may communicate with other devices. The network interface may employ connection protocols including, but not limited to, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 702.11a / b / g / n / x, etc.

[0047] In some embodiments, the memory 208 may be communicatively coupled to the processor 204. The memory 208 may be configured to store data, and instructions executable by the processor 204. In one embodiment, the memory 208 may be provided within the device 150. In an embodiment, the memory 208 may be provided within the system 110 being remote from the device 150. In an embodiment, the memory 208 may communicate with the processor 204 via a bus within the system 110. In an embodiment, the memory 208 may be located remotely from the processor 204 and may be in communication with the processor 204 via a network. The memory 208 may include, but is not limited to, a non-transitory computer-readable storage media, such as various types of volatile and non-volatile storage media including, but not limited to, random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media and the like.

[0048] In one example, the memory 208 may include a cache or random-access memory for the processor 204. In alternative examples, the memory 208 is separate from the processor 204, such as a cache memory of a processor, the system memory, or other memory. The memory 208 may be an external storage device or database for storing data. The memory 208 may be operable to store instructions executable by the processor 204. The functions, acts, or tasks illustrated in the figures or described may be performed by the programmed processor 204 for executing the instructions stored in the memory 208. The functions, acts, or tasks are independent of the particular type of instruction set, storage media, processor, or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code, and the like, operating alone or in combination. Likewise, processing strategies may include multiprocessing, multitasking, parallel processing, and the like.

[0049] In some embodiments, the plurality of modules 200 may be included within the memory 208. The plurality of modules 200 may include a set of instructions that may be executed by the processor 204 to cause the system 110 to perform any one or more of the methods / processes disclosed herein. The plurality of modules 200 may be configured to perform the steps of the present disclosure using the data stored in the database within the memory 208. For instance, the plurality of modules 200 may be configured to perform the steps disclosed with reference to Figure 6.

[0050] In an embodiment, each of the plurality of modules 200 may be a hardware unit that may be outside the memory 208. Further, the memory 208 may include an operating system for performing one or more tasks of the system 110, as performed by a generic operating system. Each of the modules 200 may be in communication with one another and the processor 204. The working and functioning of the plurality of modules 200 of the system 110 have been described in detail with reference to the following Figures.

[0051] Figure 3 illustrates a process flow 300 of the band mask generating module 210 of the system 110, in accordance with an embodiment of the present disclosure. The band mask generating module 210 is configured for generating band-masks 310 for a plurality of buffer frames 390 of the video 190 being captured. As an example, Figure 3 shows a band-mask 310 being generated for a current frame 390N having the band 392B from the buffer frames 390. At least one row of elements of the band-mask 310 may indicat band-locations 310B of one or more bands corresponding the band 392B appearing in the current frame 390N. And at least one row of elements of the band-mask 310 may indicate a non-band region 310NB corresponding to the region in the current frame 390N excluding the band 392B. The generated band mask 310 may be represented as a binary image wherein a value of "0" corresponds to a band region and a value of "1" corresponds to a non-band region. The black region (or band-location 310B) in the generated band mask 310 may depict black band region in input image. The white region(or non-band region 310NB) in the generated band mask 310 may depict non band region in input image.

[0052] The band mask generating module 210 is configured to generate an illumination map indicating an illumination level for each pixel in the buffer frames 390. The generated illumination map may be mean normalized based on a mean and a variance of the generated illumination map. As an example, a filter such as a Savitzky-Golay (SavGol) filter may be applied for polynomial smoothening of the generated illumination map to remove irregularities. Further, the band mask generating module 210 is configured for generating correlation maps corresponding to generated illumination maps by fitting a sinusoidal curve on each row in the generated illumination maps. In an embodiment, the correlation maps are generated based on a correlation between the sinusoidal filter and the generated illumination map.The band mask generating module 210 is configured for applying a predefined threshold to the generated correlation maps to generate corresponding band-masks 310, wherein values greater than the predefined threshold may be marked as the band region 310B by "0", and the values less than the predefined threshold may be marked as "1" to indicate the non-band region 310NB.

[0053] In an embodiment, the band-mask generating module 210 is configured for applying the predefined threshold based on at least one of a shutter speed of the capturing device 150 and a frame-rate of capturing associated with the video 190. Specifically, the predefined threshold value may be pre-determined based on the shutter speed and the frame rate of capturing such as the FPS of the device 150.

[0054] Figure 4A illustrates a process flow 400A of the frame selecting module 220 of the system 110, in accordance with an embodiment of the present disclosure. The frame selecting module 220 is configured for determining the band-locations in the plurality of buffer frames by correlating a rate of band movement across the plurality of frames and the frame-rate of capturing the video (e.g., 190). The frame selecting module 220 is configured for identifying frames for which a band-region corresponding to the band-location in the current frame is not occluded

[0055] In an embodiment, the frame selecting module 220 is configured for generating, based on pixel values, band-lines corresponding to each of the bands (e.g., band-locations 310B) in the band-masks 310 by using a fitting-algorithm. The frame selecting module 220 is configured for correlating the frame-rate and locations of the band-lines to determine a velocity of the band movement across the plurality of buffer frames (e.g., 390).

[0056] In an embodiment, the frame selecting module 220 includes a band distortion analyzer 410 configured for generating band lines such as Mean Squared Error (MSE) fitted lines for each band mask corresponding to the buffer frames 390 such as the band mask 310. The MSE fitted lines are generated based on the pixel values in the band mask 310 by applying a fitting algorithm such as a mean-square-fitting algorithm.

[0057] Each band mask 310 may be transformed into an MSE fitted line. The slope and interception of the MSE line represents the location and slope of the band 310B in corresponding band masks 310. The band distortion analyzer 410 is configured to use an external Application Programming Interface (API) to generate the band lines such as the MSE fitted lines taking into account the random nature of the bands. For generation of the MSE fitted lines, each "0" value in band mask 310 is treated as a data point. A line is fitted on the data points using a known fitting algorithm such as the mean square fitting algorithm. Thus, in step 420, a slope and location intercept of the bands 310B in the band mask 310 may be determined as the respective slope and location of the MSE fitted line.

[0058] Figure 4B illustrates a process flow 400B of the frame selecting module 220 of the system 110, in accordance with an embodiment of the present disclosure.

[0059] In an embodiment, the frame selecting module 220 is configured for determining a correlation between a measured movement and corresponding associated time-stamps of the capturing device. The frame selecting module 220 is configured for updating the band-locations based on the determined correlation.

[0060] Accordingly, in an embodiment, the frame selecting module 220 includes a defect-less frame detector 430 configured for determining a count of a previous frame from the buffer frames 390 based on the MSE fitted lines.

[0061] The defect-less frame detector 430 is configured for correlating data associated with the device 150, and the location and slope of the generated MSE fitted lines to determine a velocity associated with the movement of the band 310B across the plurality of buffer frames 390. The data associated with the device 150 may include sensor data (such as frame rate of capturing (fps), a shutter speed, a measured movement of the device 150 and associated timestamps such as a camera movement metadata) of the device 150 associated with the video. The plurality of buffer frames 390 may include frames before (N-pthframes) and after (N+1th, N+2th, N+pthframes) the current frame (the Nthframe, the frame 390N).

[0062] Due to the movement of the device 150 and change in capturing parameters, the location of the band 310B in the buffer frames 390 may change at a non-uniform rate. The frame selecting module 220 uses the sensor data (meta data) of the device 150 to account for such movement while determining the location of the band 310B in each of the buffer frames 390. In an embodiment, a minimum of two frames from the buffer frames 390 is used to determine the velocity of band movement using positions of bands in two consecutive frames and the fps of video 190.

[0063] Accordingly, the frame selecting module 220 selects the first set of frames from the buffer frames 390 having location of the band 310B different from location of the band 310B in the current frame 390N. The frame selecting module 220 is configured for determining a relative index (the count of the previous frame) of a frame with respect to the current frame 390N where the band 310B is not overlapping with the band 310B in the current frame 390N.

[0064] The defect-less frame detector 430 determines velocity of the band 310B for each of the buffer frames 390 using the intercepts of corresponding MSE fitted lines. Based on the determined velocity and the fps value, the defect-less frame detector 430 calculates the position of the band 310B for all previous frames (N-pthto N-1thframes) before the current frame 390N. Further, using the movement metadata (degree of rotation and motion) and fps of the device 150, the defect-less frame detector 430 updates the location of band 310B in the previous frames by incorporating the negation of motion relative to the current frame 390N. The defect-less frame detector 430 identifies the first set of frames from the previous frames for which the band position is different from the band position in the current frame 390N.

[0065] In an embodiment, the frame selecting module 220 is configured for determining a width of the generated band-lines as the widths of the bands 310B. The frame selecting module 220 is configured for identifying frames for which the determined band widths are not overlapping for the entire determined width of the band in the current frame.

[0066] In an embodiment, the defect-less frame detector 430 is configured for determining a width of the generated MSE lines as the width of the bands 310B and selecting the band masks 310 for which the determined band positions are not overlapping for the entire determined width of the band. Accordingly, the defect-less frame detector 430 may use the width of the band 310B as a measure for identifying the first set of frames. Specifically, the band position in the previous frames should be at least one band-width distance apart from the band position in the current frame 390N. Further, the index or the count of a nearest to the current frame 390N from the generated first set of frames is identified by defect-less frame detector 430.

[0067] In an embodiment, the frame selecting module 220 is configured for determining frame limits of the current frame (390N) corresponding to the scene 100S being captured. The frame selecting module 220 is configured for identifying frames for which the band-locations are within the determined frame limits.

[0068] In an embodiment, the defect-less frame detector 430 is configured for determining frame limits of the current frame 390N and identifying, from the generated band-masks 310, the band-masks for which the band position is within the determined frame limits.

[0069] Figure 5A illustrate the working of the frame regeneration module 230 of the system 110, in accordance with an embodiment of the present disclosure. The frame regeneration module 230 is configured for removing the band 392B from the current frame 390 to create a regenerated current frame 390G using the generated first set of frames and corresponding band-masks 310. The frame regeneration module 230 includes a spatial sync map (SSM) block 532, a kinematic map (KM) block 534, and a synergetic fusion block 536.

[0070] Figure 5B illustrates the working of the spatial sync map block 532 to determine common regions between a pair of current frame 390N and a frame 510. The frame 510 may be a frame from the first set of frames and temporal frames in the plurality of buffer frames 390 corresponding to the first set of frames. The spatial sync map block 532 is configured for generating spatial-sync maps 520 between each of the first set of frames (such as the frame 510) and the current frame 390N. The spatial sync map block 532 determines common regions between the current frame 390N, and the first set of frames and the temporal frames. Further, a homographic map is generated based on each of the generated spatial-sync map 520. The common region (e.g., the remaining region in the spatial-sync map 520 excluding the uncommon region 520U) may be cropped and transposed to the position of the current frame 390N and the uncommon region 520U may be assigned "0" values.

[0071] Figure 5B further illustrates the working of the kinematic map block 534 to generate a kinematic map 530 depicting velocity of the pixel movement in the current frame 390N with respect to the buffer frames 390. The kinematic map block 534 is configured for generating, for each frame in the first set of frames, the kinetic map 530 for indicating the movement of pixels by generating an occlusion map of the first set of frames and the one or more temporal frames. The kinetic map 530 may include a kinetic map for horizontal velocity(e.g., horizontal kinetic map 530H in Fig. 5C) and a kinetic map for vertical velocity(e.g., vertical kinetic map 530V in Fig. 5C). Further, the kinematic map block 534 is configured for generating an optical-flow map for the pixels of the first set of frames excluding pixels in the generated occlusion map. The generated kinetic map 530 indicates a movement of each pixel in the current frame 390N and the buffer frames 390.

[0072] Figure 5C illustrates the working of the synergetic fusion block 536 configured to create a regenerated current frame 390G having the band removed from the current frame 390. The synergetic fusion block 536 is configured for generating, based on the kinetic maps 530V and 530H, weight maps for the generated homographic maps. The synergetic fusion block 536 regenerates the current frame 390N to create the regenerated current frame 390G based on the generated weight maps. In an embodiment, the frame regeneration module 230 is configured for using machine learning models for regenerating the current frame 390N based on the generated weight maps. In an embodiment, one previous (N-1thframe) and one next (N+1th) frames may be sufficient to determine the movement of the pixels of the band 392B in the current frame 390N.

[0073] In an embodiment, the frame regeneration module 230 is configured for correcting the illumination of the regenerated current frame 390G based on the band mask 310 (e.g., 310b) corresponding to the current frame 390N.

[0074] Figure 6 is a flowchart illustrating a method 600 for removing banding artefacts while capturing the video 190 using the device 150, in accordance with an embodiment of the present disclosure. Referring to Figures 1-5 together, the method 600 may be performed by the device 150 such as a camera device, e.g., a camcorder, a mobile device, a tab with image capturing capabilities, and the like, based on instructions retrieved from non-transitory computer-readable media. A computer-readable media may include machine-executable or computer-executable instructions to perform all or portions of the described method. The computer-readable media may be, for example, digital memories, magnetic storage media, such as magnetic disks and magnetic tapes, hard drives, or optically readable data storage media. The method 600 includes a series of operations shown at step 602 through step 606 of Figure 6. The method 600 may be performed by the system 110 in conjunction with one or more modules 200, the details of which are explained in conjunction with Figures 1-5, and the same are not repeated here for the sake of brevity. The method 600 begins at step 602.

[0075] At step 602, the method 600 includes generating band-masks for a plurality of buffer frames 390 of the video 190 being captured. The band-masks are generated for indicating band-locations of the bands appearing in the plurality of buffer frames 390.

[0076] In an embodiment, the method 600 at step 602 further includes generating an illumination map indicating an illumination level for each pixel in the buffer frames 390. Further, the method 600 at step 602 includes generating correlation maps corresponding to generated illumination maps by fitting a sinusoidal curve on each row in the generated illumination maps. In an embodiment, the correlation maps are generated based on a correlation between the sinusoidal filter and the generated illumination map. The method 600 at step 602 includes applying a predefined threshold to the generated correlation maps to generate corresponding band-masks 310, wherein values greater than the predefined threshold may be marked as the band region 310B by "0", and the values less than the predefined threshold may be marked as "1" to indicate the non-band region 310NB. The predefined threshold value may be pre-determined based on the shutter speed and the frame rate of capturing such as the FPS of the device 150.

[0077] At step 604, the method 600 includes selecting a first set of frames having band-location different from a band-location in the current frame 390N from the plurality of buffer frames 390. The band-location in the buffer frames 390 of the video 190 is determined based upon the indicated band-locations in the corresponding generated band-masks 310.

[0078] In an embodiment, the method 600 at step 604 further includes generating, based on pixel values, band-lines corresponding to each of the bands 310B in the band-masks 310 by using a fitting-algorithm. The method 600 at step 604 further includes correlating the frame-rate and locations of the band-lines to determine a velocity of the band movement across the plurality of buffer frames 390.

[0079] In an embodiment, the method 600 at step 604 further includes generating band lines such as MSE fitted lines for each band mask 310 corresponding to the buffer frames 390. The MSE fitted lines are generated based on the pixel values in the band mask 310 by applying a fitting algorithm such as a mean-square-fitting algorithm. Each band mask 310 may be transformed into an MSE fitted line. The slope and interception of the MSE line represents the location and slope of the band 310B in corresponding band masks 310. The method 600 at step 604 further includes using an external Application Programming Interface (API) to generate the band lines such as the MSE fitted lines taking into account the random nature of the bands. For generation of the MSE fitted lines, each "0" value in band mask 310 is treated as a data point. A line is fitted on the data points using the known mean square fitting algorithm. Thus, a slope and location intercept of the bands 310B in the band mask 310 may be determined as the respective slope and location of the MSE fitted line.

[0080] In an embodiment, the method 600 at step 604 further includes determining a correlation between a measured movement and corresponding associated time-stamps of the capturing device such as the device 150. The method 600 at step 604 includes updating the band-locations based on the determined correlation.

[0081] In an embodiment, the method 600 at step 604 further includes correlating data associated with the device 150 such as measured movement of the device 150 and associated time-stamps corresponding to the measured movement. The data associated with the device 150 may include sensor data (e.g., frame rate of capturing (fps), shutter speed, a measured movement of the device 150 and associated timestamps such as a camera movement metadata) of the device 150 associated with the video.

[0082] In an embodiment, the method 600 at step 604 further includes determining a width of the generated band-lines as the widths of the bands 310B. The method 600 at step 604 includes identifying frames for which the determined band widths are not overlapping for the entire determined width of the band in the current frame.

[0083] In an embodiment, the method 600 at step 604 further includes determining a width of the generated MSE lines as the width of the bands 310B and selecting the band masks 310 for which the determined band positions are not overlapping for the entire determined width of the band. Accordingly, the width of the band 310B may be used as a measure for identifying the first set of frames. Specifically, the band position in the previous frames should be at least one band-width distance apart from the band position in the current frame 390N.

[0084] Further, the method 600 at step 604 includes determining frame limits of the current frame 390N corresponding to the scene 100S being captured. The method 600 at step 604 includes identifying frames for which the band-locations are within the determined frame limits.

[0085] In an embodiment, the method 600 at step 604 includes determining frame limits of the current frame 390N and identifying, from the generated band-masks 310, the band-masks for which the band position is within the determined frame limits.

[0086] At step 606, the method 600 includes fusing, by applying machine learning models such as the machine learning model 230AI, the current frame 390N with at least one frame from the first set of frames to remove the banding artefacts from the video 190. The method 600 includes using the first set of frames and corresponding generated band-masks 310 of the first set of frames. The method 600 at step 606 may include using AI or a machine learning model such as the AI model 230AI for removing the bands.

[0087] In an embodiment the method 600 at step 606 includes determining common regions between a pair of current frame 390N and the frame 510 such as a frame from the first set of frames and temporal frames in the plurality of buffer frames 390. Further, the method 600 at step 606 includes generating spatial-sync maps 520 between each of the first set of frames (such as the frame 510) and the current frame 390. Furthermore, the method 600 at step 606 includes generating a homographic map based on each of the generated spatial-sync map 520. The common region may be cropped and transposed to the position of the current frame 390 and the uncommon region may be assigned "0" values.

[0088] In an embodiment the method 600 at step 606 includes generating a kinematic map 530 depicting velocity of the pixel movement in the current frame 390N with respect to the buffer frames 390. Further, the method 600 at step 606 includes generating, for each frame in the first set of frames, the kinetic map 530 for indicating the movement of pixels by generating an occlusion map of the first set of frames and the one or more temporal frames. Further, the method 600 includes generating an optical-flow map for the pixels of the first set of frames excluding pixels in the generated occlusion map. The generated kinetic map 530 indicates a movement of each pixel in the current frame 390N and the buffer frames 390.

[0089] In an embodiment the method 600 at step 606 includes correcting the illumination of the regenerated current frame 390G based on the band mask 310 corresponding to the current frame 390.

[0090] The system and method of the invention take advantage of the previous captured frames and merge information from the previous frames to the missing pixels in the current frame due to the appearance of a band. Accordingly, the invention provides a solution to remove appearance of bands in any video format and under any light condition without the need of manual intervention or without dependence upon human skills. The invention removes the appearance of bands irrespective of the type of light (flickering, natural, artificial, etc.) and shutter speed of the camera. Therefore, the invention is able to provide real-time anti-banding of the video being recorded without the need of any post-production techniques. Since the appearance of the band and movement of the band is random, the invention is advantageous as the solution provided does not depend upon input parameters and fills in missing information in the current using previous captured frames of the same scene being captured. Further, the invention enables performance of subsequent models / engines such as denoising and HDR solution as the invention provides anti-banded video / frames for further processing. Hence, using the method and placing the system of the invention before such models / engines will further improve the video pipeline and increase the efficiency of the overall system.

[0091] While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein.

[0092] The drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.

[0093] Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.

[0094] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.

Claims

1.A method (600) for removing banding artefacts while capturing a video (190), the method (600) comprising:generating (602), for a plurality of buffer frames of the video (190), band-masks (310) for indicating band-locations of one or more bands (310B) appearing in the plurality of buffer frames (390);determining (604), from the plurality of buffer frames (390), a first set of frames having first band-locations different from a current band-location in a current frame (390N) among the plurality of buffer frames, the first band-locations and the current band-location being determined based upon the indicated band-locations in the corresponding generated band-masks (310); andfusing, by applying machine learning models (230AI), the current frame (390N) with at least one frame from the first set of frames to remove the banding artefacts from the video (190).2.The method (600) as claimed in claim 1, wherein generating the band-masks (310) comprises:generating, for each of the plurality of the buffer frames (390), an illumination map indicating an illumination level for each pixel;generating correlation maps corresponding to generated illumination maps by fitting a sinusoidal curve on each row in the generated illumination maps; andapplying a predefined threshold to the generated correlation maps to generate corresponding band-masks (310), the predefined threshold being based on at least one of a shutter speed of a capturing device capturing the video (190) and a frame-rate of capturing associated with the video (190).3.The method (600) as claimed in any one of claims 1 and 2, wherein determining the first set of frames comprises:determining band-locations of the one or more bands in the plurality of buffer frames by correlating a rate of band movement across the plurality of buffer frames and the frame-rate of capturing the video (190); andidentifying frames for which a band-region corresponding to a current band of the one or more bands in the current frame is not occluded.4.The method (600) as claimed in claim 3, wherein determining the band-locations comprises:generating, based on pixel values, band-lines corresponding to each of the one or more bands (310B) in the band-masks (310) using a fitting-algorithm; andcorrelating the frame-rate and locations of the band-lines to determine the rate of band movement across the plurality of buffer frames (390).5.The method (600) as claimed in any one of claims 3 and 4, wherein determining the band-locations comprises:determining a correlation between a measured movement and corresponding associated time-stamps of the capturing device ; andupdating the band-locations based on the determined correlation.6.The method (600) as claimed in any one of claims 4 and 5, wherein determining the first set of frames comprises:determining a width of the generated band-lines as the widths of the one or more bands (310B); andidentifying frames for which the determined band widths are not overlapping for the entire determined width of the current band in the current frame.7.The method (600) as claimed in any one of claims 4 to 6, wherein determining the first set of frames comprises:determining frame limits, corresponding to a scene being captured, of the current frame (390N); andidentifying frames for which the band-locations are within the determined frame limits.8.A system (110) for removing banding artefact while capturing a video (190), the system (110) comprising:a band-mask generating module (210) configured for generating, for a plurality of buffer frames (390) of the video (190), band-masks (310) for indicating band-locations of one or more bands (310B) appearing in the plurality of buffer frames (390);a frame selecting module (220) configured for determining (604), from the plurality of buffer frames (390), a first set of frames having first band-locations different from a current band-location in a current frame (390N) among the plurality of buffer frames, the first band-locations and the current band-location being determined based upon the indicated band-locations in the corresponding generated band-masks (310); anda frame regeneration module (230) configured for fusing, by applying machine learning models (230AI), the current frame (390N) with at least one frame from the first set of frames to remove the banding artefacts from the video (190).9.The system (110) as claimed in claim 8, wherein the band-mask generating module (210) is configured for:generating, for each of the plurality of the buffer frames (390), an illumination map indicating an illumination level for each pixel;generating correlation maps corresponding to generated illumination maps by fitting a sinusoidal curve on each row in the generated illumination maps; andapplying a predefined threshold to the generated correlation maps to generate corresponding band-masks (310), the predefined threshold being based on at least one of a shutter speed of a capturing device capturing the video (190) and a frame-rate of capturing associated with the video (190).10.The system (110) as claimed in any one of claims 8 and 9, wherein the frame selecting module is configured for:determining the band-locations of the one or more bands in the plurality of buffer frames by correlating a rate of band movement across the plurality of buffer frames and the frame-rate of capturing the video (190); andidentifying frames for which a band-region corresponding to a current band of the one or more bands in the current frame is not occluded.11.The system (110) as claimed in claim 10, wherein the frame selecting module (220) is configured for:generating, based on pixel values, band-lines corresponding to each of the one or more bands (310B) in the band-masks (310) by using a fitting-algorithm; andcorrelating the frame-rate and locations of the band-lines to determine the rate of band movement across the plurality of buffer frames (390).12.The system (110) as claimed in any one of claims 10 and 11, wherein the frame selecting module (220) is configured for:determining a correlation between a measured movement and corresponding associated time-stamps of the capturing device; andupdating the band-locations based on the determined correlation.13.The system (110) as claimed in any one of claims 11 and 12, wherein the frame selecting module (220) is configured for:determining a width of the generated band-lines as the widths of the one or more bands (310B); andidentifying frames for which the determined band widths are not overlapping for the entire determined width of the current band in the current frame.14.The system (110) as claimed in any one of claims 11 to 13, wherein the frame selecting module (220) is configured for:determining frame limits, corresponding to a scene being captured, of the current frame (390N); andidentifying frames for which the band-locations are within the determined frame limits.15.A non-transitory machine readable medium containing instructions that when executed cause at least one processor of an electronic device to:generate (602), for a plurality of buffer frames of the video (190), band-masks (310) for indicating band-locations of one or more bands (310B) appearing in the plurality of buffer frames (390);determine (604), from the plurality of buffer frames (390), a first set of frames having first band-locations different from a current band-location in a current frame (390N) among the plurality of buffer frames, the first band-locations and the current band-location being determined based upon the indicated band-locations in the corresponding generated band-masks (310); andfuse, by applying machine learning models (230AI), the current frame (390N) with at least one frame from the first set of frames to remove the banding artefacts from the video (190).