Electronic device and method for generating a panning image

The system uses neural networks to analyze single frames for panning photography, estimating speed and direction to automatically generate high-quality panning images, addressing manual setup challenges and frame-based inefficiencies.

WO2026058203A1PCT designated stage Publication Date: 2026-03-19SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing panning photography techniques require manual input and expertise to set camera parameters, are time-intensive, and struggle with accurately estimating object speed and direction from single frames, leading to inefficient and often inaccurate panning images.

Method used

A system and method using a neural network to analyze a single image frame, identify a salient object, generate a gradient map, and estimate speed and direction parameters to apply blur to the background, creating a panning image.

Benefits of technology

Enables efficient generation of panning images with accurate foreground motion and background blur, eliminating the need for manual parameter setting and improving image quality from single frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025059174_19032026_PF_FP_ABST
    Figure IB2025059174_19032026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method for generating a panning image. The method includes receiving an image frame comprising one or more objects and a background portion, identifying, in the received image frame, a first area corresponding to a salient object amongst the one or more objects, generating a histogram map based on the received image frame and the first area, the histogram map indicates a distribution and orientation of edge intensities corresponding to the salient object in the received image frame, estimating, using a set of models, a speed parameter and a direction parameter corresponding to the salient object based on the received image frame, the first area, and the histogram map, and generating the panning image by applying a blur to a second area corresponding to the background portion of the received image frame based on the estimated speed parameter and the estimated direction parameter.
Need to check novelty before this filing date? Find Prior Art

Description

DescriptionTitle of Invention: ELECTRONIC DEVICE AND METHOD FOR GENERATING A PANNING IMAGETechnical Field

[0001] The present disclosure relates to applications related to images, and more particularly relates to a system and a method for generating a panning image.Background Art

[0002] Photography has always been a preferred means to capture moments, landscapes, wildlife, etc., by a user. There have been many technological developments in a field of photography, for example, panning photography, which enables the user to capture a panning image, i.e. , a moving object in a static image. Panning photography, a widely embraced method, is employed to seize moments of motion of the object within static images as traditional still photographs do not convey the dynamic essence of the movement of the object adequately. Panning photography conveys a real sense of motion to objects in action photos. The uniqueness of panning photography lies in relatively sharp foreground objects combined with a motion-blurred background, resulting in capturing the objects more realistically.

[0003] Panning photography facilitates the production of static, yet dynamic, images that vividly portray both the direction and velocity of motion of the object. Generally, in panning photography, the object-in-motion in the foreground may be at different speeds of motion. The users may pan a camera based on his / her estimation of the speed of the object. Particularly, the user controls the rotation of the camera with respect to an angle of the object, with a degree of control. Further, several automated systems compute the different speeds of motion of the object using different types of motion cues. The automated system computes the angle of blur through the scene understanding / motion of the object. Further, for multiple frames, dense optical flow techniques are applied to measure the speed accurately. Thus, this configuration matches the direction of panning / background blur with the direction of movement of the object.

[0004] Currently, panning photography has garnered attention in cinematography and social media circles for its capacity to captivate audiences throughcompelling visual narratives. Further, panning photography provides one or more distinctive advantages. For example, the objects which are moving are brought into focus while disguising the unappealing background. Further, panning photography adds a sense of movement to the still images by keeping the foreground objects sharper and creating a streaky and blurred background. Panning photography emphasizes the object’s speed and creates a dynamic visual effect. Lastly, panning photography adds a narrative element to the image by capturing the objects in action / motion.

[0005] Further, the demand for panning photography has increased, particularly in domains like sports, wildlife, and dynamic daily situations. Panning photography, known for producing immersive visuals, has captured interest across users on a large scale.

[0006] However, various existing processes for the panning photography have some limitations, that is, manual input and proficiency are required in setting the equipment of the camera for capturing moving objects, manually, in panning photography. Further, setting up the camera equipment, manually, to synchronize with a moving object’s speed proves cumbersome and time-intensive. Particularly, manual capturing of moving objects in panning photography requires the expertise of the user, for example, professional photographers apply their niche expertise and substantial time in carefully capturing / creating the panning photograph. Further, the manual capturing involves setting different parameters manually, for example, adjusting the camera parameters like shutter speed, body positions, and controlled movement of the camera and body to produce the motion blur in the appropriate direction of the objects as shown in Figure 1 A (i-ii). Additionally, other parameters also need to be taken care of at the time of panning photography, for example, focus on the foreground objects, setting the focus of the camera, continuous capture @6FPS, panning the camera to match / follow the moving objects, accurate distance of the camera from the objects, correct state of flash of the camera, and video stabilization.

[0007] In one example, if the light is medium, the shutter speed is set at a lower speed. However, if the light is too bright, a polarizer or natural density filter is added. If the object is moving fast, the shutter speed is set at a higher speed. Thus, the setting of parameters becomes tedious when the movement of theobjects is irregular. In another example, when the shutter speed is high, the foreground objects are precise, however, there is no sense of speed due to limited background blur as shown in Figure 1 B(i). When the shutter speed is slow and the camera is moving along the objects, this results in a blurred background, and partially blurred foreground objects. Further, the user has to take multiple takes to capture a perfect image as shown in Figure 1 B (ii-iii). Thus, this capturing technique becomes tedious for the user. Additionally, the slow speed of the shutter speed increases manual efforts in panning photography as shown in Figure 1 C. Therefore, if the setting of the parameters of the camera is not accurate, or the photographer has less experience, or there is a lag in the tracking of the objects, then the moving objects in the image as captured in the panning photography are not appropriate as shown in Figures 1 D (i-iii). This results in discomfort for the user.

[0008] Further, apart from the manual setting of the parameters, to perform panning photography, a plurality of software solutions has been proposed. The software solutions consider multiple frames to generate the panning photograph effect, and successive frames are used to detect the direction and estimate the speed of the object. However, the existing solutions use multiple frames for speed calculation, thus the same solution is not as effective for a single frame input as that of multiple frames. Additionally, the existing solutions also increase the difficulty in precisely estimating the real speed and direction of the foreground objects from the single frame. Further, none of the solutions are compatible with panning photo creation from a single image input that is agnostic to capture settings. Particularly, none of the solutions convert the single image captured under any camera settings, for example, shutter speed, ISO, and autofocus in a panning image.

[0009] Additionally, the existing solutions for panning photography, consider the overall image’s structural characteristics to obtain the direction of the foreground object’s motion direction. This results in an incorrect representation of the foreground object’s motion direction as the foreground object’s motion direction is influenced by background scene characteristics. Particularly, the computed direction of motion changes when the background scene changes, which is not correct.

[0010] Thus, in order to overcome the above mentioned problems, many known techniques have been developed. However, the known techniques as disclosed have limitations that the known techniques require the need to track a moving object. The technique is multi-frame based and no intelligence is provided on the suitable scene to apply the effect. Further, if the objects are not tracked appropriately, there is a possibility of boundary leakage. Additionally, the known techniques support only manual panning shots, operating only in the bright light and when High Dynamic Range (HDR)is switched off. The known techniques require setting the manual exposure and the shutter speed based on the object’s speed. The known techniques require the application of de-filter to correct the overexposure due to the slow shutter and capture multiple frames using “Rapid Fire”. The known techniques require moving the camera along the moving objects, thus, becomes tedious for the user.

[0011] Thus, the summarized problems associated with the existing solution for panning photography are provided below:

[0012] The existing solutions as mentioned above discloses that when the user, in panning photography, captures the image, the post-processing filter often produces undesirable effects, such as color leakage at object boundaries.

[0013] The existing solutions depend on the direction information of the objects and do not consider the object’s speed for creating a panning effect. However, the existing solutions that compute speed for the creation of the panning effect, perform edge computation for photos having objects with higher speeds and thus encounter vanishing edges. Thus, these solutions lead to the computation of the amount of speed that becomes challenging and inefficient.

[0014] Further, to overcome the problems mentioned above, many solutions have been disclosed in known arts as discussed earlier. However, the solutions as disclosed in the known arts, fall short of accurately estimating object speed and direction for a single frame / photograph. Most prevailing methods rely on capturing a minimum of a 2-second video clip to gauge the speed and direction. Thus, the solutions do not support applications that consider only a single frame (gallery).

[0015] Therefore, in view of the above-mentioned problems, it is advantageous to provide a system and a method that can overcome the issues associated with generating panning images.Summary of InventionSolution to Problem

[0016] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention and nor is it intended for determining the scope of the invention.

[0017] The present disclosure discloses a method for generating a panning image, the method comprising: receiving an image frame comprising one or more objects and a background portion, identifying, in the received image frame, a first area corresponding to a salient object among the one or more objects, generating a gradient map based on the received image frame and the first area, wherein the gradient map indicates a distribution and orientation of edge intensities corresponding to the salient object in the received image frame, estimating, using a set of models, a speed parameter and a direction parameter corresponding to the salient object based on the received image frame, the first area, and the gradient map and generating the panning image by applying a blur to a second area corresponding to the background portion of the received image frame based on the estimated speed parameter and the estimated direction parameter.

[0018] Estimating the speed parameter and the direction parameter comprises: generating, using a neural network model among the set of models, feature representation maps based on the received image frame, the first area, and the gradient map, the feature representation maps comprising information on motion cues and information on direction cues corresponding to the salient object, estimating, using a first classification model among the set of models, the speed parameter based on the generated feature representation maps comprising information on the motion cues and estimating, using a second classification model among the set of models, the direction parameter based on the generated feature representation maps comprising information on the direction cues.

[0019] Estimating the speed parameter using the first classification model comprises: estimating a speed of the salient object in the received image frame, classifying the speed of the salient object in at least one of a set of speed class labels, wherein the set of speed class labels includes a slow-motion label, fast motion label, medium motion label, and still label and determining the speed parameter based on the classified speed.

[0020] Estimating the direction parameter using the second classification model comprises: estimating a direction of motion of the salient object in the received image frame, classifying the direction of motion of the salient object in at least one of a set of direction class labels, wherein the set of direction class labels includes angles among a set of pre-defined directional angles and determining the direction parameter based on the classified direction of motion.

[0021] The neural network model, the first classification model, and the second classification model are trained on one or more reference frames, one or more reference salient areas, and one or more reference gradient maps, and wherein training the neural network model comprises: constructing a future frame based on the future representation maps generated by the neural network model, wherein the future frame indicates future motion predictions associated with the one or more reference frames and backpropagating an error associated with the future frame for the training of the neural network model.

[0022] Training the neural network model, the first classification model, and the second classification model comprises: training the neural network model to generate the feature representation maps by providing the one or more reference frames, one or more reference salient areas, and one or more reference gradient maps as inputs, training the first classification model and the second classification model to determine the speed parameter and the direction parameter by providing the feature representation maps as inputs and backpropagating an error associated with the first classification model and the second classification model for training the neural network model.

[0023] Generating the panning image comprises: constructing a blur kernel based on the estimated speed parameter and the estimated direction parameter corresponding to the salient object, based on the estimated speed parameterbeing greater than a threshold and generating the panning image by applying a blur to the second area corresponding to a second area corresponding to the background portion of the received image frame using the blur kernel.

[0024] Generating the panning image comprises performing a convolution operation between the received image frame and the constructed blur kernel, wherein a point spread factor (PSF) associated with the blur kernel is defined based on the direction parameter.

[0025] A size of the blur kernel is proportional to the speed parameter, and wherein an angle to apply the blur is defined based on the direction parameter.

[0026] In another embodiment, also disclosed herein is an electronic device for generating a panning image, the electronic device comprising: a memory and a processor communicatively coupled with the memory, the processor being configured to: receive an image frame comprising one or more objects and a background portion, identify, in the received image frame, a first area corresponding to a salient object among the one or more objects, generate a gradient map based on the received image frame and the fist area, the gradient map indicates a distribution and orientation of edge intensities corresponding to the salient object in the received image frame, estimate, using a set of models, a speed parameter and a direction parameter corresponding to the salient object based on the received image frame, the fist area, and the gradient map and generate the panning image by applying a blur to a second area corresponding to the background portion of the received image frame based on the estimated speed parameter and the estimated direction parameter.

[0027] To estimate the speed parameter and the direction parameter, the processor is configured to: generate, using a neural network model among the set of models, feature representation maps based on the received image frame, the first area, and the gradient map, the future representation maps comprising information on motion cues and information on direction cues corresponding to the salient object, estimate, using a first classification model among the set of models, the speed parameter based on the generated feature representation maps comprising information on the motion cues and estimate, using a second classification model among the set of models, the direction parameter based onthe generated feature representation maps comprising information on the direction cues .

[0028] To estimate the speed parameter using the first classification model, the processor is configured to: estimate a speed of the salient object in the received image frame, classify the speed of the salient object in at least one of a set of speed class labels, wherein the set of speed class labels includes slow-motion label, fast motion label, medium motion label, and still label and determine the speed parameter based on the classified speed.

[0029] To estimate the direction parameter using the second classification model, the processor is configured to: estimate a direction of motion of the salient object in the received image frame, classify the direction of motion of the salient object in at least one of a set of direction class labels, wherein the set of direction class labels includes angles among a set of pre-defined directional angles and determine the direction parameter based on the classified direction of motion.

[0030] The neural network model, the first classification model, and the second classification model are trained on one or more reference frames, one or more reference salient areas, and one or more reference gradient maps, and wherein to train the neural network model, the processor is configured to: construct a future frame based on the future representation maps generated by the neural network model, wherein the future frame indicates future motion predictions associated with the one or more reference frames and backpropagate an error associated with the future frame for the training of the neural network model .

[0031] To train the neural network model, the first classification model, and the second classification model, the processor is configured to: train the neural network model to generate the feature representation maps by providing the one or more reference frames, one or more reference salient areas, and one or more reference gradient maps as inputs, train the first classification model and the second classification model to determine the speed parameter and the direction parameter by providing the feature representation maps as inputs and backpropagate an error associated with the first classification model and the second classification model for training of the neural network model .

[0032] To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which is illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting of its scope. The invention will be described and explained with additional specificity and detail with the accompanying drawings.Brief Description of Drawings

[0033] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:

[0034] Figures 1 A-1 D illustrate a scenario depicting challenges faced while executing a plurality of processes for panning photography as per known art;

[0035] Figure 2 illustrates an environment of a system communicably coupled with a User Equipment (UE), in accordance with an embodiment of the present disclosure;

[0036] Figure 3 illustrates a block diagram of the system in connection with the UE, in accordance with an embodiment of the present disclosure;

[0037] Figure 4 illustrates a block diagram depicting the operations of the system, in accordance with an embodiment of the present disclosure;

[0038] Figure 5A illustrates an identification of a salient Region of Interest (ROI) in an image frame by a salient object detector, in accordance with an embodiment of the present disclosure;

[0039] Figure 5B illustrates an operation of a deep feature extraction module of the salient objection detector, in accordance with an embodiment of the present disclosure;

[0040] Figure 5C illustrates an operation of a union attention module of the salient objection detector, in accordance with an embodiment of the present disclosure;

[0041] Figure 6 illustrates a HoG map generated by the system, in accordance with an embodiment of the present disclosure;

[0042] Figure 7A illustrates a training of a MOS estimator, in accordance with an embodiment of the present disclosure;

[0043] Figure 7B illustrates an inference, after training, of the MOS estimator, in accordance with an embodiment of the present disclosure;

[0044] Figure 8 illustrates a blur kernel generated by the system, in accordance with an embodiment of the present disclosure;

[0045] Figure 9A illustrates the operations of a background motion filtering module of the system, in accordance with an embodiment of the present disclosure;

[0046] Figure 9B illustrates different conditions in which a panning image may be generated by the background motion filtering module, in accordance with an embodiment of the present disclosure;

[0047] Figure 9C illustrates the background panning speed matching with a foreground salient object’s motion speed, in accordance with an embodiment of the present disclosure;

[0048] Figure 9D illustrates a low-level motion feature and a high-level scene element in the system, in accordance with an embodiment of the present disclosure;

[0049] Figure 9E illustrates a relationship between a filter of the background motion filtering module 434 and a panning effect, in accordance with an embodiment of the present disclosure;

[0050] Figure 9F illustrates a multi-directional MOS (Motion filter with Orientation and Speed) Filter Design in the system, in accordance with an embodiment of the present disclosure;

[0051] Figure 10 illustrates the panning image generated by the system, in accordance with an embodiment of the present disclosure;

[0052] Figure 11 A illustrates a block diagram depicting multi-frame-based motion cue prediction by the system, in accordance with another embodiment of the present disclosure;

[0053] Figure 11 B illustrates an optical flow computation engine in the system, in accordance with another embodiment of the present disclosure; and

[0054] Figure 12 illustrates a method performed by the system to generate the panning image, in accordance with an embodiment of the present disclosure.

[0055] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. Furthermore, in terms of the construction of the device, a plurality of components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.Description of Embodiments

[0056] For the purpose of promoting an understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the invention as illustrated therein being contemplated as would normally occur to one skilled in the art to which the invention relates. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skilled in the art to which invention belongs. The system and examples provided herein are illustrative only and not intended to be limiting.

[0057] For example, the term “some” as used herein may be understood as “none” or “one” or “more than one” or “all.” Therefore, the terms “none,” “one,” “more than one,” “more than one, but not all” or “all” would fall under the definition of “some.” It should be appreciated by a person skilled in the art that the terminology and structure employed herein is for describing, teaching, and illuminating some embodiments and their specific features and elements and therefore, should not be construed to limit, restrict, or reduce the spirit and scope of the present disclosure in any way.

[0058] For example, any terms used herein, such as “includes,” “comprises,” “has,” “consists,” and similar grammatical variants do not specify an exact limitation orrestriction, and certainly do not exclude the possible addition of a plurality of features or elements, unless otherwise stated. Further, such terms must not be taken to exclude the possible removal of the plurality of the listed features and elements, unless otherwise stated, for example, by using the limiting language including, but not limited to, "must comprise” or “needs to include.”

[0059] Whether or not a certain feature or element was limited to being used only once, it may still be referred to as “plurality of features” or “plurality of elements” or “at least one feature” or “at least one element.” Furthermore, the use of the terms “plurality of” or “at least one” feature or element do not preclude there being none of that feature or element, unless otherwise specified by limiting language including, but not limited to, “there needs to be plurality of...” or “plurality of elements is required.”

[0060] Unless otherwise defined, all terms and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by a person ordinarily skilled in the art.

[0061] Reference is made herein to some “embodiments.” It should be understood that an embodiment is an example of a possible implementation of any features and / or elements of the present disclosure. Some embodiments have been described for the purpose of explaining plurality of the potential ways in which the specific features and / or elements of the proposed disclosure fulfil the requirements of uniqueness, utility, and non-obviousness.

[0062] Use of the phrases and / or terms including, but not limited to, “a first embodiment,” “a further embodiment,” “an alternate embodiment,” “one embodiment,” “an embodiment,” “multiple embodiments,” “some embodiments,” “other embodiments,” “further embodiment”, “furthermore embodiment”, “additional embodiment” or other variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, plurality of particular features and / or elements described in connection with plurality of embodiments may be found in one embodiment, or may be found in more than one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although plurality of features and / or elements may be described herein in the context of only a single embodiment, or in the context of more than one embodiment, or inthe context of all embodiments, the features and / or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and / or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.

[0063] Any particular and all details set forth herein are used in the context of some embodiments and therefore should not necessarily be taken as limiting factors to the proposed disclosure.

[0064] In an embodiment, a system and a method for generating a panning image are disclosed. The panning image may be generated by the system using a single frame, where the system derives probable motion direction and speed of foreground moving subject / object from a single still photograph to create a panning effect. Further, in an exemplary embodiment, the system detects eight different directions and three different speeds at which an accurate and pleasing panning photograph may be generated. Further, the system also utilizes predictive models that understand moving edges verses structural edges.

[0065] Embodiments of the present invention will be described below in detail with reference to the accompanying drawings.

[0066] Figure 2 illustrates an environment 200 of a system 204 communicably coupled with a User Equipment (UE) 202, in accordance with an embodiment of the present disclosure. Figure 3 illustrates a block diagram 300 of the system 204 in connection with the user equipment 202, in accordance with an embodiment of the present disclosure.

[0067] In an embodiment, the user equipment 202 may be a smartphone, or any other electronic device having a camera, without departing from the scope of the present disclosure. In an embodiment, the system 204 may be communicatively coupled with the UE 202. In another embodiment, the system 204 may be deployed within the UE 202, without departing from the scope of the present disclosure. In an embodiment, the system 204 may be configured to provide a panning image based on an image frame received by the UE 202. In an embodiment, the system 204 may be implemented as an electronic device for generating a panning image.

[0068] In an embodiment, the system 204 may include, but is not limited to, a processor 304, a memory 308, and a plurality of modules 312 among other examples which are explained in detail in subsequent paragraphs. Further, the system 204 may include an Input / Output (I / O) interface 336 and a transceiver 334. Further, in some embodiments where the system 204 is implemented as a standalone entity at a server / cloud architecture, the system 204 may be in communication with multiple user equipment to receive data from each of the multiple user equipment, and the details provided below with respect to the system 204 and the user equipment 202 are applicable for the system 204 and the multiple user devices as well.

[0069] In an exemplary embodiment, the processor 304 may be operatively coupled to each of the I / O interface 336, the plurality of modules 312, the transceiver 336, and the memory 308. In one embodiment, the processor 304 may include a graphical processing unit (GPU) and / or an artificial intelligence engine (AIE). In one embodiment, the processor 304 may include at least one data processor for executing processes in a virtual storage area network. The processor 304 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc. In one embodiment, the processor 304 may include a central processing unit (CPU), a graphics processing unit (GPU), or both. The processor 304 may be one or more general processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, servers, networks, digital circuits, analog circuits, combinations thereof, or other now-known or later developed devices for analyzing and processing data. The processor 304 may execute a software program, such as code generated manually (i.e. , programmed) to perform the desired operation.

[0070] The processor 304 may be disposed in communication with one or more input / output (I / O) devices via the I / O interface 336. In some embodiments, the processor 304 may communicate with the UE 202 using the I / O interface 336. In some embodiments, the I / O interface 336 may be implemented within the user equipment 202. The I / O interface 336 may employ communication code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, or the like. Inan embodiment, the I / O interface 336 may enable input and output to and from the system 204 using suitable devices such as, but not limited to, display, keyboard, mouse, touch screen, microphone, speaker, and so forth.

[0071] Using the I / O interface 336, the system 204 may communicate with one or more I / O devices, specifically, the user equipment 202, to which the system 204 generates and provides the panning image. For example, the input device may be an antenna, microphone, touch screen, touchpad, storage device, transceiver, video device / source, etc. The output devices may be a video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, Plasma Display Panel (PDP), Organic light-emitting diode display (OLED) or the like), audio speaker, etc.

[0072] The processor 304 may be disposed in communication with a communication network via a network interface. In an embodiment, the network interface may be the I / O interface 336. The network interface may connect to the communication network to enable the connection of the system 204 with the UE 202. The network interface may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc. The communication network may include, without limitation, a direct interconnection, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, etc. Using the network interface and the communication network, the system 204 may communicate with other devices. The network interface may employ connection protocols including, but not limited to, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc.

[0073] The transceiver 336 may be configured to receive and / or transmit signals to and from the UE 202. In one embodiment, the database may be configured to store the information as required by the plurality of modules 312 and the processor 304 to perform one or more functions for generating the panning image, on the UE 202.

[0074] In some embodiments, the memory 308 may be communicatively coupled to the processor 304. The memory 308 may be configured to store data, and instructions executable by the processor 304 to perform the one or more methods disclosed herein throughout the present disclosure. In one embodiment, the memory 308 may be provided within the UE 202. In another embodiment, the memory 308 may be provided within the system 204 being remote from the UE 202. In yet another embodiment, the memory 308 may communicate with the processor 304 via a bus within the system 204. In yet another embodiment, the memory 308 may be located remote from the processor 304 and may be in communication with the processor 304 via a network. The memory 308 may include, but is not limited to, a non-transitory computer-readable storage media, such as various types of volatile and non-volatile storage media including, but not limited to, random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable readonly memory, flash memory, magnetic tape or disk, optical media and the like.

[0075] In one example, the memory 308 may include a cache or random-access memory for the processor 304. In alternative examples, the memory 308 is separate from the processor 304, such as a cache memory of a processor, the system memory, or other memory. The memory 308 may be an external storage device or database for storing data. The memory 308 may be operable to store instructions executable by the processor 304. The functions, acts, or tasks illustrated in the figures or described may be performed by the programmed processor 304 for executing the instructions stored in the memory 308. The functions, acts, or tasks are independent of the particular type of instruction set, storage media, processor, or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code and the like, operating alone or in combination. Likewise, processing strategies may include multiprocessing, multitasking, parallel processing, and the like.

[0076] In some embodiments, the plurality of modules 312 may be included within the memory 308. The memory 308 may further include a database to store data. The plurality of modules 312 may include a set of instructions that may be executed to cause the system 204, in particular, the processor 304 of the system 204, to perform any one or more of the methods / processes disclosed herein. Theplurality of modules 312 may be configured to perform the steps of the present disclosure using the data stored in the database. For instance, the plurality of modules 312 may be configured to perform the steps disclosed in Figures 4 to 10.

[0077] In an embodiment, each of the plurality of modules 312 may be a hardware unit which may be outside the memory 308. Further, the memory 308 may include an operating system for performing one or more tasks of the system 204, as performed by a generic operating system.

[0078] In one example, the modules 312 may include a receiving module 314, an identifying module 316, a generating module 318, an estimating module 320, a classifying module 322, a determining module 324, a constructing module 326, a backpropagating module 328, a training module 330 and a performing module 332. Each of the receiving module 314, the identifying module 316, the generating module 318, the estimating module 320, the classifying module 322, the determining module 324, the constructing module 326, the backpropagating module 328, the training module 330, and the performing module 332 may be in communication with each other. Further, each of the receiving module 314, the identifying module 316, the generating module 318, the estimating module 320, the classifying module 322, the determining module 324, the constructing module 326, the backpropagating module 328, the training module 330, and the performing module 332 may be in communication with the processor 304.

[0079] Further, the present disclosure contemplates a computer-readable medium that includes instructions or receives and executes instructions responsive to a propagated signal. Further, the instructions may be transmitted or received over the network via a communication port or interface or using a bus (not shown). The communication port or interface may be a part of the processor 304 or may be a separate component. The communication port may be created in software or may be a physical connection in hardware.

[0080] The communication port may be configured to connect with a network, external media, the display, or any other components in the system, or combinations thereof. The connection with the network may be a physical connection, such as a wired Ethernet connection, or may be established wirelessly. Likewise, the additional connections with other components of thesystem 204 may be physical or may be established wirelessly. The network may alternatively be directly connected to a bus. For the sake of brevity, the architecture and standard operations of the memory 308, the processor 304, the transceiver 336, and the I / O interface 336 are not discussed in detail.

[0081] Further, in an embodiment, the working of the system 204 to generate the panning image, on the UE 202 is explained in detail. The processor 304, in conjunction with the receiving module 314, the identifying module 316, the generating module 318, the estimating module 320, the classifying module 322, the determining module 324, the constructing module 326, the backpropagating module 328, the training module 330, and the performing module 332 may be configured to perform specific operations explained in paragraphs in conjunction with Figure 3 to Figure 10 in subsequent paragraphs.

[0082] Figure 4 illustrates a block diagram depicting operations of the system 204, in accordance with an embodiment of the present disclosure. Figure 5A illustrates an identification of a salient Region of Interest (ROI) in the image frame 418 by a salient object detector 420, in accordance with an embodiment of the present disclosure. Figure 5B illustrates an operation of a deep feature extraction module 502 of the salient objection detector 420, in accordance with an embodiment of the present disclosure. Figure 5C illustrates an operation of a union attention module 504 of the salient objection detector 420, in accordance with an embodiment of the present disclosure. Figure 6 illustrates a HoG map 602 generated by the system 204, in accordance with an embodiment of the present disclosure. Figure 7A illustrates a training of a MOS estimator 424, in accordance with an embodiment of the present disclosure. Figure 7B illustrates an inference, after training, of the MOS estimator 424, in accordance with an embodiment of the present disclosure. Figure 8 illustrates a blur kernel generated by the system 204, in accordance with an embodiment of the present disclosure. Figure 9A illustrates an operation of a background motion filtering module 432 of the system 204, in accordance with an embodiment of the present disclosure. Figure 9B illustrates different conditions in which the panning image 434 may be generated by the background motion filtering module 432, in accordance with an embodiment of the present disclosure. Figure 9C illustrates the background panning speed matching 434 with the foreground salient object’s motion speed,in accordance with an embodiment of the present disclosure. Figure 9D illustrates a low-level motion feature and a high-level scene element in the system 204, in accordance with an embodiment of the present disclosure. Figure 9E illustrates a relationship between a filter of the background motion filtering module 434 and a panning effect, in accordance with an embodiment of the present disclosure. Figure 9F illustrates a multi-directional MOS (Motion filter with Orientation and Speed) Filter Design in the system 204, in accordance with an embodiment of the present disclosure. Figure 10 illustrates the panning image 434 generated by the system 204, in accordance with an embodiment of the present disclosure.

[0083] In an embodiment, Figure 4 is explained in conjunction with Figure 5 to Figure 10, without departing from the scope of the present disclosure.

[0084] In an embodiment, referring to Figure 4, at block 402, the receiving module 314 may be configured to receive the image frame 418 comprising one or more objects and a background portion. In such an embodiment, the image frame 418 is a single image frame associated with an image or a video.

[0085] In an embodiment, at block 404, the identifying module 316 may be configured to identify a salient Region of Interest (ROI) in the received image frame 418 (referred to here as the image frame 418) with the assistance of the salient object detector 420 of the system 204. The identified salient ROI (referred to here as the salient ROI) may be associated with a salient object amongst the one or more objects. The salient ROI may be referred to as a first area corresponding to a salient object. The first area may be an area corresponding to a salient object in the received image frame 418.

[0086] Referring to Figure 5A, the salient object detector 420 may include a deep learning model. Further, the salient object detector 420 along with the identifying module 316 may be configured for identifying the salient object in the image frame 418, where the salient object is an object that draws more attention from the users than the surrounding areas. The salient object acts as a foreground around which a blur (explained later in the subsequent paragraphs) may be applied. In an embodiment, the salient object detector 420 may include a deep feature extraction module 502, and a union attention module 504. Further, the deep feature extraction module 502 may be configured to capture low-levelsemantics in the image frame 418 such as the spatial location of the salient object by employing a convolutional neural network (CNN), such as EfficientNet.

[0087] In such an embodiment, referring to Figures 5A and 5B, the deep feature extraction module 502 may be configured to extract deep features from the image frame 418 using a scalable CNN. Further, the deep extraction module 502 may perform a plurality of operations, i.e. , feature extraction from EfficientNet and multi-level feature aggregation. The EfficientNet may be used to extract information from the image frame 418 which may include the characteristics of the one or more objects present in the image frame 418.

[0088] Further, the features obtained from the CNN stages / levels 3, 5, and 7 of the EfficientNet may be combined to give multi-level features (collected from different levels of the CNN), which help in effectively locating the spatial position of the salient object in the image frame 418. Further, features from stages 3, 5, and 7 may be observed to help in capturing the spatial position of the salient object. For instance, deeper in the network, the high-level information of the image frame 418 is gradually lost, thus increasing difficulty in locating the salient object precisely. Thus, stages 3 and 5 relatively hold the high-level information (edges of the salient object) better compared to stage 7, which would be low-level information (the presence of human, the contrastive difference of subject to background), and help in giving a clear boundary in an output mask 506. Therefore, the deep feature extraction module 502 may be configured to generate deep multi-level features, for example, the presence of people, animals, and contrastive differences which may be intermediate representations of the salient object in the image frame 418 as the output with the assistance of the intermediate deep feature maps 508.

[0089] In an embodiment, referring to Figure 5C, the union attention module 504 may be configured to emphasize the significant features like body parts or edges to get clear boundaries in the output mask 506. In such an embodiment, the union attention module 504 may be configured to receive the deep multi-level features generated with the assistance of the intermediate deep feature maps 508. The union attention module 504 enhanced the performance of deep learning models, by giving relative importance to the features captured from the image frame 418, which helps determine the output mask 506. For example, the presence of a faceis relatively more important than the presence of legs in identifying a person in the image frame 418. This importance is captured through a variant of attention known as self-attention, where the features obtained from the same stage of the model are checked, and further, is decided which feature is more important than the others. The ability to make this decision is learned by the model during the training process. The union attention model 504 may include a channel attention block and a spatial attention block.

[0090] The channel attention block may be used to discriminate the significant channels (deep features like body parts of a person, or the objects in the background) from the received deep multi-level features / feature representation. For instance, from the image frame 418 of the horse, this block highlights the presence of limbs through the prominence of the head, while assigning lower importance to features like a pole in the background which may be considered less significant due to the contrastive difference. The channel attention may be determined by the following operations:

[0091] The spatial information X is mean-pooled globally to obtain a representative value for each channel represented together as Self-Attention. Further, the sigmoid o function is employed on the descriptor obtained after pooling to compute the attention score ac, which is then applied to the input feature map X to obtain channel-attended feature map Xc. Attention score ac is a weight vector equal to the size of the input channels / features, adding up to 1 . Significant features are assigned higher weights than others. Further, the attention score may be determined by Equation 1 :Sigmoid o(z)): o(z) = 1 / (1 + e-z) (1 )where sigmoid is a mathematical function, whose function is to map input values to an output range between 0 and 1 , which can be interpreted as probabilities or activation levels, here considered as attention scores / weights denoted ac.

[0093] The spatial attention block may be applied in complimentary to the channel attention block, to focus on the informative regions in the deep multi-level features. This helps in filtering any background information captured through the channel attention block. For instance, the pole in the background may be filtered once the context of the presence of the horse is evident, obtained by combining the weighted features from the channel attention model. The spatial attention may be determined by the following operations:

[0094] Self-Attention may be used to capture the inter-spatial relationship of features and further, the input features are reduced to a single output feature map Xs, which when passed through a sigmoid layer gives the final fixation output. The single output feature may be provided by Equation 2:

[0095] Therefore, the deep feature extraction module 502, and the union attention module 504 may be configured to capture the properties such as contrast difference, and foreground-background check, that make the salient object attentive, through end-to-end training. The salient object detector 420 along with the identifying module 316 may be configured to identify a mask of the Region of Interest (ROI) associated with the salient object.

[0096] In an embodiment, after identifying the salient ROI, referring to Figure 4, at block 406, the generating module 318 may be configured to generate the Histogram of Gradients (HOG) map 602 based on the image frame 418 and the identified salient ROI, with the assistance of a HOG generator 422. The HOG map 602 may indicate the distribution and orientation of edge intensities associated with the salient object amongst the one or more objects. Edge intensity refers to the magnitude of brightness change across neighboring pixels in an image, representing how prominent an edge appears. The HOG map may be referred to as a histogram map.

[0097] In such an embodiment, referring to Figures 4 and 6, the HOG generator 422 may be configured to receive the image frame 418 and the identified salient ROI as the input. Further, the HOG generator 422 along with the generating module 318 may be configured to generate the HOP map 602. The HOG generator 422 may capture information about the local gradient directions in the image frame 418 and along with the generating module 318 generate the HOG map 602.While the image frame 418 contains semantic and speed-related cues, the salient ROI localizes the characteristics of the salient object, which is sufficient for speed estimation. Further, the network is additionally fed with a HOG channel to emphasize the edge-based learning which is crucial for direction computation. Further, the HOG map 602 may be important to determine direction as the HOG map 602 provides the object’s edge characteristics which may be further used as a cue to predict the direction of the object. The HOG map 602 may be generated by performing the following operations:

[0098] Gradients of pixel intensities are calculated in an image of the image frame 418. The image is divided into cells and histograms of gradient orientations are computed within each cell. Further, adjacent cells are grouped into blocks and their histograms are normalized for illumination invariance. Normalized block histograms are concatenated to create a compact representation, the HOG descriptor. Lastly, a sliding window is applied to extract HOG features across different regions of the image.

[0099] In an embodiment, after generating the HOG map 602, at blocks 408 and 410, the estimating module 320 may be configured to estimate, using a set of models, a speed parameter, and a direction parameter associated with the salient object based on the image frame 418, the salient ROI, and the HOG map 602. Further, this configuration also constructs a next frame / future frame in an image format with the estimated speed parameter and the direction parameter.

[0100] The speed parameter may refer to a parameter indicating the movement speed of the object. For example, the speed parameter may include slow speed, medium speed, fast speed, and still object, but is not limited thereto. The speed parameter may further be subdivided into additional levels or categories of speed depending on the implementation. Also, the speed parameter may be expressed as a parameter value corresponding to a speed.

[0101] The direction parameter may refer to a parameter indicating the movement direction of the salient object. The direction parameter may include, for example, North, South, East, West, North-East, North-West, South-East, and South-West, but is not limited thereto. The direction parameter may be expressed as a parameter value corresponding to a direction.

[0102] In such an embodiment, to estimate the speed parameter and the direction parameter, the generating module 318 may be configured to generate, using a neural network model amongst the set of models, feature representation maps based on the image frame 418, the salient ROI, and the HOG map 602. The feature representation maps may be configured to indicate motion cues and direction cues associated with the salient object. The feature representation maps may comprise information on motion cues and information on direction cues corresponding to the salient object. Further, the estimating module 320 may be configured to estimate, using a first classification model (also referred to as classifier 1 ) amongst the set of models, the speed parameter based on the generated feature representation maps indicating the motion cues. Alternatively, the estimating module 320 may be configured to estimate, using a first classification model (also referred to as classifier 1 ) amongst the set of models, the speed parameter based on the generated feature representation maps comprising information on the motion cues. Furthermore, the estimating module 320 may be configured to estimate, using a second classification model (also referred to as classifier 2) amongst the set of models, the direction parameter based on the generated feature representation maps indicating the direction cues. Alternatively, the estimating module 320 may be configured to estimate, using a second classification model (also referred to as classifier 2) amongst the set of models, the direction parameter based on the generated feature representation maps comprising information on the direction cues.

[0103] Particularly, the MOS estimator 424 of the system 204 along with the estimating module 320 may be configured to estimate the speed parameter and the direction parameter associated with the salient object in the single frame. Further, the MOS estimator 424 may be a parallel network that enforces the neural network model (interchangeably referred to here as a feature representation module 702 as shown in Figure 7A), to learn both spatial contextand motion cues through two complementary tasks of a future / temporal frame construction 704 (as shown in Figure 7A) and a motion cue classification module 706 (as shown in Figure 7A). Thus, the MOS estimator 424 becomes a unique network with surpassing final results.

[0104] In such an embodiment, the feature representation module 702 may be configured to learn the features for constructing the future frame (temporal frame construction module 704). The feature representation module 702 may be a deep CNN network based on ResNet-like architecture. The feature representation module 702 may receive the image frame 418, the salient ROI, and the HOG map 602 as an input. Each module may contain a combination of convolutional layers, batch normalization, and activation functions such as ReLU. Further, the feature representation module 702 may generate the feature representation maps based on the image frame 418, the salient ROI, and the HOG map 602. Further, the aim of the feature representation module 702 is to extract and learn the relevant context and motion features from the inputs, which may help in both the temporal frame construction module 704 and the motion cue classification module 706.

[0105] Further, in the temporal frame construction module 704, the output from the feature representation module 702 may be used to construct the future frame (Ft+1 ). The future frame may be constructed with contextually suitable changes in the direction cues and the motion cues. Particularly, the construction process of the future frame involves gradually up-sampling the features to regain the original spatial dimensions of the frame. This process ensures that the network implies a good understanding of the speed and direction characteristic features of the salient object. Further, the up-sampling may be achieved using various techniques, for example, transposed convolutions or nearest-neighbor interpolation followed by regular convolutions. Further, using this module, the feature representation module 702 learns to give the features that are necessary or important for constructing the next frame / future frame.

[0106] The motion cue classification module 706 considers feature representation maps as an input from the feature representation module 702. The motion cue classification module 706 along with the estimating module 320 estimate the speed parameter and the direction parameter associated with the salient object.This module contains two separate small classification sub-modules which are small CNN networks. The first classification module is dedicated to acquiring knowledge about the speed parameter associated with the salient object (as explained in earlier paragraphs and will be explained later in subsequent paragraphs). Further, the second classification module specializes in determining the direction parameter associated with the salient object (as explained in earlier paragraphs and will be explained later in subsequent paragraphs).

[0107] The back propagating from the motion cue classification module 706 helps to refine the feature representation maps to learn motion predictive features for perfect feature frame construction and accurate classification of speed and directions of selected salient objects in the image frame (Ft) 418. Further, a weighted loss function for the single frame-based motion prediction training may be provided as follows:L (Ft, Ft+1 ) = w1 *a + w2*p, a = MSE(Ft,Ft+1),[3 = CCE(Speed) + CCE(Direction) (3)

[0108] Where: Ft: The frame under consideration, Ft+1 : The future frame at time t+1 , a: MSE loss between the reconstructed image and Ft+1 (Ground Truth), [3: Categorical Cross Entropy (CCE) loss between the predicted direction and speed values to the actual labels, w1 ,w2: weightage given to each of the loss functions.

[0109] Further, to estimate the speed parameter using the first classification model, the estimating module 320 may be configured to estimate a speed of the salient object in the image frame 418. The classifying module 322 may be configured to classify the speed of the salient object in at least one of a set of speed class labels. The set of speed class labels may include a slow-motion label, fast-motion label, medium-motion label, and still label Further, the determining module 324 may be configured to determine the speed parameter based on the classified speed.

[0110] Further, to estimate the direction parameter using the second classification model, the estimating module 320 may be configured to estimate a direction of motion of the salient object in the image frame 418. The classifying module 322 may be configured to classify the direction of the motion of the salient object in atleast one of a set of direction class labels. The set of direction class labels includes angles among a set of pre-defined directional angles. In an embodiment, the set of pre-defined directional angles may include 8 angles, without departing from the scope of the present disclosure. Further, the determining module 324 may be configured to determine the direction parameter based on the classified direction of motion.

[0111] In such an embodiment, the neural network model 702, the first classification model, and the second classification model are trained on one or more reference frames, one or more reference salient ROI, and one or more reference HOG maps. Further, to train the neural network model 702, the constructing module 326 may be configured to construct the future frame based on the future representation maps generated by the neural network model 702. The future frame indicates future motion predictions associated with the one or more reference frames. The backpropagating module 328 may be configured to backpropagate an error associated with the future frame for the training of the neural network model 702.

[0112] Further, to train the neural network model 702, the first classification model, and the second classification model, the training module 330 may be configured to train the neural network model 702 to generate the feature representation maps by providing the one or more reference frames, one or more reference salient ROI, and one or more reference HOG maps as inputs. The training module 330 may be configured to train the first classification model and the second classification model to determine the speed parameter and the direction parameter by providing the feature representation maps as inputs. Further, the backpropagating module 328 may be configured to backpropagate an error associated with the first classification model and the second classification model for training of the neural network model 702.

[0113] In such an embodiment, referring to Figure 7A, the training of the neural network model / feature representation module 702, the first classification model, and the second classification model by the training module 330 includes following steps:

[0114] Pre-processing: For training the MOS estimator 424, especially, the neural network model / feature representation module 702, the first classification model and the second classification model, a video dataset is used. Further, at least two frames are extracted from the videos which are predetermined sets of frames apart. In such an embodiment, the at least two frames are considered which are 15 frames apart, without departing from the scope of the present disclosure. Appropriate subsampling, normalization, scaling, cropping, and flipping operations are applied to these frames before being fed into the model.

[0115] Generating the salient ROI from the image frame 418: The image frame 418 after pre-processing is an RGB image of size 256x256. Further, a pre-trained saliency model / salient object detector 420 is loaded. The images are normalized using ImageNet mean and standard deviation values and passed the saliency model for the prediction of salient regions. After obtaining the predictions from the saliency model, normalization is applied to the output image to bring the intensity of pixel values to the (0-1) range, helping in improving the convergence of this model. Appropriate resizing is done to the image based on the use case, resizing an image to a smaller size requires lesser computational resources and helps significantly discard the redundant information from the image.

[0116] Calculating Histogram of Gradients from an input RGB frame: The image frame 418 after pre-processing is an RGB image of size 256x256. Further, the gradients of pixel intensities in an image are calculated. The image is divided into cells and further, histograms of gradient orientations are computed within each cell. Adjacent cells are grouped into blocks and normalizing their histograms for illumination invariance. The normalized block histograms are concatenated to create the compact representation, the HOG descriptor. Lastly, the sliding window is applied to extract HOG features across different regions of the image.

[0117] Initialization of the MOS estimator 424, especially, the neural network model / feature representation module 702, the first classification model and the second classification model, and other parameters before training: All the models are initialized for training by setting the external configuration settings, for example, batch size, learning rate, optimizer parameters. Further, epoch sizes are not learned from the data but are set before the training process begins. The epoch data significantly influences the learning process and the performance ofthe models. Further, this network architecture is driven by two losses: Mean Squared Error (a statistical measure of the average squared difference between the estimated values and the actual value) on the reconstructed image and the categorical cross-entropy on the classification results for both speed and direction. Both these losses play a crucial role in training the MOS Estimator 424 by guiding the learning process and shaping the model’s ability to make accurate predictions on unseen data.

[0118] Further, an Adam optimizer (Adaptive Moment Estimation) is utilized initially which is an optimization algorithm used in Deep neural networks. The Adam optimizer is efficient in training the MOS estimator 424. This optimizer adjusts the model parameters during training to minimize the difference between predicted and actual outputs directly impacting how effectively the neural network learns from the training data.

[0119] Training the MOS Estimator 424: The original input RGB image, HoG features, and the salient ROI are passed as the input to the models. A complete End-to-End training is performed jointly to optimize all the models for better performance. The neural network model / feature representation module 702 is trained with RGB images, HOG representations, and the salient ROI as an input. The temporal frame construction module 704 uses the output of the feature representation module 702 for training. The training process is supervised to reconstruct the future frame (Ft+1) from the learned features. Based on the loss, the error is back propagated, from two parallel streams to the feature representation module 702 for efficient learning. Further, the feature representation maps from the feature representation module 702 may be configured to train the motion cue classification module 706 having the first classification model and the second classification model. Further, the second classification model is trained to acquire knowledge about the direction of the object’s motion. Further, the first classification model is trained to specialize in determining the speed of the motion. The error from the motion cue classification module 706 is also back-propagated to the feature representation module 702 for efficient learning. Based on the gradients, backpropagated model parameters are updated. Further, the trained checkpoint model files are saved at regular intervals.

[0120] Model checkpointing: Once the acceptable accuracy is achieved, all the models are retrained with a Stochastic Gradient Descent (SGD) optimizer instead of Adam for it to converge better.

[0121] Further, after training the MOS estimator 424, especially, the neural network model / feature representation module 702, the first classification model, and the second classification model (referred to here as models), the MOS estimator 424 may be configured for inference operations 700 as shown in Figure 7B. The inference operations 700 of the MOS estimator 424 are as follows:

[0122] The image frame 418, after pre-processing, is the RGB image of size 256x 256(Ft). Further, the Histogram of Gradients for the image frame (HOG(Ft)) and the salient ROI are computed. Further, Ft, HOG(Ft), and the salient ROI are passed to the trained MOS estimator 424. The feature representation maps from the feature representation module 702 are passed to a set of convolution blocks and then to the first classification model and the second classification model for estimating the speed parameter and the direction parameter. For example, the first classification model estimates the speed parameter, including but not limited to, slow speed, fast speed, and still object. Further, the second classification model estimates the direction parameter, including but not limited to, North, South, East, West, North-East, North-West, South-East, and South-West. These classification models work by learning patterns from labelled data to make predictions or categorize new, unseen data. Once trained, these classification models may generalise the knowledge to classify new data.

[0123] Further, in an embodiment, after estimating the speed parameter and the direction parameter by the MOS estimator 424 along with the estimating module 320, the motion blur angle estimator 426 is configured to receive the estimated direction parameter and generate a corresponding operational value denoted by theta, without departing from the scope of the present disclosure. Further, the motion blur speed estimator 428 is configured to receive the estimated speed parameter and generate a corresponding operational value denoted by sigma, without departing from the scope of the present disclosure

[0124] In an embodiment, at block 410, the constructing module 326 may be configured to construct the blur kernel based on the estimated directionparameter of the salient object with a MOS filter 430. The blur kernel may be constructed upon a determination that the estimated speed parameter is greater than a threshold as shown in block 412. In such an embodiment, the size of the blur kernel is proportional to the speed parameter. Further, an angle to apply the blur is defined based on the direction parameter.

[0125] In such an embodiment, the operational value denoted by theta and the operational value denoted by sigma are used to construct the blur kernel, where a higher intensity implies a greater speed of the salient object. The size of the blur kernel may be directly proportional to the operational value denoted by sigma. Further, the operational value denoted by theta may be utilized to determine the angle at which the blur kernel should be applied.

[0126] Particularly, referring to Figure 8, in order to enhance the perception of the movement of the salient object, a streaky effect using the blur kernel is applied to the background in the image frame 418. Motion blur on a digital image may be modelled as a convolution between the image and the blur kernel having a point spread factor (PSF) distribution equals to the angle(9) of the blur. Further, the point spread factor (PSF) associated with the blur kernel may be defined based on the direction parameter. Further, to construct the blur kernel based on the operational value corresponding to the estimated speed parameter and the estimated direction parameter following operation is performed:

[0127] An ideal line segment with the length(intensity) and angle is constructed which is centered at the center coefficient of h.

[0128] For each coefficient location (i,j), the nearest distance between that location and the ideal line segment is computed.

[0129] Then, the value at the location is computed by the given formula

[0130] h(i,j) = max(1 - nearest_distance, 0); and

[0131] the filter h is normalised to maintain the sum as 1 i.e. h = h / (sum(h(:)))

[0132] Further, an example corresponding to the blur kernel with length (intensity )= 7 and angle= 45 degrees is provided below:

[0133] Example 1

[0134] In an embodiment, at blocks 414 and 416, the generating module 318 may be configured to generate the panning image 434 by applying the blur, to the background portion of the image frame 418 using the constructed blur kernel, with the help of background motion filtering module 432. In such an embodiment, the performing module 332 may be configured to perform a convolution operation between the received image frame and the constructed blur kernel.

[0135] In such an embodiment, the background motion filtering module 432 may be configured to apply the blur to the background portion of the image frame 418 which is determined by the estimated speed parameter and the estimated direction parameter associated with the salient object. Further, the background motion filtering module 432 may be configured to edit the background portion using a custom and adaptive 2D filter application based on the estimated speed parameter and the estimated direction parameter associated with the salient object. Particularly, a streaky effect is applied to blur the background in the image frame 418, using the background motion filtering module 432 and the MSO filter 430 which specifies the magnitude and direction of the blur.

[0136] In an embodiment, the background motion filtering module 432 may be configured to apply the blur to a second area corresponding to the background portion of the image frame 418 based on the estimated speed parameter and the estimated direction parameter.

[0137] Referring to Figure 9A, at block 902, the salient object with the estimated speed parameter and the estimated direction parameter, the salient ROI, the blur, may be provided as an input to the background motion filtering module 432. Further, at block 904, the generating module 318 may be configured to generate the output mask 506 based on the salient object detector 420 after alpha matting. In an embodiment, the alpha matting may be utilized to generate a highly accurate mask for the salient object. The alpha matting is a technique in an image processing that involves estimating and creating an alpha channel denoting the transparency to accurately represent the capacity of each pixel in the salient object, enabling precise blending of foreground and background elements. At block 906, the generated mask may undergo inpainting using cv2.ipaint() where each pixel of the determined mask may be 1 . The in-paint technique is applied to conceal the foreground based on the determined mask, effectively removing unwanted elements. Inpainting is performed by extrapolating information from surrounding areas. At block 908, the blur is applied on the generated mask undergone the inpainting, in a direction opposite to the direction suggested by the MOS estimator 424 to create a panning effect. Lastly, at block 910, pixels are added according to the value of the generated mask to obtain the final panning image 434.

[0138] The background motion filtering module 432 ensures that the blur applied on the background portion aligns with the motion of the salient object, leading to mode visually appealing and contextually relevant output. Further, Figure 9B specifically refers to the appearance of the blur applied on the background portion when applied with varying operational value of the estimated speed parameter (sigma) and varying operational value of the estimated direction parameter (theta). Particularly, referring to Figure 9B (i), the panning image 434 is generated with minimum effect when the background motion filtering module 432 may be applied with 0 degrees and low speed. Referring to Figure 9B (ii), the panning image 434 is generated, with optimum effect, when the background motion filtering module 432 may be applied at 45 degrees and mid-speed.

[0139] Referring to Figure 9C, a plurality of examples indicating a relationship between the parameter of the background motion filtering module 432 and the panning effect is visualized. These figures specifically depict that the backgroundpanning speed matches with foreground salient object’s motion speed. Herein, a background filter kernel is an adaptive sigma for changing speed. Further, the Figure 9C depicts that the increasing speed is mapped with a size of the background filter kernel to create a streaky background. In an embodiment, the speed level of the foreground salient object-in-motion may be predicted through scene semantic learning. In an embodiment, the foreground salient object-in- motion may be any moving objects captured by the camera exhibiting certain image characteristics. Further, a filter kernel width is inversely mapped to the speed of the salient object that helps in adding the notion of the speed.

[0140] Referring to Figure 9D, the low-level motion feature 912, and the high-level scene element 914 may be depicted. In the low-level motion feature 912, motion blur is a common indicator of the movement of the object in the image. For example, if the object, for example, a car appears blurred in day light (especially around its edge structures) or streaky head / tail-lights in the night, the car is perceived to be moving.

[0141] Further, in the high-level scene element 914, the objects, for example, the car on a road, an airplane in the sky, etc., have background motion cues such as dust trails behind the cars and background context like Road, Sky, etc.

[0142] Thus, the initial layer of the MOS estimator 424 generally learns the low-level motion features 912, for example, blur gradients, while the later layers of the MSO estimator 424 learn the high-level scene elements 914 like body-parts and background context. Therefore, with accurate and efficient training data including objects like moving car, the MOS estimator 424 learns the speed cues from both low-level and scene elements to estimate the speed parameter at quantized levels (Still, Slow, Medium, High).

[0143] Referring to Figure 9E, the direction of the background blur matches with the foreground salient object’s motion direction (angle) to generate the desired output. Minimum of 8 quantized levels of angles may be computed to create streaky and directionally homogeneous background blur.

[0144] Referring to Figure 9F, the multi-directional MOS filter design addresses more complex motion paths, paths which cannot be described with a single linear direction. By considering the multi-directional MOS filter, the expressiveness ofthe generated panning photograph is improved. Further, each pixel is provided with an individual direction vector that is sampled in the continuum of values for smooth blurring.

[0145] Thus, after considering all the operations as discussed above, the panning image 434 may be generated as shown in Figure 10.

[0146] Further, in an embodiment, when the estimated speed parameter is smaller than the threshold as shown at block 412, then the panning image is not generated as shown at block 418.

[0147] Figure 11A illustrates a block diagram depicting multi-frame-based motion cue prediction by the system 204, in accordance with another embodiment of the present disclosure. Figure 11 B illustrates an optical flow computation engine 1106 in the system 204, in accordance with another embodiment of the present disclosure.

[0148] In another embodiment, when the camera captures more than or equal to two images, underlying temporal information is available with the image sequences to be utilized. This ensures the advantage of estimating the speed parameter and the direction parameter through an optical flow computation engine 1106.

[0149] Referring to Figure 11 , at blocks 402 and 1104, a set of image frames or a video clip may be passed through a frame quality detector module 1102. Further, the frame quality detector module 1102 may be configured to select the two best frames in the video clip. Thereafter, the salient object detector 420 and the HoG generator 422 are operated at blocks 404 and 406, respectively to generate the salient object and the corresponding HOG map. Further, an optical flow map 1108 (as shown in Figure 11 B) may be incorporated within the MOS estimator 424 to estimate the speed parameter and the direction parameter of the salient object.

[0150] Referring to Figure 11 B, the optical flow computation engine 1106 incorporated within the MOS estimator 424 is RAFT. RAFT is a deep learning method that computes the optical flow of the image frame 418. Further, the optical flow is a vector field between two images, showing the way of movement of the pixel of the salient object of the first image to form the same object in the second image. Further, the engine is a pre-trained open-source model and iskept frozen during training and testing. In one example, the salient object and the HOG map are generated corresponding to the two image frames. Further, the salient object and the HOG map are provided as an input to the MOS estimator 424, where the optical flow computation engine 1106 computes the salient object and the HOG map. Further, the optical flow computation engine generates the optical flow map 1108 based on the computation. Further, the optical flow map 1108 is passed from the speed and direction classifier, i.e. , the classifier 1 and the classifier 2 to determine the speed parameter and the direction parameter.

[0151] Referring back to Figure 11A, when the motion is detected, the direction parameter of the salient object and the speed parameter of the salient object pass through the MOS filter 430. Further, the remaining operations of the present embodiment to generate the panning image 434 are the same as the operation explained in conjunction with Figures 3 to 10. Thus, for the sake of brevity, the same has not been discussed here.

[0152] The method 1200 includes a series of operations shown at step 1202 through step 1212 of Figure 12. The method 1200 may be performed by the system 204 in conjunction with modules 312, the details of which are explained in conjunction with Figures 3 to 10 and the same are not repeated here for the sake of brevity in the present disclosure. The method 1200 begins at step 1202.

[0153] At step 1202, the method 1200 includes receiving the image frame 418 including the one or more objects and the background portion. The image frame 418 may be the single image frame.

[0154] At step 1204, the method 1200 includes identifying the salient ROI in the image frame 418. The salient ROI may be associated with a salient object amongst the one or more objects. In an embodiment, the method 1200 includes identifying, in the received image frame 418, a first area corresponding to a salient object among the one or more objects.

[0155] At step 1206, the method 1200 includes generating the HOG map 602 based on the image frame 418 and the identified salient ROI. The HOG map indicates the distribution and orientation of edge intensities associated with the salient object in the image frame 418. In an embodiment, the method 1200 includes generating a gradient map based on the received image frame and the first area,wherein the gradient map indicates a distribution and orientation of edge intensities corresponding to the salient object in the received image frame.

[0156] At step 1208, the method 1200 includes estimating, using the set of models, the speed parameter and the direction parameter associated with the salient object based on the image frame 418, the identified salient ROI, and the HOG map 602. For estimating the speed parameter and the direction parameter, the method 1200 includes generating using the neural network model 702 amongst the set of models, the feature representation maps based on the image frame 418, the salient ROI, and the HOG map 602. The feature representation maps may indicate the motion cues and direction cues associated with the salient object. The method 1200 includes estimating, using the first classification model amongst the set of models, the speed parameter based on the generated feature representation maps indicating the motion cues. The method 1200 includes estimating, using the second classification model amongst the set of models, the direction parameter based on the generated feature representation maps indicating the direction cues.

[0157] In an embodiment, the method 1200 includes estimating, using a set of models, a speed parameter and a direction parameter corresponding to the salient object based on the received image frame, the first area, and the gradient map. For estimating the speed parameter and the direction parameter, the method 1200 includes generating, using a neural network model among the set of models, feature representation maps based on the received image frame, the first area, and the gradient map. The feature representation maps comprising information on motion cues and information on direction cues corresponding to the salient object. The method 1200 includes estimating, using a first classification model among the set of models, the speed parameter based on the generated feature representation maps comprising information on the motion cues. The method 1200 includes estimating, using a second classification model among the set of models, the direction parameter based on the generated feature representation maps comprising information on the direction cues.

[0158] Further, for estimating the speed parameter using the first classification model, the method 1200 includes estimating the speed of the salient object in the image frame 418. The method 1200 includes classifying the speed of the salient objectin at least one of the set of speed class labels. The set of speed class labels include the slow-motion label, fast motion label, medium motion label, and still label. The method 1200 includes determining the speed parameter based on the classified speed.

[0159] Further, for estimating the direction parameter using the first classification model, the method 1200 includes estimating the direction of motion of the salient object in the image frame 418. The method 1200 includes classifying the direction of motion of the salient object in at least one of the set of direction class labels. The set of direction class labels includes angles among the set of predefined directional angles. The method 1200 includes determining the direction parameter based on the classified direction of motion.

[0160] Furthermore, the neural network model 702, the first classification module, and the second classification module may be trained on the one or more reference frames, the one or more salient object ROI, and the one or more reference HOG maps. In an embodiment, the neural network model, the first classification model, and the second classification model are trained on one or more reference frames, one or more reference salient areas, and one or more reference gradient maps. For training the neural network, the method 1200 includes constructing the future frame based on the feature representation maps generated by the neural network model 702. The future frame indicates future motion predictions associated with the one or more reference frames. The method 1200 includes backpropagating the error associated with the future frame for training of the neural network model 702.

[0161] For training the neural network model 702, the first classification model, and the second classification model, the method 1200 includes training the neural network model 702 to generate the feature representation maps by providing the one or more reference frames, one or more reference salient ROI, and one or more reference HOG maps as inputs. The method 1200 includes training the first classification model and the second classification model to determine the speed parameter and the direction parameter by providing the feature representation maps as inputs. The method 1200 includes backpropagating the error associated with the first classification model and the second classification model for training of the neural network model 702.

[0162] In an embodiment, the method 1200 includes generating the panning image by applying a blur to a second area corresponding to the background portion of the received image frame based on the estimated speed parameter and the estimated direction parameter.

[0163] At step 1210, the method 1200 includes constructing the blur kernel based on the estimated speed parameter and the estimated direction parameter of the salient object, upon the determination that the estimated speed may be greater than the threshold. The size of the blur kernel may be defined based on the direction parameter.

[0164] At step 1212, the method 1200 includes generating the panning image by applying the blur to the background portion of the received image frame 418 using the constructed blur kernel. For generating the panning image 434, the method 1200 includes performing the convolution operation between the received image frame and the constructed blur kernel. The point spread factor associated with the blur kernel may be defined based on the direction parameter.

[0165] Further, the use cases of the system 204 and the method 1200 are provided in the subsequent table:

[0166] Table 1

[0167] Referring to use cases, additionally, the system 204 and the method 1200 as disclosed may be used in the different implementations, for example, motion phone (photo mode), single take (photo mode), gallery, portrait etc., without departing from the scope of the present disclosure.

[0168] As would be gathered, the system 204 and the method 1200 as disclosed provide a lightweight model ensuring a comprehensive approach to generate the panning image 434. The present configuration provides the single-frame-based, completely automated, end-to-end system 204 that gives stunning action-panning image based on the gallery image under any setting of parameters of the camera, for example, shutter speed, ISO, and autofocus. Further, the modules ensure obtaining the motion speed and motion direction of the salient object by using only the single frame / single image, that has the capability of understanding the characteristics of the salient object including structural and moving edges. Further, the design and application of the salient object segment, and the single frame-based speed / direction computation to identify the action of the sharp foreground subject-in-motion ensures the generation of the panning image 434. The filter 430 and the filter 432 guided by the foreground object’s speed and direction of motion, are applied to create blurry and streaky backgrounds. The streaky backgrounds introduced by the filter 432 elevate the object’s speed and direction of movement through its streakiness (its width and length), such that the final output photo is at par with professional shots.

[0169] The present system 204 and the method 1200 generate the panning image 434 while eliminating manual overload, expertise dependency, overlooked moments, and imprecise estimation, thus seamlessly simplifying and improving the art of panning photography without the need for specialized hardware or demanding user exertion.

[0170] While specific language has been used to describe the present disclosure, any limitations arising on account thereto, are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein. The drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elementsmay be split into multiple functional elements. Elements from one embodiment may be added to another embodiment.

Claims

Claims

1. A method for generating a panning image, the method comprising: receiving an image frame comprising one or more objects and a background portion; identifying, in the received image frame, a first area corresponding to a salient object among the one or more objects; generating a gradient map based on the received image frame and the first area, wherein the gradient map indicates a distribution and orientation of edge intensities corresponding to the salient object in the received image frame; estimating, using a set of models, a speed parameter and a direction parameter corresponding to the salient object based on the received image frame, the first area, and the gradient map; and generating the panning image by applying a blur to a second area corresponding to the background portion of the received image frame based on the estimated speed parameter and the estimated direction parameter.

2. The method as claimed in claim 1 , wherein estimating the speed parameter and the direction parameter comprises: generating, using a neural network model among the set of models, feature representation maps based on the received image frame, the first area, and the gradient map, the feature representation maps comprising information on motion cues and information on direction cues corresponding to the salient object; estimating, using a first classification model among the set of models, the speed parameter based on the generated feature representation maps comprising information on the motion cues; and estimating, using a second classification model among the set of models, the direction parameter based on the generated feature representation maps comprising information on the direction cues.

3. The method as claimed in claim 2, wherein estimating the speed parameter using the first classification model comprises: estimating a speed of the salient object in the received image frame; classifying the speed of the salient object in at least one of a set of speed class labels, wherein the set of speed class labels includes a slow-motion label, fast motion label, medium motion label, and still label; and determining the speed parameter based on the classified speed.

4. The method as claimed in claim 2, wherein estimating the direction parameter using the second classification model comprises: estimating a direction of motion of the salient object in the received image frame; classifying the direction of motion of the salient object in at least one of a set of direction class labels, wherein the set of direction class labels includes angles among a set of pre-defined directional angles; and determining the direction parameter based on the classified direction of motion.

5. The method as claimed in claim 2, wherein the neural network model, the first classification model, and the second classification model are trained on one or more reference frames, one or more reference salient areas, and one or more reference gradient maps, and wherein training the neural network model comprises: constructing a future frame based on the future representation maps generated by the neural network model, wherein the future frame indicates future motion predictions associated with the one or more reference frames; andbackpropagating an error associated with the future frame for the training of the neural network model.

6. The method as claimed in claim 5, wherein training the neural network model, the first classification model, and the second classification model comprises: training the neural network model to generate the feature representation maps by providing the one or more reference frames, one or more reference salient areas, and one or more reference gradient maps as inputs; training the first classification model and the second classification model to determine the speed parameter and the direction parameter by providing the feature representation maps as inputs; and backpropagating an error associated with the first classification model and the second classification model for training the neural network model.

7. The method as claimed in claim 1 , wherein generating the panning image comprises: constructing a blur kernel based on the estimated speed parameter and the estimated direction parameter corresponding to the salient object, based on the estimated speed parameter being greater than a threshold; and generating the panning image by applying a blur to the second area corresponding to a second area corresponding to the background portion of the received image frame using the blur kernel.

8. The method as claimed in claim 1 , wherein generating the panning image comprises performing a convolution operation between the received image frame and the constructed blur kernel, wherein a point spread factor (PSF) associated with the blur kernel is defined based on the direction parameter.

9. The method as claimed in claim 7, wherein a size of the blur kernel is proportional to the speed parameter, and wherein an angle to apply the blur is defined based on the direction parameter.

10. An electronic device for generating a panning image, the electronic device comprising: a memory; and a processor communicatively coupled with the memory, the processor being configured to: receive an image frame comprising one or more objects and a background portion; identify, in the received image frame, a first area corresponding to a salient object among the one or more objects; generate a gradient map based on the received image frame and the fist area, the gradient map indicates a distribution and orientation of edge intensities corresponding to the salient object in the received image frame; estimate, using a set of models, a speed parameter and a direction parameter corresponding to the salient object based on the received image frame, the fist area, and the gradient map; and generate the panning image by applying a blur to a second area corresponding to the background portion of the received image frame based on the estimated speed parameter and the estimated direction parameter.

11. The electronic device as claimed in claim 10, wherein to estimate the speed parameter and the direction parameter, the processor is configured to:generate, using a neural network model among the set of models, feature representation maps based on the received image frame, the first area, and the gradient map, the future representation maps comprising information on motion cues and information on direction cues corresponding to the salient object; estimate, using a first classification model among the set of models, the speed parameter based on the generated feature representation maps comprising information on the motion cues; and estimate, using a second classification model among the set of models, the direction parameter based on the generated feature representation maps comprising information on the direction cues.

12. The electronic device as claimed in claim 11 , wherein to estimate the speed parameter using the first classification model, the processor is configured to: estimate a speed of the salient object in the received image frame; classify the speed of the salient object in at least one of a set of speed class labels, wherein the set of speed class labels includes slow-motion label, fast motion label, medium motion label, and still label; and determine the speed parameter based on the classified speed.

13. The electronic device as claimed in claim 11 , wherein to estimate the direction parameter using the second classification model, the processor is configured to: estimate a direction of motion of the salient object in the received image frame; classify the direction of motion of the salient object in at least one of a set of direction class labels, wherein the set of direction class labels includes angles among a set of pre-defined directional angles; and determine the direction parameter based on the classified direction of motion.

14. The electronic device as claimed in claim 11 , wherein the neural network model, the first classification model, and the second classification model are trained on one or more reference frames, one or more reference salient areas, and one or more reference gradient maps, and wherein to train the neural network model, the processor is configured to: construct a future frame based on the future representation maps generated by the neural network model, wherein the future frame indicates future motion predictions associated with the one or more reference frames; and backpropagate an error associated with the future frame for the training of the neural network model.

15. The electronic device as claimed in claim 14, wherein to train the neural network model, the first classification model, and the second classification model, the processor is configured to: train the neural network model to generate the feature representation maps by providing the one or more reference frames, one or more reference salient areas, and one or more reference gradient maps as inputs; train the first classification model and the second classification model to determine the speed parameter and the direction parameter by providing the feature representation maps as inputs; and backpropagate an error associated with the first classification model and the second classification model for training of the neural network model.

Citation Information

Patent Citations

  • Image processing method, apparatus and program thereof

    JP2006050070A

  • Electronic camera

    JP2006080844A

  • Modification of Post-Viewing Parameters for Digital Images Using Image Region or Feature Information

    US20090179998A1

  • Method for obtaining panning shot image and electronic device supporting the same

    US20180176470A1

  • Systems and methods for generating panning images

    US20220044376A1