A method and a device for capturing a desired event
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2026-08-13
Smart Images

Figure KR2025005509_13082026_PF_FP_ABST
Abstract
Description
A METHOD AND A DEVICE FOR CAPTURING A DESIRED EVENT
[0001] The disclosure relates to the field of camera control. More particularly, the disclosure relates to a method and a device for capturing a desired event.
[0002] Camera control in user devices, such as smartphones, has become highly restrictive and relies significantly on manually set fixed modes that limit the user's ability to customize or dynamically interact with the camera. Typically, users are offered a limited range of pre-set controls, such as a timer with options for less than a 10-second delay, predefined frame-per-second (FPS) modes such as a normal mode, a slow motion mode, and an ultra slow motion mode, and predefined gesture-based triggers like palm or smile poses. While the pre-set controls are useful for specific scenarios, such controls often fall short when users want to capture a "perfect shot" of a unique moment or pose. For example, when trying to photograph a jump in mid-air, users often need to take multiple burst-mode photos or record a long video to later apply post-processing effects, like slow motion, in order to capture a desired image.
[0003] Figure 1A and Figure 1B illustrate a problem scenario associated with capturing a specific moment, in accordance with an existing art. As shown in Figure 1A, a user with a smartphone 102 having an integrated image-capturing device (e.g., a camera) wants to take a mid-air photo of himself. To achieve this, the user records a video of himself jumping, as shown in Figure 1A. Thereafter, as shown in Figure 1B, the user plays the video on the smartphone 102, searches for a desired frame (e.g., a frame in which the user is in mid-air), and then the user takes a screenshot of the desired frame. If the user does not get the desired frame in the video, the user may be required to keep repeating the operations 102 and 104 until the desired frame is captured. An approach of recording the video and taking screenshot of the desired frame from the video, while functional, is highly inefficient as it consumes excessive processing power, drains the smartphone's battery, and fills the smartphone's storage with unnecessary files.
[0004] However, the users need interactively control the camera in user devices (e.g., smartphones) in a more dynamic way so that the user experience becomes enhanced without wasting the resources of the user devices.
[0005] Therefore, it is required to perform the automatic capturing of specific moments or specific poses.
[0006] This is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the disclosure. This summary is neither intended to identify key or essential inventive concepts of the disclosure nor is it intended for determining the scope of the disclosure.
[0007] According to an embodiment, a device for capturing a desired event is disclosed. The device includes a camera, at least one processor, memory storing computer program code, where the memory and the computer program code are configured, with the at least one processor, to cause the apparatus to obtain a user prompt indicating the desired event to be captured in a media stream. In an embodiment, the device is caused to generate first visual features corresponding to the user prompt. In an embodiment, the device is caused to generate second visual features from the media stream. In an embodiment, the device is caused to correlate the first visual features with the second visual features of a current frame in the media stream. In an embodiment, the device is caused to determine, based on the correlation, whether the current frame includes the desired event. In an embodiment, the device is caused to control the camera to capture the current frame at a target FPS(frame per second) based on the determination.
[0008] According to an embodiment, a method of capturing a desired event is provided. In an embodiment, the method includes obtaining a user prompt indicating the desired event to be captured in a media stream. In an embodiment, the method includes generating first visual features corresponding to the user prompt. In an embodiment, the method includes generating second visual features from the media stream. In an embodiment, the method includes correlating the first visual features with the second visual features of a current frame in the media stream. In an embodiment, the method includes determining, based on the correlation, whether the current frame includes the desired event. In an embodiment, the method includes capturing, by a camera, the current frame at a target FPS based on the determination.
[0009] The foregoing and other features of embodiments will become more apparent from the following detailed description of embodiments when read in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements.
[0010] Figure 1A and Figure 1B illustrate a scenario associated with capturing a specific moment.
[0011] Figure 2 illustrates a schematic block diagram of a system for capturing a desired event, in accordance with an embodiment of the disclosure.
[0012] Figure 3 illustrates a schematic block diagram depicting an operational flow for capturing the desired event, in accordance with an embodiment of the disclosure.
[0013] Figure 4 illustrates a block diagram of a scene-aware multimodal (SAM) encoding module, in accordance with an embodiment of the disclosure.
[0014] Figure 5 illustrates a block diagram of a temporally aligned frame predictor (TAFP) module, in accordance with an embodiment of the disclosure.
[0015] Figure 6 illustrates a block diagram of a click classifier module, in accordance with an embodiment of the disclosure.
[0016] Figure 7 illustrates a block diagram of a frame-estimator (F-estimator) module, in accordance with an embodiment of the disclosure.
[0017] Figure 8 illustrates a block diagram of a frame-per-second (FPS) predictor module, in accordance with an embodiment of the disclosure.
[0018] Figure 9A and Figure 9B illustrate a flow chart of capturing the desired event, in accordance with an embodiment of the disclosure.
[0019] Figure 10 illustrates an exemplary scenario associated with capturing the desired event, in accordance with an embodiment of the disclosure.
[0020] Figure 11 depicts an exemplary scenario associated with modification of a timer value of a camera when a user wants to capture an image using a camera timer, in accordance with an embodiment of the disclosure.
[0021] Figure 12 depicts an exemplary scenario associated with capturing a slow motion video at variable frame rates, in accordance with an embodiment of the disclosure.
[0022] Figure 13 depicts an exemplary scenario associated with capturing a desired pose, in accordance with an embodiment of the disclosure.
[0023] Figure 14 depicts an exemplary scenario associated with capturing multiple images in multiple poses, in accordance with an embodiment of the disclosure. and
[0024] Figure 15 depicts an exemplary scenario associated with capturing a pose similar to a person in a reference image that features multiple people, in accordance with an embodiment of the disclosure.
[0025] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help improve understanding of aspects of the present invention. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
[0026] For the purpose of promoting an understanding of the principles of the disclosure, reference will now be made to an embodiment and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the disclosure as illustrated therein being contemplated as would normally occur to one skilled in the art to which the disclosure relates.
[0027] Whether or not a certain feature or element was limited to being used only once, it may still be referred to as "one or more features" or "one or more elements" or "at least one feature" or "at least one element." Furthermore, the use of the terms "one or more" or "at least one" feature or element do not preclude there being none of that feature or element, unless otherwise specified by limiting language including, but not limited to, "there needs to be one or more" or "one or more elements is required."
[0028] Reference is made herein to an embodiment. It should be understood that an embodiment is an example of a possible implementation of any features and / or elements of the disclosure. Some embodiments have been described for the purpose of explaining one or more of the potential ways in which the specific features and / or elements of the proposed disclosure fulfil the requirements of uniqueness, utility, and non-obviousness.
[0029] The embodiments described in this disclosure can be combined and merged with each other.
[0030] Any particular and all details set forth herein are used in the context of an embodiment and therefore should not necessarily be taken as limiting factors to the proposed disclosure.
[0031] The terms "comprises", "comprising" or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components proceeded by "comprises... a" does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.
[0032] Embodiments of the disclosure will be described below in detail with reference to the accompanying drawings.
[0033] An objective of the disclosure is to provide techniques of automatically capturing of precise moments or specific poses using camera-equipped devices, thereby eliminating the need for manual video recording and then taking screenshots, and ensuring high-quality images while optimizing battery and storage usage.
[0034] The disclosure achieves the above-mentioned objective by providing a method and a device for capturing a desired event, as described below in the forthcoming paragraphs with reference to the accompanying drawings. The desired event may refer to a specific pose or a precise moment. For example, the desired event may include a user sitting in a meditation pose, the user inverted in the air, etc. Particularly, the disclosure provides a method for predicting, for a desired user pose, from live preview frames when the desired user pose is expected to happen in a future frame. The disclosure also provides a method for capturing the future frame where the desired user pose will happen, and estimating an optimum frame rate to which a camera needs to be changed. The optimum frame rate is estimated based, for example, on a latency expected in a video frame encoding. According to an embodiment of the disclosure, the future frame to be captured is identified by matching visual features in the live preview frames with a visual features embedding associated with the desired user pose. The method are described in detail in the following.
[0035] Figure 2 illustrates a schematic block diagram of a device 200 for capturing a desired event, in accordance with an embodiment of the disclosure. The device 200 may include at least one processor 202, memory 204, a user interface 230, a camera 240 and one or more modules 206. In an embodiment, the device 200 may be implemented in a camera-equipped user device (for example, a smartphone). Throughout the disclosure, the one or more modules 206 may be part of the at least one processor 202, or functions performed by the one or more modules 206 may be performed by the at least one processor 202.
[0036] In an embodiment, the at least one processor 202 (hereinafter referred to as "processor 202") may be operatively coupled to the memory 204 and the one or more modules 206. The processor 202 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc. In an embodiment, the processor 202 may include a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), or both. The processor 202 may be one or more general processors, Digital Signal Processors (DSPs), application-specific integrated circuits, Field-Programmable Gate Arrays (FPGAs), servers, networks, digital circuits, analog circuits, combinations thereof, or other now known or later developed devices for analyzing and processing data. The processor 202 may execute a software program, such as code generated manually (i.e., programmed) to perform the desired operation. The processor 202 may implement various techniques such as, but not limited to, data extraction, Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and so forth to achieve the desired objective. The processor 202 may include various hardware processing circuitry and / or multiple processors. For example, as used herein, including the claims, the term "processor" may include various processing circuitry, including at least one processor, wherein one or more of at least one processor, individually and / or collectively in a distributed manner, may be configured to perform various functions described herein. As used herein, when "a processor", "at least one processor", and "one or more processors" are described as being configured to perform numerous functions, these terms cover situations, for example and without limitation, in which one processor performs some of recited functions and another processor(s) performs other of recited functions, and also situations in which a single processor may perform all recited functions. Additionally, the at least one processor may include a combination of processors performing various of the recited / disclosed functions, e.g., in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.
[0037] The memory 204 may store data and instructions executable by the processor 202. In an embodiment, the memory 204 may communicate via a bus within the device 200. The memory 204 may include, but is not limited to, a non-transitory computer-readable storage media, such as various types of volatile and non-volatile storage media including, but not limited to, random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media and the like. For example, the memory 204 may include a cache or random-access memory for the processor 202. In alternative examples, the memory 204 may be separate from the processor 202, such as a cache memory of a processor 202, the system memory, or other memory. The memory 204 may be an external storage device or database for storing data. The memory 204 may be operable to store instructions executable by the processor 202. The functions, acts, or tasks illustrated in the figures or described may be performed by the processor 202 for executing the instructions stored in the memory 204. The functions, acts, or tasks are independent of the particular type of instruction set, storage media, processor, or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code, and the like, operating alone or in combination. Likewise, processing strategies may include multiprocessing, multitasking, parallel processing, and the like. The memory 204 may further include a database to store the data. Further, the memory 204 may include an operating system for performing one or more tasks of the device 200, as performed by a generic operating system in the communications domain.
[0038] The user interface 230 may include an output interface and an input interface. The output interface is for outputting an audio signal or a video signal, and may include a display unit, an audio output unit, and the like. In a case in which the display unit and a touch pad form a layered structure to be configured as a touch screen, the display unit may be used as an input interface in addition to an output interface. The display may include at least one of a liquid-crystal display, a thin-film-transistor liquid-crystal display, a light-emitting diode (LED) display, an organic LED display, a flexible display, a three-dimensional (3D) display, and an electrophoretic display. The input interface is for receiving an input from the user. The input interface may be, but is not limited to, at least one of a key pad, a tact switch, a touch pad (e.g., a touch-type capacitive touch pad, a pressure-type resistive overlay touch pad, an infrared sensor-type touch pad, a surface acoustic wave conduction touch pad, an integration-type tension measurement touch pad, a piezoelectric effect-type touch pad), a jog wheel, and a jog switch. The input interface 2520 may include a speech recognition module.
[0039] The camera 240 may include a capturing device. The camera 240 may include any image capturing device capable of capturing still images and / or recording videos. The camera 240 may be implemented as a standalone device or may be integrated into various electronic devices such as smartphones, tablets, laptops, wearable devices, or home appliances. The camera 240 may include one or more image sensors, lenses, and related circuitry for acquiring image data. In an embodiment, the camera 240 may capture a single frame (still image) and / or continuously capture a sequence of frames (video) based on user input or automated control. The term "capturing" as used herein encompasses both taking photographs and recording videos unless otherwise specified.
[0040] The one or more modules 206, among other things, may include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement data types. The one or more modules 206 may also be implemented as, (signal) processor(s), state machine(s), logic circuitries, and / or any other device or component that manipulates signals based on operational instructions. Further, the one or more modules 206 can be implemented in hardware, instructions executed by a processing unit, or by a combination thereof. The processing unit can comprise a computer, the processor 202, a state machine, a logic array, or any other suitable devices capable of processing instructions. The processing unit can be a general-purpose processor that executes instructions to cause the general-purpose processor to perform the required tasks, or the processing unit can be dedicated to performing the required functions. In an embodiment of the disclosure, the one or more modules 206 may be machine-readable instructions (software) which, when executed by the processor 202 / processing unit, perform any of the described functionalities. Further, the data serves, among other things, as a repository for storing data processed, received, and generated by one or more of the modules 206.
[0041] The one or more modules 206 may include a set of instructions that may be executed to cause the device 200 to capture the desired event. Throughout the disclosure, capturing the desired event or capturing a frame including the desired event may include capturing a scene including the desired event. The one or more modules 206 may include a text tokenizer module 208, an image encoding module 210, a scene-aware multimodal (SAM) encoding module 212, a camera frame-per-second (FPS) controller module 214, a video frame encoding module 216, a temporally aligned frame predictor (TAFP) module 218, a classifier module 220, a frame-estimator (F-estimator) module 222, a camera timer controller module 224, and a FPS predictor module 226. Various operations performed by each of the one or more modules 206 in communication with each other are described in the forthcoming paragraphs in conjunction with Figure 3.
[0042] In an embodiment, the processor 202 may be configured to receive a user prompt via a user interface 230. The user prompt includes a user input indicating the desired event intended to be captured in a media stream. The media stream may include a real-time media stream. Further, the user prompt may include a user input with a form of one or more text command, an image, or a timer value. The user prompt may include the user input with a combination of one or more text command, an image and a timer value In an embodiment, the media stream may be captured by a camera. The media stream may be captured by a camera when a user activates the camera 240 on the device 200. The processor 202 may generate first visual features corresponding to the user prompt. The first visual features represent the desired event corresponding to the user prompt. The desired event may be described in the user prompt. The processor 202 may generate second visual features corresponding to the media stream. In an embodiment, the second visual features may include a numerical representation of frames of the media stream. The processor 202 may correlate the first visual features with the second visual features of the media stream. In an embodiment, correlating the features may include comparing the features and / or associating the features. In an embodiment, the processor 202 may determine, as correlating the two visual features, whether the second visual features in the media stream may correspond to the first visual features corresponding to the user prompt. For example, if the user prompt includes "capture when I jump up in mid-air", the processor 202 may determine whether the a jumping pose, as second visual features, obtained from the real-time media stream corresponds to the first visual features derived from the user prompt which includes "capture when I jump up in mid-air". In an embodiment, the processor 202 may correlate the first visual features with the second visual features of a current frame in the media stream. The processor 202 may determine, by the classifier module 220, whether the current frame represents the desired event based on the correlation. The desired event may include, but not limited to, at least one of a specific pose and / or a precise moment associated with, for example, an action. In response to the determination that the current frame is not representing the desired event, the processor 202 may predict an estimated time until the desired event appears in the media stream. The prediction may be performed by the F-estimator module 222. In response to the determination that the current frame or several frames to come is(will) not contain(ing) the desired event, the processor 202 may predict or determine, a target FPS for the media stream based on the prediction. In response to the determination that the current frame or several frames to come is(will) not represent(ing) the desired event, the processor 202 may apply the target FPS to a camera for capturing the desired event in the media stream. Alternatively, in response to the determination that the current frame represents the desired event, the processor 202 may control the camera 240 to capture the current frame. The target FPS may indicate a required frame rate associated with the camera 240 for capturing the desired event in the real-time media stream.
[0043] In an embodiment, the processor 202 may dynamically modify the target FPS based on the prediction. For instance, when the desired event is far-off from the current frame in the real-time media stream, the target FPS may become lower to optimize usage of the user device's resources. For instance, it is predicted that the desired event may be at least 3 seconds away the current frame in the real-time media stream, the target FPS may become lower. The above-noted "3 seconds" is merely an example and thus, the first predetermined time (ex) 3 seconds as shown above) indicating how far-off from the current frame may be changed according to a situation and / or a user setting. Alternatively, when the desired event is close to - for example, 500 ms away from - the current frame, the target FPS may become higher to ensure successful capturing of the desired event. Likewise, the second predetermined time (ex) 500 ms) indicating how close to the current frame may be changed according to a situation and / or a user setting. The modification of the target FPS is explained in detail later in the disclosure.
[0044] In an embodiment, the processor 202 may determine whether the predicted estimated time exceeds a predefined value. Thereafter, the processor 202 may transmit a stop instruction to the camera 240 in response to determining that the predicted estimated time exceeds a predefined value.
[0045] In an embodiment, after predicting the estimated time, the processor 202 may adjust a timer associated with the camera 240 based on the estimated time and a remaining timer value. The remaining timer value may refer to a time left on the timer when the time is adjusted.
[0046] In an embodiment, the processor 202 may generate an embedding based on the first visual features and / or the second visual features for correlating the first visual features with the second visual features of the current frame. The embedding may indicate an alignment between the current frame and at least one frame including the desired event. The alignment may indicate how close between the current frame and the at least one frame including the desired event.
[0047] In an embodiment, the processor 202 may predict, based on the embedding, whether to capture the current frame.
[0048] Figure 3 illustrates a schematic block diagram 300 depicting an operational flow for capturing the desired event, in accordance with an embodiment of the disclosure.
[0049] Referring to Figure 3, the text tokenizer module 208 and the image encoding module 210 may receive the user prompt. The text tokenizer module 208 may receive the text command. For example, the user prompt including the text command may be "Click my photo when I am mid-air". The text tokenizer module 208 may be a pre-trained text tokenizer configured to tokenize each word or phrase of the text command to generate a tokenized text command. The tokenized text command may help the device 200 for further processing (e.g., for understanding the desired event). Optionally, along with the text tokenizer module 208 obtaining the text command, the image encoding module 210 may obtain an image for reference("a reference image" hereinafter). When the text command and the reference image are obtained together, at least one text included in the textual command may refer to at least one object included in the reference image. For example, the text command may be "Capture my photo when I am in the air, like the man on left, with both my legs folded and hands straight up in the air", and the reference image may be an image showing a man as an object in air with both legs folded and hands straight up in the air, the man may be present at the left in the image. The image encoding module 210 may be a pre-trained image encoder configured to generate visual embeddings based on the reference image.
[0050] In an embodiment, the SAM encoding module 212 may generate the first visual features corresponding to the received user prompt. The SAM encoding module 212 may generate the first visual features based on the tokenized text command and the visual embeddings. The SAM encoding module 212 may use the tokenized text command and the visual embeddings to understand the text command and the reference image respectively to generate the first visual features representing the desired event included in the received user prompt. The first visual features may be scene-aware multimodal encodings representing the text command as well as the reference image. The SAM encoding module 212 is explained in detail in conjunction with Figure 4.
[0051] Figure 4 illustrates a block diagram 400 of the SAM encoding module 212, in accordance with an embodiment of the disclosure. Referring to Figure 4, the SAM encoding module 212 may receive the tokenized text command from the text tokenizer module 208 and the visual embeddings from the image encoding module 210 as inputs. Thereafter, one or more first embeddings 402 may be added to the tokenized text command to generate a first representation. The one or more first embeddings 402 may include token embeddings, positional embeddings, and text reference embeddings. Additionally and optionally, one or more second embeddings 404 may be added to the visual embedding to generate a second representation. The one or more second embeddings 404 may include patch embeddings, the positional embeddings, and image reference embeddings. The first representation and the second representation may then be combined to form a combined representation, which may be fed into the encoder layer1 406-1 for further processing. The combined representation may include one or more tokens.
[0052] In an embodiment, the SAM encoding module 212 may include a plurality of encoder layers. In an embodiment, the SAM encoding module 212 may include three encoder layers 406-1, 406-2, and 406-3 for simplicity, however, any number of encoder layers may be used for the SAM encoding module 212 based on specific requirements. The SAM encoding module 212 may have a transformer-based architecture, with each encoder layer 406-1, 406-2, or 406-3 representing a transformer block. As shown in Figure 4, each encoder layer 406-1, 406-2, or 406-3 may include a multi-headed attention layer 408, which may take the combined representation as input and generate an output representing relationships and dependencies between different tokens in the combined representation. Each of the encoder layers 406-1, 406-2, or 406-3 may also include a first Add & Norm layer 410, where the output of the multi-headed attention layer 408 is passed through a residual connection (i.e., the output of the multi-headed attention layer 408 is added to the input), followed by layer normalization. Each of the encoder layers 406-1, 406-2, or 406-3 may also include a feedforward neural network (FFN) layer 412, where each token representation is processed independently. Furthermore, after the FFN layer 412. Also, each of the encoder layers 406-1, 406-2, or 406-3 may include a second Add & Norm layer 414, where another residual connection is added, followed by the layer normalization.
[0053] In an embodiment, the SAM encoding module 212 may be a pre-trained model for generating the scene-aware multimodal encodings. The SAM encoding module 212 may be trained such that each of the encoder layers 406-1, 406-2, or 406-3 may extract information associated with the desired event. For example, the first encoder layer 406-1 may extract information related to a simple scene understanding associated with the desired event. The simple scene understanding may refer to a basic spatial relationship between key elements (e.g., scenes or actions) in the reference image. For example, the simple scene understanding may include information about outlines, shapes, individual elements (e.g., sky, road, etc.), brightness, etc. associated with the reference image. Similarly, the second encoder layer 406-2 may extract information related to complex scenes involving multi-object and contextual relationships. For example, the complex scenes may include human actions, object interactions, movements, etc. Finally, the third encoder layer 406-3 may aggregate and refine the information extracted by the previous encoder layers 406-1 and 406-2 to generate the scene-aware multimodal encodings (e.g., the first visual features). The SAM encoding module 212 may be progressively trained by taking outputs of intermediate encoder layers 406-1 and 406-2 to condition the TAFP module 218 based on the extracted information, as explained in detail in the forthcoming paragraphs.
[0054] Referring to Figure 3, in an embodiment, upon receiving the user prompt, one or more sensors associated with the camera 240 installed in the device 200 may be activated and may initiate the real-time media stream of an environment surrounding the camera 240. The camera FPS controller module 214 may be connected to the camera 240, and configured to set a required frame rate (i.e., the FPS) of the camera 240 at which the camera 240 should be processing. Initially, the frame rate may be set to a default FPS based on a camera mode. Thereafter, the frame rate may be adjusted based on an output of the FPS predictor module 226, as explained in detail in the forthcoming paragraphs.
[0055] In an embodiment, the video frame encoding module 216 may generate the second visual features corresponding to the real-time media stream. In an embodiment, the video frame encoding module 216 may encode each frame of the real-time media stream. The video frame encoding module 216 may be a pre-trained video frame encoder.
[0056] In an embodiment, the TAFP module 218 may take the first visual features from the SAM encoding module 212 and the second visual features from the video frame encoding module 216 as inputs. The TAFP module 218 may correlate the first visual features with the second visual features of the current frame in the real-time media stream. In an embodiment, for correlating the first visual features with the second visual features of the current frame, the TAFP module 218 may generate an embedding (i.e., a TAFP embedding) based on the first visual features and the second visual features. The embedding may indicate an alignment between the current frame and the frame including the desired event. The TAFP module 218 is explained in detail in conjunction with Figure 5 in the forthcoming paragraphs.
[0057] Figure 5 illustrates a block diagram 500 of the TAFP module 218, in accordance with an embodiment of the disclosure. The TAFP module 218 may include two parts i.e., a long short-term memory (LSTM) model 506 and a transformer decoder 503.
[0058] In an embodiment, the LSTM model 506 may be configured to continuously analyze the real-time media stream to understand how a scene in the real-time media stream evolves. The LSTM model 506 may take video frame encodings (i.e., the second visual features) of the current frame and a previous frame as inputs and process the video frame encodings to generate an output that represents a scene evolution (i.e., a current scene prediction). Particularly, the LSTM model 506 may capture one or more temporal dependencies between consecutive frames of the real-time media stream, thereby ensuring smooth transitions. The LSTM model 506 maintains a memory of the scene evolution for better prediction.
[0059] In an embodiment, the transformer decoder 503 may align the current frame with the frame including the desired event. The transformer decoder 503 may include a plurality of decoder layers. Although Figure 5 shows three decoder layers 504-1, 504-2, and 504-3 for simplicity, any number of decoder layers may be used based on specific requirements. In an embodiment, the transformer decoder 503 may have a transformer-based architecture, with each decoder layer 504-1, 504-2, and 504-3 representing a transformer block. Each of the decoder layers 504-1, 504-2, or 504-3 includes a multi-headed attention layer 508, a first Add & Norm layer 510, the FFN layer 512, and a second Add & Norm layer 514. Since the layers have already been explained in conjunction with Figure 4, a description thereof will be omitted in the context of Figure 5 for the sake of brevity.
[0060] In an embodiment, the output of the LSTM model 506 may be added with the positional embeddings and fed to the first decoder layer 504-1. The outputs of the intermediate encoding layers of the SAM encoding module 212 may be used to condition the TAFP module 218. For example, the output of the LSTM model 506 may undergo cross-attention with an output of the first encoder layer 406-1 of the SAM encoding module 212. The cross-attention may ensure that the first decoder layer 504-1 focuses on same information related to the simple scene understanding associated with the desired event as the first encoder layer 406-1. Similarly, the output of the first decoder layer 504-1 may undergo cross-attention with an output of the second encoder layer 406-2 of the SAM encoding module 212. The cross-attention ensures that the current scene prediction aligns closely with the desired event. Finally, the third decoder layer 504-3 generates the embedding (i.e., the TAFP embedding) indicating the alignment between the current frame and the frame including the desired event. In other words, the TAFP embedding represents an optimal frame for a capture based on scene alignment.
[0061] In an embodiment, the LSTM model 506 and the transformer decoder 503 may be trained jointly to minimize the difference between the current scene prediction and the desired event. The joint training also ensures optimization for both temporal coherence and scene similarity.
[0062] Referring to Figure 3, In an embodiment, the classifier module 220 may determine whether the current frame represents the desired event based on the correlation by the TAFP module 218. In an embodiment, to determine whether the current frame represents the desired event, the classifier module 220 may predict, based on the embedding (i.e., the TAFP embedding), whether to capture the current frame. In response to determining that the current frame represents the desired event, the classifier module 220 may transmit a click instruction to the camera 240 to capture the current frame. The classifier module 220 is explained in detail in conjunction with Figure 6 in the forthcoming paragraphs.
[0063] Figure 6 illustrates a block diagram 600 of the classifier module 220, in accordance with an embodiment of the disclosure. The classifier module 220 may be a binary classifier model that uses a fully connected neural network. The classifier module 220 may take the TAFP embedding as input. Particularly, the classifier module 220 may take a classification token (CLS) embedding associated with the TAFP embedding as the input. The CLS embedding may include a summary of the entire TAFP embedding. The classifier module 220 may classify whether to capture the current frame. In an embodiment, an output of the classifier module 220 may be either 'Yes' or 'No'. If the output of the classifier module 220 is 'Yes', the classifier module 220 may transmit the click command to the camera 240 to capture the current frame. Alternatively, if the output of the classifier module 220 is 'No', the classifier module 220 may not transmit the click command to the camera 240.
[0064] Referring to Figure 3, if the current frame does not represent the desired event, the F-estimator module 222 may predict the estimated time until the desired event appears in the real-time media stream. In an embodiment, the F-estimator module 222 may determine whether the predicted estimated time exceeds the predefined value. Thereafter, the F-estimator module 222 may transmit the stop instruction to the camera in response to determining that the predicted estimated time exceeds the predefined value. The F-estimator module 222 is explained in detail in conjunction with Figure 7 in the forthcoming paragraphs.
[0065] Figure 7 illustrates a block diagram 700 of the F-estimator module 222, in accordance with an embodiment of the disclosure. The F-estimator module 222 may predict the estimated time based on the TAFP embedding and a current FPS (for example, the default FPS based on a camera mode) of the camera 240. In an embodiment, the F-estimator module 222 may be a transformer encoder including a multi-headed attention layer 702, a first Add & Norm layer 704, a FFN layer 706, and a second Add & Norm layer 708. Since the layers have already been explained in conjunction with Figure 4, a description thereof will be omitted in the context of Figure 7 for the sake of brevity.
[0066] Referring to Figure 3, in an embodiment, the camera timer controller module 224 may adjust the timer associated with the camera 240 based on the predicted estimated time and a remaining timer value. In some cases, the user prompt may include the timer value. Particularly, the camera timer controller module 224 may increase or decrease the timer value based on the predicted estimated time and the remaining timer value. For example, the user may provide the user prompt including the text command for capturing a desired image and the timer value of 5 seconds. At a specific point in time, the F-estimator module 222 may predict that the frame including a desired pose is far off (e.g., more than 5 seconds away) from the current frame, however, at that point (5 seconds away from the desired event) the timer value may be reaching 0. In such cases, the camera timer controller module 224 may modify the timer associated with the camera 240 by extending the timer value, for example, by another 5 seconds to ensure that the frame including the desired pose is captured successfully. Similarly, if the F-estimator module 222 predicts that the desired pose is near (for example, the desired pose is just 2 seconds away from the current frame), however, the remaining timer value is 3 seconds. In such a case, the camera timer controller module 224 may modify the timer associated with the camera 240 by reducing the timer value. Thus, the disclosure provides an efficient process for capturing the desired pose without wasting any memory and processing resources.
[0067] In an embodiment, the FPS predictor module 226 may predict the target FPS for the real-time media stream based on the estimated time predicted by the F-estimator module 222. In an embodiment, the target FPS indicates the required frame rate associated with the camera 240 for capturing the desired event in the real-time media stream. Further, the FPS predictor module 226 may dynamically modify the target FPS based on the prediction by the F-estimator module 222 and the remaining timer value (if any). The FPS predictor module 226 is explained in detail in conjunction with Figure 8 in the forthcoming paragraphs.
[0068] Figure 8 illustrates a block diagram 800 of the FPS predictor module 226, in accordance with an embodiment of the disclosure. In an embodiment, the FPS predictor module 226 may be a regression-based FPS predictor that uses the fully connected neural network. Referring to Figure 8, the FPS predictor module 226 may take the current FPS from the camera FPS controller module 214 as one input, and a minimum of the estimated time from the F-estimator module 222 and the remaining modified timer value from the camera timer controller module 224 as another input. Based on the inputs, the FPS predictor module 226 may generate the target FPS, which is communicated to the camera FPS controller module 214. The camera FPS controller module 214 may then set the FPS of the camera 240. In an embodiment, the camera FPS controller module 214 may reduce or increase the FPS of the camera 240 based on the user prompt, the closeness of the current frame with the frame including the desired event, and the remaining timer value (if any).
[0069] Therefore, the disclosure provides a seamless pipeline that allows the users to control the camera-equipped user devices using the user prompt. The disclosure primarily focuses on determining the desired event of which the user wants to see or obtain a photo using the user prompt and capturing the desired event using minimal memory, processing, and latency overheads.
[0070] In an exemplary scenario, a user sets a camera-equipped user device on a tripod stand and inputs a user prompt that includes a text command - "Capture a photo of me when I throw Lucy perfectly in the air." The camera starts processing, however, at this stage (let's say at time=t1), the device 200 predicts that a desired event (when the user throws Lucy in the air) is quite in the future as the user has just started walking towards Lucy. Thus, the device 200 slows down the processing of the camera considerably. That is, the FPS predictor module 226 predicts the target FPS which is very low (e.g., 2 FPS). Therefore, instead of recording and processing everything, the camera now processes at the low target FPS (e.g., 2 FPS). Thus, the disclosure ensures minimal battery usage and reduced processing power wastage. At this point, the device 200 is only trying to predict how close the user is to the desired event. Further, no frame is being stored at this point due to which no additional memory is used. After a while (let's say at time = t2), the device 200 (using the TAFP embeddings) detects that both the user and Lucy are now correctly aligned in the frame and are ready to attempt the desired event. The device 200 also detects that the user prompt possibly means that the desired event would be within a second. Thus, the device 200 (using the FPS predictor module 226 and the camera FPS controller module 214) iteratively increases the FPS of the camera, trying to detect a time of throw (let's say at time=t3). Within a time window of (t2-t3), the camera is processing at a higher FPS (for example, 15 FPS) than at time=t1, thereby ensuring detection of the start of throwing. At this stage, none of the frames are getting stored, hence there is no memory requirement, and the processing power being used, though higher than earlier processing power, is still lower than a normal real-time camera process (which typically runs at 24 FPS). At time t3, the device 200 detects that the user is ready to throw and the desired event will now be within the next couple of seconds. The device 200 (using the camera FPS controller module 214) triggers the camera to run at the fastest possible FPS (for example, >200 FPS), actively processing each frame of the real-time media stream to match the alignment with the user prompt. The camera runs at such an FPS till the device 200 detects an end of the throw (let's say at time = t4). Within the time window of t3 to t4, the camera processes each frame of the real-time media stream and stores only the current frame that represents the desired event, thereby ensuring minimal memory usage. At time t4, the camera detects an end of the throw and the user walks back to the camera and stops processing completely. In total, the camera processes at a high FPS only for a couple of seconds, and for all other time windows, the camera processes at a much slower rate, thereby optimizing battery usage even when completing complex commands. Additionally, throughout the capturing process, only a required number of frames from the real-time media stream are captured and stored. For example, in the above-mentioned exemplary scenario, the camera captures and stores only one frame, thus taking up only about 5-8MB of space.
[0071] Figure 9A and Figure 9B illustrate a flow chart of capturing the desired event, in accordance with an embodiment of the disclosure. The method 900 includes a series of operations executed by one or more components of the device 200, for example, the processor 202.
[0072] In operation 902, the device 200 may obtain a user prompt. The user prompt may indicate the desired event intended to be captured in a real-time media stream. In an embodiment, the user prompt may include one or more of the text command, the reference image, and the timer value. In an embodiment, the desired event may include a specific pose or a precise moment.
[0073] In operation 904, the device 200 may generate the first visual features corresponding to the received prompt.
[0074] In operation 906, the device 200 may generate the second visual features from the real-time media stream. In an embodiment, the real-time stream may be captured by the camera 240.
[0075] In operation 908, the device 200 may correlate the first visual features with the second visual features of the current frame in the real-time media stream. In an embodiment, the device 200 may determine, as correlating the two visual features, whether the visual features obtained from the current frame correspond to the visual features derived from the user prompt. In other words, the device 200 may determine, as correlating, whether the current frame includes, as the second visual features, a "jumping pose" which corresponds to the first visual features derived from the user prompt including "capture when I jump up in mid-air." In an embodiment, for correlating the first visual features with the second visual features of the current frame, the device 200 may generate the embedding (i.e., a TAFP embedding) based on the first visual features and the second visual features. The embedding may indicate the alignment between the current frame and one or more frames including the desired event.
[0076] In operation 910, the device 200 may determine, by the click-classifier module 220, whether the current frame represents the desired event based on the correlation. In an embodiment, for determining whether the current frame represents the desired event, the device 200 may predict, based on the embedding, whether to capture the current frame.
[0077] In response to the determination that the current frame is not representing the desired event, the device 200 may, in operation 912, predict the estimated time until the desired event appears in the real-time media stream. In operation 914, the device 200 may predict the target FPS for the real-time media stream based on the prediction. In an embodiment, the target FPS may indicate the required frame rate associated with the camera 240 for capturing the desired event in the real-time media stream. Thereafter, in operation 916, the device 200 may apply the target FPS to the camera 240 for capturing the desired event in the real-time media stream.
[0078] Alternatively, in response to the determination that the current frame represents the desired event, the device 200 may capture, using the camera, the current frame in operation 918.
[0079] In an embodiment, the device 200 may also dynamically modify the target FPS based on the prediction.
[0080] In an embodiment, the device 200 may determine whether the predicted estimated time exceeds the predefined value. Further, the device 200 may transmit the stop instruction to the camera 240 in response to determining that the predicted estimated time exceeds the predefined value.
[0081] In an embodiment, after predicting the estimated time, the device 200 may adjust the timer associated with the camera 240 based on the predicted estimated time and the remaining timer value.
[0082] Figure 10 illustrates an exemplary scenario 1000 associated with capturing the desired event, in accordance with an embodiment of the disclosure. The scenario 1000 may be implemented in the device 200. In the scenario 1000, the user aims to capture an image of himself jumping while in the air. For this, the user provides a user prompt that includes a text command - "Capture my photo when I am in air, like the man on left, with my leg folded and hands straight up in air", and a reference image featuring a man in the air on the left side with his leg folded and hands straight up in the air. Initially, the camera 240 begins processing at low speed (i.e., at low FPS). At block 1002, the device 200 detects that the user has begun jumping and may reach a desired pose soon. Accordingly, the device 200 increases the FPS of the camera 240 at this stage to ensure that the desired pose can be captured successfully. At block 1004, the device 200 predicts that the user has achieved the desired pose (mid-air with legs folded and hands straight up in the air) and accordingly, controls the camera 240 to capture the desired event.
[0083] Figure 11 depicts an exemplary scenario 1100 associated with modification of a timer value of the camera when the user wants to capture an image using the camera timer, in accordance with an embodiment of the disclosure. The scenario 1100 may be implemented in the device 200. In the scenario 1100, the user aims to capture an image of himself while being in a meditation pose. For this, the user sets the camera included his user device (e.g., the smartphone) to be operated to capture an event desired by the user. The user provides the user prompt that includes a text command 1101 (i.e., "Capture when I am in meditation pose") and initiates the camera timer with the timer value 1103 of 5 seconds, as shown in 1102. As the camera begins processing, the device 200 predicts that a desired pose (e.g., the meditation pose) is still far off from the present moment and there is ample time remaining on the camera timer (i.e., the remaining timer value is 2 seconds), as shown in block 1104. Consequently, the device 200 significantly slows down the processing speed of the camera. Instead of continuously recording and processing every frame of the real-time media stream, the camera operates at a reduced FPS, e.g., 2 FPS. This approach minimizes battery consumption and avoids unnecessary wastage of the device's processing power. At this stage, the device 200 focuses on predicting how close the user is to achieving the meditation pose. However, no frames are being saved, thereby also conserving memory usage and saving power usage. After some time, at block 1106, the device 200 identifies that the user is progressing towards the meditation pose, but the timer value 1103 is reaching 0 and there may not be sufficient time to capture the desired pose. In response, the device 200 slightly extends the timer value 1103 (e.g., by 2 seconds) to allow the user more time to complete the pose. Despite the camera timer adjustment, the FPS remains relatively low, for example, around 4 FPS. At block 1108, the device 200 detects that the user is nearly finished with the meditation pose, and the timer value is approaching zero. Recognizing such a critical moment, the device 200 increases the FPS substantially, for example, reaching 30 FPS. An increased processing speed ensures that the best possible pose is captured, resulting in a successful capturing of the desired pose. Thus, the disclosure improves the accuracy and efficiency of capturing the specific poses or the precise moments with camera-equipped devices.
[0084] Figure 12 depicts an exemplary scenario 1200 associated with capturing a slow motion video at variable frame rates, in accordance with an embodiment of the disclosure. The scenario 1200 may be implemented in the device 200. In the scenario 1200, the user aims to capture the slow motion video of himself at different frame rates while playing basketball. For this, the user provides a user prompt including a text command 1201 ("Capture a slow motion video of me throwing a ball and then a super slow motion when I dunk"), as shown in block 1202. At block 1204, the device 200 detects that the user is attempting to throw the ball. Accordingly, the device 200 determines a first frame rate for capturing the slow motion video of the user throwing the ball, and the camera is triggered to capture the slow motion video at the determined first frame rate. Next, at block 1206, the device 200 predicts that a dunk shot is in a pipeline. Accordingly, the device 200 determines a second frame rate (higher than the first frame rate) for capturing the super slow motion video of the user dunking, and the camera is triggered to capture the dunk shot at the determined second frame rate. Thus, the disclosure enhances user experience by improving the efficiency of capturing the specific poses or the precise moments with the camera-equipped devices.
[0085] Figure 13 depicts an exemplary scenario 1300 associated with capturing a desired pose, in accordance with an embodiment of the disclosure. The scenario 1300 may be implemented in the device 200. In the scenario 1300, the user wants to capture an image of himself while being inverted during a somersault and has his user device (e.g., the smartphone). For this, the user provides a user prompt that includes a text command 1301 ("Capture when I am inverted in the air while doing somersault"), as shown in block 1302. At block 1304, the device 200 predicts that the desired pose is still far off from the present moment. Consequently, the device 200 significantly slows down the processing speed of the camera. Instead of continuously recording and processing every frame of the real-time media stream, the camera operates at a reduced FPS, e.g., 2 FPS. This approach minimizes battery consumption and avoids unnecessary wastage of the device's power. At this stage, the device 200 focuses on predicting how close the user is to completing or achieving the desired pose. However, no frames are being saved, thereby also conserving memory usage. After some time, at block 1306, the device 200 identifies that the user is progressing towards the desired pose (e.g., the user has started somersaulting). At this stage, the FPS remains relatively low, for example, around 4 FPS. At block 1308, the device 200 detects that the user has almost approached or achieved the desired pose and accordingly, the device 200 increases the FPS substantially, for example, reaching 30 FPS to ensure that the best possible pose is captured. When the device 200 detects a frame that exactly matches a description provided by the user (i.e., when the user is inverted in the air while somersaulting), the device 200 triggers the camera, and the desired pose is captured.
[0086] Figure 14 depicts an exemplary scenario 1400 associated with capturing multiple images in multiple poses, in accordance with an embodiment of the disclosure. The scenario 1400 may be implemented in the device 200. In the scenario 1400, the user aims to capture multiple pictures of himself and his partner in front of monoliths in multiple poses. For this, the user provides a user prompt including a text command 1401 ("Capture multiple pictures of us together in multiple poses"), as shown in block 1402. At block 1404, the device 200 detects that the user and his partner have achieved a first pose and accordingly, triggers the camera to capture the first pose. Next, at block 1406, the device 200 detects that the user and his partner have achieved a second pose and accordingly, the device 200 triggers the camera to capture the second pose. Thus, the disclosure simplifies the process of capturing specific multiple moments or multiple poses and allows for smarter camera control.
[0087] Figure 15 depicts an exemplary scenario 1500 associated with capturing a pose similar to a person in a reference image 1503 that features multiple people, in accordance with an embodiment of the disclosure. The scenario 1500 may be implemented in the device 200. In the scenario 1500, the user aims to capture an image of himself in a pose similar to a pose made in the reference image 1503. The user prompt may designate, among a plurality of subjects, a subject making the pose in the reference image 1503. For this, the user provides a user prompt that includes a text command 1501 ("Capture my picture when I pose similar to the monk on right side in the given image") and the reference image 1503 featuring the monk on the right side throwing a flying kick and other people, as shown in block 1502. At block 1504, the device 200 detects that the user has begun jumping and may reach the desired pose soon. Accordingly, the system increases the FPS of the camera at this stage to ensure that the desired pose is captured successfully. At block 1506, the device 200 predicts that the user has achieved the desired pose (of the flying kick similar to the monk in the reference image 1503) and accordingly, triggers the camera to capture the desired pose. Thus, the disclosure simplifies the process of capturing specific moments or poses and improves user experience.
[0088] At least by virtue of the aforesaid, the disclosure provides various advantages. The disclosure automates the capture of precise moments or specific poses using the camera-equipped devices. By eliminating the need for manual video recording and taking screenshots, the disclosure streamlines the process of capturing high-quality images. The disclosure also optimizes battery and storage usage, thereby ensuring efficient utilization of resources. The disclosure enhances convenience, accuracy, and efficiency in capturing desired moments or poses using the camera-equipped devices. Further, the disclosure simplifies the process of capturing specific moments or poses while maintaining optimal performance and resource management. Furthermore, the disclosure ensures prolonged device performance and reduced operational costs while improving the user experience.
[0089] The disclosure allows for intuitive fine-grained control to capture a perfect image with minimal processing, latency, and memory overheads. Instead of a constant video shot at a pre-set FPS, which consumes more battery and processing power as well as significant memory resources, the disclosure allows for a very low processing speed for most of the time, allowing for a focus on only a target time period. The disclosure provides a fully interactive cross-modal (image / text / voice) based camera control system that intelligently decides when to "click / record" and when to wait, thus there is no wastage of memory resources. The disclosed techniques may have various potential applications, for example, in the field of photography, social media, and content creation. Further, the disclosure allows for smart control of camera FPS processing to reduce redundant processing overheads, thereby significantly saving on power consumption. The disclosure enables an intelligent understanding of the user prompt, thus there is no requirement for post-processing or any user intervention. Moreover, the disclosure allows the users a scene-aware camera feature control, thereby capturing moments / scenes in accordance with the user commands.
[0090] According to an embodiment of the disclosure, disclosed herein is a method for capturing a desired event. The method includes receiving a user prompt. The prompt indicating the desired event intended to be captured in a real-time media stream. The method also includes generating first visual features corresponding to the received prompt. Further, the method includes generating second visual features corresponding to the real-time media stream. Thereafter, the method includes correlating the first visual features with the second visual features of a current frame in the real-time media stream. Furthermore, the method includes determining, by a click-classifier module, whether the current frame represents the desired event based on the correlation. Thereafter, the method includes performing one of: predicting, using a frame-estimator (F-estimator) module, an estimated time until the desired event appears in the real-time media stream, predicting, a target frame per second (FPS) for the real-time media stream based on the prediction by the F-estimator module, wherein the target FPS indicates a required frame rate associated with the camera for capturing the desired event in the real-time media stream, and applying the target FPS to a camera for capturing the desired event in the real-time media stream in response to the determination that the current frame is not representing the desired event; or capturing, using the camera, the current frame in response to the determination that the current frame represents the desired event.
[0091] According to an embodiment of the disclosure, disclosed herein is a system for capturing a desired event. The system includes a memory, and at least one processor in communication with the memory. The at least one processor is configured to receive a user prompt, wherein the prompt indicates the desired event intended to be captured in a real-time media stream. The at least one processor is further configured to generate first visual features corresponding to the received prompt. Further, the at least one processor is configured to generate second visual features corresponding to the real-time media stream. Thereafter, the at least one processor is configured to correlate the first visual features with the second visual features of a current frame in the real-time media stream. Furthermore, the at least one processor is configured to determine, by a click classifier module, whether the current frame represents the desired event based on the correlation. Thereafter, the at least one processor is configured to perform one of: predict, using a frame-estimator (F-estimator) module, an estimated time until the desired event appears in the real-time media stream, predict, a target frame per second (FPS) for the real-time media stream based on the prediction by the F-estimator module, wherein the target FPS indicates a required frame rate associated with the camera for capturing the desired event in the real-time media stream, and apply the target FPS to a camera for capturing the desired event in the real-time media stream in response to the determination that the current frame is not representing the desired event; or capture, using the camera, the current frame in response to the determination that the current frame represents the desired event.
[0092] According to an embodiment, a device for capturing a desired event is disclosed. The device includes a camera, at least one processor, memory storing computer program code, where the memory and the computer program code are configured, with the at least one processor, to cause the apparatus to obtain a user prompt indicating the desired event to be captured in a media stream. In an embodiment, the device is caused to generate first visual features corresponding to the user prompt. In an embodiment, the device is caused to generate second visual features from the media stream. In an embodiment, the device is caused to correlate the first visual features with the second visual features of a current frame in the media stream. In an embodiment, the device is caused to determine, based on the correlation, whether the current frame includes the desired event. In an embodiment, the device is caused to control the camera to capture the current frame at a target FPS(frame per second) based on the determination.
[0093] In an embodiment, at least one processor is configured to predict an estimated time until the desired event appears in the media stream.
[0094] In an embodiment, the at least one processor is configured to predict the target FPS for the media stream based on the prediction of the estimated time, wherein the target FPS indicates a required frame rate associated with the camera for capturing the desired event in the media stream.
[0095] In an embodiment, the at least one processor is configured to apply the target FPS to the camera for capturing the desired event in the media stream based on the determination that the current frame is not representing the desired event.
[0096] In an embodiment, the at least one processor is configured to modify the target FPS based on the prediction of the estimated time.
[0097] In an embodiment, the at least one processor is configured to lower the target FPS based on the prediction that the desired event is a first predetermined time away from the current frame, or to raise the target FPS based on the prediction that the desired event appears within a second predetermined time from the current frame.
[0098] In an embodiment, the user prompt comprises at least one selected from a group of a text command, a reference image, and a timer value.
[0099] In an embodiment, the text command comprises at least one text referring to an object included in the reference image. In an embodiment, the desired event includes at least one selected from a group of a specific pose and a specific moment.
[0100] In an embodiment, the at least one processor is configured to transmit a stop instruction to the camera in response to determining that the predicted estimated time exceeds a predefined value.
[0101] Disclosed is a method of capturing a desired event. According to an embodiment, the method includes obtaining a user prompt indicating the desired event to be captured in a media stream. In an embodiment, the method includes generating first visual features corresponding to the user prompt. In an embodiment, the method includes generating second visual features from the media stream. In an embodiment, the method includes correlating the first visual features with the second visual features of a current frame in the media stream. In an embodiment, the method includes determining, based on the correlation, whether the current frame includes the desired event. In an embodiment, the method includes capturing, by a camera, the current frame at a target FPS based on the determination.
[0102] In an embodiment, the method includes predicting an estimated time until the desired event appears in the media stream.
[0103] In an embodiment, the method includes predicting the target FPS for the media stream based on the prediction of the estimated time, wherein the target FPS indicates a required frame rate associated with the camera for capturing the desired event in the media stream.
[0104] In an embodiment, the method includes modifying the target FPS based on the prediction of the estimated time.
[0105] In an embodiment, the method includes lowering the target FPS based on the prediction that the desired event is a first predetermined time away from the current frame.
[0106] In an embodiment, the method includes raising the target FPS based on the prediction that the desired event appears within a second predetermined time from the current frame.
[0107] In an embodiment, the method includes transmitting a stop instruction to the camera in response to determining that the predicted estimated time exceeds a predefined value.
[0108] In this application, unless specifically stated otherwise, the use of the singular includes the plural, and the use of "or" means "and / or." Furthermore, the use of the terms "including" or "having" is not limiting. Any range described herein will be understood to include the endpoints and all values between the endpoints. Features of the disclosed embodiments may be combined, rearranged, omitted, etc., within the scope of the invention to produce additional embodiments. Furthermore, certain features may sometimes be used to advantage without a corresponding use of other features.
[0109] The method according to an embodiment of the disclosure may be implemented in program instructions which are executable by various computing means and recorded in computer-readable media. The computer-readable media may include program instructions, data files, data structures, etc., separately or in combination. The program instructions recorded on the computer-readable media may be designed and configured specially for the disclosure, or may be well-known to those of ordinary skill in the art of computer software. Examples of the computer readable recording medium include a magnetic medium such as a hard disk, a floppy disk and a magnetic tape, an optical medium such as a compact disc read-only memory (CD-ROM) and a digital versatile disc (DVD), a magneto-optical medium such as a floptical disk, and a hardware device specially configured to store and perform program instructions, such as a read-only memory (ROM), a random-access memory (RAM), a flash memory, etc. Examples of the program instructions include not only machine language codes but also high-level language codes which are executable by various computing means using an interpreter.
[0110] The embodiment of the disclosure may be implemented in the form of a computer-readable recording medium that includes computer-executable instructions such as the program modules executed by the computer. The computer-readable medium may be an arbitrary available medium that may be accessed by the computer, including volatile, non-volatile, removable, and non-removable mediums. The computer-readable recording medium may also include a computer storage medium and a communication medium. The computer-readable medium includes all the volatile, non-volatile, removable, and non-removable mediums implemented by an arbitrary method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. The communication medium generally includes computer-readable instructions, data structures, program modules, or other data or other transmission mechanism for modulated data signals like carrier waves, and include arbitrary information delivery medium. Furthermore, some embodiments of the disclosure may be implemented in a computer program or a computer program product including computer-executable instructions.
[0111] The machine-readable storage medium may be provided in the form of a non-transitory storage medium. The term 'non-transitory storage medium' may mean a tangible device without including a signal, e.g., electromagnetic waves, and may not distinguish between storing data in the storage medium semi-permanently and temporarily. For example, the non-transitory storage medium may include a buffer that temporarily stores data.
[0112] In an embodiment of the disclosure, the aforementioned method may be provided in a computer program product. The computer program product may be a commercial product that may be traded between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a CD-ROM) or distributed directly between two user devices (e.g., smart phones) or online (e.g., downloaded or uploaded). In the case of the online distribution, at least part of the computer program product (e.g., a downloadable app) may be at least temporarily stored or arbitrarily created in a storage medium that may be readable to a device such as a server of the manufacturer, a server of the application store, or a relay server.
[0113] While at least one exemplary embodiment has been presented in the foregoing detailed description, it should be appreciated that a vast number of variations exist.
Claims
1.An apparatus for capturing a desired event, the device comprising:a camera;at least one processor;memory storing computer program code, where the memory and the computer program code are configured, with the at least one processor, to cause the apparatus to:obtain a user prompt indicating the desired event to be captured in a media stream,generate first visual features corresponding to the user prompt,generate second visual features from the media stream,correlate the first visual features with the second visual features of a current frame in the media stream,determine, based on the correlation, whether the current frame includes the desired event, andcontrol the camera to capture the current frame at a target FPS(frame per second) based on the determination.2.The apparatus of claim 1,wherein the at least one processor is configured to predict an estimated time until the desired event appears in the media stream.3.The apparatus of claim 2,wherein the at least one processor is configured to predict the target FPS for the media stream based on the prediction of the estimated time, wherein the target FPS indicates a required frame rate associated with the camera for capturing the desired event in the media stream.4.The apparatus of claim 3,wherein the at least one processor is configured to apply the target FPS to the camera for capturing the desired event in the media stream based on the determination that the current frame is not representing the desired event.5.The apparatus of any one of claims 2 to 4,wherein the at least one processor is configured to modify the target FPS based on the prediction of the estimated time.6.The apparatus of any one of claims 2 to 5,wherein the at least one processor is configured to lower the target FPS based on the prediction that the desired event is a first predetermined time away from the current frame, or to raise the target FPS based on the prediction that the desired event appears within a second predetermined time from the current frame.7.The apparatus of any one of claims 1 to 6,wherein the user prompt comprises at least one selected from a group of a text command, a reference image, and a timer value.8.The apparatus of claim 7,wherein the text command comprises at least one text referring to an object included in the reference image.9.The apparatus of any one of claims 1 to 8,wherein the desired event comprises at least one selected from a group of a specific pose and a specific moment.10.The apparatus of any one of claims 1 to 9,wherein the at least one processor is configured to transmit a stop instruction to the camera in response to determining that the predicted estimated time exceeds a predefined value.11.A method of capturing a desired event, the method comprising:obtaining a user prompt indicating the desired event to be captured in a media stream;generating first visual features corresponding to the user prompt;generating second visual features from the media stream;correlating the first visual features with the second visual features of a current frame in the media stream;determining, based on the correlation, whether the current frame includes the desired event; andcapturing, by a camera, the current frame at a target FPS based on the determination.12.The method of claim 11, further comprising:predicting an estimated time until the desired event appears in the media stream; andpredicting the target FPS for the media stream based on the prediction of the estimated time, wherein the target FPS indicates a required frame rate associated with the camera for capturing the desired event in the media stream.13.The method of claim 12, further comprisingmodifying the target FPS based on the prediction of the estimated time.14.The method of any one of claims 12 to 13,lowering the target FPS based on the prediction that the desired event is a first predetermined time away from the current frame, orraising the target FPS based on the prediction that the desired event appears within a second predetermined time from the current frame.15.The method of any one of claims 12 to 14,transmitting a stop instruction to the camera in response to determining that the predicted estimated time exceeds a predefined value.