Electronic device, method, and non-transitory computer-readable storage medium for determining execution timing of high-speed capturing function
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-11-26
- Publication Date
- 2026-07-30
Smart Images

Figure KR2025019721_30072026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transient computer-readable storage medium for determining the execution timing of a high-speed shooting function
[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for determining the execution timing of a high-speed shooting function.
[0002] An electronic device can analyze an image through an artificial intelligence model. The image can be acquired through a camera. The image can represent a real environment. By performing object recognition through an artificial intelligence model, the electronic device can identify the types of visual objects included within the image. By performing image classification through an artificial intelligence model, the electronic device can identify the types of images.
[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure.
[0004] No claim or determination is made as to whether any of the foregoing can be applied as prior art related to the present disclosure.
[0005] An electronic device is described. The electronic device may include a memory comprising one or more storage media for storing instructions. The electronic device may include a camera. The electronic device may include at least one processor comprising a processing circuit. The instructions may cause the electronic device to determine the type of the first video through a first trained model based on acquiring a first video at a first frame rate through the camera, when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to identify the state of a visual object included in the first video using a third trained model corresponding to the determined type among second trained models for identifying the state of a visual object, when executed individually or collectively by the at least one processor. When the above instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to acquire a second video through the camera at a second frame rate higher than the first frame rate, based on the determination that the identified state is a target state.
[0006] A method is provided. The method may be executed within an electronic device having a camera. The method may include an operation of determining the type of the first video through a first trained model based on acquiring a first video at a first frame rate through the camera. The method may include an operation of identifying the state of a visual object included in the first video using a third trained model corresponding to the determined type among second trained models for identifying the state of a visual object. The method may include an operation of acquiring a second video through the camera at a second frame rate higher than the first frame rate, based on a determination that the identified state is a target state.
[0007] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to determine the type of the first video through a first trained model based on acquiring a first video at a first frame rate through the camera when executed by the electronic device having a camera. The one or more programs may include instructions that cause the electronic device to identify the state of a visual object included in the first video using a third trained model corresponding to the determined type among second trained models for identifying the state of a visual object when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire a second video through the camera at a second frame rate higher than the first frame rate based on a determination that the identified state is a target state when executed by the electronic device.
[0008] An electronic device is described. The electronic device may include a memory comprising one or more storage media for storing instructions. The electronic device may include a camera. The electronic device may include at least one processor comprising a processing circuit. The instructions may cause the electronic device to identify states of visual objects included in the first video based on acquiring a first video at a first frame rate through the camera when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine the timing for acquiring a second video at a second frame rate higher than the first frame rate by using transition information to determine the timing according to the transition between the states based on identifying the states when executed individually or collectively by the at least one processor. When the above instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to acquire the second video through the camera at the second frame rate based on the determined timing.
[0009] A method is provided. The method may be executed within an electronic device having a camera. The method may include an operation of identifying states of visual objects included in the first video based on acquiring a first video at a first frame rate through the camera. The method may include an operation of determining a timing for acquiring a second video at a second frame rate higher than the first frame rate, using transition information for determining timing according to transitions between the states based on identifying the states. The method may include an operation of acquiring the second video at the second frame rate through the camera based on the determined timing.
[0010] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to identify the states of a visual object included in the first video based on acquiring a first video at a first frame rate through the camera when executed by the electronic device having a camera. The one or more programs may include instructions that cause the electronic device to determine the timing for acquiring a second video at a second frame rate higher than the first frame rate by using transition information to determine the timing according to the transition between the states based on identifying the states when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the second video at the second frame rate through the camera based on the determined timing when executed by the electronic device.
[0011] Figure 1 illustrates an example of an electronic device that performs a high-speed shooting function.
[0012] Figure 2 is a simplified block diagram of an exemplary electronic device.
[0013] Figure 3 is a flowchart illustrating the operation of an electronic device that acquires video based on trained models.
[0014] FIG. 4 illustrates an exemplary operation of an electronic device that determines the type of video based on a first trained model.
[0015] FIG. 5 illustrates an exemplary operation of an electronic device that acquires state data based on a third trained model.
[0016] FIG. 6 illustrates an exemplary operation of an electronic device that calculates a cumulative value based on a fourth trained model.
[0017] FIG. 7 illustrates an exemplary operation of an electronic device that calculates an accumulated value based on transitions between states.
[0018] Figure 8 is a flowchart showing the operation of an electronic device that performs a high-speed shooting function through a Region of Interest (ROI).
[0019] FIG. 9 illustrates an exemplary operation of an electronic device that performs a high-speed shooting function based on identifying movement within an ROI.
[0020] FIG. 10 is a flowchart illustrating the operation of an electronic device that determines the execution timing of a high-speed shooting function using transition information.
[0021] FIG. 11 is a block diagram of an electronic device in a network environment according to various embodiments.
[0022] Figure 12 is a schematic diagram of an exemplary AI (Artificial Intelligence) system.
[0023] Figure 1 illustrates an example of an electronic device that performs a high-speed shooting function.
[0024] Referring to FIG. 1, an electronic device (100) may be used to acquire a video (140) of an environment (150). For example, the electronic device (100) may include a camera (e.g., camera (203) in FIG. 2). For example, the camera may be used to acquire a video (140) representing the environment (150). For example, the video (140) may be described as a video acquired through the camera and displayed through a display included in the electronic device (100). For example, the video (140) may be (temporarily) stored in the memory of the electronic device (100) (e.g., memory (206) in FIG. 2). For example, the video (140) may be referred to as a preview image or a preview video.
[0025] The environment (150) may include a user (120). For example, the user (120) may enjoy sports (e.g., golf). For example, the electronic device (100) may acquire a video (140) of the user (120) playing sports (e.g., golf) through a camera (e.g., camera (203) of FIG. 2).
[0026] The electronic device (100) can acquire video at a first frame rate (e.g., 24 to 30 fps (frames per second)) through a camera (e.g., camera (203) of FIG. 2). For example, the electronic device (100) can provide or implement a high-speed shooting function or high-speed shooting feature. For example, the high-speed shooting function may be described as a function for acquiring video at a second frame rate (e.g., 960 fps) higher than the first frame rate (or a function for acquiring video having a frame rate higher than a frame rate threshold). For example, the second frame rate may be changed according to user input. For example, the high-speed shooting function may be described as a function for displaying frames of video (e.g., 960 frames) acquired through the camera during a first time interval (e.g., 1 second) at a preset frame rate (e.g., 30 fps) during a second time interval (e.g., 32 seconds) that is longer than the first time interval. For example, the high-speed shooting function may be described as a function for playing back a high-speed video (e.g., at 30 fps) acquired by shooting an environment (150) containing the electronic device (100) at high speed (e.g., 480 fps or 960 fps). For example, the high-speed shooting function may be referred to as a slow motion function, a super slow motion function, a slow motion mode, and a super slow motion mode.
[0027] The electronic device (100) may provide a super slow motion function only during the first time interval. For example, when performing high-speed shooting, the electronic device (100) may perform high-speed shooting only for a short time (e.g., 1 second or less) because it is required to acquire many frames (e.g., 200 frames or more). For example, the electronic device (100) may perform high-speed shooting only for a short time because an image sensor and storage, etc., are required to acquire many high-quality frames. For example, the electronic device (100) may identify the motion of a visual object through images (or video) acquired through a camera (e.g., camera (203) of FIG. 2) in order to provide a super slow motion function for a section desired by the user (120). For example, the electronic device (100) may change or switch the mode of the camera (e.g., camera (203) of FIG. 2) from normal mode to super slow mode based on identifying motion of a visual object corresponding to a part of the user (120) or an external object within the video acquired through the camera. For example, the normal mode may be described as a mode for acquiring video according to a frame rate lower than the frame rate provided in super slow mode (e.g., 30 fps).
[0028] An electronic device (100) may be required to identify the motion of a visual object (e.g., a golf ball or a golf club) included within a preview image based on a preview video (e.g., a video (140)) obtained through a camera (e.g., a camera (203) in FIG. 2). For example, the electronic device (100) may identify the motion of the visual object based on the position of the visual object moved within the frames obtained through the camera. For example, the electronic device (100) may identify the motion of the visual object by calculating a motion vector for the visual object. To identify the motion of the visual object, the electronic device (100) may be required to distinguish another visual object (e.g., a background) that is different from the visual object. The electronic device (100) may be required to identify complex motion of the visual object within the preview image. For example, since the time interval (e.g., 1 second) at which the electronic device (100) can provide a super slow motion mode is very short, it may be required for the user (120) to identify the motion of the desired visual object.
[0029] The electronic device (100) can identify the state of a visual object included in the preview image through trained models (e.g., the second trained models of FIG. 5 (510-1, 510-2, ..., 510-N)) to provide a super slow motion mode for a section desired by the user (120). For example, the electronic device (100) can predict the next state of the visual object based on identifying the state of the visual object. For example, the electronic device (100) can determine the timing to execute the super slow mode based on the determination that the next state is the state desired by the user (120).
[0030] For example, the electronic device (100) may include hardware components used to perform or execute the above operations. The hardware components are described and illustrated with reference to FIG. 2.
[0031] Figure 2 is a simplified block diagram of an exemplary electronic device.
[0032] Referring to FIG. 2, the electronic device (100) may include a camera (203), a display (208), at least one processor (207), and a memory (206).
[0033] At least one processor (207) may include a hardware component for processing data using instructions stored in memory (206). The hardware component for processing data may include a CPU (central processing unit) (e.g., including processing circuits). The hardware component for processing data may include a GPU (graphic processing unit) (e.g., including processing circuits). The hardware component for processing data may include a DPU (display processing unit) (e.g., including processing circuits). The hardware component for processing data may include a NPU (neural processing unit) (e.g., including processing circuits).
[0034] At least one processor (207) may include one or more cores. For example, at least one processor (207) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.
[0035] Memory (206) may include a hardware component for storing data and / or instructions that are input to and / or output from at least one processor (207). Memory (206) may include, for example, volatile memory such as RAM (random-access memory) and / or non-volatile memory such as ROM (read-only memory). Volatile memory may include, for example, at least one of DRAM (dynamic RAM), SRAM (static RAM), cache RAM, and PSRAM (pseudo SRAM). Non-volatile memory may include, for example, at least one of PROM (programmable ROM), EPROM (erasable PROM), EEPROM (electrically erasable PROM), flash memory, hard disk, compact disk, and EMMC (embedded multimedia card).
[0036] The display (208) can output visualized information. For example, the display (208) can output visualized information to a user under the control of at least one processor (207). The display (208) may include hardware components of an electronic device (100) used to display a screen. For example, the display (208) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements may include an organic light-emitting diode (OLED) or a micro LED. However, it is not limited thereto. For example, the display (208) may include a liquid crystal display (LCD). The display (208) may not be an essential hardware component. The display (208) may be an optional hardware component.
[0037] The camera (203) may include one or more light sensors (e.g., a CCD (charged coupled device) sensor, a CMOS (complementary metal oxide semiconductor) sensor) that generate an electrical signal indicating the color and / or brightness of light. For example, the camera (203) may be described as one or more image sensors. For example, the camera (203) may be available to acquire an image of the environment surrounding the electronic device (100).
[0038] At least one processor (207) can determine the type of the first video (e.g., the video (410) of FIG. 4) based on acquiring the first video at a first frame rate through the camera (203) using a first trained model (e.g., the first trained model (420) of FIG. 4). At least one processor (207) can identify the state of a visual object included in the first video using second trained models for identifying the state of a visual object (e.g., the second trained models (510-1, 510-2, ..., 510-N) of FIG. 5) (N is a natural number). At least one processor (207) can acquire the second video at a second frame rate higher than the first frame rate through the camera (203) based on the determination that the identified state is a target state (e.g., the fourth state (740) of FIG. 7). For example, at least one processor (207) can display or play the second video through a display (208).
[0039] FIG. 3 is a flowchart illustrating the operation of an electronic device for acquiring video based on trained models. This method may be executed by the electronic device (100) illustrated in FIG. 2 or by at least one processor (207) of the electronic device (100). At least some of the operations described below may be executed in parallel.
[0040] Referring to FIG. 3, in operation 310, at least one processor (207) can determine the type of the first video (e.g., the video (410) of FIG. 4) based on acquiring the first video (e.g., the video (410) of FIG. 4) through the camera (203) at a first frame rate (e.g., 30 fps) using a first trained model (e.g., the first trained model (420) of FIG. 4). For example, the first video can be described as a video captured according to the first frame rate. For example, the first video can be described as a video acquired according to normal mode. For example, at least one processor (207) can determine the type of the first video using a first trained model (420) trained to determine the type of the video among a plurality of types by performing image classification on the frames of the video. For example, the first trained model (420) can determine the type of the first video based on identifying visual objects (e.g., golf club or grass) included in the frames of the first video.
[0041] At least one processor (207) can determine the type of video (e.g., video (410) of FIG. 4) based on acquiring a video (e.g., 30 fps) through a camera (203) using a first trained model (e.g., first trained model (420) of FIG. 4). The frame rate may be referred to as FPS (frames per second). For example, the video (e.g., video (410) of FIG. 4) at the first frame rate (or first FPS) may be described as a video acquired according to normal mode. For example, at least one processor (207) can determine the type of the video by performing image classification and / or object recognition on the frames of the video (e.g., video (410) of FIG. 4). For example, at least one processor (207) may determine the type as a first type in which the state of the frames of the video (e.g., the video (410) of FIG. 4) cannot be identified, or a second type in which the state of the frames can be identified. For example, the determination of the type will be described later with reference to FIG. 4.
[0042] In operation 320, at least one processor (207) can identify the state of a visual object included in a first video (e.g., video (410) of FIG. 4) by using a third trained model (e.g., third trained model (520) of FIG. 5) corresponding to a determined type among second trained models (e.g., second trained models (510-1, 510-2, ..., 510-N) of FIG. 5) for identifying the state of a visual object.
[0043] At least one processor (207) can identify the state of a visual object included in a video (e.g., video (410) of FIG. 4) acquired through a camera (203) by using a third trained model (e.g., third trained model (520) of FIG. 5) corresponding to a determined type among second trained models (e.g., second trained models (510-1, 510-2, ..., 510-N) of FIG. 5) for identifying the state of a visual object included in a video (e.g., video (410) of FIG. 4). For example, at least one processor (207) can identify or determine the state (e.g., motion or pose) of a visual object included in the video by providing the video (e.g., video (410) of FIG. 4) to the third trained model (e.g., third trained model (520) of FIG. 5). For example, at least one processor (207) may calculate an accumulated value to determine the timing for providing a super slow motion mode based on identifying the state of a visual object. For example, at least one processor (207) may switch the mode of the camera (203) from normal mode to super slow motion mode based on the determination that the accumulated value exceeds a threshold value. For example, the identification of the state of the visual object and the calculation of the accumulated value will be described later with reference to FIGS. 5 through 7.
[0044] In operation 330, at least one processor (207) may acquire a second video through a camera (203) at a second frame rate (e.g., 960 fps) higher than a first frame rate (e.g., 30 fps) based on a determination that the identified state is a target state using a third trained model (520). For example, at least one processor (207) may perform high-speed shooting based on a determination that the identified state is a target state. For example, the second video may be described as a video acquired while performing high-speed shooting. For example, the frame rate of the second video may be described as a second frame rate higher than the first frame rate. For example, the second video may be acquired through the camera (203) for a first time interval (e.g., 1 second) while performing high-speed shooting and played back through a display (208) for a second time interval (e.g., 32 seconds) longer than the first time interval. For example, the first time interval may be changed according to the resources of the electronic device (100) (e.g., the performance of the image processor or the capacity of the storage). For example, at least one processor (207) may play or display frames of the second video acquired during the first time interval through the display (208) during the second time interval. For example, the shooting speed at which the second video is acquired may be faster than the playback speed at which the second video is played. For example, if the second video contains 960 frames acquired during 1 second, the 960 frames may be displayed through the display (208) at a rate of 30 frames per second for 32 seconds.
[0045] At least one processor (207) can calculate an accumulated value corresponding to an identified state (e.g., an accumulated value (632) in FIG. 6) by using transition information (e.g., transition information (625) in FIG. 6) for calculating an accumulated value according to a transition between states based on identifying a state of a visual object included in the first video. At least one processor (207) can acquire a second video at a second frame rate through a camera (203) based on a determination that the accumulated value calculated according to identifying a target state (e.g., a fourth state (740) in FIG. 7) exceeds a threshold value. The accumulated value can be calculated based on a reference value (e.g., 0) corresponding to a reference state (e.g., a first state (710) in FIG. 7) and a transition value (e.g., a transition value (712) in FIG. 7) corresponding to a transition between states indicated by the transition information (625).
[0046] At least one processor (207) can acquire a video at a second frame rate (e.g., 960 fps) higher than a first frame rate (e.g., 30 fps) through a camera (203) based on a determination that the state of a visual object included in the video (e.g., video (410) of FIG. 4) is a target state (e.g., fourth state (740) of FIG. 7). For example, at least one processor (207) can acquire a video in super slow motion mode based on a determination that an accumulated value calculated based on identifying the target state exceeds a threshold value. For example, at least one processor (207) can display the video acquired in super slow motion mode through a display (208). For example, at least one processor (207) can display or play frames (e.g., 960 frames) acquired during a first time interval (e.g., 1 second) through a display (208) at a third frame rate (e.g., 30 fps) during a second time interval (e.g., 32 seconds) longer than the first time interval.
[0047] The electronic device (100) can determine one of a plurality of types of a video acquired through a camera (203) (e.g., the video (410) of FIG. 4) through a first trained model (e.g., the first trained model (420) of FIG. 4). For example, the operation of determining the type of the video is described and illustrated in more detail with reference to FIG. 4.
[0048] FIG. 4 illustrates an exemplary operation of an electronic device that determines the type of video based on a first trained model.
[0049] Referring to FIG. 4, at least one processor (207) can determine the type of video (410) through a first trained model (420). The video (410) can be described as a video acquired through a camera (203) and (temporarily) stored in memory (206). For example, each frame of the video (410) can be referred to as a preview image or a preview frame. For example, the video (410) can be referred to as a preview video. For example, when the electronic device (100) acquires the video (410) through the camera (203), it can display the video (410) through a display (208).
[0050] The electronic device (100) may include a first trained model (420). For example, the first trained model (420) may be described as a model trained to determine the type of video using video. For example, the first trained model (420) may be described as a model trained through machine learning techniques (or deep learning techniques). For example, the first trained model (420) may include a convolutional neural network (CNN). For example, the first trained model (420) may include a large multimodal model (LMM). For example, the first trained model (420) may be referred to as a classification model, a type model, or a type determination model.
[0051] At least one processor (207) can determine the type of video (410) among a plurality of types through the first trained model (420). For example, at least one processor (207) can determine the video (410) as a first type through the first trained model (420). For example, the video (430) of the first type can be described as a video in which the state of a visual object included in the video (430) is not identified as at least one of the predefined states. For example, the video (430) of the first type can be described as a video in which states different from the predefined states are identified. For example, at least one processor (207) can determine the video (410) as a second type of video (440) through the first trained model (420). For example, the video (440) of the second type can be described as a video in which predefined states are identified. For example, the state of a visual object included in the second type of video (440) can be described as one of the predefined states.
[0052] The first trained model (420) may include information about states that can be identified within the video. For example, at least one processor (207) may determine the type of video by performing at least one of object recognition and image classification on the video (410) (or frames of the video (410)). For example, at least one processor (207) may determine the type of the video (410) as a first type and one of the remaining types that are different from the first type among a plurality of types. For example, the video of the first type (430) may be described as a video in which states different from predefined states are identified. Each video of the remaining type (440) included in the remaining types (e.g., sports such as golf, soccer, baseball, tennis, basketball, etc.) may be described as a video in which at least some of the predefined states for one (a) type (e.g., golf) are identified.
[0053] According to one embodiment, at least one processor (207) may receive user input to determine the type of video (410). For example, at least one processor (207) may display a User Interface (UI) object through a display (208) to inquire about the type of video (410) acquired through a camera (203). For example, at least one processor (207) may determine the type of video (410) based on receiving user input regarding the UI object.
[0054] According to one embodiment, at least one processor (207) may provide a super slow motion mode by identifying motion within a region of interest (ROI) of the video (410) acquired through the camera (203) based on a determination that the type of the video (410) is a first type. For example, the operation of providing a super slow motion mode by identifying motion within the ROI will be described later with reference to FIGS. 8 and 9.
[0055] At least one processor (207) can identify the state of the video (410) by using one of the second trained models (e.g., the second trained models of FIG. 5 (510-1, 510-2, ..., 510-N)) based on the determination that the type of the video (410) is one of the remaining types that are different from the first type among a plurality of types. For example, the operation of identifying the state of the video (410) using the trained model will be described later with reference to FIG. 5.
[0056] FIG. 5 illustrates an exemplary operation of an electronic device that acquires state data based on a third trained model.
[0057] Referring to FIG. 5, the electronic device (100) may include a second trained model (510-1), a second trained model (510-2) to a second trained model (510-N). The electronic device (100) may determine a third trained model (520) corresponding to a type of video (440) among the second trained models (510-1), the second trained models (510-2) to a second trained model (510-N). For example, at least one processor (207) may determine a third trained model (520) corresponding to a second type among the second trained models (510-1, 510-2, ..., 510-N) based on the determination that the type of video (410) is a second type (e.g., golf). For example, each second trained model may be mapped to each type. For example, each second trained model may include data for predefined states (e.g., address, backswing, top of swing, etc.) for each type (e.g., golf, soccer, baseball, basketball, or tennis). For example, the second trained model may be referred to as a definition model or a state definition model.
[0058] At least one processor (207) can reduce the power required to determine the execution timing of the high-speed shooting function by determining the third trained model (520) among the second trained models (510-1, 510-2, ..., 510-N). For example, the power consumed to acquire state data (530) using a large model including the second trained models (510-1, 510-2, ..., 510-N) may be greater than the power consumed to acquire state data (530) using the third trained model (520). At least one processor (207) can reduce the time required to determine the execution timing of the high-speed shooting function by determining the third trained model (520) among the second trained models (510-1, 510-2, ..., 510-N). For example, the time required to obtain state data (530) using the above-mentioned large model may be longer than the time required to obtain state data (530) using the third trained model (520).
[0059] Each second trained model can be described as a model trained to output state data using a video. For example, each second trained model can be described as a model trained to output data (e.g., state data (530) of FIG. 5) representing the state of a visual object included in the video, among states predefined for one type, using a video (e.g., video (410) of FIG. 4). For example, each second trained model can be described as a model trained through machine learning (or deep learning) techniques. For example, each second trained model may include a model trained through supervised learning techniques. However, it is not limited thereto. For example, each second trained model may be trained through unsupervised learning techniques. Each second trained model may be used to identify the state of a visual object included in the video (e.g., movement, motion, pose, etc.) by performing at least one of image classification and object recognition on the frames of the video.
[0060] At least one processor (207) can obtain state data (530) representing the state of the video (440) by providing the video (440) to a third trained model (520). For example, at least one processor (207) can obtain state data (530) representing the state of a visual object included in the video (440) based on providing the video (440) to the third trained model (520). For example, the visual object may correspond to a user (120). For example, the state data (530) may be based on text. However, it is not limited thereto.
[0061] At least one processor (207) can sequentially identify the states of visual objects included in the video by providing the video (440) to a third trained model (520). For example, at least one processor (207) can calculate an accumulation value based on the sequentially identified states using transition information (e.g., transition information (625) of FIG. 6). For example, at least one processor (207) can calculate an accumulation value based on identifying a transition from a first state to a second state. For example, at least one processor (207) can determine the timing for providing a super slow motion mode based on calculating the accumulation value. For example, the calculation of the accumulation value is described and illustrated in more detail with reference to FIG. 6.
[0062] FIG. 6 illustrates an exemplary operation of an electronic device that calculates a cumulative value based on a fourth trained model.
[0063] Referring to FIG. 6, at least one processor (207) can obtain a set of state data (e.g., state data (612), state data (614), and state data (616)) as it sequentially identifies states. For example, at least one processor (207) can calculate cumulative values (sequentially) through a fourth trained model (620) based on obtaining the set.
[0064] The fourth trained model (620) can be described as a trained model for calculating a cumulative value using state data (e.g., state data (612), state data (614), or state data (616)). For example, the fourth trained model (620) can be described as a model trained through machine learning (or deep learning) techniques.
[0065] At least one processor (207) can obtain state data representing the state of a visual object included in the video (440) through a third trained model (520). At least one processor (207) can obtain an accumulated value (632) by providing state data (612) to a fourth trained model (620). At least one processor (207) can obtain an accumulated value (634) by providing state data (614) obtained after state data (612) to the fourth trained model (620). For example, at least one processor (207) can obtain an accumulated value (634) through the fourth trained model (620) based on the accumulated value (632) and state data (614). At least one processor (207) can obtain an accumulated value (636) by providing the state data (616) obtained after the state data (614) to the fourth trained model (620). For example, at least one processor (207) can obtain an accumulated value (636) through the fourth trained model (620) based on the accumulated value (634) and the state data (616).
[0066] According to one embodiment, executing a high-speed shooting function using an electronic device (100) comprising at least one of second trained models (e.g., second trained model (510-1) to second trained model (510-N)) and / or a fourth trained model (620) can be performed through a memory of a relatively small capacity (e.g., memory (206)). However, the present disclosure is not limited thereto. For example, the high-speed shooting function may be performed based on communication between an external electronic device (e.g., a cloud server) comprising at least one of second trained models (e.g., second trained model (510-1) to second trained model (510-N)) and / or a fourth trained model (620) and the electronic device (100). For example, the electronic device (100) may include a communication circuit (not shown). For example, the electronic device (100) may receive a signal from the external electronic device through the communication circuit to execute the high-speed shooting function. For example, when the electronic device (100) executes a high-speed shooting function based on a signal received from the external electronic device, the external electronic device may determine the timing for executing the high-speed shooting function by using a large model comprising at least one of the second trained models (e.g., second trained model (510-1) to second trained model (510-N)) and / or a fourth trained model (620). For example, since the external electronic device may include a relatively large capacity memory, the large model may be used to determine the timing for executing the high-speed shooting function.
[0067] At least one processor (207) may use transition information (625) included in the fourth trained model (620) when obtaining accumulated values (632, 634, 636) through the fourth trained model (620). For example, the transition information (625) may represent transition values (e.g., transition values (712, 714) of FIG. 7) based on transitions between states represented by a set of state data (e.g., state data (612), state data (614), and state data (616)). For example, the operation of obtaining accumulated values (632, 634, 636) using the transition values represented by the transition information (625) is described and illustrated in more detail with reference to FIG. 7.
[0068] FIG. 7 illustrates an exemplary operation of an electronic device that calculates an accumulated value based on transitions between states.
[0069] Referring to FIG. 7, when at least one processor (207) identifies states (e.g., first state (710), second state (720), third state (730), and fourth state (740)), it may identify or calculate cumulative values (632, 634, 636) using transition information (625) representing transition values (transition value (712), transition value (714), etc.). Examples of golf states are described below to explain the operation of calculating the cumulative values, but embodiments of the present disclosure are not limited thereto. For example, in a video (440) of the second type (e.g., golf), the first state (710) may represent an address state. For example, the second state (720) may represent a backswing state. For example, the third state (730) may represent a top of swing state. For example, the fourth state (740) may represent a downswing state. Transition information (625) may represent transition values corresponding to transitions between states. For example, transition information (625) may include a transition matrix based on a state transition diagram. For example, the state transition diagram may be described as a diagram representing the relationship between states that can be transitioned, as in FIG. 7. For example, the transition matrix may be described as a matrix representing transition values to be calculated as a transition from one state (a) to another state. For example, the transition value may be described as a value that is the subject of calculation when transitioning from one state to another. For example, the transition value may be described as a value that is added to or subtracted from the accumulated value when transitioning from one state to another. For convenience of explanation, it is described as an operation of adding or subtracting transition values, but the embodiments are not limited thereto. For example, when at least one processor (207) identifies a transition from one state to another, it can obtain an accumulated value through complex mathematical calculations.
[0070] According to one embodiment, at least one processor (207) may determine a target state from among predefined states of one type (a) based on receiving user input for determining a target state. For example, the target state may be described as a trigger state for providing a super slow motion mode. For example, in FIG. 7, the target state may be a fourth state (740). At least one processor (207) may identify the state of a visual object in the video (440). For example, at least one processor (207) may provide a super slow motion mode based on identifying the target state. For example, at least one processor (207) may set a reference value based on the determination that the state of a visual object included in the video (440) is a first state (710). For example, the reference value corresponding to the first state (710) may be used to calculate a cumulative value.
[0071] According to one embodiment, at least one processor (207) can sequentially identify a first state (710), a second state (720), a third state (730), and a fourth state (740). For example, at least one processor (207) can obtain a first accumulated value (e.g., 0) which is a reference value based on identifying the first state (710). For example, at least one processor (207) can obtain a second accumulated value (e.g., 1) by calculating a transition value (712) (e.g., 1) for the first accumulated value based on identifying the second state (720) following the first state (710). For example, at least one processor (207) can obtain a third accumulated value (e.g., 2) by calculating a transition value (722) for the second accumulated value based on identifying the third state (730) following the second state (720). For example, at least one processor (207) can obtain a fourth accumulated value (e.g., 3) by calculating a transition value (732) to the third accumulated value based on identifying a fourth state (740) following the third state (730). For example, at least one processor (207) can switch the mode of the camera (203) from normal mode to super slow motion mode based on a determination that the calculated accumulated value (e.g., 3) exceeds a threshold value (e.g., 2.5). For example, at least one processor (207) can obtain video of a second frame rate (or a second FPS (frames per second)) higher than the first frame rate through the camera (203) based on a determination that the accumulated value exceeds a threshold value.
[0072] According to one embodiment, at least one processor (207) can sequentially identify a first state (710), a third state (730), and a fourth state (740). For example, based on identifying the first state (710), at least one processor (207) can determine a reference value corresponding to the first state (710) as a first accumulated value (e.g., 0). For example, based on identifying the third state (730) following the first state (710), at least one processor (207) can obtain a second accumulated value (e.g., 2) by calculating a transition value (736) on the first accumulated value. For example, based on identifying the fourth state (740) following the third state (730), at least one processor (207) can obtain a third accumulated value (e.g., 3.0) by calculating a transition value (732) on the second accumulated value. For example, at least one processor (207) may switch the mode of the camera (203) from normal mode to super slow motion mode based on a determination that the calculated accumulated value (e.g., 3) exceeds a threshold value (e.g., 2.5). For example, at least one processor (207) may acquire a video of a second frame rate higher than a first frame rate through the camera (203) based on a determination that the accumulated value exceeds a threshold value.
[0073] According to one embodiment, at least one processor (207) can sequentially identify a first state (710), a second state (720), a second state (720) (identifying the second state (720) sequentially), a third state (730), and a fourth state (740). At least one processor (207) can determine a reference value as a first accumulated value (e.g., 0) based on identifying the first state (710). For example, at least one processor (207) can obtain a second accumulated value (e.g., 1) by calculating a transition value (712) to the first accumulated value based on identifying the second state (720) following the first state (710). For example, at least one processor (207) can obtain a third accumulated value (e.g., 1) by calculating a transition value (e.g., 0) on the second accumulated value based on identifying the second state (720) following the second state (720). For example, at least one processor (207) can obtain a fourth accumulated value (e.g., 2) by calculating a transition value (722) on the third accumulated value based on identifying the third state (730) following the second state (720). For example, at least one processor (207) can obtain a fifth accumulated value (e.g., 3) by calculating a transition value (732) on the fourth accumulated value based on identifying the fourth state (740) following the third state (730). For example, at least one processor (207) may switch the mode of the camera (203) from normal mode to super slow motion mode based on a determination that the calculated accumulated value (e.g., 3) exceeds a threshold value (e.g., 2.5). For example, at least one processor (207) may acquire a video of a second frame rate higher than a first frame rate through the camera (203) based on a determination that the accumulated value exceeds a threshold value.
[0074] According to one embodiment, at least one processor (207) can sequentially identify a first state (710), a second state (720), a third state (730), and a fourth state (740). At least one processor (207) can determine a reference value as a first accumulated value (e.g., 0) based on identifying the first state (710). For example, at least one processor (207) can obtain a second accumulated value (e.g., 1) by calculating a transition value (712) on the first accumulated value based on identifying the second state (720) following the first state (710). For example, at least one processor (207) can obtain a third accumulated value (e.g., 0) by calculating a transition value (714) (e.g., -1) on the second accumulated value based on identifying the first state (710) following the second state (720). For example, at least one processor (207) can obtain a fourth accumulated value (e.g., 1) by calculating a transition value (712) on the third accumulated value based on identifying the second state (720) following the first state (710). For example, at least one processor (207) can obtain a fifth accumulated value (e.g., 2) by calculating a transition value (722) on the fourth accumulated value based on identifying the third state (730) following the second state (720). For example, at least one processor (207) can obtain a sixth accumulated value (e.g., 3) by calculating a transition value (732) to the fifth accumulated value based on identifying a fourth state (740) following the third state (730). For example, at least one processor (207) can switch the mode of the camera (203) from normal mode to super slow motion mode based on the determination that the calculated accumulated value (e.g., 3) exceeds a threshold value (e.g., 2.5).For example, at least one processor (207) can acquire a video of a second frame rate higher than a first frame rate through a camera (203) based on a determination that the accumulated value exceeds a threshold value.
[0075] According to one embodiment, a Markov chain (or model) can be described as a model in which the next state depends on the current state. In a Markov chain, a first state can transition to one of a second state and a third state. For example, in a Markov chain, a first state can transition to a second state or a third state depending on a first probability of transitioning from a first state to a second state and a second probability of transitioning from a first state to a third state. For example, the sum of the probabilities of transitioning from the first state to the next state can be 1.
[0076] At least one processor (207) can identify that the state of a visual object transitions from a reference state to the next state. For example, the state of a visual object may transition from a first state (710) to one of a second state (720), a third state (730), and a fourth state (740). The state of a visual object may transition from a first state (710) to a fourth state (740). The state of a visual object may transition in the order of a first state (710), a third state (730), and a fourth state (740). The state of a visual object may transition in the order of a first state (710), a second state (720), a third state (730), and a fourth state (740). A first accumulated value corresponding to the sequential order of the first state (710) and the fourth state (740) of the state of the visual object may be the same as a second accumulated value corresponding to the sequential order of the first state (710), the third state (730), and the fourth state (740) of the state of the visual object. The second accumulated value may be the same as a third accumulated value corresponding to the sequential order of the first state (710), the second state (720), the third state (730), and the fourth state (740) of the state of the visual object. For example, since the first accumulated value, the second accumulated value, and the third accumulated value are the same as each other, at least one processor (207) may determine the timing for executing a super slow motion mode according to a single threshold value (e.g., 2.5).
[0077] According to one embodiment, in a Markov chain, a first probability corresponding to the sequential order of the first state (710) and the fourth state (740) may differ from a second probability corresponding to the sequential order of the first state (710), the third state (730), and the fourth state (740). For example, in a Markov chain, the probability of reaching the current state may be based on the path of previous states. For example, the path may be described as the states identified (e.g., the first state (710) and the third state (730)) to reach the current state (e.g., the fourth state (740)). For example, because the first probability differs from the second probability, it may be difficult to determine the threshold probability for executing a super slow motion mode. For example, when the threshold probability is greater than the first probability and less than the second probability, at least one processor (207) may refrain from, bypass, or block the acquisition of video based on super slow motion mode along the sequential path of the first state (710) and the fourth state (740). For example, when the threshold probability is greater than the first probability and less than the second probability, at least one processor (207) may acquire video based on super slow motion mode along the sequential path of the first state (710), the third state (730), and the fourth state (740). For example, at least one processor (207) can calculate cumulative values using transition values (712) (e.g., 1), transition values (714) (e.g., -1), transition values (716) (e.g., 3), transition values (722) (e.g., 1), transition values (724) (e.g., -1), transition values (726) (e.g., 2), transition values (732) (e.g., 1), transition values (734) (e.g., -2), and transition values (736) (e.g., 2) indicated by transition information (625), so that cumulative values independent of the path can be calculated.For example, since at least one processor (207) calculates a path-independent cumulative value, it can reliably switch to super slow motion mode.
[0078] At least one processor (207) can execute a super slow motion mode by identifying motion within an ROI (e.g., ROI (945) of FIG. 9) based on a determination that the type of video (410) is a first type in which a different state from a predefined state is identified. For example, the operation of identifying motion within the ROI is described and illustrated in more detail with reference to FIG. 8.
[0079] FIG. 8 is a flowchart illustrating the operation of an electronic device that performs a high-speed shooting function through a region of interest (ROI). This method can be executed by the electronic device (100) shown in FIG. 2 or by at least one processor (207) of the electronic device (100).
[0080] Referring to FIG. 8, in operation 810, at least one processor (207) can acquire a first video at a first frame rate through a camera (203). For example, the first video may be an example of a video (410). For example, the first video may be described as a video acquired at a first frame rate in normal mode.
[0081] In operation 820, at least one processor (207) can determine whether the first video acquired through the camera (203) is of the first type. For example, at least one processor (207) can execute operation 830 based on the determination that the first video is of the first type, and execute operation 840 based on the determination that the first video is not of the first type. For example, the operation of determining whether the first video is of the first type can be described with reference to FIG. 4. For example, at least one processor (207) can determine whether the video (410) is of the first type by using the first trained model (420).
[0082] According to one embodiment, at least one processor (207) may determine whether the video (410) is of a first type based on receiving user input. For example, at least one processor (207) may receive user input to determine whether a first video (e.g., video (410)) acquired through a camera (203) is of a first type. For example, at least one processor (207) may determine that the type of the first video is of a first type based on receiving user input indicating that the first video is of a first type (e.g., user input regarding an executable object). For example, at least one processor (207) may determine that the type of the first video is of a different type (e.g., a second type) different from the first type based on receiving other user input indicating that the first video is not of a first type. For example, at least one processor (207) may refrain from, bypass, or block the identification or determination of the video type using the first trained model (420) when determining the video type based on user input.
[0083] In operation 830, at least one processor (207) can acquire a second video at a second frame rate through a camera (203) by identifying motion of a visual object within a region of interest (ROI) of the first video (e.g., ROI (945) in FIG. 9) based on a determination that the type of the first video is a first type. For example, at least one processor (207) can identify motion of a visual object included in the ROI of the first video (e.g., ROI (945) in FIG. 9) based on a determination that the type of the first video is a first type. For example, at least one processor (207) can perform high-speed shooting based on identifying motion of a visual object included in the ROI. At least one processor (207) can acquire a second video at a second frame rate by identifying motion within the ROI of the first video acquired through the camera (203). For example, the second frame rate may be greater than the first frame rate. For example, the second video may be described as a video acquired according to a super slow motion mode. For example, at least one processor (207) may identify the motion of a visual object included in the ROI by detecting changes between frames. For example, the acquisition of the second video based on the ROI will be described later with reference to FIG. 9.
[0084] In operation 840, at least one processor (207) can acquire the second video using a third trained model (520) based on the determination that the first video acquired through the camera (203) is not of the first type. For example, at least one processor (207) can use the first trained model (420) to determine the type of the first video as one of the remaining types that are different from the first type (e.g., the second type). For example, at least one processor (207) can use a trained model corresponding to the determined one (e.g., the third trained model (520)) to acquire state data (e.g., state data (530)) representing the state of a visual object included in the first video. For example, at least one processor (207) can determine the timing for the execution of the super slow motion mode by sequentially identifying the states of the visual object using the state data. For example, at least one processor (207) can calculate cumulative values using transition values (e.g., transition values (712, 714, 716, 722, 724, 726, 732, 734, 736)) by using a fourth trained model (620) that includes transition information (625). For example, at least one processor (207) can switch or change the mode for acquiring video from normal mode to super slow motion mode based on the calculated cumulative values.
[0085] At least one processor (207) can switch the mode for acquiring video from normal mode to super slow motion mode based on identifying motion within an ROI (e.g., ROI (945) of FIG. 9) for a first type of video (430) that is different from the remaining types in which predefined states are identified. For example, the operation of identifying motion within the ROI is described and illustrated in more detail with reference to FIG. 9.
[0086] FIG. 9 illustrates an exemplary operation of an electronic device that performs a high-speed shooting function based on identifying movement within an ROI.
[0087] Referring to FIG. 9, at least one processor (207) can acquire a video (940) representing a user (120) through a camera (203). For example, at least one processor (207) can acquire the video in a super slow motion mode based on identifying the motion of a visual object included in the ROI (945) of the video (940). For example, at least one processor (207) can identify the motion of a visual object included in the ROI (945) using the ROI (945) of the video (940). For example, at least one processor (207) can identify the motion of a visual object included in the ROI (945) based on the difference between frames of the video (940). For example, at least one processor (207) can acquire a vector corresponding to a pixel by calculating the position change of the pixels included in each frame of the video (940). For example, at least one processor (207) can identify the motion of a visual object by analyzing a vector corresponding to the visual object included in each frame of the video (940). For example, at least one processor (207) can determine the timing for high-speed shooting by analyzing information (e.g., vector, change in position, etc.) about the visual object included in each frame of the video (940). For example, at least one processor (207) can switch the mode for acquiring video from normal mode to super slow motion mode based on identifying the motion of the visual object within the ROI (945). For example, at least one processor (207) can acquire a video of a second frame rate through the camera (203) based on identifying the motion of the visual object within the ROI (945).
[0088] According to one embodiment, at least one processor (207) can perform high-speed shooting based on an event identifying motion within the ROI (945). For example, at least one processor (207) can perform high-speed shooting based on a motion trigger within the ROI (945).
[0089] FIG. 10 is a flowchart illustrating the operation of an electronic device that determines the execution timing of a high-speed shooting function using transition information. This method may be executed by the electronic device (100) illustrated in FIG. 2 or by at least one processor (207) of the electronic device (100).
[0090] Referring to FIG. 10, in operation 1010, at least one processor (207) can identify the states of visual objects included in the first video based on acquiring a first video at a first frame rate through a camera (203). At least one processor (207) can identify the states of visual objects included in the video (440) at the first frame rate (sequentially). For example, at least one processor (207) can identify the states of visual objects (sequentially) using one of the trained models (e.g., second trained models (510-1, 510-2, ..., 510-N)) for defining the states of visual objects (e.g., third trained model (520)). For example, at least one processor (207) can identify the states of visual objects included in the video (410) by performing at least one of object recognition and image classification on frames of the video (410) of the first frame rate.
[0091] According to one embodiment, a video (440) of a first frame rate can be described as a video whose type is determined using a first trained model (420). For example, a video (440) of a first frame rate in which the states of visual objects are identified can be described as a video determined to be one of the remaining types (e.g., a second type) different from the first type through the first trained model (420) for determining the type of video. For example, a video of the first type can be described as a video in which an undefined state is identified. For example, a video of one of the remaining types can be described as a video determined to represent one (a) theme (e.g., golf) based on at least one of object recognition and image classification.
[0092] In operation 1020, at least one processor (207) can determine the timing for acquiring a second video at a second frame rate higher than the first frame rate by using transition information (625) to determine the execution timing of a high-speed shooting function based on the transition between the states, based on identifying the states of visual objects included in the first video. At least one processor (207) can determine the timing for acquiring a video at the second frame rate by using the transition information (625). For example, at least one processor (207) can acquire or calculate cumulative values for each of the sequentially identified states by using the transition information (625). For example, at least one processor (207) can calculate or obtain an accumulated value based on a transition value (e.g., transition value (712) or transition value (714)) corresponding to a transition between a first state (710) and a second state (720) using transition information (625). At least one processor (207) can determine the timing for executing a super slow motion mode based on a determination that the calculated accumulated value exceeds a threshold value.
[0093] According to one embodiment, at least one processor (207) can calculate cumulative values according to each of the states using transition information (625) based on identifying the states of a visual object within a first video. For example, each of the states may be identified based on one of the second trained models (510-1, 510-2, ..., 510-N). For example, each second trained model may be trained to output data representing the state of a visual object included in the video among states predefined for one type. For example, at least one processor (207) may determine the timing for acquiring the second video based on acquiring a cumulative value that exceeds a threshold value among the cumulative values. For example, the calculation of the cumulative value may be described with reference to FIGS. 6 and FIGS. 7.
[0094] In operation 1030, at least one processor (207) can acquire a second video at a second frame rate through the camera (203) based on a determined timing. At least one processor (207) can acquire a video at a second frame rate (e.g., 960 fps) through the camera (203) according to the determined timing. For example, at least one processor (207) can acquire a video at a second frame rate (e.g., 960 fps) through the camera (203) according to a super slow motion mode based on a determination that a calculated accumulated value exceeds a threshold value.
[0095] According to one embodiment, the determined timing may be determined based on an accumulated value calculated according to the transition between states. For example, the accumulated value may be based on a model trained to identify whether the accumulated value calculated according to the transition between states exceeds a threshold value (e.g., a fourth trained model (620)). Although the operation of calculating the accumulated value through the model (e.g., the fourth trained model (620)) has been described above, the embodiment is not limited thereto. For example, when at least one processor (207) identifies states using the third trained model (520), it may determine the timing for high-speed shooting to be performed without using the model by using the transition value mapped to the transition between states.
[0096] FIG. 11 is a block diagram of an electronic device in a network environment according to various embodiments.
[0097] FIG. 11 is a block diagram of an electronic device (1101) in a network environment (1100) according to various embodiments. Referring to FIG. 11, in the network environment (1100), the electronic device (1101) may communicate with an electronic device (1102) through a first network (1198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (1104) or a server (1108) through a second network (1199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1101) may communicate with the electronic device (1104) through a server (1108). According to one embodiment, the electronic device (1101) may include a processor (1120), memory (1130), input module (1150), sound output module (1155), display module (1160), audio module (1170), sensor module (1176), interface (1177), connection terminal (1178), haptic module (1179), camera module (1180), power management module (1188), battery (1189), communication module (1190), subscriber identification module (1196), or antenna module (1197). In some embodiments, at least one of these components (e.g., connection terminal (1178)) may be omitted from the electronic device (1101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (1176), camera module (1180), or antenna module (1197)) may be integrated into a single component (e.g., display module (1160)).
[0098] The processor (1120) can, for example, execute software (e.g., program (1140)) to control at least one other component (e.g., hardware or software component) of the electronic device (1101) connected to the processor (1120) and perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (1120) can store commands or data received from other components (e.g., sensor module (1176) or communication module (1190)) in volatile memory (1132), process the commands or data stored in volatile memory (1132), and store the resulting data in non-volatile memory (1134). According to one embodiment, the processor (1120) may include a main processor (1121) (e.g., a central processing unit or an application processor) or an auxiliary processor (1123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (1101) includes a main processor (1121) and an auxiliary processor (1123), the auxiliary processor (1123) may be configured to use less power than the main processor (1121) or to be specialized for a specified function. The auxiliary processor (1123) may be implemented separately from the main processor (1121) or as part thereof.
[0099] The auxiliary processor (1123) may control at least some of the functions or states associated with at least one component of the electronic device (1101) (e.g., display module (1160), sensor module (1176), or communication module (1190)) on behalf of the main processor (1121) while the main processor (1121) is in an inactive (e.g., sleep) state, or together with the main processor (1121) while the main processor (1121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (1123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (1180) or communication module (1190)). According to one embodiment, the auxiliary processor (1123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (1101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (1108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0100] The memory (1130) can store various data used by at least one component of the electronic device (1101) (e.g., processor (1120) or sensor module (1176)). The data may include, for example, software (e.g., program (1140)) and input or output data for related commands. The memory (1130) may include volatile memory (1132) or non-volatile memory (1134).
[0101] The program (1140) may be stored as software in memory (1130) and may include, for example, an operating system (1142), middleware (1144), or an application (1146).
[0102] The input module (1150) can receive commands or data to be used for a component of the electronic device (1101) (e.g., processor (1120)) from outside the electronic device (1101) (e.g., user). The input module (1150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0103] The sound output module (1155) can output a sound signal to the outside of the electronic device (1101). The sound output module (1155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0104] The display module (1160) can visually provide information to an external (e.g., user) of the electronic device (1101). The display module (1160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (1160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0105] The audio module (1170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (1170) can acquire sound through the input module (1150) or output sound through the sound output module (1155) or an external electronic device (e.g., electronic device (1102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (1101).
[0106] The sensor module (1176) can detect the operating state of the electronic device (1101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (1176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0107] The interface (1177) may support one or more specified protocols that can be used for the electronic device (1101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (1102)). According to one embodiment, the interface (1177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0108] The connection terminal (1178) may include a connector through which the electronic device (1101) can be physically connected to an external electronic device (e.g., electronic device (1102)). According to one embodiment, the connection terminal (1178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0109] The haptic module (1179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (1179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0110] The camera module (1180) can capture still images and video. According to one embodiment, the camera module (1180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0111] The power management module (1188) can manage the power supplied to the electronic device (1101). According to one embodiment, the power management module (1188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0112] The battery (1189) can supply power to at least one component of the electronic device (1101). According to one embodiment, the battery (1189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0113] The communication module (1190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (1101) and an external electronic device (e.g., electronic device (1102), electronic device (1104), or server (1108)), and the performance of communication through the established communication channel. The communication module (1190) may include one or more communication processors that operate independently of the processor (1120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1190) may include a wireless communication module (1192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (1194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (1104) through a first network (1198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (1199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1192) can identify or authenticate the electronic device (1101) within a communication network such as the first network (1198) or the second network (1199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (1196).
[0114] The wireless communication module (1192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (1192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (1192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (1192) can support various requirements specified in the electronic device (1101), external electronic device (e.g., electronic device (1104)), or network system (e.g., second network (1199)). According to one embodiment, the wireless communication module (1192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0115] An antenna module (1197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (1197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (1197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (1198) or a second network (1199), may be selected from the plurality of antennas, for example, by a communication module (1190). A signal or power may be transmitted or received between the communication module (1190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (1197).
[0116] According to various embodiments, the antenna module (1197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0117] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0118] According to one embodiment, commands or data may be transmitted or received between the electronic device (1101) and an external electronic device (1104) through a server (1108) connected to a second network (1199). Each of the external electronic devices (1102, or 1104) may be the same or a different type of device as the electronic device (1101). According to one embodiment, all or part of the operations performed on the electronic device (1101) may be performed on one or more of the external electronic devices (1102, 1104, or 1108). For example, if the electronic device (1101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (1101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (1101). The electronic device (1101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (1101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (1104) may include an Internet of Things (IoT) device. The server (1108) may be an intelligent server using machine learning and / or neural networks.According to one embodiment, an external electronic device (1104) or server (1108) may be included within the second network (1199). The electronic device (1101) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0119] Figure 12 is a schematic diagram of an exemplary artificial intelligence (AI) system.
[0120] Referring to FIG. 12, the AI system (1200) may include an input / output interface (1210), an AI framework (1220), a generative AI model (1230), and / or a knowledge repository (1290).
[0121] The input / output interface (1210) can receive input. The input may include user input and / or data obtained or generated by an electronic device (e.g., the electronic device (100) or electronic device (1101) described above). The above data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (207) or processor (1120)) (e.g., illumination data around the electronic device obtained from a sensor or sensor hub (e.g., auxiliary processor (1123)), posture data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., temperature of the display (208) or at least one processor (207)), size information of the display area of the display (208), and / or images obtained through an image sensor of the electronic device (e.g., included in the camera module (1180)). The user input may include natural language, touch data obtained through a touch circuit included in the display panel (e.g., used to identify input from a finger and / or stylus), images displayed (and / or to be displayed) on the display panel, and / or videos. By example, without limitation, the user input may be provided to the input / output interface (1210) along with context information. It may be received by. The above context information may be described as additional information obtained in relation to the user input. The above context information may be related to the state at the time the user input is received (e.g., the state of the electronic device and / or the state of the surroundings of the electronic device (e.g., user state)). For example, the above context information may include information about one or more software applications executed within the electronic device at the time the user input is received.For example, the above situation information may include information about the location of the electronic device (or the location of the user of the electronic device) when the user input is received. For example, the user input may be integrated with the situation information. For example, the user input with the situation information integrated as the input may be received by the input / output interface (1210).
[0122] The input / output interface (1210) may transmit (or provide) an output. The output may include a result (or result information) generated or obtained by the AI system (1200) based on at least part of the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., media content and / or multimedia content). For example, the output may include actions related to the user of the electronic device. For example, the output may have a format according to the user settings of the electronic device.
[0123] The input / output interface (1210) can be described as a user question / response interface (1207).
[0124] The AI framework (1220) can be used to obtain information (or data) about the input from the input / output interface (1210) and to control one or more components related to the AI system (1200) using the obtained information.
[0125] For example, a prompt design component (1221) within an AI framework (1220) can generate or obtain prompts for a generative AI model (1230) (e.g., including a large language model (LLM) or a large multimodal model (LMM)) using the acquired information. For example, the prompt design component (1221) may be described as an AI component that uses a learning algorithm and / or a neural network to provide prompts that are enhanced over time. For example, the prompt design component (1221) can generate or obtain prompts by accessing a knowledge component (e.g., a knowledge repository (1290)) containing user preference data, a prompt library, and / or prompt examples using the acquired information. The generated prompts may be provided to the generative AI model (1230) (e.g., including an LLM or LMM).
[0126] For example, an API / plugin management component (1222) within the AI framework (1220) may be used to support communication for additional information requested (or induced) in relation to the prompt provided (or to be provided) to the generative AI model (1230). For example, the API / plugin management component (1222) may be used to create or establish a channel for communication with various data sources (e.g., knowledge repository (1290)). For example, the API / plugin management component (1222) may support access to at least some of the data sources. For example, the API / plugin management component (1222) may be used to request another component (e.g., application / service component (1280)) that performs feedback (or response) according to the prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1222) may be provided to the prompt design component (1221) for generating a prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1222) may be provided to the generative AI model (1230).
[0127] For example, an improvement component (1223) within the AI framework (1220) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (1230). For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) is related to the input. For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) contains biased content. For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) contains harmful content. For example, the improvement component (1223) can support or assist in performing additional processing to improve the content obtained from the generative AI model (1230). For example, the improvement component (1223) may support providing a hint to the user to improve the content.
[0128] A generative AI model (1230) can be described as an artificial intelligence neural network that generates feedback in response to a prompt. For example, the feedback may include additional data and / or information relative to the prompt, but relative to the prompt. For example, the feedback may include new content relative to the prompt. For example, the generative AI model (1230) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, the generative AI model (1230) may include an LMM that generates the feedback by recognizing characters, images, and / or voice.
[0129] As an example without limitation, the AI framework (1220) and / or generative AI model (1230) may be included within an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI module may be operatively coupled with at least one processor of the electronic device (e.g., at least one processor (207) or processor (1120)). For example, the AI module may be operatively coupled with a display driving circuit of the electronic device (e.g., a display driving circuit or DDI). For example, the AI module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.
[0130] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs.
[0131] An electronic device as described above (e.g., electronic device (100)) may include a memory (e.g., memory (206)) for storing instructions. The electronic device may include a camera (e.g., camera (203)). The electronic device may include at least one processor (e.g., at least one processor (207)). The instructions may cause the electronic device to determine the type of the first video through a first trained model based on acquiring a first video at a first frame rate through the camera, when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to identify the state of a visual object included in the first video using a third trained model corresponding to the determined type among second trained models for identifying the state of a visual object, when executed individually or collectively by the at least one processor. When the above instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to acquire a second video through the camera at a second frame rate higher than the first frame rate, based on the determination that the identified state is a target state.
[0132] According to one embodiment, the second video is acquired during a first time interval based on high-speed shooting, and can be played back through a display during a second time interval longer than the first time interval.
[0133] According to one embodiment, the first trained model can be trained to determine the type of video among a plurality of types by performing image classification on the frames of the video.
[0134] According to one embodiment, the third trained model can be trained to output data representing the state of a visual object included in the video among states predefined for one type (a) using the video.
[0135] According to one embodiment, the instructions may cause the electronic device to calculate an accumulated value corresponding to the identified state by using transition information for calculating an accumulated value according to a transition between states, based on identifying the state of the visual object when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to acquire the second video based on a determination that the accumulated value calculated according to identifying the target state exceeds a threshold value when executed individually or collectively by the at least one processor.
[0136] According to one embodiment, the accumulated value can be calculated based on a reference value corresponding to a reference state and a transition value corresponding to a transition between states indicated by the transition information.
[0137] According to one embodiment, the instructions may cause the electronic device to acquire the second video by identifying the state of the visual object using the third trained model, based on the determination that the determined type is a first type, when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to acquire the second video by identifying the motion of the visual object within the Region of Interest (ROI) of the first video, based on the determination that the determined type is a second type, when executed individually or collectively by the at least one processor.
[0138] A method performed by an electronic device (e.g., electronic device (100)) having a camera (e.g., camera (203)) as described above may include an operation of determining the type of the first video through a first trained model based on acquiring a first video at a first frame rate through the camera. The method may include an operation of identifying the state of a visual object included in the first video using a third trained model corresponding to the determined type among second trained models for identifying the state of a visual object. The method may include an operation of acquiring a second video through the camera at a second frame rate higher than the first frame rate based on the determination that the identified state is a target state.
[0139] According to one embodiment, the second video is acquired during a first time interval based on high-speed shooting, and can be played back through a display during a second time interval longer than the first time interval.
[0140] According to one embodiment, the first trained model can be trained to determine the type of video among a plurality of types by performing image classification on the frames of the video.
[0141] According to one embodiment, the third trained model can be trained to output data representing the state of a visual object included in the video among states predefined for one type (a) using the video.
[0142] According to one embodiment, the method may include an operation of calculating an accumulated value corresponding to the identified state by using transition information for calculating an accumulated value according to a transition between states, based on identifying the state of the visual object. The method may include an operation of acquiring the second video based on a determination that the accumulated value calculated according to identifying the target state exceeds a threshold value.
[0143] According to one embodiment, the accumulated value can be calculated based on a reference value corresponding to a reference state and a transition value corresponding to a transition between states indicated by the transition information.
[0144] According to one embodiment, the method may include an operation of acquiring the second video by identifying the state of the visual object using the third trained model based on a determination that the determined type is a first type. The method may also include an operation of acquiring the second video by identifying the motion of the visual object within the Region of Interest (ROI) of the first video based on a determination that the determined type is a second type.
[0145] In a computer-readable storage medium in which one or more programs as described above are stored, the one or more programs may include instructions that cause the electronic device (e.g., electronic device (100)) having a camera (e.g., camera (203)) to determine the type of the first video through a first trained model based on acquiring a first video through the camera when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to identify the state of a visual object included in the first video using a third trained model corresponding to the determined type among second trained models for identifying the state of a visual object when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire a second video through the camera at a second frame rate higher than the first frame rate based on a determination that the identified state is a target state when executed by the electronic device.
[0146] According to one embodiment, the second video is acquired during a first time interval based on high-speed shooting, and can be played back through a display during a second time interval longer than the first time interval.
[0147] According to one embodiment, the first trained model can be trained to determine the type of video among a plurality of types by performing image classification on the frames of the video.
[0148] According to one embodiment, the third trained model can be trained to output data representing the state of a visual object included in the video among states predefined for one type (a) using the video.
[0149] According to one embodiment, the one or more programs may include instructions that cause the electronic device to calculate an accumulated value corresponding to the identified state by using transition information for calculating an accumulated value according to a transition between states, based on identifying the state of the visual object when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the second video based on a determination that the accumulated value calculated according to identifying the target state exceeds a threshold value when executed by the electronic device.
[0150] According to one embodiment, the accumulated value can be calculated based on a reference value corresponding to a reference state and a transition value corresponding to a transition between states indicated by the transition information.
[0151] According to one embodiment, the one or more programs may include instructions that cause the electronic device to acquire the second video by identifying the state of the visual object using the third trained model based on the determination that the determined type is a first type when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the second video by identifying the motion of the visual object within the Region of Interest (ROI) of the first video based on the determination that the determined type is a second type when executed by the electronic device.
[0152] An electronic device as described above (e.g., electronic device (100)) may include a memory (e.g., memory (206)) for storing instructions. The electronic device may include a camera (e.g., camera (203)). The electronic device may include at least one processor (e.g., at least one processor (207)). The instructions may cause the electronic device to identify the states of a visual object included in the first video based on acquiring a first video at a first frame rate through the camera when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine the timing for acquiring a second video at a second frame rate higher than the first frame rate by using transition information to determine the timing according to the transition between the states based on identifying the states when executed individually or collectively by the at least one processor. When the above instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to acquire the second video through the camera at the second frame rate based on the determined timing.
[0153] According to one embodiment, the second video is acquired during a first time interval based on high-speed shooting, and can be played back through a display during a second time interval longer than the first time interval.
[0154] According to one embodiment, the instructions may cause the electronic device to calculate cumulative values according to each of the states using the transition information based on identifying the states when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine the timing for acquiring the second video based on obtaining a cumulative value that exceeds a threshold value among the cumulative values when executed individually or collectively by the at least one processor.
[0155] According to one embodiment, the first video may be a video of a type determined based on a first model trained to determine the type of video among a plurality of types by performing image classification on the frames of the video.
[0156] According to one embodiment, the states can be identified based on a second model trained to output data representing the state of a visual object included in the video among states predefined for one type (a) using the video.
[0157] According to one embodiment, the determined timing may be based on a third model trained to identify whether the accumulated value calculated according to the transition between the states exceeds a threshold value.
[0158] A method performed by an electronic device (e.g., electronic device (100)) having a camera (e.g., camera (203)) as described above may include an operation of identifying the states of a visual object included in the first video based on acquiring a first video at a first frame rate through the camera. The method may include an operation of determining the timing for acquiring a second video at a second frame rate higher than the first frame rate by using transition information to determine the timing according to the transition between the states based on identifying the states. The method may include an operation of acquiring the second video at the second frame rate through the camera based on the determined timing.
[0159] According to one embodiment, the second video is acquired during a first time interval based on high-speed shooting, and can be played back through a display during a second time interval longer than the first time interval.
[0160] According to one embodiment, the method may include an operation of calculating cumulative values according to each of the states using the transition information based on identifying the states. The method may include an operation of determining the timing for acquiring the second video based on obtaining a cumulative value that exceeds a threshold value among the cumulative values.
[0161] According to one embodiment, the first video may be a video of a type determined based on a first model trained to determine the type of video among a plurality of types by performing image classification on the frames of the video.
[0162] According to one embodiment, the states can be identified based on a second model trained to output data representing the state of a visual object included in the video among states predefined for one type (a) using the video.
[0163] According to one embodiment, the determined timing may be based on a third model trained to identify whether the accumulated value calculated according to the transition between the states exceeds a threshold value.
[0164] In a computer-readable storage medium in which one or more programs as described above are stored, the one or more programs may include instructions that cause the electronic device (e.g., electronic device (100)) having a camera (e.g., camera (203)) to identify states of visual objects included in the first video based on acquiring a first video at a first frame rate through the camera when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine a timing for acquiring a second video at a second frame rate higher than the first frame rate by using transition information to determine timing according to transitions between the states based on identifying the states when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the second video at the second frame rate through the camera based on the determined timing when executed by the electronic device.
[0165] According to one embodiment, the second video is acquired during a first time interval based on high-speed shooting, and can be played back through a display during a second time interval longer than the first time interval.
[0166] According to one embodiment, the one or more programs may include instructions that cause the electronic device to calculate cumulative values according to each of the states using the transition information based on identifying the states when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine the timing for acquiring the second video based on obtaining a cumulative value exceeding a threshold value among the cumulative values when executed by the electronic device.
[0167] According to one embodiment, the first video may be a video of a type determined based on a first model trained to determine the type of video among a plurality of types by performing image classification on the frames of the video.
[0168] According to one embodiment, the states can be identified based on a second model trained to output data representing the state of a visual object included in the video among states predefined for one type (a) using the video.
[0169] According to one embodiment, the determined timing may be based on a third model trained to identify whether the accumulated value calculated according to the transition between the states exceeds a threshold value.
[0170] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.
[0171] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0172] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0173] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0174] Various embodiments of this document may be implemented as software (e.g., program (#40)) comprising one or more instructions stored in a storage medium (e.g., internal memory (#36) or external memory (#38)) readable by a machine (e.g., electronic device (#01)). For example, a processor (e.g., processor (#20)) of the machine (e.g., electronic device (#01)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0175] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0176] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0177] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may continuously store a computer-executable program, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or several combined hardware, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Additionally, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.
[0178] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0179] Therefore, other implementations, other embodiments, and equivalents to the claims set forth below are also within the scope of the claims. According to one embodiment, the method according to the various embodiments disclosed herein may be provided as a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0180] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device, camera; Memory comprising one or more storage media for storing instructions; and It includes at least one processor comprising processing circuitry, and When the above instructions are executed individually or collectively by the at least one processor, Based on acquiring a first video at a first frame rate through the camera, the type of the first video is determined through a first trained model, and Among the second trained models for identifying the state of a visual object, the state of a visual object included in the first video is identified using a third trained model corresponding to the determined type, and Based on the determination that the above identified state is a target state, to acquire a second video through the camera at a second frame rate higher than the first frame rate, causing the above electronic device, Electronic device.
2. In claim 1, the second video is, Acquired during a first time interval based on performing high-speed shooting, and Played through a display during a second time interval longer than the first time interval mentioned above, Electronic device.
3. In claim 1, the first trained model is, Trained to determine the type of video among multiple types by performing image classification on video frames, Electronic device.
4. In claim 1, the third trained model is, Trained to use a video to output data representing the state of a visual object included in the video among predefined states for one (a) type, Electronic device.
5. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Based on identifying the state of the visual object, using transition information for calculating a cumulative value according to the transition between states, a cumulative value corresponding to the identified state is calculated, and Based on the determination that the accumulated value calculated by identifying the above target state exceeds a threshold value, to acquire the second video, causing the above electronic device, Electronic device.
6. In claim 5, the accumulated value is, Calculated based on a reference value corresponding to a reference state and a transition value corresponding to a transition between states indicated by the transition information, Electronic device.
7. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Based on the determination that the above-determined type is the first type, the state of the visual object is identified using the third trained model, thereby acquiring the second video, and Based on the determination that the above-determined type is the second type, the second video is acquired by identifying the motion of a visual object within the ROI (Region Of Interest) of the first video. causing the above electronic device, Electronic device.
8. In an electronic device, camera; Memory comprising one or more storage media for storing instructions; and It includes at least one processor comprising processing circuitry, and When the above instructions are executed individually or collectively by the at least one processor, Based on acquiring a first video at a first frame rate through the camera, the states of visual objects included in the first video are identified, and Based on identifying the above states, using transition information for determining timing according to the transition between the above states, determine the timing for acquiring a second video at a second frame rate higher than the first frame rate, and Based on the above-determined timing, to acquire the second video through the camera at the second frame rate, causing the above electronic device, Electronic device.
9. In claim 8, the second video is, Acquired during a first time interval based on performing high-speed shooting, and Played through a display during a second time interval longer than the first time interval mentioned above, Electronic device.
10. In claim 8, When the above instructions are executed individually or collectively by the at least one processor, Based on identifying the above states, using the above transition information, cumulative values are calculated according to the above states, respectively, and Based on obtaining an accumulated value exceeding a threshold value among the above accumulated values, to determine the timing for acquiring the second video, causing the above electronic device, Electronic device.
11. In claim 8, the first video is, A video of a type determined based on a first model trained to determine the type of video among multiple types by performing image classification on the frames of the video, Electronic device.
12. In claim 8, the above states are, Identifyed based on a second model trained to output data representing the state of a visual object included in the video among predefined states for one type (a) using the video, Electronic device.
13. In claim 8, the determined timing is, Based on a third model trained to identify whether the cumulative value calculated according to the transition between the above states exceeds a threshold value, Electronic device.
14. In a non-transient computer-readable storage medium storing one or more programs, said one or more programs are, When executed by an electronic device having a camera, Based on acquiring a first video at a first frame rate through the camera, the type of the first video is determined through a first trained model, and Among the second trained models for identifying the state of a visual object, the state of a visual object included in the first video is identified using a third trained model corresponding to the determined type, and Based on the determination that the above identified state is a target state, to acquire a second video through the camera at a second frame rate higher than the first frame rate, Including instructions that cause the above electronic device, Non-transient computer-readable storage media.
15. In claim 14, the second video is, Acquired during a first time interval based on performing high-speed shooting, and Played through a display during a second time interval longer than the first time interval mentioned above, Non-transient computer-readable storage media.