Electronic device and image generation method thereof
The electronic device addresses the challenge of identifying important frames by using tags and priority information to generate edited videos that highlight significant moments and user-defined preferences, resulting in improved content relevance and quality.
Patent Information
- Application Number
- PCT/KR2025/009985
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-07-09
- Publication Date
- 2026-02-05
AI Technical Summary
Existing electronic devices lack efficient methods for identifying and prioritizing important frames in video data based on subject motion and user input, leading to incomplete or irrelevant edited videos.
An electronic device equipped with a memory and processor that identifies tags corresponding to subject motion and determines frame importance, generating an edited video or image by including only frames with a predetermined level of importance, using priority information and user input to refine the selection process.
The solution enables the creation of edited videos that focus on meaningful frames, enhancing user experience by prioritizing significant moments and user-defined preferences, thus improving the quality and relevance of generated content.
Smart Images

Figure KR2025009985_05022026_PF_FP_ABST
Abstract
Description
Electronic device and method for generating images thereof
[0001] The present disclosure relates to an electronic device for generating an edited image and an image generating method thereof.
[0002] Advances in electronic technology have led to the emergence of various types of electronic devices in everyday life. Among these devices are those that generate edited video.
[0003] For example, there may be an electronic device that edits a video based on video data and generates an edited video.
[0004] According to at least one embodiment of the present disclosure, an electronic device includes a memory storing priority information and at least one processor. The at least one processor identifies at least one tag corresponding to a subject motion within at least one frame, and determines an importance for each of the at least one frame based on the priority information stored in the memory and the identified tag, and generates an edited video including at least one frame having an importance equal to or greater than a predetermined value based on the importance for each of the at least one frame.
[0005] According to at least one embodiment of the present disclosure, a method for generating an image of an electronic device includes the steps of: identifying at least one tag corresponding to a subject motion within at least one frame; determining an importance of the at least one frame based on priority information for each subject motion and the identified tag; and generating an edited image including at least one frame having an importance equal to or greater than a predetermined value based on the at least one importance of the frame.
[0006] A non-transitory readable recording medium according to at least one embodiment of the present disclosure has stored thereon a program for performing an image generation method of an electronic device, the method comprising: identifying at least one tag corresponding to a subject motion within at least one frame; determining an importance of the at least one frame based on priority information for each subject motion and the identified tag; and generating an edited image including at least one frame having an importance equal to or greater than a predetermined value based on the importance of the at least one frame.
[0007] FIG. 1 is a drawing for explaining the operation of an electronic device according to at least one embodiment of the present disclosure.
[0008] FIG. 2 is a drawing showing an example of a screen displayed by an electronic device according to at least one embodiment of the present disclosure.
[0009] FIG. 3 is a block diagram illustrating a configuration of an electronic device according to at least one embodiment of the present disclosure.
[0010] FIG. 4 is a detailed block diagram illustrating an electronic device according to at least one embodiment of the present disclosure.
[0011] FIG. 5 is a diagram illustrating a method for an electronic device according to at least one embodiment of the present disclosure to extract at least one frame.
[0012] FIG. 6 is a diagram illustrating a method for an electronic device according to at least one embodiment of the present disclosure to generate an edited image.
[0013] FIG. 7 is a diagram illustrating an example of an electronic device according to at least one embodiment of the present disclosure obtaining data from an external device and displaying an image.
[0014] FIG. 8 is a drawing showing an example of a screen displayed by an electronic device according to at least one embodiment of the present disclosure.
[0015] FIG. 9 is a drawing showing an example of a screen displayed by an electronic device according to at least one embodiment of the present disclosure.
[0016] FIG. 10 is a flowchart illustrating an image generation method of an electronic device according to at least one embodiment of the present disclosure.
[0017] The terms used in the various embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should be defined based on the meaning of the terms and the overall content of this disclosure, rather than simply their names.
[0018] It should be understood that the various embodiments of the present disclosure and the terminology used therein are not intended to limit the technical features described in the present disclosure to specific embodiments, but include various modifications, equivalents, or substitutes of the embodiments.
[0019] In connection with the description of the drawings, similar reference numerals may be used for similar or related components.
[0020] The singular form of a noun corresponding to an item may include one or more of said items, unless the relevant context clearly indicates otherwise.
[0021] In this disclosure, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof.
[0022] Terms such as "first," "second," or "first" or "second" may be used simply to distinguish one component from another and do not qualify the components in any other respect (e.g., importance or order).
[0023] When a component (e.g., a first component) is referred to as being “coupled” or “connected” to another component (e.g., a second component), with or without the terms “functionally” or “communicatively,” it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0024] Terms such as "include" or "have" are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in the present disclosure, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.
[0025] When a component is said to be “connected,” “coupled,” “supported,” or “in contact with” another component, this includes not only cases where the components are directly connected, coupled, supported, or in contact, but also cases where the components are indirectly connected, coupled, supported, or in contact through a third component.
[0026] When we say that a component is "on" another component, this includes not only cases where the component is in contact with the other component, but also cases where there is another component between the two components.
[0027] The term "and / or" includes any combination of a plurality of related described elements or any one of a plurality of related described elements.
[0028] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented as at least one processor, excluding any "modules" or "parts" that need to be implemented as specific hardware.
[0029] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0030] In this disclosure, the term user may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence electronic device).
[0031] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.
[0032] FIG. 1 is a drawing for explaining the operation of an electronic device according to at least one embodiment of the present disclosure.
[0033] In FIG. 1, the electronic device (100) is illustrated as a mobile phone, but the electronic device (100) can be implemented as various types of electronic devices such as a TV, a laptop, a monitor, a kiosk, a tablet PC, a smartphone, a user terminal, an electronic picture frame, a large format display (LFD), a digital signage, a digital information display (DID), a video wall, a projector display, a server device, etc.
[0034] However, the present invention is not limited thereto, and may be implemented as various devices such as an image processing device (e.g., a set-top box, a single connected box) that is connected to an electronic device and provides an image, depending on the case. In addition, the electronic devices mentioned above are merely examples, and in addition to the electronic devices mentioned above, a device that is connected to another electronic device or server and can perform the operations described below may be included in the electronic devices according to various embodiments of the present disclosure.
[0035] Referring to FIG. 1, the electronic device (100) can display an image on a display. Specifically, the electronic device (100) can display an image captured by a camera based on user input on the display.
[0036] According to FIG. 1, an electronic device (100) can acquire an image including at least one frame through an action of photographing a subject (1-1, 1-2) and display the image on a screen (101). The image being photographed can be analyzed in real time to generate at least one tag corresponding to the subject's motion within the frame.
[0037] At least one frame can be an image representing a specific moment in a video. Typically, a video can contain multiple frames per second (FPS). For example, a video shot at 24 FPS consists of 24 still images per second, which can be played back in rapid succession to create a moving image.
[0038] Specifically, the electronic device (100) can recognize the subject (1-1, 1-2) in real time for each frame while shooting an image. The electronic device (100) can recognize not only the action or expression of the subject (1-1, 1-2), but also the background atmosphere other than the subject (1-1, 1-2), the image quality and audio for each frame, and tag each piece of information.
[0039] For example, if the subjects (1-1, 1-2) are a wife and a daughter, tags may be generated in the form of "name of wife, daughter, family member, or person." A tag called "smile" may be generated by recognizing the facial expression of the subjects (1-1, 1-2).
[0040] However, this is not limited to cases where tags are generated by the electronic device (100), and users can also directly input or modify tags for each frame. That is, the tagging operation can be performed automatically by the electronic device or manually by the user. In this case, additional points can be awarded when calculating the importance of each frame for frames directly tagged by the user. At least one frame directly tagged by the user can be searched for separately or collected and confirmed by the electronic device (100).
[0041] Tags can be generated by recognizing visual and auditory elements. Visual elements can include subjects, motions / expressions, and backgrounds, while auditory elements can include speakers, specific words, and event sounds (e.g., applause, laughter, cheers, etc.).
[0042] For example, the electronic device (100) may recognize cheers and applause in the introduction of a lecture video and generate tags such as "cheers" and "applause" in the corresponding frames. Furthermore, if the lecturer is a celebrity, the electronic device (100) may recognize the lecturer's face and generate tags using his or her name. Furthermore, the electronic device (100) may recognize specific words frequently mentioned in the lecturer's voice and generate tags representing the lecture topic.
[0043] Once tags are created, they can be used to extract at least one frame to include in an edited video, and the tags created in the video can also be displayed to allow navigation, sharing, and searching by those tags.
[0044] At least one frame may be described by various terms, such as Best moment, Highlight cut, or section, but in this disclosure, it is described as at least one frame. An edited video may be described by various terms, such as Thumbnail, Section, Highlight video, Story video, or Auto Sequence, but in this disclosure, it is described as an edited video.
[0045] Meanwhile, the various embodiments of the present disclosure described below do not necessarily require that the electronic device (100) be preceded by an operation of photographing and displaying a subject (1-1, 1-2) as illustrated in FIG. 1. The electronic device (100) may also receive image data photographed by another electronic device and perform an operation according to the present embodiment. Specific details of obtaining image data from another electronic device are described in detail in FIG. 6.
[0046] FIG. 2 is a drawing showing an example of a screen displayed by an electronic device according to at least one embodiment of the present disclosure.
[0047] Referring to FIG. 2, at least one processor (120) can recognize the action, expression, voice, etc. of the subject in real time while shooting to generate tags, and display all or part of the generated tags on the image displayed on the screen.
[0048] For example, when a video button (21) is pressed to start shooting a subject (1), at least one processor (120) can analyze the subject's actions, facial expressions, etc. in real time to generate tags (22) of 'laughter, applause, smile, heart' and display them on the screen (101) of the electronic device (100). The user can quickly identify the subject, characters, key points, etc. of the selected video through the tags (22) displayed on the screen (101) of the electronic device (100).
[0049] FIG. 3 is a block diagram illustrating a configuration of an electronic device according to at least one embodiment of the present disclosure.
[0050] Referring to FIG. 3, the electronic device (100) includes a memory (110) and at least one processor (120).
[0051] According to an embodiment, the memory (110) may store data required for various embodiments of the present disclosure. Depending on the purpose of data storage, the memory (110) may be implemented in the form of memory embedded in the electronic device (100) or may be implemented in the form of memory that is attachable to the electronic device (100).
[0052] For example, data for driving an electronic device (100) may be stored in a memory embedded in the electronic device (100), and data for an extended function of the electronic device (100) may be stored in a memory that can be attached or detached to the electronic device (100).
[0053] In the case of memory embedded in an electronic device (100), it may be implemented in the form of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)).
[0054] In the case of a memory that can be attached or detached to an electronic device (100), it can be implemented in the form of a memory card (e.g., CF (compact flash), SD (secure digital), Micro-SD (micro secure digital), Mini-SD (mini secure digital), xD (extreme digital), MMC (multi-media card), etc.), an external memory that can be connected to a USB port (e.g., USB memory), etc.
[0055] According to an embodiment, the memory (110) may store a computer program including at least one instruction or instructions for controlling the electronic device (100).
[0056] For example, the memory (110) may store an application capable of executing various operations of the present disclosure, such as an AI application capable of generating an edited image.
[0057] According to an embodiment, the memory (110) may store information on priorities for each subject motion, preset data, user information, etc. Specifically, priority information may be stored for subject motions based on changes in the subject's actions, facial expressions, sounds, and mood, and user motions based on zoom-in motions during shooting, actions for changing shooting options, actions for creating tags, and actions for giving voice commands.
[0058] Subject motion can refer to actions other than user input, such as the subject's actions, facial expressions, sounds, and moods. Conversely, user motion can refer to user input actions, such as zooming in while shooting, changing shooting options, creating tags, and giving voice commands.
[0059] According to an embodiment, the memory (110) may store information about the user. Specifically, the memory may store information about the user, such as the user's relationship information, the user's interest information, the user's video editing pattern, the user's personal information, and the user's calendar information.
[0060] According to an embodiment, the memory (110) may store information about the identified tag for each of at least one frame. Specifically, if at least one tag corresponding to the subject motion within the frame is identified for each of at least one frame, the memory (110) may store information about the identified tag for each of at least one frame.
[0061] According to an embodiment, at least one processor (120) controls the overall operation of the electronic device (100). Specifically, at least one processor (120) may be connected to each component of the electronic device (100) to control the overall operation of the electronic device (100).
[0062] At least one processor (120) can perform operations of the electronic device (100) according to various embodiments by executing at least one instruction stored in memory.
[0063] According to an embodiment, at least one processor (120) may be implemented as a digital signal processor (DSP), a microprocessor, or a timing controller (TCON) for processing a digital signal. However, the present invention is not limited thereto, and may include one or more of a central processing unit (CPU), a micro controller unit (MCU), a micro processing unit (MPU), a controller, an application processor (AP), a communication processor (CP), an ARM processor, or an artificial intelligence (AI) processor, or may be defined by the terms thereof.
[0064] At least one processor (120) may be implemented as a SoC (System on Chip), an LSI (Large Scale Integration) with a built-in processing algorithm, or may be implemented in the form of an FPGA (Field Programmable Gate Array). At least one processor (120) may perform various functions by executing computer executable instructions stored in memory.
[0065] At least one processor (120) may include one or more of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an APU (Accelerated Processing Unit), a MIC (Many Integrated Core), a DSP (Digital Signal Processor), an NPU (Neural Processing Unit), a hardware accelerator, or a machine learning accelerator. The at least one processor (120) may control one or any combination of other components of the electronic device, and may perform operations related to communication or data processing.
[0066] At least one processor (120) may execute one or more programs or instructions stored in memory. For example, at least one processor (120) may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in memory.
[0067] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-dedicated processor).
[0068] At least one processor (120) may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicores or heterogeneous multicores).
[0069] When at least one processor (120) is implemented as a multi-core processor, each of the plurality of cores included in the multi-core processor may include internal processor memory such as cache memory and on-chip memory, and a common cache shared by the plurality of cores may be included in the multi-core processor. In addition, each of the plurality of cores (or some of the plurality of cores) included in the multi-core processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the plurality of cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.
[0070] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among a plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.
[0071] According to an embodiment, at least one processor (120) may identify, for each of at least one frame, at least one tag corresponding to a subject motion within the frame.
[0072] According to an embodiment, at least one frame may include data representing the subject's motion, such as the subject's action, facial expression, and mood. For example, if the subject is in a pose expressing a heart with their arms within the frame, the processor may generate a "heart" tag corresponding to the subject's motion, and identify the "heart" tag for the frame.
[0073] The electronic device (100) can acquire an image including at least one frame using at least one processor (120), analyze the image in real time while shooting, and generate at least one tag corresponding to the motion of a subject within the frame. Specifically, the at least one processor (120) can analyze the image in real time while shooting, identify motion of the subject, such as action, expression, and mood of the subject within the frame, and generate a tag corresponding to the identified motion of the subject.
[0074] According to an embodiment, at least one processor (120) may determine at least one frame-specific importance based on priority information stored in memory (110) and identified tags.
[0075] For example, at least one processor (120) may give a higher priority to a frame containing a person when comparing a frame containing a person with a frame containing only a background.
[0076] At least one processor (120) may determine that among the frames containing a person, a frame with a “static” tag is more important than a frame with a “jumping” tag.
[0077] According to an embodiment, at least one processor (120) may generate an edited video including at least one frame having an importance equal to or greater than a predetermined value based on at least one frame-by-frame importance. If the frame does not have an importance equal to or greater than the predetermined value, the processor (120) may store the frame in the memory (110) without including it in the edited video.
[0078] For example, at least one processor (120) may determine at least one frame-by-frame importance and generate an edited video including at least one frame having an importance greater than a predetermined value. A detailed description of determining at least one frame-by-frame importance and generating an edited video will be described later in FIG. 6.
[0079] FIG. 4 is a detailed block diagram illustrating an electronic device according to at least one embodiment of the present disclosure.
[0080] Referring to FIG. 4, an electronic device (100) according to an embodiment of the present disclosure may include a memory (110), at least one processor (120), a camera (130), a communication interface (140), and a display (150). Parts that overlap with the above description are omitted or abbreviated.
[0081] The camera (130) is configured to photograph a subject, background, etc. In FIG. 4, one camera (130) is illustrated, but the camera (130) may include multiple cameras such as a stereo camera, a 3D camera, a TOF (Time of Flight) camera, a depth camera, a multi-lens array camera, a stereo vision system, a fused lidar camera, etc.
[0082] At least one processor (120) can acquire an image including at least one frame through a camera (130). At least one processor (120) can analyze the image being captured in real time and generate at least one tag corresponding to the motion of a subject within the frame.
[0083] However, the present invention is not limited thereto, and at least one frame according to an example may include at least one frame obtained from at least one external device via a communication interface (140).
[0084] The communication interface (140) can communicate with various electronic devices according to various types of communication methods. To this end, the communication interface (140) may include a communication module such as a short-range wireless communication module (not shown) or a wireless LAN communication module (not shown). Here, the short-range wireless communication module (not shown) is a communication module that performs data communication wirelessly with an electronic device located in close proximity.
[0085] For example, it can be a Bluetooth module, a ZigBee module, an NFC (Near Field Communication) module, etc. In addition, a wireless LAN communication module (not shown) can be a module that performs communication by connecting to an external network according to a wireless communication protocol such as WiFi, IEEE, etc.
[0086] The communication interface (140) may include a mobile communication module that connects to a mobile communication network and performs communication according to various mobile communication standards such as 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evoloution), 5G (5th Generation), etc.
[0087] The communication interface (140) may include at least one of wired communication modules (not shown) such as USB (Universal Serial Bus), IEEE (Institute of Electrical and Electronics Engineers) 1394, RS-232, etc., and may also include a broadcast reception module for receiving TV broadcasts.
[0088] The display (150) can display various screens and edited images under the control of the processor (120). The display (150) can be implemented as a display including self-luminous elements or a display including non-luminous elements and a backlight. Additionally, the display (150) can be implemented as an LFD display.
[0089] For example, it can be implemented as various types of displays such as LCD (Liquid Crystal Display), OLED (Organic Light Emitting Diodes) display, LED (Light Emitting Diodes), micro LED, Mini LED, PDP (Plasma Display Panel), QD (Quantum dot) display, QLED (Quantum dot light-emitting diodes), etc. The display (150) may also include a driving circuit, a backlight unit, etc., which can be implemented in a form such as a-si TFT, LTPS (low temperature poly silicon) TFT, OTFT (organic TFT), etc.
[0090] For example, at least one processor (120) may control the display (150) to generate tags for each frame and display at least one tag on the display (150). In addition, at least one processor (120) may generate an edited video including at least one frame having an importance equal to or greater than a predetermined value based on at least one importance for each frame, and display the edited video on the display (150). In addition, at least one processor may identify an unnecessary section in the video, skip the portion, and display the skipped portion on the display (150).
[0091] Unnecessary sections may include sections where obstacles or unknown people appear, the camera is obstructed, the camera is shaken, out of focus, the image quality is poor, the brightness level is outside the normal range, or the camera is being set up.
[0092] FIG. 5 is a diagram illustrating a method for an electronic device according to at least one embodiment of the present disclosure to extract at least one frame.
[0093] According to FIG. 5, at least one processor (120) can 'identify a priority for each subject motion based on the subject motion' (510) and 'adjust a priority for each subject motion based on the user motion' (520).
[0094] Here, "subject motion" can refer to actions other than user input, such as the subject's actions, facial expressions, sounds, and mood. Conversely, "user motion" can refer to user input actions, such as zooming in while shooting, changing shooting options, creating tags, and giving voice commands.
[0095] The action of 'identifying the priority of each subject action based on the subject action' can be subdivided into the actions of 'assigning priority by category' (511) and 'setting a high priority' (512). Here, the categories can include actions (511-1), expressions (511-2), moods (511-3), and sounds (511-4), and high priorities can be assigned in the order of actions (511-1), expressions (511-2), moods (511-3), and sounds (511-4).
[0096] However, this is not limited to this, and the detailed criteria for assigning priorities may vary depending on various embodiments. For example, high priorities may be assigned in the following order: action (511-1), expression (511-2), sound (511-4), and atmosphere (511-3).
[0097] Specifically, at least one processor (120) can identify subject motion based on changes in the subject's action, expression, sound, and mood, classify at least one frame into categories for the subject's action (511-1), expression (511-2), mood (511-3), and sound (511-4), and assign priorities to the subject motions according to the classified categories.
[0098] At least one processor (120) may assign different priorities within the same category. For example, within the action (511-1) category, a higher priority may be assigned to "dynamic action" among "dynamic action" and "static action," and a higher priority may be assigned to "special action" among "special action" and "routine action."
[0099] Similarly, within the expression (511-2) category, "special" expressions may be given higher priority than "ordinary" expressions. The term "special" here may be replaced with terms such as "unusual" or "peculiar." The above description can also be applied to the mood (511-3) category.
[0100] For the Sound (511-4) category, specific words, core conversation, positive voice, and background music can be prioritized in that order. Core conversation can be identified based on the frequency of recognized conversation, specific words, and conversation segments. Positive voice can be identified based on voice patterns, tone, speed, and specific words.
[0101] At least one processor (120) may 'set a high priority' (512) based on the amount of change in the subject motion. For example, among the 'section with change' (512-1) and the 'section that is maintained continuously' (512-2), the section with change (512-1) may be given a high priority. In other words, a frame with a greater amount of change in the subject motion may be given a high priority.
[0102] At least one processor (120) can subdivide the action of 'adjusting the priority of each subject action based on the user action' (520) into 'the action of creating a tag and giving a voice command' (520-1), 'the action of zooming in' (520-2), and 'the action of changing the shooting option' (520-3).
[0103] At least one processor (120) may be configured to prioritize 'action of creating a tag and giving a voice command' (520-1), 'action of zooming in' (520-2), and 'action of changing a shooting option' (520-3) in that order.
[0104] Specifically, at least one processor (120) can identify user actions based on the user's 'action of creating a tag and giving a voice command' (520-1), 'action of zooming in' (520-2), and 'action of changing a shooting option' (520-3).
[0105] At least one processor (120) can identify a priority for each subject motion based on the subject motion, and can adjust the priority information for each subject motion based on the user motion.
[0106] According to one embodiment of the present disclosure, it is possible to extract frames that are meaningful to the user by reflecting the user's intention by adjusting priorities by considering not only the subject motion identified based on image data but also the user motion identified based on user input.
[0107] FIG. 6 is a diagram illustrating a method for an electronic device according to at least one embodiment of the present disclosure to generate an edited image.
[0108] Referring to FIG. 6, the electronic device (100) can 'take a video by executing the Camera App' (610) based on user input.
[0109] The electronic device (100) can generate tags in real time while recording a video. The specific details of generating tags in real time have been described in FIG. 1, so a redundant description will be omitted.
[0110] At least one processor (120) can 'identify tags per frame and determine their importance' (620). Specifically, it can 'recognize a subject' (621) within a frame as one of Human, Pet, Object, and Landscape.
[0111] At least one processor (120) may assign different importance levels to at least one frame containing a subject, depending on the type of subject recognized. For example, a frame containing a human or a pet may be assigned a higher importance level than a frame containing an object or a landscape.
[0112] At least one processor (120) can recognize the "Moment" (622) of each subject and adjust the importance per frame. For example, the importance per frame can be adjusted based on the subject's motion and the user's motion. Since the specific details of adjusting the importance per frame based on the subject's motion and the user's motion have been described above, redundant details will be omitted.
[0113] At least one processor (120) can adjust the importance of each frame based on the amount of change in subject motion, image quality, and user information. This is described in detail below.
[0114] At least one processor (120) can identify the image quality for each frame based on at least one frame. If the image quality is determined to be lower than a preset image quality, the importance for each frame can be set low. At least one processor (120) can identify the rate of change in subject motion based on at least one frame. If the rate of change in subject motion is high, the importance for each frame can be set high.
[0115] However, frames included in the middle of a video that may contain intentional shaking may not be penalized. In cases where there are shaky frames, the regularity of the shaking can be identified based on at least one frame, and only sections with irregular shaking can be penalized, while sections with regular shaking can be exempted from penalization.
[0116] In another embodiment, at least one processor (120) may adjust the frame-by-frame importance based on user information when setting the frame-by-frame importance. The user information may include the user's relationship information, the user's interest information, the user's video editing pattern, the user's personal information, and the user's calendar information.
[0117] Normally, when extracting frames to create a video, there may be a problem where sections that the user considers important are omitted. However, this problem can be solved by adjusting the importance of each frame by reflecting user information.
[0118] At least one processor (120) can identify sections to be included and sections to be excluded for each frame, and thereby ‘extract at least one frame’ (630). For example, sections in which actions / dialogue / atmosphere are emphasized, sections with high contrast (or change) in the video, sections containing key content during a dialogue / lecture, and sections that the user focused on when filming can be identified as ‘sections to be included’ (631) and included in the frames to be extracted.
[0119] On the other hand, at least one processor (120) may identify sections of the image that are not smooth, sections containing disturbing elements, sections that are overlapping or do not have significant changes, and sections that induce negative emotions as 'sections to be excluded' (632) and exclude them from the frames to be extracted.
[0120] However, the above-described content is only an example and is not limited to the listed examples, and various examples may be included, such as including a frame from which a large section of audio is extracted.
[0121] At least one processor (120) can extract (640) a Best moment including at least one extracted frame and generate (650) an edited video based on the extracted Best moment.
[0122] FIG. 7 is a diagram illustrating an example of an electronic device according to at least one embodiment of the present disclosure obtaining data from an external device and displaying an image.
[0123] Referring to FIG. 7, the electronic device (100) is illustrated in the form of a monitor, but may be implemented as a variety of electronic devices capable of receiving and displaying images from an external device.
[0124] At least one processor (120) may include a communication interface, and may acquire image data including at least one frame from at least one external device through the communication interface. Specifically, when an electronic device (62) to be shared is selected by pressing a video sharing button (61) on the external device, the electronic device (100) may acquire the image data and extract at least one frame (63) according to the various embodiments described above.
[0125] For example, when sharing a video containing a child performing a dance, the electronic device (100) may extract at least one frame of the child dancing from among at least one acquired frame to create an edited video. In this case, the electronic device (100) may also display only the edited section of the acquired video on the screen in the form of an edited video.
[0126] However, the shared data can be shared not only as a video but also as a single image. For example, the importance of each frame can be determined, and the frame identified as the most important can be extracted and shared as a "best moment."
[0127] An edited video can be a video that simply displays at least one extracted frame in succession, or it can be a video with an adjusted playback speed. For example, if there is a section that is recognized as a single action, such as a person hitting a golf ball with a golf club, the playback speed for that section may be adjusted to a slower speed than the preset playback speed.
[0128] As another example, if there is a section that is perceived as a repetitive motion, such as a person performing repetitive movements in a factory, the playback speed of that section may be adjusted to be faster than the preset playback speed.
[0129] Edited videos may include regenerated videos with user input such as video length, aspect ratio, filters, text insertion, and image insertion.
[0130] FIG. 8 is a drawing showing an example of a screen displayed by an electronic device according to at least one embodiment of the present disclosure.
[0131] The electronic device (100) may further include a display, and at least one processor (120) may control the display to display at least one tag on the display.
[0132] For example, if the electronic device (100) obtains video data of a child clapping and laughing while looking at a kangaroo, it can analyze at least one frame in the video to generate tags such as 'kangaroo, clapping, smile, laughter'.
[0133] When the generated tags are displayed on the display, at least one processor (120) can select a 'smile' tag (71) by user input. When the user long-presses the selected tag, at least one processor (120) can extract at least one frame containing the tag and generate an edited video. In this case, when the user presses the 'share' button (72), at least one processor (120) can transmit the generated edited video to another electronic device.
[0134] The method of transmitting the edited video may include creating a multi-clip and an individual clip. If a multi-clip (73) is selected by the user, at least one processor (120) may share the file in the form of a multi-clip. Here, a multi-clip may refer to a video that is a continuous composite of at least one extracted frame. On the other hand, an individual clip may refer to images of each of at least one extracted frame.
[0135] FIG. 9 is a drawing showing an example of a screen displayed by an electronic device according to at least one embodiment of the present disclosure.
[0136] Referring to FIG. 9, when an image is selected by a user, at least one processor (120) may provide another image with a high degree of relevance to the selected image as a recommended image based on a tag identified in the image.
[0137] Specifically, at least one processor (120) can generate a plurality of thumbnails based on a selected image. When a thumbnail (101-1) including a frame of the 'playful' tag is selected from among the plurality of thumbnails, buttons such as 'Find in Gallery' (81), 'Trash Can', and 'Edit Tag' are displayed, and when 'Find in Gallery' (81) is selected, another edited image including a frame of the 'playful' tag can be displayed as a recommended image (82, 83) on the screen (101).
[0138] Recommended videos may be described in various terms, such as similar sequence, but in this disclosure, recommended videos are described.
[0139] FIG. 10 is a flowchart illustrating an image generation method of an electronic device according to at least one embodiment of the present disclosure.
[0140] Referring to FIG. 10, the electronic device can identify at least one tag corresponding to a subject motion within at least one frame (S1010).
[0141] Specifically, when a subject is smiling and running within a frame, the electronic device can identify the tags “smiling” and “running” in response to the subject’s facial expression and actions.
[0142] The electronic device can determine at least one frame-by-frame importance based on the subject motion-specific priority information and the identified tag (S1020).
[0143] For example, subject motion-specific priority information may be information that assigns a higher priority to dynamic motion when comparing static and dynamic motion among subject motions. Accordingly, based on subject motion-specific priority information, a frame containing dynamic motion may be determined to have a higher priority than a frame containing static motion.
[0144] An edited video including at least one frame having an importance greater than a certain value based on at least one frame-by-frame importance can be generated (S1030).
[0145] Specifically, the electronic device can extract at least one frame having a preset importance level based on the importance level of at least one frame. Furthermore, the electronic device can generate an edited video including the extracted at least one frame and display it on the screen of the electronic device or transmit it to another external device.
[0146] The image generation method described in Fig. 10 can be performed by devices having various configurations such as those of Figs. 3 and 4 described above, but is not necessarily limited thereto, and can also be performed by devices having various configurations.
[0147] The various embodiments described above may be implemented as a single embodiment, or at least one embodiment may be combined with each other in whole or in part and implemented together in one device.
[0148] According to the various embodiments described above, by generating an edited video that includes frames with a certain importance value or higher from a video containing at least one frame, the video can be presented to the user centered on meaningful scenes. Ultimately, this can enhance the user experience.
[0149] Various embodiments of the present disclosure may be implemented as software stored in a machine-readable storage media that can be installed or connected to a smartphone, a user terminal device, or other various electronic devices (e.g., a computer).
[0150] Specifically, a non-transitory readable recording medium may be provided having software stored thereon for sequentially performing the steps of: identifying at least one tag corresponding to a subject motion within a frame for at least one frame; determining an importance of at least one frame based on priority information for each subject motion and the identified tag; and generating an edited image including a frame having an importance equal to or greater than a predetermined value based on the importance of at least one frame.
[0151] A device equipped with such a non-transitory readable medium can perform various operations, such as tag identification corresponding to the subject motion described in the various embodiments described above, confirmation of importance for at least one frame, and creation of an edited video.
[0152] In the context of non-transitory readable recording media, 'non-transitory' means that the recording medium does not contain signals and is tangible, but does not distinguish between whether data is stored semi-permanently or temporarily on the recording medium.
[0153] Alternatively, a program for performing the method according to the various embodiments described above may be distributed online through an application store. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated on a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0154] Each component (e.g., a module or a program) according to various embodiments may be composed of one or more entities, and some of the aforementioned sub-components may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., a module or a program) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the respective components prior to integration. Operations performed by a module, program, or other component according to various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0155] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person skilled in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In electronic devices, Memory that stores priority information for each subject motion; and comprising at least one processor; At least one processor, Identifying at least one tag corresponding to the subject motion for at least one frame, Based on the priority information stored in the memory and the identified tag, the importance of at least one frame is determined, An electronic device that generates an edited image including at least one frame having an importance greater than a certain value based on the importance of at least one frame.
2. In paragraph 1, At least one processor above Identifying subject motion based on at least one of the subject's actions, facial expressions, sounds, or mood changes; Classifying at least one frame into a category for at least one of the subject's action, expression, sound, or mood, and setting a priority for each subject's motion for each category, An electronic device that sets a high priority for each subject motion for each frame when the amount of change in the subject motion is large for each of the at least one frame.
3. In paragraph 2, Including more cameras; At least one processor above Identifying a user action based on at least one of the following actions: zooming in while shooting, changing shooting options, creating a tag, or giving a voice command; An electronic device that sets a priority for each subject motion based on the subject motion, and adjusts the priority for each subject motion based on the user motion.
4. In paragraph 1, Including more cameras; At least one processor above Obtaining an image including at least one frame through the camera, An electronic device that analyzes the video in real time while shooting and identifies at least one tag corresponding to the motion of a subject within the frame.
5. In paragraph 1, At least one processor, An electronic device, wherein, based on at least one frame, if the image quality is identified as being lower than a preset image quality, the importance per frame is set low, and if the rate of change in the subject motion is large, the importance per frame is set high.
6. In paragraph 1, At least one processor, When determining the importance per frame, adjust the importance per frame based on user information, The above user information is: An electronic device comprising at least one of the user's relationship information, the user's interest information, the user's video editing pattern, the user's personal information, or the user's calendar information.
7. In paragraph 1, further comprising a communication interface; At least one frame above, An electronic device comprising at least one frame obtained from at least one external device via the communication interface.
8. In paragraph 1, including display; At least one processor, Controlling the display to display at least one tag on the display; An electronic device that generates an edited video including frames corresponding to the selected tags when one or more tags are selected by a user from among the tags displayed above.
9. In paragraph 1, At least one processor, An electronic device that, when a video is selected by a user, provides other videos with a high degree of relevance to the selected video as recommended videos based on the identified tags and user information.
10. In a method for generating an image of an electronic device, A step of identifying, for each of at least one frame, at least one tag corresponding to a subject motion within the frame; A step of determining the importance of at least one frame based on the priority information for each subject motion and the identified tag; and A method for generating an image, comprising: generating an edited image including at least one frame having an importance equal to or greater than a predetermined value based on the importance per at least one frame.
11. In paragraph 10, A step of identifying a subject's motion based on at least one of a change in the subject's action, facial expression, sound, or mood; A step of classifying at least one frame into a category for at least one of the subject's action, expression, sound, or mood, and setting a priority for each subject's action for each category; and An image generation method further comprising a step of setting a high priority for each subject motion for each frame when the amount of change in the subject motion is large for each of the at least one frame.
12. In paragraph 11, A step of identifying a user action based on at least one of a zoom-in action during shooting, an action of changing shooting options, an action of creating a tag, or an action of giving a voice command; and An image generation method further comprising: a step of adjusting the priority of each subject motion based on the user motion when identifying the priority of each subject motion based on the subject motion.
13. In paragraph 10, An image generation method further comprising: a step of acquiring an image including at least one frame, analyzing the image in real time while shooting, and identifying at least one tag corresponding to the motion of a subject within the frame; 14. In paragraph 10, An image generation method further comprising: a step of setting the importance per frame low when the image quality is identified as being lower than a preset image quality based on at least one frame, and setting the importance per frame high when the rate of change in the subject motion is large.
15. A non-transitory computer-readable recording medium including a program for executing a method for generating an image of an electronic device, The image generation method of the above electronic device is: A step of identifying, for each of at least one frame, at least one tag corresponding to a subject motion within the frame; A step of determining the importance of at least one frame based on the priority information for each subject motion and the identified tag; and A computer-readable recording medium comprising: a step of generating an edited video including a frame having an importance equal to or greater than a predetermined value based on at least one frame-by-frame importance;
Citation Information
Patent Citations
Extract keyframe candidates from video clips
JP2009539273A
Anti-rust bolt-cap
KR1020220049649A
Cell guide for jig formation of secondary battery
KR1020250155762A
Training corpus generating method, apparatus, device and storage medium
KR102345156B1
Method and apparattus for outputting motor imagery result from brain wave signal
KR102916004B1