Electronic device for generating highlight video on basis of keyword and operation method thereof
The electronic device addresses limitations in existing highlight video technologies by obtaining keywords, determining importance, and generating highlight videos based on user input, enabling versatile and context-aware real-time processing.
Patent Information
- Application Number
- PCT/KR2025/001706
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-24
- Filing Date
- 2025-02-05
- Publication Date
- 2025-10-30
AI Technical Summary
Existing technologies for generating highlight videos are limited to specific fields and struggle with real-time operation in mobile environments, and users face difficulty interpreting the context of generated highlight videos.
An electronic device that obtains keywords from video frames, determines importance distribution, and generates highlight videos based on user input, allowing for context interpretation and real-time processing.
Enables the generation of highlight videos for various types of content with user-friendly context interpretation and real-time capability.
Smart Images

Figure KR2025001706_30102025_PF_FP_ABST
Abstract
Description
Electronic device for generating highlight videos based on keywords and method for operating the same
[0001] The present disclosure relates to an electronic device for generating a highlight video and a method for operating the same. Specifically, the present disclosure relates to an electronic device for obtaining keywords from a video and generating the video based on the keywords, and a method for operating the same.
[0002] Related technologies identify highlight segments based on video analysis, such as frame differences, or audio analysis to generate highlight videos. Therefore, these technologies can only generate highlight videos for specific fields, such as sports. Furthermore, even if these technologies generate highlight videos, it can be difficult for users to interpret the context of the generated highlight videos. For example, users may have difficulty identifying the criteria used to determine whether a highlight video is a highlight. Furthermore, because these technologies process the entire video, they may struggle to operate in real-time mobile environments.
[0003] There may be a need for a method to generate highlight videos for all types of videos and provide descriptions related to the highlight videos so that users can understand and edit the highlight videos.
[0004] The above information is provided solely as background information to aid in understanding the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above constitutes prior art in connection with the present disclosure.
[0005] In one embodiment of the present disclosure, a method for generating a highlight video based on keywords by an electronic device is provided. The method may include obtaining one or more keywords included in a plurality of frames of a video. The method may include obtaining importance distribution information corresponding to each of the one or more keywords. The method may include determining at least one highlight keyword from among the one or more keywords based on the importance distribution information. The method may include displaying at least one highlight keyword. The method may include obtaining an input corresponding to at least some of the displayed at least one highlight keyword. The method may include generating a highlight video based on the highlight keyword corresponding to the input.
[0006] In one embodiment of the present disclosure, an electronic device is provided that generates a highlight video based on keywords. The electronic device may include at least one processor including a processing circuit and a memory that stores at least one instruction. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device obtains one or more keywords included in a plurality of frames of a video. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device obtains importance distribution information corresponding to each of the one or more keywords. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device determines at least one highlight keyword from among the one or more keywords based on the importance distribution information. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device displays at least one highlight keyword. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device obtains an input corresponding to at least some of the at least one highlighted keyword displayed. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to generate a highlight video based on a highlight keyword corresponding to an input.
[0007] In one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing an operation of an electronic device, and executing any one of the methods described above and below may be provided.
[0008] FIG. 1 is a diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to generate a highlight video based on a keyword.
[0009] FIG. 2 is a flowchart illustrating a method for generating a highlight video based on keywords according to one embodiment of the present disclosure.
[0010] FIG. 3 is a diagram for explaining keywords of a video according to one embodiment of the present disclosure.
[0011] FIG. 4 is a diagram showing importance distribution information corresponding to a keyword according to one embodiment of the present disclosure.
[0012] FIG. 5 is a drawing for explaining a highlight section according to one embodiment of the present disclosure.
[0013] FIG. 6 is a drawing for explaining a highlight section according to one embodiment of the present disclosure.
[0014] FIG. 7 is a diagram illustrating a process for generating a highlight video according to one embodiment of the present disclosure.
[0015] FIG. 8A is a drawing for explaining modules included in an electronic device (100) according to one embodiment of the present disclosure.
[0016] FIG. 8b is a drawing for explaining modules included in an electronic device (100) according to one embodiment of the present disclosure.
[0017] FIG. 9 is a diagram showing a plurality of groups into which a video is divided according to one embodiment of the present disclosure.
[0018] FIG. 10 is a diagram illustrating a process for generating a highlight video according to one embodiment of the present disclosure.
[0019] FIG. 11 is a diagram illustrating a process for determining highlight keywords and generalization keywords according to one embodiment of the present disclosure.
[0020] FIG. 12 is a diagram showing a UI in which an electronic device (100) displays a highlight keyword according to one embodiment of the present disclosure.
[0021] FIG. 13 is a drawing showing a UI for editing a video by an electronic device (100) according to one embodiment of the present disclosure.
[0022] FIG. 14 is a diagram illustrating a UI that displays search results of images and / or videos according to one embodiment of the present disclosure.
[0023] FIG. 15 is a drawing showing a UI for editing multiple videos by an electronic device (100) according to one embodiment of the present disclosure.
[0024] FIG. 16 is a diagram showing a UI for sharing a highlight video based on a keyword by an electronic device (100) according to one embodiment of the present disclosure.
[0025] FIG. 17 is a diagram illustrating a UI including search results of images and / or videos according to one embodiment of the present disclosure.
[0026] FIG. 18 is a block diagram showing the configuration of an electronic device according to one embodiment of the present disclosure.
[0027] FIG. 19 is a block diagram showing the configuration of an electronic device according to one embodiment of the present disclosure.
[0028] In one embodiment of the present disclosure, the expression “at least one of a, b, or c” may refer to “a,” “b,” “c,” “a and b,” “a and c,” “b and c,” “all of a, b, and c,” or variations thereof.
[0029] In one embodiment of the present disclosure, the expression "a, b, and / or c" may be interpreted identically to the expression "at least one of a, b, or c." For example, the expression "a, b, and / or c" may refer to "a," "b," "c," "a and b," "a and c," "b and c," "all of a, b, and c," or variations thereof.
[0030] The terms used in one embodiment of the present disclosure are selected from widely used, general terms, taking into account the functions of one embodiment of the present disclosure. However, these terms may vary depending on the intentions of engineers working in the field, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings can be understood through the relevant description. Therefore, the terms used in one embodiment of the present disclosure should be defined based on the meaning of the term and the overall content of the present disclosure, rather than simply the name of the term.
[0031] In one embodiment of the present disclosure, a singular expression may include a plural expression unless the context clearly indicates otherwise. For example, the description "a constituent surface" may also refer to one or more of such surfaces. While terms including ordinal numbers, such as "first" or "second," used in one embodiment of the present disclosure may be used to describe various components, the components should not be limited by these terms. These terms are used solely to distinguish one component from another.
[0032] In one embodiment of the present disclosure, when a part is said to "include" a certain component, this does not exclude other components, but rather may include other components, unless otherwise specifically stated. In one embodiment of the present disclosure, terms such as "part" and "module" refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.
[0033] The expression "configured to" used in one embodiment of the present disclosure can be used interchangeably with, for example, "suitable for", "having the capacity to", "designed to", "adapted to", "made to", or "capable of", depending on the context. The term "configured to" does not necessarily mean only "specifically designed to" in terms of hardware. Alternatively, in some contexts, the expression "a system configured to" can include that the system is "capable of" in conjunction with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" can include a dedicated processor for performing the operations (e.g., an embedded processor), or a general-purpose processor (e.g., a CPU or an application processor) that can perform the operations by executing one or more software programs stored in a memory.
[0034] In one embodiment of the present disclosure, when a component is referred to as being “connected” or “connected” to another component, it should be understood that the component may be directly connected or connected to the other component, but may also be connected or connected via another component in between, unless specifically stated otherwise.
[0035] In describing the present disclosure, descriptions of technical details that are well known in the technical field to which the present disclosure pertains and are not directly related to the present disclosure may be omitted. This is to convey the gist of the present disclosure more clearly without obscuring unnecessary explanation. In the drawings, parts irrelevant to the description are omitted for clarity in describing the present disclosure, and similar parts are designated with similar reference numerals throughout the specification. The size of each component does not entirely reflect the actual size. The same or corresponding components in each drawing are given the same reference numerals.
[0036] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure is complete and to fully inform those skilled in the art of the disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined in accordance with the claims.
[0037] In one embodiment of the present disclosure, combinations of each block in the flowchart and the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing equipment, and the instructions executed by the processor of the computer or other programmable data processing equipment can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to perform a function in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing equipment. It should be understood that each combination of blocks in the flowchart and the flowchart diagrams can be performed by one or more computer programs that include computer-executable instructions. One or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories. In one embodiment, the memory may include one or more storage media that store instructions.
[0038] All functions or operations described in one embodiment of the present disclosure may be processed by a single processor or a combination of processors. A single processor or a combination of processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), an Integrated Chip (IC), etc.
[0039] At least one processor according to an embodiment of the present disclosure may include various processing circuits and / or multiple processors. For example, the term "processor" as used in an embodiment of the present disclosure, including in the claims, may include various processing circuits including at least one processor, one or more of which are configured to individually and / or collectively perform the various functions described herein in a distributed manner. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms may include, for example, without limitation, a single processor performing some of the recited functions, other processor(s) performing other of the recited functions, and still other situations where a single processor can perform all of the recited functions. Additionally, the at least one processor may include a combination of processors that perform the various functions enumerated / disclosed, for example, in a distributed manner. The at least one processor may execute program instructions to achieve or perform the various functions.
[0040] In one embodiment of the present disclosure, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment of the present disclosure, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.
[0041] The term '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the '~ unit' may perform a specific role. Meanwhile, the '~ unit' is not limited to software or hardware. The '~ unit' may be configured to be on an addressable storage medium and may be configured to play one or more processors. In one embodiment of the present disclosure, the '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided through a specific component or a specific '~ unit' may be combined to reduce the number of components or separated into additional components. In one embodiment of the present disclosure, the '~ unit' may include one or more processors.
[0042] The artificial intelligence-related function according to one embodiment of the present disclosure operates via a processor and memory. The processor may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a Digital Signal Processor (DSP), a graphics-only processor such as a GPU, a Vision Processing Unit (VPU), or an artificial intelligence-only processor such as an NPU. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in the memory. Alternatively, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0043] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is learned by a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to an embodiment of the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0044] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks.
[0045] In a method for generating a highlight video of an electronic device according to an embodiment of the present disclosure, a frame of a video may be used as input data of an artificial intelligence model as a method for recognizing a keyword, and output data in which a keyword corresponding to at least a portion of the frame is recognized may be obtained. The artificial intelligence model may be created through learning. Here, being created through learning means that a basic artificial intelligence model is learned using a plurality of learning data by a learning algorithm, thereby creating a predefined operation rule or artificial intelligence model set to perform a desired characteristic (or purpose). The artificial intelligence model may be composed of a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weight values, and performs a neural network operation through an operation between the operation result of the previous layer and the plurality of weight values.
[0046] Visual understanding is a technology that recognizes and processes objects like human vision, and includes object recognition, object tracking, image retrieval, human recognition, scene recognition, spatial understanding (3D reconstruction / localization), and image enhancement.
[0047] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings so that those skilled in the art can easily practice them. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In addition, in order to clearly describe the present disclosure in the drawings, parts that are not related to the description are omitted, and similar parts are designated with similar reference numerals throughout the specification. In addition, the reference numerals used in each drawing are only for describing each drawing, and different reference numerals used in different drawings do not indicate different elements. The present disclosure will be described in detail below with reference to the attached drawings.
[0048] In the present disclosure, a "video" may include multiple still images. A video may be composed of multiple still images in a time series. For example, a video may represent the movement of an object included in the images by being composed of sequentially captured images in a time series. In one embodiment of the present disclosure, a still image included in a video may be referred to as a frame. In one embodiment of the present disclosure, a video may further include audio data, such as voice or music.
[0049] In the present disclosure, a 'keyword' may include a word or phrase indicating a feature related to a video or a frame of a video. For example, a keyword may include a word or phrase containing information regarding at least one of an emotion, an object, an action, a behavior, a situation, a place, or an event related to a video or a frame of a video. For example, a keyword may indicate at least one of an emotion, an object, an action, a behavior, a situation, a place, or an event contained in a video or a frame of a video. A keyword may be stored in memory as information related to a video or a frame of a video.
[0050] In the present disclosure, "importance distribution information" may include a distribution of the importance of keywords in a video. The importance score may indicate the importance of a keyword in the video. In one embodiment of the present disclosure, the importance score may be determined to be high if the keyword is frequently identified in a portion of the video. The importance score may be determined to be low if the keyword is frequently identified throughout the entire video. For example, in a 1-minute video, the importance score for the keyword "pet" when a pet is identified for 50 seconds may be determined to be lower than the importance score when a pet is identified for 10 seconds. The importance distribution information may include time-series data. The importance distribution information may include an importance score corresponding to time. For example, the importance distribution information may include an importance score according to the playback time of the video.
[0051] In the present disclosure, a "highlight keyword" may include a keyword identified in a small section of a video. Highlight keywords may be determined based on importance distribution information. Keywords whose importance scores in the importance distribution information satisfy predetermined conditions may be determined as highlight keywords. For example, a keyword with a maximum or average importance score greater than a threshold may be determined as a highlight keyword.
[0052] In the present disclosure, a "highlight video" may refer to a video generated using a portion of an original video. For example, a highlight video may include a frame corresponding to a highlight section among multiple frames of the original video. In one embodiment of the present disclosure, a highlight section may be determined based on the importance score of a highlight keyword. For example, a certain section centered around the highest point of a highlight keyword may be determined as a highlight section.
[0053] FIG. 1 is a diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to generate a highlight video based on a keyword.
[0054] Referring to FIG. 1, the electronic device (100) may be a mobile device such as a smart phone, a tablet PC, a laptop computer, a digital camera, an e-book terminal, a digital broadcasting terminal, a PDA (Personal Digital Assistant), a PMP (Portable Multimedia Player), a navigation device, or an MP3 player. In one embodiment of the present disclosure, the electronic device (100) may be a home appliance such as a TV, an air conditioner, a robot vacuum cleaner, or a clothes manager. However, the present disclosure is not limited thereto, and in another embodiment of the present disclosure, the electronic device (100) may be implemented as a wearable device such as a smart watch, an eyeglass-type augmented reality device (e.g., AR (Augmented Reality) glasses), a head-mounted device (HMD), or a body-attached device (e.g., a skin pad).
[0055] Referring to FIG. 1, an electronic device (100) may obtain keywords (110) for a video. In one embodiment of the present disclosure, the electronic device (100) may obtain one or more keywords included in multiple frames of the video. The electronic device (100) may analyze frames of the video and obtain keywords corresponding to the frames. For example, the electronic device (100) may obtain keywords indicating emotions, objects, actions, behaviors, places, situations, or events included in frames of the video. For example, keywords may include kiss, beach, pet, hug, woman, smiling, and love. The electronic device (100) may obtain keywords (110) for the video based on keywords for each frame of the video.
[0056] The electronic device (100) can obtain importance distribution information (120) corresponding to a keyword (110). In one embodiment of the present disclosure, the electronic device (100) can obtain importance distribution information (120) corresponding to each of one or more keywords. For example, the electronic device (100) can obtain importance distribution information (120) corresponding to 'hug' and importance distribution information (120) corresponding to 'smile'.
[0057] In one embodiment of the present disclosure, the importance distribution information (120) may include an importance score corresponding to time. The importance score may be determined to be high if a keyword is identified in a portion of the video. For example, if a video includes frames related to hugging in the 130- to 310-second interval, the importance score for "hug" may be determined to be high in the 130- to 310-second interval. The importance score may be determined to be low if a keyword is identified throughout the entire video. For example, if a video includes frames related to pets throughout the entire interval, the importance score for "pets" may be determined to be low throughout the entire interval.
[0058] The electronic device (100) may determine a highlight keyword (130). In one embodiment of the present disclosure, the electronic device (100) may determine the highlight keyword (130) based on the importance distribution information (120). If the importance score of the importance distribution information (120) satisfies a predetermined condition, the electronic device (100) may determine a keyword corresponding to the importance distribution information (120) as a highlight keyword. For example, the electronic device (100) may determine 'hug' as a highlight keyword based on the fact that the maximum value of the importance score of 'hug' is greater than a threshold value. For example, 'pet' may not be determined as a highlight keyword based on the fact that the maximum value of the importance score of 'pet' is less than a threshold value.
[0059] In one embodiment of the present disclosure, the electronic device (100) can determine at least some of a plurality of keywords of a video as highlight keywords based on importance distribution information (120).
[0060] The electronic device (100) can generate a highlight video (150) based on a highlight keyword (130). The electronic device (100) can generate the highlight video (150) to include a frame corresponding to the highlight keyword (130).
[0061] In one embodiment of the present disclosure, the electronic device (100) may display a highlight keyword (130). The electronic device (100) may display a user interface (UI) (140) corresponding to the highlight keyword (130). For example, the electronic device (100) may display a UI (140) corresponding to 'kiss', 'hug', and 'smile'. The electronic device (100) may omit and display a portion of the highlight keyword (130) based on the size of the UI (140) and the size of the displayed area.
[0062] The electronic device (100) can obtain an input corresponding to the UI (140). For example, the electronic device (100) can obtain a touch input for an area where the UI (140) is displayed. For example, the electronic device (100) can obtain a voice input corresponding to a highlight keyword (130). In one embodiment of the present disclosure, the electronic device (100) can generate a highlight video (150) based on the highlight keyword (130) corresponding to the input. The electronic device (100) can determine a highlight section based on the highlight keyword (130) corresponding to the input. The highlight section can be determined based on the importance score of the highlight keyword (130). For example, the highlight section can be determined as a section in which the importance score is greater than a threshold value. For example, the highlight section for 'hug' can be determined to be 130 seconds to 310 seconds. The electronic device (100) can generate a highlight video to include frames included in the highlight section.
[0063] FIG. 2 is a flowchart illustrating a method for generating a highlight video based on keywords according to one embodiment of the present disclosure.
[0064] In one embodiment of the present disclosure, a method for generating a highlight video may be performed by an electronic device (100). For example, the electronic device (100) may perform each step of a video decoding method by having a processor of the electronic device (100) execute at least one instruction contained in a memory.
[0065] In step S210, the method may include a step of acquiring one or more keywords included in a plurality of frames of a video. The electronic device (100) may acquire one or more keywords included in a plurality of frames of the video. For example, the electronic device (100) may extract keywords for each frame of the video. The keywords may include words representing at least a portion of the frame. In one embodiment of the present disclosure, the electronic device (100) may select a keyword for a frame from among words stored in a database (DB).
[0066] In one embodiment of the present disclosure, the electronic device (100) can obtain keywords included in a frame using an artificial intelligence model. The artificial intelligence model can be trained to output keywords included in the frame. The number of keywords output by the artificial intelligence model is not limited, and the artificial intelligence model can output one or more keywords. The artificial intelligence model may not output keywords if the frame has no features. The electronic device (100) can input frames of a video into the artificial intelligence model to obtain keywords and / or confidence scores corresponding to the keywords. The confidence score may refer to the probability that the artificial intelligence model determines that the frame includes an output keyword.
[0067] In one embodiment of the present disclosure, the electronic device (100) can acquire keywords for each frame while capturing a video. For example, the electronic device (100) can acquire keywords for each frame captured in real time while capturing a video. The electronic device (100) can extract keywords from frames independently of capturing the video.
[0068] In one embodiment of the present disclosure, the electronic device (100) can acquire frame-by-frame keywords after the video recording is completed. For example, the electronic device (100) can acquire frame-by-frame keywords during a post-processing process after the video recording is completed.
[0069] In step S220, the method may include a step of obtaining importance distribution information corresponding to each of one or more keywords. The electronic device (100) may obtain importance distribution information corresponding to each of one or more keywords.
[0070] In one embodiment of the present disclosure, the importance distribution information may include an importance score corresponding to time. The importance score may be determined to be high if the keyword is identified only in a portion of the video. The importance score may be determined to be low if the keyword is identified throughout the entire video or if the keyword is identified only a small number of times.
[0071] In step S230, the method may include a step of determining at least one highlight keyword from among one or more keywords based on importance distribution information. The electronic device (100) may determine at least one highlight keyword from among one or more keywords based on importance distribution information.
[0072] In one embodiment of the present disclosure, the electronic device (100) may determine a keyword corresponding to the importance distribution information as a highlight keyword when the importance score of the importance distribution information satisfies a predetermined condition.
[0073] In step S240, the method may include a step of displaying at least one highlight keyword. The electronic device (100) may display at least one highlight keyword. In one embodiment of the present disclosure, the electronic device (100) may display a UI corresponding to the highlight keyword through the display.
[0074] In step S250, the method may include a step of obtaining an input corresponding to at least some of the at least one highlighted keyword displayed. The electronic device (100) may obtain an input corresponding to at least some of the at least one highlighted keyword displayed. In one embodiment of the present disclosure, the electronic device (100) may obtain a touch input or a voice input corresponding to the highlighted keyword. For example, the electronic device (100) may obtain a touch input for an area displaying a UI corresponding to the highlighted keyword or a voice input indicating the highlighted keyword.
[0075] In step S260, the method may include a step of generating a highlight video based on a highlight keyword corresponding to the input. The electronic device (100) may generate a highlight video based on a highlight keyword corresponding to the input.
[0076] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight section corresponding to a highlight keyword corresponding to an input. The electronic device (100) can generate a highlight video based on frames included in the highlight section.
[0077] In one embodiment of the present disclosure, the electronic device (100) can generate a highlight video that reflects the user's intent by generating a highlight video corresponding to an input. The electronic device (100) can extract keywords from the video and generate a highlight video based on the keywords, thereby providing the user with a reason why the highlight video was generated, and allowing the user to change the highlight video using the keywords. For example, the user can request the electronic device (100) to generate a highlight video for a keyword of interest.
[0078] FIG. 3 is a diagram for explaining keywords of a video according to one embodiment of the present disclosure.
[0079] In one embodiment of the present disclosure, the electronic device (100) can determine keywords based on a video. The electronic device (100) can obtain keywords corresponding to frames of the video.
[0080] Referring to FIG. 3, for example, a frame of a video may be a scene of a girl and a boy feeding animals outdoors. In one embodiment of the present disclosure, keywords corresponding to the frame may include 'child,' 'girl,' 'joy,' 'park,' 'summer,' 'person,' 'playing,' and 'cute.'
[0081] In one embodiment of the present disclosure, multiple predetermined keywords may be correlated. For example, if the frame includes a girl, "child," "girl," and "person" may be selected. The electronic device (100) may determine multiple keywords corresponding to a single object included in the frame. However, the present invention is not limited thereto, and the number of keywords corresponding to a single object may be limited. For example, the number of keywords corresponding to a single object may be limited to one.
[0082] In one embodiment of the present disclosure, keywords may include characteristic information expressed in natural language about a video. In one embodiment of the present disclosure, the electronic device (100) may obtain characteristic information indicating at least one of an emotion, an object, an action, an action, a situation, a place, or an event included in a frame of the video. The electronic device (100) may determine keywords of the video to include keywords corresponding to the characteristic information. For example, the electronic device (100) may determine keywords of the video to include a first keyword included in a first frame and a second keyword included in a second frame.
[0083] In one embodiment of the present disclosure, the electronic device (100) may determine a keyword using an artificial intelligence model. The artificial intelligence model may be an artificial intelligence model trained to input a frame of a video and output a keyword corresponding to at least one of an emotion, object, action, behavior, situation, location, or event included in the frame. If there is no keyword corresponding to the frame, the electronic device (100) may not determine a keyword or may determine that there is no keyword. In one embodiment of the present disclosure, the artificial intelligence model may output a confidence score for the corresponding keyword along with the keyword. The confidence score may refer to the probability that the artificial intelligence model determines that the output keyword is included in the frame. For example, a keyword with a high confidence score may be determined by the artificial intelligence model to have a high probability of being included in the frame.
[0084] In one embodiment of the present disclosure, the electronic device (100) may select a keyword corresponding to at least a portion of a frame from among a plurality of predetermined keywords. For example, the electronic device (100) may select a keyword indicated by at least a portion of a frame from among a plurality of keywords included in a keyword database.
[0085] FIG. 4 is a diagram showing importance distribution information corresponding to a keyword according to one embodiment of the present disclosure.
[0086] Referring to Figure 4, importance distribution information for each keyword is illustrated. For example, the importance distribution information illustrated in Figure 4 may each correspond to a keyword included in a video. In one embodiment of the present disclosure, the importance distribution information may include a distribution of the importance of keywords in the video.
[0087] In one embodiment of the present disclosure, the importance distribution information may include information regarding importance scores based on time or frame of a video. The number of frames may indicate the display order. For example, the number of frames may be displayed in increasing order starting from 0. The importance distribution information may be expressed as information regarding importance scores based on time.
[0088] In one embodiment of the present disclosure, an importance score may indicate the importance of a keyword in a video. The importance score may be determined based on the frequency of the keyword. For example, the importance score may be determined based on the number of times a keyword is identified in a certain section of the video.
[0089] In one embodiment of the present disclosure, a high importance score may be determined when a keyword is frequently identified in a certain section of a video. The first distribution (410) represents importance distribution information corresponding to the keyword "kiss." The first distribution (410) has a high importance score between 150 and 250 seconds. The first distribution (410) may indicate that the keyword "kiss" is frequently identified in the 150- to 250-second section of the video.
[0090] In one embodiment of the present disclosure, an importance score may be determined to be low if a keyword is frequently identified throughout the entire video. For example, a keyword frequently identified throughout the entire video may be determined to have low importance for generating a highlight video. The second distribution (420) represents importance distribution information corresponding to the keyword "walking." The second distribution (420) has low importance scores across all frames of the video. The second distribution (420) may indicate that the keyword "walking" is frequently identified throughout the entire video.
[0091] In one embodiment of the present disclosure, the importance score may be determined to be low if a keyword is rarely identified throughout the entire video. For example, a keyword that is rarely identified throughout the entire video may be determined to have low importance for generating a highlight video. The second distribution (420) may indicate that the keyword "walking" is rarely identified throughout the entire video.
[0092] In one embodiment of the present disclosure, the electronic device (100) can obtain keywords included in frames of a video and obtain importance distribution information corresponding to the keywords. The electronic device (100) can determine highlight keywords based on the importance distribution information. In one embodiment of the present disclosure, the electronic device (100) can determine highlight keywords of the video from among keywords included in frames of the video. In one embodiment of the present disclosure, highlight keywords can refer to keywords that are relatively rare and indicate highlight moments.
[0093] In one embodiment of the present disclosure, the electronic device (100) may determine a highlight keyword based on an importance score. In one embodiment of the present disclosure, the electronic device (100) may determine a keyword whose importance score satisfies a predetermined condition as a highlight keyword. For example, the electronic device (100) may determine a keyword whose maximum importance score is greater than a threshold value as a highlight keyword. For example, the electronic device (100) may determine a keyword whose average importance score is greater than a threshold value as a highlight keyword. In one embodiment of the present disclosure, the electronic device (100) may determine only a predetermined number of keywords in descending order of importance score as highlight keywords. In one embodiment of the present disclosure, the electronic device (100) may determine a keyword whose importance score does not satisfy a predetermined condition as not a highlight keyword.
[0094] FIG. 5 is a drawing for explaining a highlight section according to one embodiment of the present disclosure.
[0095] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight section based on importance distribution information. The importance distribution information may include importance scores of keywords over frames or time in a video. Referring to FIG. 5 , example keyword importance distribution information is illustrated. Frame F1 may denote the frame with the highest importance score.
[0096] In one embodiment of the present disclosure, the electronic device (100) may determine a highlight section based on importance distribution information. The electronic device (100) may determine the highlight section based on an importance score. The electronic device (100) may determine the highlight section based on the highest value of the importance score. The electronic device (100) may determine a certain section including the highest value of the importance score as a highlight section. For example, the electronic device (100) may determine a certain section including frame F1 as a highlight section. The electronic device (100) may determine the highlight section to include a frame having the highest importance score and a frame having an importance score greater than a threshold value. For example, the electronic device (100) may determine the highlight section to include frame F1 and a frame having an importance score greater than 40% of the highest value.
[0097] In one embodiment of the present disclosure, the electronic device (100) may determine the start and end frames of a highlight section. For example, the importance score of frame F2 may be 40% of the importance score of frame F1, and frame F3 may be the last frame of the video. The electronic device (100) may store information regarding the start and end frames of the highlight section.
[0098] In one embodiment of the present disclosure, the electronic device (100) may determine a plurality of highlight sections based on importance distribution information. For example, the plurality of highlight sections may be determined to include local maximum points of the importance distribution information. For example, the importance distribution information has local maximum points of importance scores in frames F1 and F4. Therefore, in this example, the electronic device (100) may determine a first highlight section including F1 and a second highlight section including F4. In one embodiment of the present disclosure, the electronic device (100) may determine a plurality of highlight sections using an artificial intelligence model. For example, the electronic device (100) may determine a plurality of highlight sections using a Gaussian Mixture Model (GMM).
[0099] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight section only for a highlight keyword. The electronic device (100) can determine a highlight keyword and a highlight section corresponding to the highlight keyword based on importance distribution information.
[0100] FIG. 6 is a drawing for explaining a highlight section according to one embodiment of the present disclosure.
[0101] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight section based on importance distribution information. The importance distribution information may include importance scores of keywords based on video frames or time. Referring to FIG. 6 , importance distribution information for exemplary keywords is illustrated.
[0102] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight section based on importance distribution information. The electronic device (100) can determine a highlight section based on an importance score. The electronic device (100) can determine a highlight section based on an importance score. The electronic device (100) can determine a certain section including a frame having an importance score greater than a threshold value as a highlight section.
[0103] In one embodiment of the present disclosure, the electronic device (100) may determine the start frame and the end frame of the highlight section. The start frame and the end frame of the highlight section may be the first frame of the video, the last frame of the video, or a frame whose importance score is equal to a threshold value. For example, the start frame F1 of the highlight section has an importance score equal to the threshold value, and the end frame F2 of the highlight section is the last frame of the video. The electronic device (100) may store information regarding the start frame and the end frame of the highlight section.
[0104] In one embodiment of the present disclosure, the electronic device (100) may determine a plurality of highlight sections based on importance distribution information. For example, the electronic device (100) may determine a plurality of highlight sections having importance scores greater than a threshold value.
[0105] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight section only for a highlight keyword. The electronic device (100) can determine a highlight keyword and a highlight section corresponding to the highlight keyword based on importance distribution information.
[0106] FIG. 7 is a diagram illustrating a process for generating a highlight video according to one embodiment of the present disclosure.
[0107] In one embodiment of the present disclosure, the electronic device (100) can generate a highlight video (730) corresponding to a highlight keyword (710). The electronic device (100) can display the highlight keyword (710). For example, the electronic device (100) can output a UI corresponding to the highlight keyword (710) using a display.
[0108] In one embodiment of the present disclosure, the electronic device (100) can obtain an input corresponding to a highlight keyword (710). The user can select at least some of the highlight keywords (710) displayed on the display. The electronic device (100) can obtain an input for selecting at least some of the highlight keywords (710). For example, the user can select 'kiss' from the highlight keywords (710) displayed on the display, and the electronic device (100) can obtain an input for selecting 'kiss' from the highlight keywords (710). The input can include at least one of a voice input or a touch input.
[0109] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight section based on a highlight keyword (710) corresponding to an input. The electronic device (100) can identify the highlight keyword (710) corresponding to the input, and determine a highlight section based on importance distribution information (720) of the highlight keyword (710). For example, the electronic device (100) can identify 'kiss' corresponding to the input among the highlight keywords (710), and determine a highlight section based on importance distribution information (720) of 'kiss'.
[0110] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight section based on a plurality of highlight keywords (710) corresponding to an input.
[0111] The electronic device (100) can obtain inputs for multiple highlight keywords. In one embodiment of the present disclosure, the electronic device (100) can obtain multiple inputs indicating highlight keywords (710). For example, the electronic device (100) can obtain a first input indicating a first highlight keyword 'kiss' and a second input indicating a second highlight keyword 'love'. In one embodiment of the present disclosure, the electronic device (100) can obtain inputs indicating multiple highlight keywords (710). For example, the electronic device (100) can obtain one input indicating a first highlight keyword 'kiss' and a second highlight keyword 'love'.
[0112] The electronic device (100) may determine a plurality of highlight sections based on importance distribution information (720) corresponding to each highlight keyword (710). For example, the electronic device (100) may determine a first highlight section based on first importance distribution information corresponding to a first highlight keyword 'kiss', and may determine a second highlight section based on second importance distribution information corresponding to a second highlight keyword 'love'. In this example, the electronic device (100) may determine a highlight section to include a first highlight section and a second highlight section.
[0113] In one embodiment of the present disclosure, the electronic device (100) can generate a highlight video (730) corresponding to a highlight section. The electronic device (100) can generate the highlight video (730) to include a plurality of frames of the video corresponding to the highlight section.
[0114] FIG. 8A is a drawing for explaining modules included in an electronic device (100) according to one embodiment of the present disclosure.
[0115] In one embodiment of the present disclosure, the electronic device (100) may include a keyword extraction module (810), an interval sampling module (820), an importance distribution determination module (830), a highlight keyword determination module (840), and a highlight video generation module (850).
[0116] The keyword extraction module (810), the interval sampling module (820), the importance distribution determination module (830), the highlight keyword determination module (840), and the highlight video generation module (850) included in the electronic device (100) of FIG. 8A may be components classified based on their functions or roles. The keyword extraction module (810), the interval sampling module (820), the importance distribution determination module (830), the highlight keyword determination module (840), and the highlight video generation module (850) included in the electronic device (100) of FIG. 8A may be a software configuration implemented by a processor of the electronic device (1000) executing a program or instruction stored in a memory, or may be a virtual configuration in which an actual matching hardware device exists. In other words, the operations performed by the processor of the electronic device (1000) by executing a program or instruction stored in the memory may be classified into a plurality of groups by function or purpose, and the entities performing the operations included in each classified group may be expressed as the keyword extraction module (810), the interval sampling module (820), the importance distribution determination module (830), the highlight keyword determination module (840), and the highlight video generation module (850) included in the electronic device (100) of FIG. 8A. The operations and functions described as being performed by the keyword extraction module (810), the interval sampling module (820), the importance distribution determination module (830), the highlight keyword determination module (840), and the highlight video generation module (850) of FIG. 8A may be performed by the electronic device (100) or the processor of the electronic device (100) by executing a program or instruction stored in the memory.
[0117] In one embodiment of the present disclosure, the keyword extraction module (810) can obtain keywords based on a video. The keyword extraction module (810) can obtain keywords corresponding to each frame of the video. The keyword extraction module (810) can include an artificial intelligence model that inputs frames and outputs keywords. The keywords can represent at least a portion of a frame of the video.
[0118] In one embodiment of the present disclosure, the interval sampling module (820) can divide a video into multiple groups. The interval sampling module (820) can divide the video into multiple groups based on a division interval. For example, the interval sampling module (820) can determine a group of the video in units of 30 frames. The division interval can refer to a value determined to divide the video evenly. In one embodiment of the present disclosure, the process of dividing the video into multiple groups is described in detail with reference to FIG. 9.
[0119] In one embodiment of the present disclosure, the keyword extraction module (810) and the interval sampling module (820) are sequentially connected, but can be processed in parallel, and the order between the keyword extraction module (810) and the interval sampling module (820) can be exchanged.
[0120] In one embodiment of the present disclosure, the importance distribution determination module (830) may obtain importance distribution information corresponding to a keyword. In one embodiment of the present disclosure, the importance distribution determination module (830) may obtain importance distribution information including importance scores according to time or frame. In one embodiment of the present disclosure, the importance distribution determination module (830) may determine an importance score based on the number of keywords included in multiple groups of a video.
[0121] In one embodiment of the present disclosure, the importance distribution determination module (830) can identify the number of frames including the first keyword among the plurality of frames included in the first group among the plurality of groups. For example, when the segmentation interval is 30, the importance distribution determination module (830) can identify how many frames include the first keyword among the 30 frames. The importance distribution determination module (830) can identify the number of groups including the first keyword among the plurality of groups. For example, when the segmentation interval is 30 and the total number of frames of the video is 300 frames, the importance distribution determination module (830) can identify the number of groups including the first keyword among 10 groups. The importance distribution determination module (830) can obtain importance distribution information corresponding to the first keyword based on the number of frames including the first keyword and the number of groups including the first keyword. For example, the importance distribution determination module (830) may determine a higher importance score as the number of frames containing the first keyword increases, and may determine a lower importance score as the number of groups containing the first keyword increases. In one embodiment of the present disclosure, the importance score may be determined based on mathematical expression 1.
[0122]
[0123] Here, k represents the keyword index, d represents the group index, and TF represents the k,d is the number of keywords included in the group, N is the total number of groups in the video, df k can mean the number of groups containing the keyword.
[0124] In one embodiment of the present disclosure, the importance distribution determination module (830) may obtain importance distribution information based on the ratio between the number of groups including the first keyword and the total number of groups included in the video. The importance distribution determination module (830) may determine the ratio between the number of groups including the first keyword and the total number of groups included in the video. For example, the importance distribution determination module (830) may determine N / df of mathematical expression 1. k Determine N / df k Based on this, importance distribution information can be obtained.
[0125] In one embodiment of the present disclosure, the importance distribution determination module (830) may obtain importance distribution information based on the ratio between the number of frames of the first group including the first keyword and the total number of keywords included in the first group. The importance distribution determination module (830) may determine the ratio between the number of frames of the first group including the first keyword and the total number of keywords included in the first group. For example, the importance distribution determination module (830) may determine TF of Mathematical Formula 1. k,d The ratio between the total number of keywords included in the first group d can be determined, and importance distribution information can be obtained based on the determined ratio.
[0126] In one embodiment of the present disclosure, the importance score may be determined based on Equation 2 or Equation 3.
[0127]
[0128] Here, c k,d represents the confidence score for the keyword, and the remaining variables can have the same meaning as in mathematical expression 1.
[0129] In one embodiment of the present disclosure, the importance score may be determined based on mathematical expression 3.
[0130]
[0131] In one embodiment of the present disclosure, the importance distribution determination module (830) is not limited to mathematical expressions 1 to 3, and may determine the importance score such that the importance score increases as the number of keywords included in a group increases, and the importance score decreases as the number of groups including keywords increases.
[0132] In one embodiment of the present disclosure, the importance distribution determination module (830) may obtain a confidence score indicating whether each frame included in the first group includes the first keyword. For example, the importance distribution determination module (830) may obtain c in Equation 2. k,d can be obtained. In one embodiment of the present disclosure, the reliability score can be determined from the artificial intelligence model of the keyword extraction module (810). For example, the artificial intelligence model of the keyword extraction module (810) can determine the reliability score corresponding to the first keyword for each frame, and ck,d in mathematical expression 2 can mean the average of the reliability scores of the first keywords included in the first group. In one embodiment of the present disclosure, the importance distribution determination module (830) can obtain the importance distribution information corresponding to the first keyword based on a plurality of reliability scores, the number of frames including the first keyword, and the number of groups including the first keyword.
[0133] In one embodiment of the present disclosure, the importance distribution determination module (830) can obtain importance distribution information based on the section-specific importance scores. The importance distribution determination module (830) can obtain the frame-specific importance scores by interpolating the section-specific importance scores.
[0134] In one embodiment of the present disclosure, the highlight keyword determination module (840) can determine highlight keywords based on importance distribution information. The highlight keyword determination module (840) can determine highlight keywords based on importance distribution information using the highlight keyword determination method described in the present disclosure.
[0135] In one embodiment of the present disclosure, the highlight keyword determination module (840) can determine highlight keywords based on the highlight keyword priority DB (860). In one embodiment of the present disclosure, the highlight keyword priority DB (860) can store weights for each keyword. In one embodiment of the present disclosure, the highlight keyword priority DB (860) can be stored in the memory of the electronic device (100) or obtained from an external electronic device. In one embodiment of the present disclosure, the highlight keyword determination module (840) can determine only keywords included in the highlight keyword priority DB (860) as highlight keywords.
[0136] The highlight keyword determination module (840) can determine highlight keywords by applying keyword-specific weights to the importance distribution information of each keyword. According to one embodiment of the present disclosure, the highlight keyword determination module (840) can determine highlight keywords based on the weighted importance distribution information using the highlight keyword determination method described in the present disclosure. For example, a keyword with a low weight may not be selected as a highlight keyword even if it has a high importance score in the keyword distribution information.
[0137] In one embodiment of the present disclosure, the highlight video generation module (850) can generate a highlight video corresponding to a highlight keyword. In one embodiment of the present disclosure, the highlight video generation module (850) can generate a highlight video including video frames corresponding to a highlight section.
[0138] In one embodiment of the present disclosure, the highlight section may be determined by at least one of the importance distribution determination module (830), the highlight keyword determination module (840), or the highlight video generation module (850). The highlight section may be determined based on an importance score. For example, the highlight section may include a section with a high importance score.
[0139] In one embodiment of the present disclosure, the highlight video generation module (850) can obtain a selection input for highlight keywords. The highlight keywords can be displayed through the display of the electronic device (100). The highlight keyword determination module (840) can obtain an input for selecting at least some of the displayed keywords. For example, the highlight keyword determination module (840) can obtain a selection input through an input interface. The highlight video generation module (850) can generate a highlight video corresponding to the highlight keyword corresponding to the input. The electronic device (100) can provide a highlight video that meets the user's intention by generating a highlight video based on the input for selecting a highlight keyword, and can provide a function for the user to modify the highlight video according to the user's intention.
[0140] FIG. 8b is a drawing for explaining modules included in an electronic device (100) according to one embodiment of the present disclosure.
[0141] In one embodiment of the present disclosure, the electronic device (100) may include a keyword extraction module (810), an interval sampling module (820), an importance distribution determination module (830), a generalized distribution determination module (835), a highlight keyword determination module (840), a generalized keyword determination module (845), and a highlight video generation module (850). With respect to the keyword extraction module (810), the interval sampling module (820), the importance distribution determination module (830), the highlight keyword determination module (840), and the highlight video generation module (850), the description made with reference to FIG. 8A is omitted.
[0142] In one embodiment of the present disclosure, the generalization distribution determination module (835) may obtain generalization distribution information corresponding to a keyword. In one embodiment of the present disclosure, the generalization distribution determination module (835) may obtain generalization distribution information including a generalization score according to time or frame. In one embodiment of the present disclosure, the generalization distribution determination module (835) may determine a generalization score based on the number of keywords included in multiple groups of a video.
[0143] In one embodiment of the present disclosure, the generalization distribution determination module (835) can identify the number of frames including the first keyword among the plurality of frames included in the first group among the plurality of groups. The generalization distribution determination module (835) can identify the number of groups including the first keyword among the plurality of groups. The generalization distribution determination module (835) can obtain generalization distribution information corresponding to the first keyword based on the number of frames including the first keyword and the number of groups including the first keyword. For example, the generalization distribution determination module (835) can determine a large generalization score as the number of frames including the first keyword increases, and can determine a large generalization score as the number of groups including the first keyword increases. In one embodiment of the present disclosure, the generalization score can be determined based on mathematical expression 4.
[0144]
[0145] Here, k represents the keyword index, d represents the group index, and TF represents the k,d is the number of keywords included in the group, N is the total number of groups in the video, df k can mean the number of groups containing the keyword.
[0146] In one embodiment of the present disclosure, the generalization distribution determination module (835) may obtain generalization distribution information based on the number of groups containing the first keyword. For example, the generalization distribution determination module (835) may determine a higher generalization score as the number of groups containing the first keyword increases, and may determine a lower generalization score as the number of groups containing the first keyword decreases.
[0147] In one embodiment of the present disclosure, the generalization distribution determination module (835) may obtain generalization distribution information based on the ratio between the number of frames of the first group including the first keyword and the total number of keywords included in the first group. The generalization distribution determination module (835) may determine the ratio between the number of frames of the first group including the first keyword and the total number of keywords included in the first group. For example, the generalization distribution determination module (835) may determine TF of Equation 1. k,d The ratio between the total number of keywords included in the first group d can be determined, and generalized distribution information can be obtained based on the determined ratio.
[0148] In one embodiment of the present disclosure, the generalization score can be determined based on Equation 5.
[0149]
[0150] Here, c k,d represents the confidence score for the keyword, and the remaining variables can have the same meaning as in mathematical expression 4.
[0151] In one embodiment of the present disclosure, the generalization distribution determination module (835) is not limited to Equations 4 and 5, and may determine the generalization score such that the generalization score increases as the number of keywords included in a group increases, and the generalization score increases as the number of groups including keywords increases.
[0152] In one embodiment of the present disclosure, the generalized distribution determination module (835) may obtain a confidence score indicating whether each frame included in the first group includes the first keyword. For example, the generalized distribution determination module (835) may obtain c in Equation 5. k,dIn one embodiment of the present disclosure, the reliability score can be determined by the method described in FIG. 8A. In one embodiment of the present disclosure, the generalization distribution determination module (835) can obtain generalization distribution information corresponding to the first keyword based on a plurality of reliability scores, the number of frames including the first keyword, and the number of groups including the first keyword.
[0153] In one embodiment of the present disclosure, the generalization distribution determination module (835) can obtain generalization distribution information based on interval-specific generalization scores. The generalization distribution determination module (835) can obtain frame-specific generalization scores by interpolating interval-specific generalization scores.
[0154] In one embodiment of the present disclosure, the generalized keyword determination module (845) may determine a generalized keyword based on generalized distribution information. In one embodiment of the present disclosure, a 'generalized keyword' may include a keyword identified in many sections of a video. The generalized keyword determination module (845) may determine a generalized keyword if a generalized score satisfies a predetermined condition. For example, the generalized keyword determination module (845) may determine a generalized keyword if the maximum or average of the generalized scores is greater than a threshold.
[0155] In one embodiment of the present disclosure, the importance distribution determination module (830) and the generalization distribution determination module (835) may be implemented as a single module. For example, a module including the importance distribution determination module (830) and the generalization distribution determination module (835) may determine importance distribution information and generalization distribution information based on the frequency of keywords. In one embodiment of the present disclosure, the highlight keyword determination module (840) and the generalization keyword determination module (845) may be implemented as a single module. For example, a module including the importance distribution determination module (830) and the generalization distribution determination module (835) may determine highlight keywords and generalized keywords based on the importance distribution information and the generalization distribution information.
[0156] FIG. 9 is a diagram showing a plurality of groups into which a video is divided according to one embodiment of the present disclosure.
[0157] In one embodiment of the present disclosure, the electronic device (100) can divide a video into a plurality of groups. The electronic device (100) can divide the video into a plurality of groups based on a division interval. For example, the electronic device (100) can divide the video into a first group (912) corresponding to section 1, a second group (914) corresponding to section 2, and a third group (916) corresponding to section 3. The first group (912), the second group (914), and the third group (916) can all include the same number of frames. The electronic device (100) can obtain importance distribution information and / or generalized distribution information based on the number of keywords included in the first group (912), the second group (914), and the third group (916).
[0158] In one embodiment of the present disclosure, the electronic device (100) can divide a video into a plurality of groups using a plurality of division intervals. For example, referring to FIG. 9, the electronic device (100) can divide the video into a first group (912), a second group (914), and a third group (916) using a first division interval. The electronic device (100) can divide the video into a plurality of groups including a fourth group (922) and a fifth group (924) using a second division interval. The electronic device (100) can divide the video into a plurality of groups including a sixth group (932) and a seventh group (934) using a third division interval.
[0159] In one embodiment of the present disclosure, the second split section may be half of the first split section, and the third split section may be half of the second split section. However, this is not limited thereto, and multiple split sections may be independently determined.
[0160] In one embodiment of the present disclosure, the electronic device (100) can obtain importance distribution information and / or generalized distribution information based on the number of keywords included in a plurality of groups divided using a plurality of division intervals. The electronic device (100) can obtain importance distribution information and / or generalized distribution information for each division interval.
[0161] In one embodiment of the present disclosure, the electronic device (100) can determine highlight keywords for each segmentation interval based on importance distribution information for each segmentation interval. The electronic device (100) can determine a final highlight keyword based on the highlight keywords for each segmentation interval.
[0162] In one embodiment of the present disclosure, the electronic device (100) may determine final importance distribution information and / or final generalized distribution information based on importance distribution information and / or generalized distribution information for each segmentation interval. For example, the electronic device (100) may determine final importance distribution information and / or final generalized distribution information based on an average or sum of importance scores of importance distribution information for each segmentation interval and / or generalized scores of generalized distribution information.
[0163] In one embodiment of the present disclosure, the electronic device (100) can determine a highlight keyword using final importance distribution information. By determining the highlight keyword using multiple segmentation intervals, the electronic device (100) can improve the reliability of highlight keyword determination.
[0164] FIG. 10 is a diagram illustrating a process for generating a highlight video according to one embodiment of the present disclosure.
[0165] In one embodiment of the present disclosure, the electronic device (100) can generate a highlight video according to a set of rules. For example, instead of operations S240 to S260 of FIG. 2 , the electronic device (100) can determine a highlight section based on a highlight keyword and generate a highlight video based on the highlight section.
[0166] In one embodiment of the present disclosure, the importance distribution determination module (830) can determine importance distribution information (1010) for each keyword. The importance distribution determination module (830) can determine importance distribution information (1010) corresponding to each keyword of a video. For example, the importance distribution determination module (830) can obtain importance distribution information (1010) related to kiss, love, cuteness, daily life, hug, laughter, and walking, respectively.
[0167] In one embodiment of the present disclosure, the highlight keyword determination module (840) can determine highlight keywords. The highlight keyword determination module (840) can select some of a plurality of video keywords as highlight keywords. For example, the highlight keyword determination module (840) can determine "kiss," "love," and "daily life" as highlight keywords.
[0168] In one embodiment of the present disclosure, the highlight video generation module (850) can generate a highlight video based on highlight keywords. The highlight video generation module (850) can generate a highlight video based on importance distribution information (1020) corresponding to the selected highlight keywords. For example, the highlight video generation module (850) can select all of the highlight keywords "kiss," "love," and "daily life," and generate a highlight video based on the importance distribution information (1020) of "kiss," "love," and "daily life."
[0169] In one embodiment of the present disclosure, the highlight video generation module (850) can generate a highlight video according to predetermined criteria. The highlight video generation module (850) can select at least some highlight keywords according to predetermined criteria. The highlight video generation module (850) can generate a highlight video using the selected highlight keywords.
[0170] In one embodiment of the present disclosure, the highlight video generation module (850) may select at least some of the highlight keywords based on user preference information. The user preference information may include keywords that the user has previously expressed interest in and / or keywords used in previously generated highlight videos.
[0171] In one embodiment of the present disclosure, the highlight video generation module (850) may select at least some of the highlight keywords based on priority information. The priority information may include predetermined keyword priorities. Based on the priority information, the highlight video generation module (850) may select some of the highlight keywords according to their priorities.
[0172] In one embodiment of the present disclosure, the highlight video generation module (850) can determine a highlight section corresponding to a highlight keyword. The highlight section can be determined based on the importance score of the importance distribution information (1020). The highlight video generation module (850) can combine highlight sections to determine a final highlight section (1030). For example, the highlight video generation module (850) can determine a final highlight section (1030) by combining highlight sections of highlight keywords selected to generate a highlight video.
[0173] The highlight video generation module (850) can generate a highlight video based on the final highlight section (1030). The highlight video generation module (850) can generate a highlight video so as to include frames of the final highlight section.
[0174] FIG. 11 is a diagram illustrating a process for determining highlight keywords and generalization keywords according to one embodiment of the present disclosure.
[0175] In one embodiment of the present disclosure, the keyword extraction module (810) can acquire (e.g., extract or determine) keywords from a video. Referring to FIG. 11 , for example, the video may include a frame of a woman walking on a beach with her pet. The keywords may include "woman," "pet," and "beach."
[0176] The highlight keyword determination module (840) can determine highlight keywords from among the keywords extracted by the keyword extraction module (810). For example, highlight keywords may include "kiss," "hug," "laughter," and "love." Highlight keywords may be keywords that frequently appear in certain sections of a video. For example, highlight keywords may be keywords included only in frames of certain sections of a video. Highlight keywords may be used to determine highlight sections or generate highlight videos.
[0177] The generalized keyword determination module (845) can determine generalized keywords from among the keywords extracted by the keyword extraction module (810). For example, generalized keywords may include "beach," "pet," and "woman." Generalized keywords may be keywords that frequently appear throughout the entire frame of a video. For example, the video may include scenes of a woman and a pet on a beach. Generalized keywords can be used to search a video. For example, generalized keywords may be suitable for video search because they appear throughout the entire video.
[0178] FIG. 12 is a diagram showing a UI in which an electronic device (100) displays a highlight keyword according to one embodiment of the present disclosure.
[0179] In one embodiment of the present disclosure, the electronic device (100) may display an indicator of a video. For example, the electronic device (100) may display a highlight video as an indicator representing the video.
[0180] In one embodiment of the present disclosure, an electronic device (100) may display stored videos. The electronic device (100) may display a UI (1210) representing the stored videos. Referring to FIG. 12 , the UI (1210) may include an area (1215) that displays one or more stored videos. For example, the electronic device (100) may display the UI (1210) including an area that displays an image representing a video or a portion of the video (e.g., a highlight video). The UI (1210) may include an area that displays the playback time of each video.
[0181] In one embodiment of the present disclosure, the electronic device (100) may obtain an input for selecting one of the videos. The electronic device (100) may display the selected video. In one embodiment of the present disclosure, the electronic device (100) may display an image representing the video or a highlight video. For example, the electronic device (100) may display a UI (1220) representing the selected video in response to the input. The UI (1220) may include an area (1225) for displaying an image of the selected video or a highlight video and an area (1230) for displaying keywords corresponding to the displayed image or highlight video. For example, the electronic device (100) may display a highlight video and display highlight keywords corresponding to highlight sections or frames of the displayed highlight video.
[0182] In one embodiment of the present disclosure, an electronic device (100) can transmit a highlight video to another electronic device. The UI (1220) may include an indicator indicating that the highlight video is being transmitted to another electronic device. For example, the UI (1220) may include a share icon (1235). A user can view the highlight video displayed on the electronic device (100) and share it with another electronic device.
[0183] FIG. 13 is a drawing showing a UI for editing a video by an electronic device (100) according to one embodiment of the present disclosure.
[0184] In one embodiment of the present disclosure, an electronic device (100) can transmit a highlight video to another electronic device based on a keyword. The electronic device (100) can display a UI (1300) for editing the video. The electronic device (100) can display keywords of the video. For example, referring to FIG. 13 , the UI (1300) can include an area (1310) for displaying keywords. The electronic device (100) can display a video corresponding to the keyword. For example, the UI (1300) can include an area (1320) for displaying a video corresponding to the keyword. The electronic device (100) can display a section of the highlight video in which the video corresponding to the keyword is included. For example, the UI (1300) can include an area (1330) for displaying a section of the highlight video in which the video corresponding to the keyword is included.
[0185] The electronic device (100) can generate a highlight video based on an input corresponding to a keyword. For example, the electronic device (100) can generate a highlight video including a highlight section corresponding to 'congratulations' based on an input for a keyword corresponding to 'congratulations'. The electronic device (100) can generate a highlight video including a highlight section corresponding to 'food' after a highlight section corresponding to 'congratulations' based on an input for 'food' after an input for 'congratulations'. In one embodiment of the present disclosure, the electronic device (100) can transmit a highlight video generated using one or more keywords to another electronic device.
[0186] FIG. 14 is a diagram illustrating a UI including search results of images and / or videos according to one embodiment of the present disclosure.
[0187] In one embodiment of the present disclosure, the electronic device (100) may display search results for images and / or videos based on keywords. The electronic device (100) may display a UI (1400) that displays search results for images and / or videos. For example, referring to FIG. 14 , the UI (1400) may include an area (1410) that displays searched keywords and an area (1420) that displays videos corresponding to the keywords.
[0188] In one embodiment of the present disclosure, the electronic device (100) can search for images and / or videos based on highlight keywords and / or generalized keywords. The electronic device (100) can display images and / or videos searched for based on highlight keywords and / or generalized keywords. For example, the electronic device (100) can display only images and / or videos that have the searched keyword as a highlight keyword, only images and / or videos that have the searched keyword as a generalized keyword, or only images and / or videos that have the searched keyword as a highlight keyword or generalized keyword.
[0189] In one embodiment of the present disclosure, the electronic device (100) may display the type of keyword corresponding to the displayed image and / or video. For example, the electronic device (100) may display an indicator indicating that the image and / or video has the searched keyword as a highlighted keyword. For example, the electronic device (100) may display an indicator indicating that the image and / or video has the searched keyword as a generalized keyword. Here, the indicator may be displayed by an icon, text, color, or a distinction in the displayed area.
[0190] FIG. 15 is a drawing showing a UI for editing multiple videos by an electronic device (100) according to one embodiment of the present disclosure.
[0191] In one embodiment of the present disclosure, the electronic device (100) can generate a highlight video based on a plurality of videos. In one embodiment of the present disclosure, the electronic device (100) can display stored videos. The electronic device (100) can display a UI (1510) representing the stored videos. Referring to FIG. 15, the UI (1510) can include an area (1512) that displays a plurality of stored videos. In one embodiment of the present disclosure, the UI (1510) can correspond to the UI (1210) of FIG. 12.
[0192] In one embodiment of the present disclosure, the electronic device (100) may obtain an input for selecting multiple videos. The electronic device (100) may display an indicator indicating the selected multiple videos. For example, the electronic device (100) may display a checkmark icon on the selected multiple videos. The electronic device (100) may generate a highlight video based on the selected multiple videos.
[0193] The electronic device (100) may display a first video among the selected videos. For example, the electronic device (100) may display a UI (1520) for editing the first video. The electronic device (100) may display keywords of the first video. For example, referring to FIG. 15 , the UI (1520) may include an area (1522) for displaying keywords. The electronic device (100) may display a first video corresponding to the keyword. For example, the UI (1520) may include an area (1524) for displaying the first video corresponding to the keyword. The electronic device (100) may display a section including the first video corresponding to the keyword. For example, the UI (1520) may include an area (1526) for displaying a section including the first video corresponding to the keyword.
[0194] The electronic device (100) can generate a highlight video based on an input corresponding to a keyword. For example, the electronic device (100) can generate a highlight video including a highlight section corresponding to 'congratulations' based on an input for a keyword corresponding to 'congratulations'. The electronic device (100) can generate a highlight video based on one or more keywords of the first video. The electronic device (100) can display another video. For example, the UI (1520) can include an area (1528) for proceeding to an editing UI of another video.
[0195] The electronic device (100) may display a second video from among the selected videos. For example, the electronic device (100) may display a UI (1530) for editing the second video. The electronic device (100) may display keywords of the second video. For example, the UI (1530) may include an area (1532) for displaying keywords. The electronic device (100) may display a second video corresponding to the keyword. For example, the UI (1530) may include an area (1534) for displaying the second video corresponding to the keyword. The electronic device (100) may display a section including the second video corresponding to the keyword. For example, the UI (1530) may include an area (1536) for displaying a section including the second video corresponding to the keyword.
[0196] The electronic device (100) can generate a highlight video based on an input corresponding to a keyword. For example, the electronic device (100) can generate a highlight video including a highlight section corresponding to "high five" based on an input for a keyword corresponding to "high five." The electronic device (100) can generate a highlight video based on one or more keywords of a second video.
[0197] FIG. 16 is a diagram showing a UI for sharing a highlight video based on a keyword by an electronic device (100) according to one embodiment of the present disclosure.
[0198] In one embodiment of the present disclosure, an electronic device (100) can transmit a highlight video to another electronic device based on a keyword. The electronic device (100) can display a UI (1600) for editing the video. The electronic device (100) can display keywords of the video. For example, referring to FIG. 16 , the UI (1600) can include an area (1610) for displaying keywords. The electronic device (100) can display a video corresponding to the keyword. The electronic device (100) can display a section of the highlight video in which the video corresponding to the keyword is included.
[0199] The electronic device (100) can obtain inputs corresponding to multiple keywords. For example, the electronic device (100) can obtain inputs corresponding to "celebration" and "food." The electronic device (100) can transmit a highlight video generated based on the multiple keywords corresponding to the inputs to another electronic device. For example, the UI (1600) can include, for example, a share icon (1620). The user can check the highlight video displayed on the electronic device (100) and share it with another electronic device.
[0200] FIG. 17 is a diagram illustrating a UI including search results of images and / or videos according to one embodiment of the present disclosure.
[0201] In one embodiment of the present disclosure, the electronic device (100) may display search results of images and / or videos based on keywords. The electronic device (100) may display a UI (1700) that displays search results of images and / or videos. For example, referring to FIG. 17 , the UI (1700) may include an area (1710) that displays searched keywords and an area (1720) that displays videos corresponding to the keywords.
[0202] In one embodiment of the present disclosure, the electronic device (100) can search for images and / or videos based on highlight keywords and / or generalized keywords. The electronic device (100) can display images and / or videos searched for based on highlight keywords and / or generalized keywords. The electronic device (100) can display videos that have the searched keywords as highlight keywords and / or generalized keywords, or videos that have keywords similar to the searched keywords as highlight keywords and / or generalized keywords.
[0203] In one embodiment of the present disclosure, the electronic device (100) may display the type of keyword corresponding to the displayed image and / or video. For example, the electronic device (100) may display an indicator indicating that the image and / or video has the searched keyword as a highlighted keyword. For example, the electronic device (100) may display an indicator indicating that the image and / or video has the searched keyword as a generalized keyword. For example, the electronic device (100) may display an indicator indicating that the image and / or video has a keyword similar to the searched keyword as a generalized keyword. For example, the electronic device (100) may display a video having the keyword "love", which is similar to the searched keyword "heart," as a highlighted keyword.
[0204] FIG. 18 is a block diagram showing the configuration of an electronic device according to one embodiment of the present disclosure.
[0205] In one embodiment of the present disclosure, an electronic device (100) may include a processor (1810) and a memory (1820).
[0206] The processor (1810) can control the overall operations of the electronic device (100). For example, the processor (1810) can control the overall operations of the electronic device (100) for generating a highlight video by executing one or more instructions of a program stored in the memory (1820). In one embodiment of the present disclosure, the processor (1810) can be configured with multiple processors. For example, the multiple processors can individually or collectively execute one or more instructions of the program stored in the memory (1820) to control the overall operations of the electronic device (100) for generating a highlight video.
[0207] In one embodiment of the present disclosure, the processor (1810) may include a configuration that controls a series of processes so that the electronic device (100) operates according to the embodiments described in the present disclosure. The processor (1810) may be composed of one or more processors. The one or more processors included in the processor (1810) may include circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), and the like. The processor (1810) may be one or more processors including, but not limited to, a central processing unit, a microprocessor unit, an application processor, a digital signal processor (DSP), a graphic processing unit, a vision processing unit (VPU), an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), a neural processing unit, a communication processor, and / or an artificial intelligence processor designed as a hardware structure for processing an artificial intelligence model.
[0208] Meanwhile, although not illustrated in FIG. 18, the electronic device (100) may further include additional components to perform the operations described in the aforementioned embodiments. For example, the electronic device (100) may further include a display, a camera, a microphone, a speaker, an input / output interface, and the like.
[0209] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one processor or multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment of the present disclosure, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an AI-dedicated processor). Here, an AI-dedicated processor, which is an example of the second processor, may perform operations for training / inference of an AI model. However, the embodiments of the present disclosure are not limited thereto.
[0210] One or more processors (1810) according to the present disclosure may be implemented as a single-core processor or as a multi-core processor.
[0211] When a method according to one embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one core or may be performed by multiple cores included in one or more processors.
[0212] At least one processor (1810) according to an embodiment of the present invention may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include various processing circuits including at least one processor, one or more of which are configured to individually and / or collectively perform the various functions described herein in a distributed manner. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms encompass, for example, without limitation, a single processor performing some of the recited functions, other processor(s) performing other of the recited functions, and even a single processor performing all of the recited functions. Additionally, the at least one processor may include a combination of processors that perform the various functions enumerated / disclosed, for example, in a distributed manner. The at least one processor may execute program instructions to achieve or perform the various functions.
[0213] The memory (1820) may store instructions, data structures, and program codes that can be read by the processor (1810). Operations performed by the processor (1810) may be implemented by executing instructions or codes of a program stored in the memory (1820).
[0214] The memory (1820) may include at least one of volatile memory and non-volatile memory. For example, the memory (1820) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, an optical disk, a RAM (Random Access Memory), or a SRAM (Static Random Access Memory).
[0215] The memory (1820) may store one or more instructions and / or programs that cause the electronic device (100) to operate to generate a highlight video. For example, the memory (1820) may store instructions and / or programs for implementing operations for generating a highlight video. The processor (1810) may write data to the memory (1820) or read data stored in the memory (1820). The processor (1810) may process data according to predefined operation rules or artificial intelligence models by executing the program or at least one instruction stored in the memory (1820). The processor (1810) may perform operations described in embodiments of the present disclosure.
[0216] Meanwhile, the memory (1820) may further store instructions and / or programs for implementing at least one of the functions of a keyword extraction module, an interval sampling module, an importance distribution information generation module, a generalized distribution information generation module, a highlight keyword determination module, a generalized keyword determination module, or a highlight video generation module.
[0217] In one embodiment of the present disclosure, at least one instruction stored in a memory (1820) may be individually or collectively executed by at least one processor (1810) to enable an electronic device to obtain one or more keywords included in a plurality of frames of a video. The at least one instruction may be individually or collectively executed by at least one processor (1810) to enable the electronic device to obtain importance distribution information corresponding to each of the one or more keywords. The at least one instruction may be individually or collectively executed by at least one processor (1810) to enable the electronic device to determine at least one highlight keyword from among the one or more keywords based on the importance distribution information. The at least one instruction may be individually or collectively executed by at least one processor (1810) to enable the electronic device to display at least one highlight keyword. The at least one instruction may be individually or collectively executed by at least one processor (1810) to enable the electronic device to obtain an input corresponding to at least some of the at least one highlighted keyword displayed. At least one instruction may be individually or collectively executed by at least one processor (1810) to enable the electronic device to generate a highlight video based on a highlight keyword corresponding to an input.
[0218] In one embodiment of the present disclosure, the electronic device (100) may include other components in addition to the processor (1810) and the memory (1820). Components that the electronic device (100) may include in one embodiment of the present disclosure are described in detail with reference to FIG. 19.
[0219] FIG. 19 is a block diagram showing the configuration of an electronic device according to one embodiment of the present disclosure.
[0220] As illustrated in FIG. 19, an electronic device (100) according to one embodiment of the present disclosure may further include a camera (1910), a sensor unit (1920), a communication interface (1930), and / or a user interface (1940) in addition to a processor (1810) and a memory (1820).
[0221] The processor (1810) controls the operation of the electronic device (100). The processor (1810) can control the camera (1910), the sensor unit (1920), the communication interface (1930), the user interface (1940), and the memory (1820) by executing programs stored in the memory (1820).
[0222] The memory (1820) may store programs for processing and controlling the processor (1810), and may also store input / output data (e.g., video keywords, highlight keywords, generalized keywords, etc.). The memory (1820) may also store an artificial intelligence model. For example, the memory (1820) may store an artificial intelligence model for keyword acquisition.
[0223] A camera (1910) may refer to a device that acquires at least one frame. Here, the at least one frame may be expressed as an image (still image or moving image) or a photograph.
[0224] The camera (1910) may include a wide-angle camera and / or a telephoto camera capable of capturing the front or rear of the electronic device (100). The camera (1910) may also include a miniature camera or a pinhole camera. In one embodiment of the present disclosure, the electronic device (100) may include multiple cameras (1910).
[0225] The camera (1910) may include an image sensor (1911) and / or an image signal processor (ISP) (1912). The image sensor (1911) may include a device that identifies light transmitted through a lens of the camera (1910). The image signal processor (1912) may refer to a processor that processes a signal identified by the image sensor (1911). In one embodiment of the present disclosure, the image signal processor (1912) may perform primary correction for distortion generated by the camera (1910).
[0226] The sensor unit (1920) may include, but is not limited to, a depth sensor (1921) and / or an infrared sensor (1922). The depth sensor (1921) may include a sensor that measures depth information of an object. The infrared sensor (1922) may include a sensor that measures numerical values using infrared rays. The infrared sensor (1922) may measure depth information of an object using infrared rays.
[0227] The sensor unit (1920) can obtain depth information using a UWB sensor. The UWB sensor can measure the round-trip time of a UWB signal. The sensor unit (1920) can determine depth information based on the round-trip time of the UWB signal.
[0228] The communication interface (1930) may include one or more components that enable communication between the electronic device (100) and a server device, or between the electronic device (100) and a mobile terminal. For example, the communication interface (1930) may include a short-range communication unit (1931), a long-range communication unit (1932), etc.
[0229] The short-range communication unit (1931) may include a Bluetooth communication unit, a BLE (Bluetooth Low Energy) communication unit, a near field communication unit (NFC, Near Field Communication unit), a WLAN (Wi-Fi) communication unit, a Zigbee communication unit, an infrared (IrDA, infrared Data Association) communication unit, a WFD (Wi-Fi Direct) communication unit, an UWB (ultra wideband) communication unit, or an Ant+ communication unit, but the present disclosure is not limited thereto.
[0230] The remote communication unit (1932) may include the Internet, a computer network (e.g., a LAN or WAN), and a mobile communication unit. The mobile communication unit transmits and receives wireless signals with at least one of a base station, an external terminal, and a server on the mobile communication network. For example, the wireless signals may include various types of data according to transmission and reception of voice call signals, video call call signals, or text / multimedia messages. The mobile communication unit may include a 3G module, a 4G module, a 5G module, an LTE module, an NB-IoT module, an LTE-M module, and the like, but the present disclosure is not limited thereto.
[0231] The user interface (1940) may include an output interface (1941) and an input interface (1942). The output interface (1941) is for outputting an audio signal or a video signal and may include, for example, a display and / or an audio output unit.
[0232] In one embodiment of the present disclosure, the display may be configured as a touch screen by forming a layer structure with a touchpad. When the display and the touchpad are configured as a touch screen by forming a layer structure, the display may be used as an input interface (1942) in addition to an output interface (1941). The display may include at least one of a liquid crystal display, a thin film transistor-liquid crystal display, a light-emitting diode (LED), an organic light-emitting diode (OLED), a flexible display, a 3D display, or an electrophoretic display. In addition, depending on the implementation form of the electronic device (100), the electronic device (100) may include two or more displays. For example, the electronic device (100) may include a front-facing display and a rear-facing display opposite to the front-facing display.
[0233] According to one embodiment of the present disclosure, the display can display and output information processed in the electronic device (100). For example, the display can display an image captured by a camera (1910) of the electronic device (100) in real time, or display images, videos, and / or highlight videos stored in the memory (1820). For example, the display can display the UI described with reference to FIGS. 12 to 17. The display can output an interface for controlling the electronic device (100), an interface for displaying the status of the electronic device (100), and the like.
[0234] The audio output unit may output audio data received from the communication interface (1930) or stored in the memory (1820). In addition, the audio output unit may output audio signals related to functions performed in the electronic device (100). For example, the audio output unit may include a speaker or a buzzer. For example, the speaker or buzzer may output signals related to functions performed in the electronic device (1000) (e.g., call signal reception sound, message reception sound, notification sound) as sound.
[0235] The input interface (1942) can receive input from a user. The input interface (1942) can include at least one of a key pad, a dome switch, a touch pad (e.g., a contact-type electrostatic capacitance type, a pressure-type resistive type, an infrared sensing type, a surface ultrasonic conduction type, an integral tension measurement type, or a piezoelectric effect type), a jog wheel, a jog switch, and a microphone, but the present disclosure is not limited thereto.
[0236] In one embodiment of the present disclosure, a method for generating a highlight video based on keywords by an electronic device is provided. The method may include obtaining one or more keywords included in a plurality of frames of a video. The method may include obtaining importance distribution information corresponding to each of the one or more keywords. The method may include determining at least one highlight keyword from among the one or more keywords based on the importance distribution information. The method may include displaying at least one highlight keyword. The method may include obtaining an input corresponding to at least some of the displayed at least one highlight keyword. The method may include generating a highlight video based on the highlight keyword corresponding to the input.
[0237] In one embodiment of the present disclosure, the step of obtaining one or more keywords included in a video may include the step of obtaining a first frame from among a plurality of frames of the video. The step of obtaining one or more keywords included in the video may include the step of obtaining feature information included in the first frame. The feature information may include information indicating at least one of an emotion, an object, an action, a behavior, a situation, a place, or an event included in the first frame. The step of obtaining one or more keywords included in the video may include the step of obtaining one or more keywords so as to include keywords corresponding to the feature information.
[0238] In one embodiment of the present disclosure, one or more keywords may include a first keyword. The step of obtaining importance distribution information corresponding to each of the one or more keywords may include a step of dividing the video into a plurality of groups based on a division interval. The step of obtaining importance distribution information corresponding to each of the one or more keywords may include a step of identifying the number of frames including the first keyword among a plurality of frames included in a first group among the plurality of groups. The step of obtaining importance distribution information corresponding to each of the one or more keywords may include a step of identifying the number of groups including the first keyword among the plurality of groups. The step of obtaining importance distribution information corresponding to each of the one or more keywords may include a step of obtaining importance distribution information corresponding to the first keyword based on the number of frames including the first keyword and the number of groups including the first keyword.
[0239] In one embodiment of the present disclosure, the step of obtaining importance distribution information corresponding to the first keyword may include the step of determining the total number of keywords included in a first group of a video. The step of obtaining importance distribution information corresponding to the first keyword may include the step of determining a first ratio between the number of frames of the first group including the first keyword and the total number of keywords included in the first group. The step of obtaining importance distribution information corresponding to the first keyword may include the step of determining a second ratio between the number of groups including the first keyword and the total number of groups included in the video. The step of obtaining importance distribution information corresponding to the first keyword may include the step of obtaining importance distribution information corresponding to the first keyword based on the first ratio and the second ratio.
[0240] In one embodiment of the present disclosure, the step of obtaining importance distribution information corresponding to the first keyword may include the step of obtaining a plurality of confidence scores indicating whether each of the plurality of frames of the first group includes the first keyword. The step of obtaining importance distribution information corresponding to the first keyword may include the step of obtaining importance distribution information corresponding to the first keyword based on the plurality of confidence scores, the number of frames, and the number of groups.
[0241] In one embodiment of the present disclosure, one or more keywords may include a second keyword. The importance distribution information corresponding to the second keyword may include an importance score of the second keyword according to a frame or time of a video. The step of determining at least one highlight keyword may include a step of determining at least one highlight keyword to include the second keyword based on whether the importance score of the second keyword satisfies a predetermined condition. The step of determining at least one highlight keyword may include a step of determining at least one highlight keyword not to include the second keyword based on whether the importance score of the second keyword does not satisfy a predetermined condition.
[0242] In one embodiment of the present disclosure, the step of generating a highlight video may include a step of determining a highlight section based on a highlight keyword corresponding to an input. The step of generating a highlight video may include a step of generating a highlight video including a plurality of frames of the video corresponding to the highlight section.
[0243] In one embodiment of the present disclosure, the input may include a first input corresponding to a first highlight keyword and a second input corresponding to a second highlight keyword. The step of determining the highlight section may include a step of determining the first highlight section based on first importance distribution information corresponding to the first highlight keyword. The step of determining the highlight section may include a step of determining the second highlight section based on second importance distribution information corresponding to the second highlight keyword. The step of determining the highlight section may include a step of determining a highlight section including the first highlight section and the second highlight section.
[0244] In one embodiment of the present disclosure, the first importance distribution information may include the importance scores of the first highlight keyword over frames or time of the video. The first highlight section may include frames with the maximum importance score of the first highlight keyword and frames with an importance score of the first highlight keyword greater than a threshold value.
[0245] In one embodiment of the present disclosure, the method may include the step of displaying a highlight video. The method may include the step of displaying highlight keywords corresponding to highlight sections of the displayed highlight video.
[0246] In one embodiment of the present disclosure, an electronic device is provided that generates a highlight video based on keywords. The electronic device may include at least one processor including a processing circuit and a memory that stores at least one instruction. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device obtains one or more keywords included in a plurality of frames of a video. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device obtains importance distribution information corresponding to each of the one or more keywords. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device determines at least one highlight keyword from among the one or more keywords based on the importance distribution information. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device displays at least one highlight keyword. The at least one instruction may be individually or collectively executed by the at least one processor, such that the electronic device obtains an input corresponding to at least some of the at least one highlighted keyword displayed. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to generate a highlight video based on a highlight keyword corresponding to an input.
[0247] In one embodiment of the present disclosure, at least one instruction may be individually or collectively executed by at least one processor to cause an electronic device to obtain a first frame from among a plurality of frames of a video. Feature information included in the first frame may be obtained, and the feature information may include information indicating at least one of an emotion, an object, an action, an action, a situation, a location, or an event included in the first frame. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to obtain one or more keywords such that the keywords correspond to the feature information.
[0248] In one embodiment of the present disclosure, one or more keywords may include a first keyword. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to divide the video into a plurality of groups based on a division interval. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to identify the number of frames including the first keyword among a plurality of frames included in a first group among the plurality of groups. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to identify the number of groups including the first keyword among the plurality of groups. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to obtain importance distribution information corresponding to the first keyword based on the number of frames including the first keyword and the number of groups including the first keyword.
[0249] In one embodiment of the present disclosure, at least one instruction may be individually or collectively executed by at least one processor to cause an electronic device to determine a total number of keywords included in a first group of a video. The at least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to determine a first ratio between a number of frames of the first group including the first keyword and a total number of keywords included in the first group. The at least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to determine a second ratio between a number of groups including the first keyword and a total number of groups included in the video. The at least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to obtain importance distribution information corresponding to the first keyword based on the first ratio and the second ratio.
[0250] In one embodiment of the present disclosure, at least one instruction may be individually or collectively executed by at least one processor to cause an electronic device to obtain a plurality of confidence scores indicating whether each of a plurality of frames of a first group includes a first keyword. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to obtain importance distribution information corresponding to the first keyword based on the plurality of confidence scores, the number of frames, and the number of groups.
[0251] In one embodiment of the present disclosure, one or more keywords may include a second keyword. The importance distribution information corresponding to the second keyword may include an importance score of the second keyword according to a frame or time of a video. At least one instruction may be individually or collectively executed by at least one processor, such that the electronic device may determine at least one highlight keyword to include the second keyword based on whether the importance score of the second keyword satisfies a predetermined condition. At least one instruction may be individually or collectively executed by at least one processor, such that the electronic device may determine at least one highlight keyword not to include the second keyword based on whether the importance score of the second keyword does not satisfy a predetermined condition.
[0252] In one embodiment of the present disclosure, at least one instruction may be individually or collectively executed by at least one processor to cause an electronic device to determine a highlight section based on a highlight keyword corresponding to an input. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to generate a highlight video comprising a plurality of frames of the video corresponding to the highlight section.
[0253] In one embodiment of the present disclosure, the input may include a first input corresponding to a first highlight keyword and a second input corresponding to a second highlight keyword. At least one instruction may be individually or collectively executed by at least one processor, such that the electronic device may determine a first highlight section based on first importance distribution information corresponding to the first highlight keyword. At least one instruction may be individually or collectively executed by at least one processor, such that the electronic device may determine a second highlight section based on second importance distribution information corresponding to the second highlight keyword. At least one instruction may be individually or collectively executed by at least one processor, such that the electronic device may determine a highlight section including the first highlight section and the second highlight section.
[0254] In one embodiment of the present disclosure, the first importance distribution information may include the importance scores of the first highlight keyword over frames or time of the video. The first highlight section may include frames with the maximum importance score of the first highlight keyword and frames with an importance score of the first highlight keyword greater than a threshold value.
[0255] In one embodiment of the present disclosure, at least one instruction may be individually or collectively executed by at least one processor to cause an electronic device to display a highlight video. At least one instruction may be individually or collectively executed by at least one processor to cause the electronic device to display a highlight keyword corresponding to a highlight section of the displayed highlight video.
[0256] In one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing an operation of an electronic device, and executing any one of the methods described above and below may be provided.
[0257] According to one embodiment of the present disclosure, the computer-executable instructions, such as program modules executed by a computer, may also be implemented in the form of a recording medium. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.
[0258] A computer-readable storage medium according to one embodiment of the present disclosure may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is stored semi-permanently in the storage medium and cases where data is stored temporarily. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0259] A method according to one embodiment of the present disclosure may be provided as a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0260] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.
[0261] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. A method for an electronic device to generate a highlight video based on keywords, A step (S210) of obtaining one or more keywords included in multiple frames of a video; A step (S220) of obtaining importance distribution information corresponding to each of the above one or more keywords; A step (S230) of determining at least one highlight keyword among the one or more keywords based on the above importance distribution information; A step of displaying at least one highlight keyword (S240); A step (S250) of obtaining an input corresponding to at least some of the at least one highlighted keyword displayed above; A method comprising a step (S260) of generating a highlight video based on a highlight keyword corresponding to the above input.
2. In paragraph 1, The step (S210) of obtaining one or more keywords included in the above video is: A step of obtaining a first frame from among a plurality of frames of the above video; A step of acquiring feature information included in the first frame, wherein the feature information includes information indicating at least one of an emotion, an object, an action, a behavior, a situation, a place, or an event included in the first frame; and A method comprising the step of obtaining one or more keywords so as to include keywords corresponding to the above characteristic information.
3. In any one of paragraphs 1 and 2, wherein said one or more keywords include a first keyword, The step (S220) of obtaining importance distribution information corresponding to each of the above one or more keywords is as follows: A step of dividing the video into a plurality of groups based on a segmentation interval; A step of identifying the number of frames including the first keyword among the plurality of frames included in the first group among the plurality of groups; A step of identifying the number of groups including the first keyword among the plurality of groups; and A method comprising the step of obtaining importance distribution information corresponding to the first keyword based on the number of frames including the first keyword and the number of groups including the first keyword.
4. In paragraph 3, The step of obtaining importance distribution information corresponding to the first keyword above is: A step of determining the total number of keywords included in the first group of the above video; A step of determining a first ratio between the number of frames of the first group including the first keyword and the total number of keywords included in the first group; a step of determining a second ratio between the number of groups including the first keyword and the total number of groups included in the video; and A method comprising a step of obtaining importance distribution information corresponding to the first keyword based on the first ratio and the second ratio.
5. In any one of paragraphs 3 to 4, The step of obtaining importance distribution information corresponding to the first keyword above is: A step of obtaining a plurality of confidence scores indicating whether each of the plurality of frames of the first group includes the first keyword; and A method comprising the step of obtaining importance distribution information corresponding to the first keyword based on the plurality of reliability scores, the number of frames, and the number of groups.
6. In any one of paragraphs 1 to 5, the one or more keywords include a second keyword, The importance distribution information corresponding to the second keyword includes the importance score of the second keyword according to the frame or time of the video, The step (S230) of determining at least one highlight keyword is: A step of determining at least one highlight keyword to include the second keyword based on whether the importance score of the second keyword satisfies a defined condition; and A method comprising a step of determining at least one highlight keyword not to include the second keyword based on the importance score of the second keyword not satisfying the set condition.
7. In any one of paragraphs 1 to 6, The step of generating the above highlight video (S260) is: A step of determining a highlight section based on a highlight keyword corresponding to the above input; and A method comprising the step of generating a highlight video, the highlight video including a plurality of frames of the video corresponding to the highlight section.
8. In the 7th paragraph, the input includes a first input corresponding to the first highlight keyword and a second input corresponding to the second highlight keyword, The step of determining the above highlighted section is: A step of determining a first highlight section based on first importance distribution information corresponding to the first highlight keyword; A step of determining a second highlight section based on second importance distribution information corresponding to the second highlight keyword; and A method comprising the step of determining the highlight section including the first highlight section and the second highlight section.
9. In paragraph 8, The above first importance distribution information includes the importance score of the first highlight keyword according to the frame or time of the video, A method wherein the first highlight section includes a frame in which the importance score of the first highlight keyword is maximum and a frame in which the importance score of the first highlight keyword is greater than a threshold value.
10. In any one of paragraphs 7 to 9, a step of displaying the above highlight video; and A method further comprising the step of displaying highlight keywords corresponding to highlight sections of the highlighted video displayed above.
11. In an electronic device (100) that generates a highlight video based on a keyword, At least one processor (1810) comprising a processing circuit; and A memory (1820) storing at least one instruction, wherein the at least one instruction is individually or collectively executed by the at least one processor (1810), so that the electronic device (100) Obtain one or more keywords contained in multiple frames of a video, Obtain importance distribution information corresponding to each of the above one or more keywords, Based on the above importance distribution information, at least one highlight keyword is determined from among the one or more keywords, Display at least one highlight keyword above, Obtaining an input corresponding to at least some of at least one of the highlighted keywords displayed above, An electronic device that generates a highlight video based on highlight keywords corresponding to the above input.
12. In paragraph 11, The at least one instruction is individually or collectively executed by the at least one processor (1810) so that the electronic device (100) Obtaining the first frame from among multiple frames of the above video, Obtaining feature information included in the first frame, wherein the feature information includes information indicating at least one of an object, action, behavior, situation, place, or event included in the first frame, An electronic device that acquires one or more keywords to include keywords corresponding to the above characteristic information.
13. In any one of paragraphs 11 to 12, wherein said one or more keywords include a first keyword, The at least one instruction is individually or collectively executed by the at least one processor (1810) so that the electronic device (100) Divide the video into multiple groups based on the segmentation interval, Identifying the number of frames including the first keyword among the plurality of frames included in the first group among the plurality of groups, Identify the number of groups that include the first keyword among the plurality of groups, An electronic device that obtains importance distribution information corresponding to the first keyword based on the number of frames containing the first keyword and the number of groups containing the first keyword.
14. In paragraph 13, The at least one instruction is individually or collectively executed by the at least one processor (1810) so that the electronic device (100) Determine the total number of keywords included in the first group of the above video, Determine a first ratio between the number of frames of the first group including the first keyword and the total number of keywords included in the first group, Determine a second ratio between the number of groups containing the first keyword and the total number of groups contained in the video, An electronic device that obtains importance distribution information corresponding to the first keyword based on the first ratio and the second ratio.
15. A computer-readable recording medium having recorded thereon a program for performing the method of any one of clauses 1 to 10 on a computer.
Citation Information
Patent Citations
Program for imparting keyword tag to scene of interest in motion picture contents, terminal, server, and method
JP2012155695A
Apparatus of extracting highlight and method of the same
KR1020090019582A
A manufacturing method of Paste containing Ginkgo nuts and Bellflower Root Extract
KR1020210051724A
narrow mask
KR1020230108623A
Touch detection module and display device including the same
KR1020240025095A