Electronic device providing emoticon generation function, or operation method thereof
Patent Information
- Application Number
- PCT/KR2025/000664
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-04
- Filing Date
- 2025-01-10
- Publication Date
- 2025-10-02
AI Technical Summary
Video editing requires significant time and effort, and creating personalized emoticons necessitates design skills and proficiency in editing tools, making it cumbersome for users to convey emotions effectively through emoticons.
An electronic device equipped with artificial intelligence models for scene understanding and image editing capabilities automatically detects objects in images, adds relevant content, and generates emoticons, enabling quick and convenient video editing and emoticon creation.
Users can quickly and easily create a variety of edited videos and emoticons, meeting their intended intent without requiring advanced editing skills.
Smart Images

Figure KR2025000664_02102025_PF_FP_ABST
Abstract
Description
Electronic device providing emoticon generation function or method of operating same
[0001] Various embodiments disclosed in this document relate to an electronic device providing an emoticon generation function or a method of operating the same.
[0002] Recently, vision systems that use artificial intelligence models to understand video scenes and identify specific objects have been developed and are being utilized in various fields.
[0003] For example, in the field of computer vision, scene understanding technology using artificial intelligence models can provide the ability to interpret various elements within an image, such as objects, environments, and situations. Scene understanding technology allows computers to infer events occurring within an image, and this inference can include object recognition, feature extraction, and situational awareness.
[0004] Furthermore, deep learning-based semantic segmentation can be considered, for example, in relation to AI models that identify objects within images. Semantic segmentation divides an image into multiple classes and assigns pixels to each class. By predicting which class each pixel belongs to, it can be used to segment object boundaries within the image. Methods utilizing convolutional neural networks (CNNs) are being extensively studied for deep learning-based semantic segmentation.
[0005] With the recent proliferation of services utilizing personal videos, video content creators are increasingly seeking to create more effective and engaging content through editing, rather than simply displaying footage captured with cameras or other devices. To achieve this, video content creators must select specific image frames from videos containing multiple image frames and then edit them one by one. Consequently, video editing requires significant time and effort.
[0006] Furthermore, as messenger application services become more active, users are increasingly seeking to convey their intentions and emotions more effectively through emoticons. Emoticons play a crucial role in communication, and users are increasingly seeking to create and use personalized emoticons. However, creating personalized emoticons requires design skills and proficiency in editing tools.
[0007] According to various embodiments of the present disclosure, a scene understanding technology can be used to automatically edit an image. For example, an electronic device can detect at least one object contained in a specific image frame within the image and add content (e.g., a visual object) related to the at least one object to obtain an edited image.
[0008] According to various embodiments of the present disclosure, emoticons can be automatically generated based on acquired images. For example, the acquired images can be edited and emoticons can be acquired based on the edited images.
[0009] According to various embodiments, an electronic device includes a communication device, a storage device storing at least one artificial intelligence model trained to generate an emoticon based on an input image, and at least one processor, wherein the at least one processor is configured to obtain a first image including a plurality of images from a user device connected to the electronic device through the communication device, input the first image to the at least one artificial intelligence model to obtain a first emoticon, and transmit the first emoticon through the communication device so that the user device outputs the first emoticon, and the at least one artificial intelligence model may be trained to obtain scene information for the first image, the scene information being information related to scene understanding of the plurality of images, obtain a second image obtained by editing the first image based on the scene information, and generate the first emoticon based on the second image.
[0010] According to various embodiments, a method of operating an electronic device may include an operation of obtaining a first image including a plurality of images from a user device connected to the electronic device, an operation of obtaining scene information about the first image through at least one artificial intelligence model, the scene information being information related to scene understanding of the plurality of images, an operation of obtaining a second image obtained by editing the first image based on the scene information through at least one artificial intelligence model, an operation of obtaining a first emoticon based on the second image through at least one artificial intelligence model, and an operation of transmitting the first emoticon so that the user device outputs the first emoticon.
[0011] According to various embodiments, an electronic device includes a display device, a storage device storing at least one artificial intelligence model trained to generate an edited image based on an input image, and at least one processor, wherein the at least one processor obtains a first image including a plurality of images, inputs the first image to the at least one artificial intelligence model to obtain a first emoticon, and outputs the first emoticon through the display device, and the at least one artificial intelligence model obtains scene information for the first image, the scene information being information related to scene understanding of the plurality of images, obtains a second image obtained by editing the first image based on the scene information, and may be trained to generate the first emoticon based on the second image.
[0012] According to various embodiments, a non-transitory computer-readable recording medium including a program for executing a method for controlling an electronic device providing a video editing function may include the steps of: obtaining a first image including a plurality of images; obtaining scene information about the first image through at least one artificial intelligence model; obtaining a second image obtained by editing the first image based on the scene information, the scene information being information related to scene understanding of the plurality of images through at least one artificial intelligence model; obtaining a first emoticon based on the second image through at least one artificial intelligence model; and transmitting the first emoticon so as to output the first emoticon.
[0013] According to various embodiments disclosed in this document, a scene understanding technology can be used to automatically edit an image. For example, an electronic device can detect at least one object within a specific image frame within the image and add content (e.g., a visual object) related to the at least one object to obtain an edited image. Therefore, a user can obtain an edited image more quickly and conveniently.
[0014] According to various embodiments of the present disclosure, a user can further edit a video automatically edited using scene understanding technology. Therefore, by performing additional video editing based on the initially edited video, the user can obtain a video edited more quickly and conveniently, while still meeting the user's intended intent.
[0015] According to various embodiments of the present disclosure, emoticons can be automatically generated based on acquired images. For example, the acquired images can be edited and emoticons can be generated based on the edited images. Therefore, users can create a wider variety of emoticons more easily and quickly.
[0016] In addition, various effects may be provided, either directly or indirectly, through this document.
[0017] FIG. 1 is a diagram illustrating a system that provides video editing functions according to various embodiments.
[0018] FIG. 2 is a block diagram of an electronic device according to various embodiments.
[0019] FIG. 3 illustrates a concept for controlling functions related to video editing in an electronic device according to various embodiments.
[0020] FIG. 4 is a flowchart illustrating an operation in which an electronic device provides information about a specific scene in a video according to a user request according to various embodiments.
[0021] FIG. 5 is a diagram illustrating providing information about a specific scene in a video using an artificial intelligence model according to various embodiments.
[0022] FIG. 6 is a flowchart illustrating an operation of an electronic device to obtain a summary video based on an input video according to various embodiments.
[0023] FIG. 7 is a diagram illustrating obtaining a summary video using an artificial intelligence model according to various embodiments.
[0024] FIG. 8 is a flowchart illustrating an operation of an electronic device providing an edited video according to various embodiments.
[0025] FIG. 9 is a diagram for explaining obtaining an edited video using an artificial intelligence model according to various embodiments.
[0026] FIG. 10 illustrates an execution screen related to video editing output through a user device according to various embodiments.
[0027] FIG. 11 is a flowchart illustrating an operation of an electronic device to obtain an emoticon based on an input video according to various embodiments.
[0028] FIG. 12 is a diagram illustrating obtaining an emoticon using an artificial intelligence model according to various embodiments.
[0029] FIG. 13 illustrates an execution screen for using emoticons through a messenger application according to various embodiments.
[0030] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0031] Specific structural or functional descriptions of various embodiments are merely illustrative for the purpose of explaining the various embodiments, and the various embodiments may be implemented in various forms and should not be construed as limited to the embodiments described in this specification or application.
[0032] Since various embodiments may have various modifications and take various forms, various embodiments are illustrated in the drawings and described in detail in this specification or application. However, the matters disclosed in the drawings are not intended to specify or limit the various embodiments, and should be understood to include all modifications, equivalents, and alternatives included within the spirit and technical scope of the various embodiments.
[0033] While terms such as "first" and / or "second" may be used to describe various components, these components should not be limited by these terms. These terms are only intended to distinguish one component from another; for example, without departing from the scope of the present disclosure, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component."
[0034] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between. Conversely, when a component is referred to as being "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Other expressions that describe the relationship between components, such as "between" and "directly between" or "adjacent to" and "directly adjacent to", should be interpreted similarly.
[0035] The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the various embodiments. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, it should be understood that the terms "comprises" or "has" specify the presence of a described feature, number, step, operation, component, part, or combination thereof, but do not exclude in advance the presence or possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0036] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this disclosure pertains. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0037] Hereinafter, the present disclosure will be described in detail by describing preferred embodiments of the present disclosure with reference to the attached drawings. The same reference numerals presented in each drawing represent the same components.
[0038] FIG. 1 is a diagram illustrating a video editing system that provides video editing functions according to various embodiments.
[0039] Referring to FIG. 1, a video editing system (100) may include a user device (102), a network (104), and an electronic device (106).
[0040] According to various embodiments, the user device (102) may be a device including a signage device, such as a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a television, a wearable device, or a head mounted device (HMD).
[0041] According to various embodiments, the user device (102) may include various output devices capable of providing video content to the user. For example, the user device (102) may include at least one of an audio device, a display device, or at least one camera capable of capturing video.
[0042] According to various embodiments, the user device (102) may include various input devices capable of obtaining input from a user. For example, it may include at least one of a keyboard, a touchpad, keys (e.g., buttons), a mouse, a microphone, and a digital pen (e.g., a stylus pen).
[0043] According to various embodiments, the network (104) may include any of a variety of wireless communication networks suitable for coupling to enable communication with the user device (102). For example, the network may include WLAN, WAN, PAN, cellular, WMN, WiMAX, GAN, 6LowPAN, and the like.
[0044] According to various embodiments, the electronic device (106) may include a standalone host computing system, an on-board computer system integrated with the user device (102), a mobile device, or any other hardware platform capable of providing video editing capabilities and video content (e.g., emoticons) to the user device (102). For example, the electronic device (106) may include a cloud-based computing architecture suitable for servicing video editing running on the user device (102). Accordingly, the electronic device (106) may include one or more servers (110) and data storage (108). For example, the electronic device (106) may include a Software as a Service (SaaS) architecture, a Platform as a Service (PaaS) architecture, an Infrastructure as a Service (IaaS) architecture, or other similar cloud-based computing architecture.
[0045] The term "emoticon" is a combination of the words "emotion" and "icon," meaning a representation of a feeling. The emoticons of the present disclosure may be understood as encompassing not only emojis, kaomojis, or symbols, but also animated emoticons and / or image emoticons. The emoticons of the present disclosure are not limited to the terms expressed and may include various digital contents displayed through a display.
[0046] According to various embodiments, the electronic device (106) and / or the user device (102) may be configured as one device performing each function, without being limited to the illustrated example.
[0047] For example, the electronic device (106) may perform functions of the user device (102), including components included in the user device (102). When the electronic device (106) provides functions of the user device (102), the electronic device (106) may provide an editing function for a video stored in the electronic device (106). For example, the electronic device (106) may acquire and store a video including a plurality of images, audio information, and / or subtitle information through a camera and a microphone. In addition, the electronic device (106) may edit the video based on a user input and output the edited video through a display device. In addition, the electronic device (106) may provide various functions related to the video (e.g., video summary, search for a specific scene within the video, and generation of emoticons based on the video).
[0048] The electronic device (106) provides various functions related to video according to various embodiments, which will be described below.
[0049]
[0050] Figure 2 is a block diagram of an electronic device (200) according to one embodiment.
[0051] Referring to FIG. 2, an electronic device (200) (e.g., the electronic device (106) of FIG. 1) may include a processor (210), a storage device (220) (e.g., the data storage (108) of FIG. 1), and / or a communication device (230). The components listed above may be operatively or electrically connected to one another. The components of the electronic device (200) illustrated in FIG. 2 may be modified, deleted, or added, for example.
[0052] According to various embodiments, the electronic device (200) may include a processor (210). The processor (210) may include hardware for executing instructions, such as instructions constituting a computer program. For example, to execute an instruction, the processor (210) may retrieve (or fetch) the instruction from an internal register, an internal cache, a storage device (220) (including a memory), decode and execute the instruction, and then store the result in the internal register, the internal cache, or the storage device (220).
[0053] In various embodiments, the processor (210) may execute software (e.g., a computer program) to control at least one other component (e.g., a hardware or software component) of the electronic device (200) connected to the processor (210) and perform various data processing or operations. According to various embodiments, as at least a part of the data processing or operations, the processor (210) may store commands or data received from another component (e.g., a communication device (230)) in volatile memory, process the commands or data stored in the volatile memory, and store result data in non-volatile memory.
[0054] According to various embodiments, the processor (210) may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), a micro controller unit (MCU), a sensor hub, a supplementary processor, a communication processor, an application processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a neural processing unit (NPU), and may have multiple cores.
[0055] According to various embodiments, the processor (210) (e.g., a neural network processing device) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the above-described examples. The artificial intelligence model may include a plurality of artificial neural network layers. The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, or a combination of two or more of the above, but is not limited to the above-described examples. In addition to the hardware structure, the artificial intelligence model may additionally or alternatively include a software structure.
[0056] According to various embodiments, the processor (210) may obtain a first image including a plurality of images from a user device (e.g., the user device (102) of FIG. 1) via a communication device (230). For example, the first image may include a plurality of images having a series of temporal flows and an image including audio information and / or subtitle information corresponding to the plurality of images.
[0057] According to various embodiments, the first image may include an image captured by at least one camera of the user device (102). For example, the user may capture an image using the user device (102) and transmit the image to the electronic device (200) for editing.
[0058] According to various embodiments, the processor (210) may input the first image into at least one artificial intelligence model to obtain scene information for a plurality of images. For example, the processor (210) may obtain scene information, which is information related to scene understanding for a plurality of images in the first image, by using a scene understanding model (e.g., a scene understanding model of the scene understanding module (303) described with reference to FIG. 3).
[0059] According to various embodiments, the processor (210) may obtain a second image by editing at least some of the plurality of images based on the scene information through at least one artificial intelligence model. For example, the processor (210) may add at least one content related to a first image among the plurality of images based on the scene information through at least one artificial intelligence model, and obtain a second image by editing the first image by adding the at least one content.
[0060] According to various embodiments, the processor (210) may transmit the second image so that the user device (102) outputs the second image. For example, the processor (210) may transmit the second image to the user device (102) via a communication device (230).
[0061] According to various embodiments, the user device (102) that has acquired the second image can output the second image through a display device (e.g., a display) included in the user device (102).
[0062] According to various embodiments, the storage device (220) may include mass storage for data or instructions. For example, the storage device (220) may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, an optical-magnetic disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more thereof.
[0063] In various embodiments, the storage device (220) may include nonvolatile, solid-state memory, read-only memory (ROM). Such ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more thereof.
[0064] Although this disclosure describes and illustrates a particular storage device, this disclosure contemplates any suitable storage device, and according to various embodiments, the storage device (220) may be internal or external to the electronic device (106).
[0065] According to various embodiments, the processor (210) may store a module related to video editing described with reference to FIG. 3 in the storage device (220).
[0066] According to various embodiments, the processor (210) may execute operations or data processing related to control and / or communication of at least one other component of the electronic device (200) using instructions stored in the storage device (220).
[0067] According to various embodiments, the electronic device (200) may include a storage device (220). According to various embodiments, the storage device (220) may store various data used by at least one component (e.g., processor (210)) of the electronic device (200). The data may include, for example, input data or output data for software (e.g., program) and commands related thereto.
[0068] According to various embodiments, the program may be stored as software in the storage device (220) and may include, for example, an operating system, middleware, or an application. According to various embodiments, the storage device (220) may store instructions that cause the processor (210) to process data or control components of the electronic device (200) to perform operations of the electronic device (200) when executed. The instructions may include code generated by a compiler or code that can be executed by an interpreter.
[0069] According to various embodiments, the storage device (220) may store various information obtained through the processor (210). For example, the storage device (220) may store at least one of a plurality of images obtained from the processor (210), an image including the plurality of images, scene information for the image, order information for each of the plurality of images, information for image groups that group the plurality of image frames, information for at least one object in each of the plurality of images, information for a point set group and a bounding box set for the at least one object output through an object recognition model, and user input information obtained from the user device (102). In addition, the storage device (220) may store identification information for the user device (102) connected to the electronic device (200).
[0070] According to various embodiments, the storage device (220) may store at least one AI model trained to provide various functions based on an input image. For example, the storage device (220) may store a model trained to output information related to a specific scene within the input image in response to a user request to search for a specific scene within the input image. Furthermore, for example, the storage device (220) may store a model trained to summarize the input image. Furthermore, the storage device (220) may store a model trained to generate an image by editing the input image (by adding at least one piece of content) based on scene-understanding of the input image. Furthermore, the storage device (220) may store a model trained to automatically generate an emoticon based on user input data (e.g., an image, text, sound, image, etc.). In this case, the emoticon may be provided in the form of a video. According to various embodiments, the electronic device (200) may store a model trained to provide various functions related to video processing in the storage device (220) without being limited to the models listed above. At this time, the model learned to provide the above-mentioned various functions may be composed of a single artificial intelligence model or may be composed of a combination of multiple models (e.g., an ensemble model).
[0071] According to various embodiments, the electronic device (200) may include a communication device (230). In various embodiments, the communication device (230) may support establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (200) and an external electronic device (e.g., the user device (102) of FIG. 1), and performance of communication through the established communication channel. The communication device (230) may operate independently from the processor (210) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication device (230) may include a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device via a first network (e.g., network (104) of FIG. 1) (e.g., a short-range communication network such as Bluetooth, WiFi Direct (wireless fidelity direct) or IrDA (infrared data association)) or a second network (e.g., network (104) of FIG. 1) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., network (104) of FIG. 1) (e.g., a LAN or WAN). The various types of communication modules may be integrated into one component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips).
[0072] According to various embodiments, the electronic device (200) can transmit and receive various data with various external devices through the communication device (230). In addition, the electronic device (200) can store the obtained data in the storage device (220). For example, the electronic device (200) can obtain an image including a plurality of images from the user device (102) through the communication device (230). In addition, for example, the electronic device (200) can obtain a user input indicating a user request for the image through the communication device (230). For example, the electronic device (200) can obtain a user input for searching for a specific scene of the image, a user input for generating a summary image of the image, a user input for generating an edited image of the image, and / or a user input for generating an emoticon based on the image through the communication device (230).
[0073] Additionally, for example, the electronic device (200) may transmit information about a specific scene, a summary video, an edited video, and / or an emoticon generated according to the execution of a function of the electronic device (200) to the user device (102) through the communication device (230). At this time, a request for modification of the information about the specific scene, the summary video, the edited video, and / or the emoticon may also be obtained from the user device (102) through the communication device (230).
[0074] According to various embodiments, the electronic device (200) may include a computer system. For example, the computer system may be at least one of an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC), a computer-on-module (COM), a system-on-module (SOM), a desktop computer system, a laptop or notebook computer system, a server, a tablet computer system, and a mobile terminal. For example, the electronic device (200) may include one or more computer systems residing in a cloud, which may include one or more cloud components.
[0075] According to various embodiments, the electronic device (200) may perform one or more operations of one or more methods described or presented in the present disclosure without substantial spatial or temporal limitations. Furthermore, the electronic device (200) may perform one or more operations of one or more methods described or presented in the present disclosure in real time or in batch mode. For example, the electronic device (200) may perform one or more operations of one or more methods described or presented in the present disclosure at different times or locations.
[0076] According to various embodiments, when the electronic device (200) described in the present disclosure provides the functions of the user device (102), the electronic device (200) may include at least one camera (not shown) and / or a display device (not shown). Hereinafter, a case in which the electronic device (200) includes at least one camera (not shown) and / or a display device (not shown) will be described.
[0077] According to various embodiments, the processor (210) may acquire a first image including a plurality of images through at least one camera. For example, the processor (210) may activate the at least one camera based on a shooting start command, and acquire the first image including a plurality of image frames, audio information, and / or subtitle information through the at least one camera.
[0078] According to various embodiments, the processor (210) may obtain scene information for a plurality of images included in the first image using at least one artificial intelligence model. For example, the processor (210) may input the first image into at least one artificial intelligence model trained to perform a scene understanding function, thereby obtaining information related to scene understanding for each of the plurality of images, and may obtain a second image by adding at least one content to the first image based on the scene information.
[0079] According to various embodiments, the processor (210) may control the display device to output the second image. For example, the processor (210) may output the second image through the display device included in the electronic device (200).
[0080] According to various embodiments, the processor (210) may obtain user input for various information (e.g., specific scene information, summary video, edited video, emoticon) regarding an image obtained through at least one artificial intelligence model. For example, the processor (210) may obtain user input through an input device (e.g., keyboard, touch pad, key (e.g., button), mouse, microphone, digital pen (e.g., stylus pen)) included in the electronic device (200).
[0081] According to various embodiments, the processor (210) may modify and provide the various information based on the user input. According to one embodiment, the processor (210) may obtain a user input for a second image, which is an edited image of the first image obtained based on scene understanding through at least one artificial intelligence model. For example, the processor (210) may obtain a user input to add at least one additional content to the second image or to delete added content. In this case, the processor (210) may obtain a third image obtained by editing the first image based on the user input.
[0082]
[0083] FIG. 3 illustrates a concept for controlling functions related to video editing in an electronic device according to various embodiments.
[0084] Referring to FIG. 3, the electronic device (200) may utilize hardware and / or software modules (300) to support various video editing-related functions. For example, the processor (210) may drive an image acquisition module (301), a scene understanding module (303), an object recognition module (305), an image search module (307), an image summary module (309), an image editing module (311), an emoticon generation module (313), an image generation module (315), and / or an evaluation module (317) by executing commands stored in a storage device (220). In various embodiments, software modules other than those illustrated in FIG. 3 may be implemented. For example, at least two modules may be integrated into one module, or one module may be split into two or more modules. In addition, work performance may be improved by having hardware and software modules share a single function. For example, the electronic device (200) may include both an encoder implemented as hardware and an encoder implemented as a software module, and some of the data acquired through at least one camera module may be processed by the hardware encoder and the remaining part may be processed by the software encoder.
[0085] According to various embodiments, the image acquisition module (301) may provide a user interface (UI) / graphical UI (GUI) related to video upload to a user through a user device (102), and may acquire an image (e.g., a first image described with reference to FIG. 2) through the user device. For example, the image may be acquired by controlling a function related to video upload in response to a user input provided through a UI / GUI output through a display device of the user device (102). In addition, the image acquisition module (301) may extract a plurality of image frames, audio information, and / or subtitle information of the image acquired through the user device (102).
[0086] According to various embodiments, the scene understanding module (303) can obtain scene information for the acquired image. The scene understanding module (303) can be a model trained to understand the scene for the input image and multiple images included in the image through various learning data.
[0087] According to various embodiments, the scene understanding module (303) may be configured as an artificial neural network model. For example, the scene understanding module (303) may be a deep neural network model trained to recognize an object included in an input image and extract information about the object, extract features from each part of the image included in the image, extract spatial information in terms of spatial aspects such as the arrangement, relative position, and depth of objects within the image, and / or generate scene information by understanding the interaction and situation of objects within the image and understanding the overall scene. For example, the scene understanding module (303) may be trained using a neural network structure such as a Transformer architecture or a Convolution Neural Network (CNN). However, the scene understanding module (303) of the present disclosure is not limited to the aforementioned neural network model and may be implemented with any other appropriate neural network model.
[0088] According to various embodiments, the scene understanding module (303) can understand and interpret the environment of the acquired image to generate scene information. For example, the scene understanding module (303) can understand and infer the image and the images contained in the image to obtain scene information through object recognition, feature extraction, situation recognition, etc.
[0089] According to one embodiment, the scene understanding module (303) may group a plurality of images included in an input image based on the scene information to create image groups. For example, the scene understanding module (303) may group a plurality of images included in an image acquired through the image acquisition module (301) into preset standard units based on the scene information to create image groups. For example, the scene understanding module (303) may recognize a scene transition of the image based on the scene information, and group a plurality of images of the image based on the scene transition to create image groups.
[0090] According to various embodiments, the object recognition module (object detection model) (305) may include an object recognition model trained to detect a point set and a bounding box of at least one object included in an image. For example, the object recognition model may be a model trained to detect a point set and a bounding box set of at least one object through various training data.
[0091] According to various embodiments, the training data for learning the object recognition model may include training data that distinguishes at least one object and a background within a plurality of image frames included in an image, and assigns a label corresponding to the background and a label corresponding to each of at least one object.
[0092] According to various embodiments, the object recognition module (305) may be configured as an artificial neural network model. For example, the object recognition module (305) may be a deep neural network model trained to identify and track objects in an image and extract a set of key points and a set of bounding boxes for at least one object in the image. For example, the object recognition model may be implemented as a Region-based Convolution Neural Network (R-CNN), a Faster Region-based Convolution Neural Network (Faster R-CNN), a Single Shot Multibox Detector (SSD), YOLO v4, CenterNet, or MobileNet. However, the object recognition model of the present disclosure is not limited to the aforementioned deep neural network model and may be implemented as any other appropriate neural network model.
[0093] According to various embodiments, the object recognition module (305) may obtain a point set by extracting skeleton data of at least one object included in an image. For example, the object recognition module (305) may extract at least one object included in an image and extract skeleton data of the extracted object to obtain a point set. For example, if the object is a human object or an animal, the joint part or specific part of the object may be detected. In addition, if the object is a human object, body parts such as the head, eyes, nose, mouth, ears, neck, shoulders, elbows, wrists, fingertips, torso, hip joints, wrists, knees, ankles, and toes may be extracted. The skeleton data may be expressed as XY coordinate values as coordinates in the image to constitute a point set.
[0094] According to various embodiments, the object recognition module (305) may utilize a joint detection algorithm, such as a Kinetics dataset or a NTU-RGB-D (Nanyang Technological University's Red Blue Green and Depth information) dataset, to extract skeletal data of joints or specific parts of the body to obtain a point set. In this case, the number of skeletal joints per at least one object may be arbitrarily defined.
[0095] According to various embodiments, the object recognition module (305) can track at least one object in an image. For example, it can track at least one object recognized from multiple images included in the image, and track changes in at least one object within multiple image frames. For example, when multiple objects are included in an image, the object recognition module (305) can extract skeleton data for each of the multiple objects to obtain a point set, and create and track a layer for each of the multiple objects. According to various embodiments, when multiple objects included in an image are recognized for a specific time interval, a layer can be created for each time interval in which each object is recognized.
[0096] According to various embodiments, the object recognition module (305) of the electronic device (200) may include an encoder and a decoder for extracting a set of bounding boxes from a plurality of image frames within an image.
[0097] According to various embodiments, the encoder and decoder may be connected as a network having a nested U-shaped structure. This nested U-shaped structure can effectively extract and combine multi-scale features of intra-stages. The encoder may extract and compress features from image frames of a video to generate context information. The decoder may be configured to expand a feature map including the context information and output a set of bounding boxes based on segmentation.
[0098] According to various embodiments, the object recognition module (305) can identify the location and outline of at least one object with higher accuracy by utilizing both the point set and the bounding box for at least one object in each of the plurality of image frames of the video using the object recognition model.
[0099] According to various embodiments, the scene understanding module (303) and the object recognition module (305) may be configured as a single module. For example, since the scene information provided by the scene understanding module (303) includes identification information about an object, the function of the object recognition module (305) described above may be performed by the scene understanding module (303).
[0100] According to various embodiments, the image search module (307) can search for a specific scene in an image acquired through the image acquisition module (301) and generate information about the specific scene. For example, the image search module (307) can provide the acquired image to the user and provide the user with a UI / GUI for the user input of the image to obtain a user input for searching for a specific scene in the image. According to various embodiments, the image search module (307) can understand a scene of the image through the scene understanding module (303) based on the user input, extract a scene (or image, section image) corresponding to the specific scene, and output information about the specific scene.
[0101] According to one embodiment, the image search module (307) may be implemented with at least one artificial intelligence model and may perform a function of outputting information about a specific scene. For example, the image search module (307) may include a first artificial intelligence model trained to identify the user's intention to search for a scene within a video based on a user input and output search information. In addition, for example, the image search module (307) may include a second artificial intelligence model trained to receive the search information and the video and output information about the specific scene the user wants to search for.
[0102] According to various embodiments, the image summary module (309) can generate information about a summary image that summarizes the image acquired through the image acquisition module (301). For example, the image summary module (309) can provide the acquired image to the user and provide the user with a UI / GUI for user input of the image to obtain a user input including a summary request and / or summary criteria for the image. According to various embodiments, the image summary module (309) can understand a scene of the image through the scene understanding module (303) based on the user input and output information about the summary image based on the user input.
[0103] According to one embodiment, the video summary module (309) may be implemented with at least one artificial intelligence model and may perform a function of outputting a summary video for an input video. For example, the video summary module (309) may automatically output a summary video for an input video, or may output a summary video according to a summary criterion by identifying the user's intention based on a user input. According to one embodiment, the video summary module (309) may include a first artificial intelligence model trained to identify the user's intention and output reference information for a video summary in order to perform the above function. In addition, for example, the video summary module (309) may include a second artificial intelligence model trained to input the reference information and the video and output a summary video that summarizes the video according to the user's intention. According to various embodiments, without being limited to the above-described example, the first artificial intelligence model and the second artificial intelligence model of the video summary module (309) may be configured as a single model.
[0104] According to various embodiments, the image editing module (311) can generate an edited image by editing the input image by adding at least one content to the image acquired through the image acquisition module (301). For example, the image editing module (311) can acquire scene information about the image acquired through the image acquisition module (301) through the scene understanding module (303).
[0105] According to one embodiment, the image editing module (311) may generate at least one content to be added to the image based on scene information acquired through the scene understanding module (303). For example, the image editing module (311) may generate at least one content to be added to the image using at least one artificial intelligence model. In one embodiment, the at least one content may be generated based on scene information about the input image and a plurality of images included in the image. In this case, the at least one content may be generated based on scene information about a scene (image) to which the at least one content is to be added.
[0106] According to various embodiments, the image editing module (311) may generate at least one content and add the at least one content to the input image to generate an edited image. For example, the image editing module (311) may generate at least one content in relation to a first image in an image input through the image acquisition module (301), and generate a second image by adding the at least one content to the first image. In addition, the image editing module (311) may generate an image by editing the input image based on the second image. In this case, the image editing module (311) may generate the image using the image generation module (315).
[0107] According to one embodiment, the image editing module (311) may obtain a point set and a bounding box set for at least one object included in the image obtained through the image acquisition module (301) using the object recognition module (305). In addition, the image editing module (311) may identify an outline of the at least one object based on the point set and the bounding box set, and obtain object information for the at least one object based on the outline of the at least one object. For example, the image editing module (311) may obtain first scene information for a first image included in the image obtained through the image acquisition module (301) using the scene understanding module (303), and obtain a point set and a bounding box set for at least one object included in the first image using the object recognition module (305). In addition, object information for the at least one object may be obtained based on the first scene information and the outline of the at least one object.
[0108] According to various embodiments, the image editing module (311) may determine a location to which at least one content is to be added in relation to the first image based on the object information. For example, when the image editing module (311) adds at least one content to a first image included in the first image, the location to which the at least one content is to be added may be determined based on a point set and a bounding box set acquired through the object recognition module (305) and first scene information acquired through the scene understanding module (303). The image editing module (311) may generate a second image by adding the at least one content to the first image based on the location.
[0109] According to various embodiments, the image editing module (311) may obtain scene information for an input image and object information for at least one object, and determine the type and location of at least one content to be added to the image based on the scene information and the object information. According to one embodiment, when the type of the at least one content is a visual object, the image editing module (311) may determine an image in the image on which the visual object is to be displayed based on the scene information and the object information, and may determine in which area of the image the visual object is to be displayed.
[0110] According to various embodiments, a user can obtain an edited video more conveniently and quickly through an electronic device (200). Furthermore, by obtaining an edited video with at least one piece of content added to a more precise location through the electronic device, the user can obtain a video that is similar to one edited directly by the user.
[0111] According to various embodiments, the image editing module (311) may provide the acquired image to the user and obtain user input requesting editing of the image by providing the user with a UI / GUI for user input of the image. For example, the image editing module (311) may initially provide the user with a completed edited image. At this time, the image editing module (311) may generate an additionally edited image of the edited image based on the user input for the edited image. In one embodiment, the image editing module (311) may also generate an additionally processed additionally edited image while reflecting the user input based on the scene information and the object information when generating the additionally edited image.
[0112] According to one embodiment, the image editing module (311) may be implemented with at least one artificial intelligence model and may perform a function of outputting an edited image for an input image. For example, the image editing module (311) may generate at least one content to be included in the first image through a first artificial intelligence model based on the input image. In addition, the image editing module (311) may output an edited image to which at least one content has been added based on scene information and object information for the input image. According to various embodiments, the function provided by the image editing module (311) may be provided through one artificial intelligence model or may be provided through multiple artificial intelligence models.
[0113] According to various embodiments, the emoticon generation module (313) may generate an emoticon that can be used in a messenger application, etc., based on an input image. According to various embodiments, the emoticon generation module (313) may obtain user input by providing a UI / GUI for user input related to generating an emoticon. According to one embodiment, the emoticon generation module (313) may obtain a user input requesting the generation of an emoticon, and generate an emoticon through at least one artificial intelligence model based on the user input. The at least one artificial intelligence model may be a model trained to identify a user's intention from a user input and output an emoticon corresponding to the user's intention.
[0114] According to one embodiment, the emoticon generation module (313) can obtain an image to be used for generating an emoticon through user input. In one embodiment, the emoticon generation module (313) can obtain scene information about the input image through the scene understanding module (303) and object information through the object recognition module (305). In one embodiment, the emoticon generation module (313) can generate an emoticon using at least one artificial intelligence model through the input image, the scene information, and / or the object information.
[0115] For example, the emoticon generation module (313) may obtain an image as a user input requesting the generation of an emoticon, and obtain a summary image to be converted into an emoticon through the image summary module (309) based on the scene information and the object information. The emoticon generation module (313) may generate an emoticon that can be used in a messenger application based on the summary image.
[0116] For example, the emoticon generation module (313) may obtain an image as a user input requesting the generation of an emoticon, and obtain an edited image to be converted into an emoticon through the image editing module (311) based on the scene information and the object information. The emoticon generation module (313) may generate an emoticon that can be used in a messenger application based on the edited image.
[0117] According to various embodiments, the emoticon generation module (313) can generate an emoticon through the functions of the scene understanding module (303), object recognition module (305), image summary module (309) and / or image editing module (311) described above.
[0118] According to various embodiments, the image generation module (315) may generate an image using information output through the image search module (307), the image summary module (309), and / or the image editing module (311). For example, when a specific scene is searched for in a first image through the image search module (307), the image generation module (315) may generate an image including the specific scene based on information about the specific scene. In addition, for example, when information about a summary image is output through the image summary module (309), the image generation module (315) may generate a summary image. In addition, when editing information about an input image is output through the image editing module (311), the image generation module (315) may generate an edited image.
[0119] According to various embodiments, the image generation module (313) may obtain a corrected image that encodes various information based on the input image based on data provided from the image search module (307), the image summary module (309), the image editing module (311), and / or the emoticon generation module (313).
[0120] According to various embodiments, the image generated by the image generation module (315) may be transmitted to the user so that it can be recognized. For example, the electronic device (200) may transmit the image to the user device (102) via the communication device (230).
[0121] According to various embodiments, the evaluation module (317) may provide a UI / GUI related to feedback on the output results (e.g., searched images, summarized images, edited images, emoticons) from the image search module (307), the image summary module (309), the image editing module (311), and / or the emoticon generation module (313) to the user through the user device (102), and obtain user feedback through the user device. Therefore, according to various embodiments, the evaluation module (317) may obtain feedback information indicating user satisfaction with the output results, and may retrain at least one of the scene understanding module (303), the object recognition module (305), the image search module (307), the image summary module (309), the image editing module (311), and / or the emoticon generation module (313) using the feedback information.
[0122] According to various embodiments, the evaluation module (317) may use the acquired feedback information to control at least one of the scene understanding module (303), the object recognition module (305), the image search module (307), the image summary module (309), the image editing module (311), and / or the emoticon generation module (313) to be customized for the user. For example, the image editing module (311) may be used to acquire feedback information indicating the user's satisfaction with the edited image obtained by editing the acquired image, and the feedback information may be used to retrain the artificial intelligence model included in the image editing module (311) to provide a user-customized image editing model. Alternatively, the evaluation module (317) may be controlled to store user input information related to the generation of emoticons through the emoticon generation module (313), and generate emoticons (e.g., using a macro method or history information) based on the stored input information without the user having to repeatedly input the information through the user device again.
[0123] In the embodiment of FIG. 3, the functions performed by the image acquisition module (301), the scene understanding module (303), the object recognition module (305), the image search module (307), the image summary module (309), the image editing module (311), the emoticon generation module (313), the image generation module (315) and / or the evaluation module (317) can be understood as being performed by the processor (210) executing commands stored in the storage device (220). In addition, in various embodiments, the electronic device (200) can utilize one or more hardware processing circuits to perform various functions and operations disclosed in this document.
[0124] Additionally, the hardware / software connection relationship illustrated in FIG. 3 is for convenience of explanation and does not limit the flow / direction of data or commands. Components included in the electronic device (200) may have various electrical / operational connection relationships.
[0125]
[0126] FIG. 4 is a flowchart (400) illustrating an operation of an electronic device to generate an image from which an area excluding at least one object is removed according to various embodiments.
[0127] FIG. 5 is a diagram illustrating providing information about a specific scene in a video using an artificial intelligence model according to various embodiments.
[0128] The operations described in this document below may be performed in combination with each other. Furthermore, the operations described in this document are not limited to the order shown and may be performed in various orders, or more or fewer operations may be performed. Among the operations described below, operations performed by an electronic device (200) (e.g., electronic device (106) of FIG. 1) may refer to operations performed by the processor (210) of the electronic device (200).
[0129] In addition, the “information” described below may be interpreted to mean “data” or “signal,” and “data” may be understood as a concept including both analog data and digital data. Descriptions that overlap or are similar to the above description among the operations of the electronic device (200) according to various embodiments may be omitted.
[0130] Referring to FIG. 4, the electronic device (200) may obtain a user input for requesting a search for a first image and a specific scene within the first image in operation 401. For example, the electronic device (200) may obtain a user input requesting a search for a specific scene within the first image through the user device (102). The user input may be obtained in various forms, such as video, text, sound, image, motion (e.g., gesture), etc.
[0131] According to various embodiments, the electronic device (200) may obtain search information by inputting user input into at least one artificial intelligence model in operation 403. For example, referring to the first flow (510) of FIG. 5, the electronic device (200) may obtain search information (515) by inputting user input (511) into the first artificial intelligence model (513). According to various embodiments, the first artificial intelligence model (513) may be a model trained to identify user intent based on user input.
[0132] According to various embodiments, the electronic device (200) may, in operation 405, input search information and a first image into at least one artificial intelligence model to obtain information about a specific scene. For example, referring to the second flow (520) of FIG. 5, the electronic device (200) may input the search information (521) (e.g., search information (515)) and the first image (522) into a second artificial intelligence model (523) to obtain information (525) about a specific scene. According to one embodiment, the information (525) about the specific scene may be output in various forms. For example, the information about the specific scene may include an image order (Num) in the first image, an image (Image) about the specific scene, and / or a specific scene (Scene). In various embodiments, the electronic device (200) may also generate an image related to the specific scene through the image search module (307) and the image generation module (315) described with reference to FIG. 3.
[0133] According to various embodiments, the electronic device (200) may acquire information (525) about a specific scene by executing the functions of the image search module (307) and the image generation module (315) described with reference to FIG. 3. Accordingly, the operations of the electronic device (200) related to searching for a specific scene described with reference to FIG. 3 may be applied in the same or similar manner.
[0134] According to various embodiments, the electronic device (200) may transmit information (525) about the specific scene to the user device (102). For example, the electronic device (200) may transmit information about the specific scene to the user device (102) via the communication device (230).
[0135]
[0136] FIG. 6 is a flowchart (600) illustrating an operation of an electronic device to obtain a summary video based on an input video according to various embodiments.
[0137] FIG. 7 is a diagram illustrating obtaining a summary video using an artificial intelligence model according to various embodiments.
[0138] Referring to FIG. 6, the electronic device (200) may obtain a user input requesting a first image and a summary image for the first image in operation 601. For example, the electronic device (200) may obtain a user input requesting a summary for the first image through the user device (102). The user input may include a video that is the subject of the summary. In addition, the user input may include additional requests that serve as a basis for the summary or are related to the summary. The additional requests may be obtained in various forms, such as text, sound, image, and motion (e.g., gesture).
[0139] According to various embodiments, the electronic device (200) can obtain image groups by grouping a plurality of images included in the first image into reference units in operation 603.
[0140] According to various embodiments, the electronic device (200) may obtain a summary image for the first image based on the image groups in operation 605.
[0141] Referring to FIG. 7, an electronic device (200) according to various embodiments may obtain a summary image (705) for a first image (701) through at least one artificial intelligence model (703). For example, the electronic device (200) may obtain image groups that group a plurality of images included in the first image (701) into standard units based on at least one artificial intelligence model (703).
[0142] According to one embodiment, the electronic device (200) can recognize a scene transition of the first image through the scene understanding module (303) and / or the object recognition module (305) described with reference to FIG. 3. For example, the electronic device (200) can obtain scene information about the first image through the scene understanding module (303), and group a plurality of images included in the first image based on the scene information to create image groups. For example, the electronic device (200) can recognize a scene transition of the first image based on the scene information, and group a plurality of images of the image based on the scene transition to create image groups.
[0143] According to various embodiments, the electronic device (200) may also, for example, identify at least one object included in each of a plurality of images of the first video through a scene understanding module (303) and / or an object recognition module (305), and recognize a scene transition based on at least one of a change in the at least one object, a change in the type of the at least one object, a change in the number of the at least one object, a change in the main color value of each of the plurality of images, audio information of the first video, subtitle information of the first video, order information of the plurality of images, shooting time information of each of the plurality of images, and a user input for image grouping. In addition, the electronic device (200) may group the plurality of images of the first video based on the scene transition to generate the image groups.
[0144] According to various embodiments, the electronic device (200) may obtain a summary image for the first image based on the image groups. For example, the electronic device (200) may obtain at least one group among the image groups generated for the first image as a summary image, and / or may select at least one image from each of the image groups and obtain the summary image.
[0145] However, without being limited to the above-described examples, according to various embodiments, the electronic device (200) may obtain a summary image (705) for the first image (701) using at least one artificial intelligence model (703). For example, even without performing a grouping operation on images included in the first image, the electronic device (200) may obtain a summary image (705) based on the first image (701) using at least one artificial intelligence model (703) trained to output a summary image. In this case, the electronic device (200) may obtain the summary image (705) based on scene information and / or object information obtained from the scene understanding module (303) and / or the object recognition module (305). In one embodiment, when an additional request related to image summary is obtained through user input, the electronic device (200) may interpret the additional user request using at least one artificial intelligence model (703) and generate a summary image (705) using scene information and / or object information based on the additional request.
[0146] In various embodiments, techniques related to natural language interpretation, which are well known to those skilled in the art in interpreting user input (text, sound, image, motion (e.g., gesture)), may be applied in various ways. Therefore, a description thereof may be omitted.
[0147] According to various embodiments, the electronic device (200) may obtain a summary image (705) (or information about the summary image) by executing the functions of the image summary module (309) and the image generation module (315) described with reference to FIG. 3. Accordingly, the operations of the electronic device (200) related to generating a summary image described with reference to FIG. 3 may be applied in the same or similar manner.
[0148] According to various embodiments, the electronic device (200) may transmit the summary image (705) to the user device (102). For example, the electronic device (200) may transmit the summary image (705) (or information about the summary image) to the user device (102) via the communication device (230).
[0149]
[0150] FIG. 8 is a flowchart (800) illustrating an operation of an electronic device providing an edited video according to various embodiments.
[0151] FIG. 9 is a diagram for explaining obtaining an edited video using an artificial intelligence model according to various embodiments.
[0152] FIG. 10 illustrates an execution screen (1000) related to video editing output through a user device according to various embodiments.
[0153] Referring to FIG. 8, in operation 801, the electronic device (200) can obtain a first image including a plurality of images.
[0154] According to various embodiments, the electronic device (200) may, in operation 803, input a first image into at least one artificial intelligence model to obtain scene information for a plurality of images.
[0155] According to various embodiments, the electronic device (200) may, in operation 805, obtain a second image having at least one content added in relation to a first image among a plurality of images based on scene information through at least one artificial intelligence model.
[0156] Referring to FIG. 9, for example, the electronic device (200) may input a first image (901) into at least one artificial intelligence model (903) to obtain a second image (905). According to various embodiments, the at least one artificial intelligence model (903) may include an artificial intelligence model of a scene understanding module (303), an object recognition module (305), and / or an image editing module (311) described with reference to FIG. 3. Therefore, redundant or similar descriptions may be omitted.
[0157] According to various embodiments, the electronic device (200) may obtain scene information and object information for the first image (901) through at least one artificial intelligence model (903) (e.g., the scene understanding module (303), the object recognition module (305), and the image editing module (311) of FIG. 3). In addition, according to one embodiment, the electronic device (200) may generate at least one content to be added to the first image (901) through the at least one artificial intelligence model (903). In one embodiment, the at least one content may be generated based on scene information for the first image (901) and a plurality of images included in the first image (901). In this case, the at least one content may be generated based on scene information for a scene (image (e.g., the first image) to which the at least one content is to be added.
[0158] According to various embodiments, the electronic device (200) may generate a second image (905) by adding at least one content to the first image (901) through at least one artificial intelligence model (903). For example, the electronic device (200) may generate at least one content in relation to a first image in the input first image (901) through at least one artificial intelligence model (903), and generate a second image by adding the at least one content to the first image. In addition, the electronic device (200) may generate a second image (905) by editing the input image based on the second image.
[0159] According to various embodiments, the electronic device (200) may add at least one content to a first image (901) to obtain a second image (905), and may add at least one content based on a point set and a bounding box set for at least one object of the first image (901). For example, the electronic device (200) may obtain first scene information for the first image through at least one artificial intelligence model (903), and may obtain object information for at least one object by obtaining a point set and a bounding box set for at least one object included in the first image. In addition, the electronic device (200) may determine a location to which at least one content is to be added in relation to the first image based on the first scene information and the object information through at least one artificial intelligence model (903), and may generate a second image in which the at least one content is added to the first image based on the location. Accordingly, a second image (905) that is an edited version of the first image (901) can be obtained based on the second image.
[0160] According to various embodiments, the electronic device (200) may transmit the second image (905) to the user device (102) to output the second image (905). For example, the electronic device (200) may transmit the second image (905) (or information about the summary image) to the user device (102) via the communication device (230).
[0161] Referring to FIG. 10, an execution screen (1000) related to video editing output through a user device is illustrated.
[0162] According to various embodiments, the user device (102) may display an execution screen related to a video editing service based on a control signal obtained through the electronic device (200). The execution screen (1000) of the video editing service displayed through the user device may include a first area (1010), a second area (1020), and a third area (1030).
[0163] In this document, icons may be replaced with representations of buttons, menus, objects, etc. In addition, the visual objects depicted in the first area (1010), the second area (1020), and / or the third area (1030) in FIG. 10 are exemplary, and other visual objects may be placed, replaced with other icons, or omitted.
[0164] According to various embodiments, the user device (102) may display an execution screen (1000) based on a control signal of the electronic device (200). According to various embodiments, the execution screen (1000) may include a first area (1010) in which a second image (905) is output by editing the first image (901) described with reference to FIGS. 8 and 9 using at least one artificial intelligence model (903). The electronic device (200) may output a screen for the edited second image (905) to the first area (1010) through the user device (102).
[0165] According to various embodiments, the first area (1010) may display a plurality of images included in the edited second image (905). At this time, when at least one image among the plurality of images in the edited second image (905) is displayed in the first area (1010), at least one content added to the first image (901) may be displayed through the second area (1020). According to various embodiments, the second area (1020) where at least one content is displayed in the execution screen (1000) may include an area corresponding to a location where the at least one content is to be displayed.
[0166] Referring to FIG. 10, the user device (102) may receive a control signal from the electronic device (200) and display an image edited as a second image by adding at least one content among a plurality of images in the first image (901) in the first area (1010). For example, the electronic device (200) may understand that the first image in the first image (901) is of a child making a surprised expression through scene understanding using at least one artificial intelligence model (903), and may identify the location of the child through object recognition using at least one artificial intelligence model (903). In addition, the electronic device (200) may generate content indicating an exclamation mark for the first image through at least one artificial intelligence model (903), and determine a location at which the exclamation mark is to be displayed. Accordingly, the electronic device (200) can generate a second image by adding an exclamation mark to the first image in the first image (901) in the second area (1020), and generate a second image (905) including the second image. At this time, the second image, which is an image edited with added content, can be displayed through the first area (1010).
[0167] According to various embodiments, the electronic device (200) may obtain a user input requesting correction of at least one content within a video through the first area (1010), the second area (1020), and / or the third area (1030) of the execution screen (1000) displayed through the user device (102). For example, the electronic device (200) may obtain a user input requesting correction of additional content (e.g., an exclamation mark) displayed for a second image through the user device (102). According to various embodiments, the constant user input may be obtained through a mouse click, a touch input, a sound input, a character input, a keyboard input, etc.
[0168] According to various embodiments, the electronic device (200) may correct (position, shape size, type, etc.) at least one content (e.g., an exclamation mark) based on a user input obtained through the user device (102), and output a third image generated based on the corrected at least one content.
[0169] According to various embodiments, the user device (102) may display information about a plurality of images and image groups included in the second image (905) through the third area (1030) based on a control signal of the electronic device (200). At this time, an image (e.g., the second image) (1031) edited by adding content to at least one of the plurality of images may be displayed through a distinct visual object.
[0170] According to the present disclosure, for convenience of explanation, an example of obtaining a second image (905) by adding at least one content to a first image in a first image (901) has been described, but the present disclosure is not limited thereto and at least one content may be added to at least one image in the first image (901).
[0171] According to various embodiments, the electronic device (200) may obtain a second image (905) (or information about the edited image) as an edited image by executing the functions of the scene understanding module (303), the object recognition module (305), the image editing module (311), and / or the image generation module (315) described with reference to FIG. 3. Accordingly, the operations of the electronic device (200) related to generating an edited image (e.g., the second image (905)) described with reference to FIG. 3 may be applied in the same or similar manner.
[0172]
[0173] FIG. 11 is a flowchart (1100) showing an operation of an electronic device to obtain an emoticon based on an input video according to various embodiments.
[0174] FIG. 12 is a diagram illustrating obtaining an emoticon using an artificial intelligence model according to various embodiments.
[0175] FIG. 13 illustrates an execution screen (1300) for using emoticons through a messenger application according to various embodiments.
[0176] Referring to FIG. 11, the electronic device (200) can obtain a basic image used to generate an emoticon based on user input data through at least one artificial intelligence model in operation 1101.
[0177] Referring to FIG. 12, the electronic device (200) may obtain user input data (1201) based on a user request to generate an emoticon. According to various embodiments, the electronic device (200) may use at least one artificial intelligence model (1203) to identify the user's intention from the user input data (1201) and output an emoticon (1205) corresponding to the user's intention. According to various embodiments, the at least one artificial intelligence model (1203) (e.g., the scene understanding module (303), the object recognition module (305), the emoticon generation module (313) and / or the image generation module (315) of FIG. 3) may include a model trained to output an emoticon (1205) based on the user input data (1201).
[0178] According to various embodiments, the electronic device (200) may obtain an emoticon based on a basic image through at least one artificial intelligence model (1203) in operation 1103. For example, the electronic device (200) may obtain a basic image to be used for generating an emoticon through a user input. In one embodiment, the electronic device (200) may obtain scene information and object information for the basic image input through at least one artificial intelligence model (1203). In one embodiment, the electronic device (200) may generate an emoticon (1205) through the basic image input through at least one artificial intelligence model (1203), the scene information, and / or the object information.
[0179] According to various embodiments, when generating an emoticon (1205) using at least one artificial intelligence model (1203), the electronic device (200) may include at least one image (1205_b) included in a basic image in the emoticon (1205). In addition, when generating an emoticon (1205) using at least one artificial intelligence model (1203), the electronic device (200) may include at least one content (1205_a) in the emoticon (1205) based on user input data (1201), scene information for the basic image, and object information for the basic image.
[0180] According to various embodiments, the electronic device (200) may transmit the emoticon (1205) to be used in a messenger application at operation 1105. For example, the electronic device (200) may transmit the emoticon (1205) to the user device (102) so that the user can use it as the messenger application stored in the user device (102) is executed.
[0181] Referring to FIG. 13, an execution screen (1300) of a messenger application executed through a user device (102) is illustrated. According to various embodiments, the user device (102) can use at least one emoticon (1320) acquired through an electronic device (200) in the messenger application. For example, the user device (102) can output a conversation content (1310) resulting from executing the messenger application in a portion of the execution screen (1300) and output at least one emoticon (1320) in a portion of the execution screen (1300). Accordingly, the user can use various emoticons generated through the electronic device (200) in the messenger application to express his or her intentions and emotions in various ways. In addition, emoticons can be generated more easily and simply and used in the messenger application.
[0182] For convenience of explanation, it is described that emoticons created through a messenger application are used, but this is not limited to this, and the user device (102) can use at least one emoticon created through the electronic device (200) in executing various applications.
[0183] According to various embodiments, the electronic device (200) may obtain an emoticon (1205) (or information about the emoticon) by executing the functions of the scene understanding module (303), the object recognition module (305), the emoticon generation module (313), and / or the image generation module (315) described with reference to FIG. 3. Accordingly, the operations of the electronic device (200) related to generating the emoticon (1205) (or information about the emoticon) described with reference to FIG. 3 may be applied in the same or similar manner.
[0184]
[0185] As described above, an electronic device (e.g., the electronic device (200) of FIG. 2) according to an embodiment includes a communication device (e.g., the communication device (230) of FIG. 2), a storage device (e.g., the electronic device (220) of FIG. 2) for storing at least one artificial intelligence model trained to generate an emoticon based on an input image, and at least one processor (e.g., the processor (210) of FIG. 2), wherein the at least one processor is configured to obtain a first image including a plurality of images from a user device connected to the electronic device through the communication device, input the first image into the at least one artificial intelligence model to obtain a first emoticon, and transmit the first emoticon so that the user device outputs the first emoticon through the communication device, wherein the at least one artificial intelligence model obtains scene information for the first image, the scene information is information related to scene understanding of the plurality of images, obtains a second image obtained by editing the first image based on the scene information, and generates the first emoticon based on the second image. can be learned to generate.
[0186] According to one embodiment, the at least one processor may be configured to obtain image groups by grouping a plurality of images included in the first image into reference units through the at least one artificial intelligence model, and obtain the second image that summarizes the first image based on the image groups.
[0187] According to one embodiment, the image groups may be obtained based on scene transitions of the first image identified based on the scene information.
[0188] According to one embodiment, the at least one processor may be configured to generate at least one content to be added to the first image based on the scene information through the at least one artificial intelligence model, and to obtain the second image having the at least one content added to the first image.
[0189] According to one embodiment, the at least one processor may be configured to obtain a point set and a bounding box set for at least one object included in the first image through the at least one artificial intelligence model, identify an outline of the at least one object based on the point set and the bounding box set, and obtain object information for the at least one object based on the scene information and the outline of the at least one object.
[0190] According to one embodiment, the at least one processor may be configured to determine, through the at least one artificial intelligence model, a location where the at least one content is to be added in relation to the first image based on the object information, and obtain the second image with the at least one content added to the first image based on the location.
[0191] According to one embodiment, the at least one processor is configured to obtain a user input for generating an emoticon, and obtain a second emoticon through the at least one artificial intelligence model based on the user input and the first image, and the at least one artificial intelligence model may be trained to identify a user intent based on the user input, obtain a third image obtained by editing the first image based on the user intent and the scene information for the first image, and generate the second emoticon based on the third image.
[0192] According to one embodiment, the at least one processor may be configured to generate at least one content to be added to the first image based on the user input and the scene information through the at least one artificial intelligence model, and to obtain the third image having the at least one content added to the first image.
[0193] According to one embodiment, the at least one processor may be configured to obtain a user input for modifying the first emoticon through the user device, obtain a third emoticon generated based on the user input and the scene information through the at least one artificial intelligence model, and transmit the third emoticon through the communication device so that the user device outputs the third emoticon.
[0194] According to one embodiment, the first emoticon can be used in a messenger application running through the user device.
[0195] As described above, an operating method of an electronic device according to an embodiment may include an operation of obtaining a first image including a plurality of images from a user device connected to the electronic device, an operation of obtaining scene information about the first image through at least one artificial intelligence model, the scene information being information related to scene understanding of the plurality of images, an operation of obtaining a second image obtained by editing the first image based on the scene information through the at least one artificial intelligence model, an operation of obtaining a first emoticon based on the second image through the at least one artificial intelligence model, and an operation of transmitting the first emoticon so that the user device outputs the first emoticon.
[0196] According to one embodiment, the operation of obtaining the second image may include an operation of obtaining image groups that group a plurality of images included in the first image into reference units, and an operation of obtaining the second image that summarizes the first image based on the image groups.
[0197] According to one embodiment, the operation of obtaining the second image may include an operation of generating at least one content to be added to the first image based on the scene information, and an operation of obtaining the second image with the at least one content added to the first image.
[0198] According to one embodiment, the method of operating the electronic device may further include an operation of obtaining a point set and a bounding box set for at least one object included in the first image, an operation of identifying an outline of the at least one object based on the point set and the bounding box set, and an operation of obtaining object information for the at least one object based on the scene information and the outline of the at least one object.
[0199] According to one embodiment, the operation of obtaining the second image may include an operation of determining a location to which the at least one content is to be added in relation to the first image based on the object information, and an operation of obtaining the second image with the at least one content added to the first image based on the location.
[0200] According to one embodiment, the method of operating the electronic device may further include an operation of obtaining a user input for generating an emoticon, an operation of identifying a user intention based on the user input through the at least one artificial intelligence model, an operation of obtaining a third image obtained by editing the first image based on the user intention and the scene information for the first image through the at least one artificial intelligence model, and an operation of generating a second emoticon based on the third image.
[0201] As described above, an electronic device (e.g., the electronic device (200) of FIG. 2) according to an embodiment includes a display device, a storage device (e.g., the storage device (220) of FIG. 2) for storing at least one artificial intelligence model trained to generate an emoticon based on an input image, and at least one processor (e.g., the processor (210) of FIG. 2), wherein the at least one processor obtains a first image including a plurality of images, inputs the first image to the at least one artificial intelligence model to obtain a first emoticon, and outputs the first emoticon through the display device, and the at least one artificial intelligence model obtains scene information for the first image, the scene information being information related to scene understanding of the plurality of images, obtains a second image obtained by editing the first image based on the scene information, and may be trained to generate the first emoticon based on the second image.
[0202] As described above, in a non-transitory computer-readable recording medium including a program for executing a control method of an electronic device providing a video editing function according to an embodiment, the control method of the electronic device may include the steps of: obtaining a first image including a plurality of images; obtaining scene information about the first image through at least one artificial intelligence model; obtaining a second image obtained by editing the first image based on the scene information, the scene information being information related to scene understanding of the plurality of images; obtaining a first emoticon based on the second image through the at least one artificial intelligence model; and transmitting the first emoticon so as to output the first emoticon.
[0203]
[0204] In this disclosure, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof.
[0205] Terms such as "first," "second," or "first" or "second" may be used simply to distinguish one component from another and do not qualify the components in any other respect (e.g., importance or order).
[0206] The term "module" used in various embodiments of the present disclosure may include a unit implemented in hardware, software, or firmware. For example, it may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions.
[0207] Various embodiments of the present disclosure may be implemented as software (e.g., a program) including one or more commands stored in a storage device (220) (e.g., built-in memory or external memory) readable by a device (e.g., an electronic device (200)). The storage device (220) may be represented as a storage medium.
[0208] According to one embodiment, the methods according to the various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a device-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices.
[0209] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the above-described components may be omitted, or one or more other components or operations may be added. Additionally or alternatively, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each component of the plurality of components in a manner identical to or similar to that performed by the corresponding component among the plurality of components prior to the integration.
[0210] According to various embodiments, the operations performed by a module, program or other component may be performed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be performed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, communication device; A storage device storing at least one artificial intelligence model trained to generate emoticons based on an input image; and Contains at least one processor, At least one processor of the above: Through the communication device, a first image including a plurality of images is obtained from a user device connected to the electronic device, Inputting the first image into at least one artificial intelligence model to obtain a first emoticon, and Through the communication device, the user device is set to transmit the first emoticon so that the first emoticon is output, At least one artificial intelligence model of the above: Obtain scene information for the first image, wherein the scene information is information related to scene understanding for the plurality of images, Obtaining a second image by editing the first image based on the above scene information, and An electronic device that is trained to generate the first emoticon based on the second image.
2. In claim 1, The at least one processor, through the at least one artificial intelligence model: Acquire image groups by grouping multiple images included in the first image into standard units, and An electronic device configured to obtain a second image summarizing the first image based on the image groups.
3. In claim 2, An electronic device wherein the image groups are acquired based on scene transitions of the first image identified based on the scene information.
4. In claim 1, The at least one processor, through the at least one artificial intelligence model: Generating at least one content to be added to the first video based on the scene information, and An electronic device configured to obtain the second image having at least one content added to the first image.
5. In claim 4, The at least one processor, through the at least one artificial intelligence model: Obtain a set of points and a set of bounding boxes for at least one object included in the first image, Identifying the outline of at least one object based on the set of points and the set of bounding boxes, and An electronic device configured to obtain object information about at least one object based on the scene information and the outline of the at least one object.
6. In claim 5 The at least one processor, through the at least one artificial intelligence model: Based on the object information, determining a location where the at least one content is to be added in relation to the first image, and An electronic device configured to acquire the second image having the at least one content added to the first image based on the location.
7. In claim 1, The at least one processor obtains user input for generating an emoticon, and Based on the user input and the first image, a second emoticon is set to be obtained through the at least one artificial intelligence model, At least one artificial intelligence model of the above: Identify user intent based on the above user input, Obtaining a third image by editing the first image based on the user's intention and the scene information for the first image, An electronic device that is trained to generate the second emoticon based on the third image.
8. In claim 7, The at least one processor, through the at least one artificial intelligence model: Generating at least one content to be added to the first image based on the user input and the scene information, and An electronic device set to obtain the third image by adding at least one content to the first image.
9. In claim 1, At least one processor, Obtaining user input for modifying the first emoticon through the user device, Obtaining a third emoticon generated based on the user input and the scene information through at least one artificial intelligence model, and An electronic device configured to transmit the third emoticon through the communication device so that the user device outputs the third emoticon.
10. In the method of operating an electronic device, An operation of obtaining a first image including a plurality of images from a user device connected to the electronic device; An operation of obtaining scene information for the first image through at least one artificial intelligence model, wherein the scene information is information related to scene understanding of the plurality of images; An operation of obtaining a second image by editing the first image based on the scene information through at least one artificial intelligence model; An operation of obtaining a first emoticon based on the second image through at least one artificial intelligence model; and An operating method of an electronic device, comprising an action of transmitting the first emoticon so that the user device outputs the first emoticon.
11. In claim 10, The operation of obtaining the above second image is: An operation of obtaining image groups by grouping a plurality of images included in the first image into standard units; and An operating method of an electronic device, comprising an operation of obtaining a second image that summarizes the first image based on the image groups.
12. In claim 10, The operation of obtaining the above second image is: An operation of generating at least one content to be added to the first image based on the scene information; and An operating method of an electronic device, comprising an operation of obtaining a second image having at least one content added to the first image.
13. In claim 12, An operation of obtaining a set of points and a set of bounding boxes for at least one object included in the first image; An operation of identifying an outline of at least one object based on the set of points and the set of bounding boxes; and An operating method of an electronic device, further comprising an operation of obtaining object information for the at least one object based on the scene information and the outline of the at least one object.
14. In claim 13, The operation of obtaining the above second image is: An operation of determining a location where at least one content is to be added in relation to the first image based on the object information; and An operating method of an electronic device, comprising an operation of obtaining the second image by adding at least one content to the first image based on the location.
15. A non-transitory computer-readable recording medium including a program for executing a control method of an electronic device providing a video editing function, The method of controlling the above electronic device is: A step of acquiring a first image including a plurality of images; A step of obtaining scene information for the first image through at least one artificial intelligence model, wherein the scene information is information related to scene understanding for the plurality of images; A step of obtaining a second image by editing the first image based on the scene information through at least one artificial intelligence model; A step of obtaining a first emoticon based on the second image through at least one artificial intelligence model; and A computer-readable recording medium comprising a step of transmitting the first emoticon to output the first emoticon.