Electronic device for generating emoticon and methods thereof
The use of AI models in electronic devices automates emoticon generation, addressing inefficiencies in manual creation, enabling timely and effective emoticon production that adapts to user emotions and content trends.
Patent Information
- Application Number
- PCT/KR2025/099007
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-05
- Filing Date
- 2025-01-15
- Publication Date
- 2025-08-14
AI Technical Summary
The inefficiency in creating emoticons due to the time and cost involved in manual drawing and development, which fails to keep pace with the rapid evolution of trends in chat and social networking services.
An electronic device equipped with artificial intelligence models to automatically generate emoticons and dynamic emoticons by extracting joint position information, applying motion files, and outputting the results, utilizing a first AI model for joint position, a second AI model for still emoticons, and a third AI model for converting to dynamic emoticons.
Enables rapid and efficient generation of emoticons and dynamic emoticons that reflect user emotions and content trends, enhancing user interaction in chat and social networking services.
Smart Images

Figure KR2025099007_14082025_PF_FP_ABST
Abstract
Description
Electronic devices for generating emoticons and methods thereof
[0001] The present invention relates to an electronic device and method for generating emoticons.
[0002] As the variety and use of electronic devices increases, so does the number of users utilizing chat and social networking services (SNS). Users can use emoticons in chat and social networking services. Recently, chat services have also been provided for real-time content, allowing users to express their emotions and communicate with others by using emoticons in chat windows.
[0003] To express rich emotions and create fun, a variety of emoticons are needed for each situation. Furthermore, it's necessary to provide users with new emoticons at a rapid pace, keeping pace with changing trends.
[0004] However, in the past, there was a problem of inefficiency in terms of time and cost because people had to draw and create emoticons themselves.
[0005] According to at least one embodiment of the present disclosure, an electronic device includes a memory, a communication unit, a display, and a processor.
[0006] In addition, the processor may control the display to obtain an emoticon to which the joint position information has been applied from a second artificial intelligence model based on joint position information of an image corresponding to the emoticon obtained from a first artificial intelligence model, obtain a dynamic emoticon converted from the emoticon based on a rigging point corresponding to the joint position information from a third artificial intelligence model for applying a motion file, and output the obtained dynamic emoticon. In addition, the method for generating an emoticon using an electronic device includes the steps of obtaining an emoticon to which the joint position information has been applied from a second artificial intelligence model based on joint position information of an image corresponding to the emoticon obtained from the first artificial intelligence model, obtain a dynamic emoticon converted from the emoticon based on a rigging point corresponding to the joint position information from a third artificial intelligence model for applying a motion file, and output the obtained dynamic emoticon.
[0007] Also, in a non-transitory computer-readable recording medium storing one or more instructions executed by a control unit of an electronic device to cause the electronic device to perform an operation, the operation includes: a step of obtaining an emoticon to which the joint position information has been applied from a second artificial intelligence model based on joint position information of an image corresponding to the emoticon obtained from a first artificial intelligence model; a step of obtaining a dynamic emoticon converted from the emoticon based on a rigging point corresponding to the joint position information from a third artificial intelligence model for applying a motion file; and a step of outputting the obtained dynamic emoticon.
[0008] FIG. 1 is a drawing for explaining the configuration of an electronic device according to at least one embodiment of the present disclosure.
[0009] FIG. 2 is a drawing for explaining joint position information of a dynamic emoticon according to at least one embodiment of the present disclosure.
[0010] FIG. 3 is a drawing for explaining the operation when an electronic device according to at least one embodiment of the present disclosure is implemented as a server device.
[0011] FIG. 4 is a block diagram illustrating a configuration when an electronic device according to at least one embodiment of the present disclosure is implemented as a server device.
[0012] FIG. 5 is a drawing for explaining each module of a program for generating emoticons according to at least one embodiment of the present disclosure.
[0013] FIG. 6 is a drawing for explaining an example of an emoticon being displayed in a chat window when a user requests display of an emoticon according to at least one embodiment of the present disclosure.
[0014] FIG. 7 is a diagram for showing a chat window UI when a user requests to create an emoticon according to at least one embodiment of the present disclosure.
[0015] FIG. 8 is a diagram illustrating another configuration of an electronic device according to at least one embodiment of the present disclosure.
[0016] FIG. 9 is a diagram illustrating a configuration when an electronic device according to at least one embodiment of the present disclosure is implemented as a display device.
[0017] FIG. 10 is a flowchart illustrating a method for generating an emoticon of an electronic device according to at least one embodiment of the present disclosure.
[0018] FIG. 11 is a flowchart illustrating a method for generating emoticons in an electronic device according to another embodiment of the present disclosure.
[0019] The terms used in the various embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should be defined based on the meaning of the terms and the overall content of this disclosure, rather than simply their names.
[0020] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0021] The expression "at least one of A and / or B" should be understood to mean either "A" or "B" or "A and B".
[0022] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0023] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).
[0024] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this disclosure, terms such as "comprise" or "comprises" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0025] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented as at least one processor (not shown), excluding any "modules" or "parts" that need to be implemented as specific hardware.
[0026] In this disclosure, the term user may refer to a person using an electronic device or a device used by the person.
[0027] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.
[0028] FIG. 1 is a drawing for explaining the configuration of an electronic device according to at least one embodiment of the present disclosure.
[0029] According to FIG. 1, an electronic device (100) includes a memory (110) and a processor (120). The electronic device (100) may be implemented as various types of electronic devices, such as a server device, a PC, a laptop PC, a smartphone, a tablet PC, a set-top box, and a display device such as a TV. When implemented as a server device, the electronic device (100) may communicate with various types of external devices. For example, when a TV or the like is connected, the electronic device (100) may provide data on emoticons that a TV user can use in a chat window or a Social Network Service (SNS) screen. The electronic device (100) may directly generate and provide emoticons or dynamic emoticons.
[0030] The memory (110) is a configuration for storing various software, commands, control codes, and data required for the operation of the electronic device (100). The memory (110) may be implemented as at least one of various memories, such as DRAM (dynamic RAM), SRAM (static RAM), SDRAM (synchronous dynamic RAM), OTPROM (one time programmable ROM), PROM (programmable ROM), EPROM (erasable and programmable ROM), EEPROM (electrically erasable and programmable ROM), mask ROM, flash ROM, flash memory, a hard drive, or a solid state drive (SSD). In various embodiments of the present disclosure, the memory (110) stores at least one artificial intelligence model. The types and operations of the artificial intelligence models will be described in detail in the following section.
[0031] The processor (120) is a configuration for controlling the overall operation of the electronic device (100).
[0032] The processor (120) may include one or more of a digital signal processor (DSP), a microprocessor, a central processing unit (CPU), a micro controller unit (MCU), a micro processing unit (MPU), a controller, an application processor (AP), a communication processor (CP), an ARM processor, and an artificial intelligence (AI) processor, or may be defined by the terms thereof. In addition, the processor (120) may be implemented as a system on chip (SoC) or large scale integration (LSI) having a built-in processing algorithm, or may be implemented in the form of a field programmable gate array (FPGA).
[0033] Additionally, the processor (120) may be implemented through a combination of a general-purpose processor such as a CPU, AP, DSP (Digital Signal Processor), a graphics-only processor such as a GPU, VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU, and software.
[0034] Alternatively, the processor (120) may include a dedicated artificial intelligence processor. In this case, the processor (120) may be designed with a hardware structure specialized for processing a specific artificial intelligence model. For example, the processor (120) may be designed as a hardware chip such as an ASIC or FPGA.
[0035] The processor (120) can perform various functions by executing computer executable instructions stored in the memory (110).
[0036] Specifically, the processor (120) can generate an emoticon or a dynamic emoticon using an artificial intelligence model stored in the memory (110). An emoticon is a compound word combining Emotion and Icon, and can be an image that can express emotions. An emoticon can be referred to by various terms such as a smiley, a graphic object, an icon, a thumbnail image, a character, a mini image, etc., but is described as an emoticon in the present disclosure. A dynamic emoticon (movable emoticon) can be an emoticon that moves. A dynamic emoticon can be referred to by motion emoticons, action emoticons, etc., but is described as a dynamic emoticon in the present disclosure.
[0037] In this disclosure, the term "emoticon" may be understood as referring to a static emoticon without movement, but is not necessarily limited thereto, and may also be used as a general term to refer to both static and dynamic emoticons. For example, the term "method for generating an emoticon" may not refer to a method for generating a static emoticon, but rather a method for generating at least one of a static emoticon and a dynamic emoticon.
[0038] Users can express their feelings while watching content using appropriate emoticons or dynamic emoticons.
[0039] To generate emoticons or dynamic emoticons, the processor (120) may use an artificial intelligence model.
[0040] For example, the memory (110) can store the first to third artificial intelligence models.
[0041] The first artificial intelligence model is an artificial intelligence model for extracting joint position information of a dynamic emoticon to be created. The first artificial intelligence model may be an artificial intelligence model that has undergone iterative learning to generate optimal joint position values by checking for errors that arise in advance while changing various rigging points during the animation process for creating dynamic emoticons. The iterative learning process may be performed in advance by the manufacturer of the electronic device (100) or a supplier providing software for the electronic device (100).
[0042] The second AI model is an AI model designed to generate still emoticons. Alternatively, the second AI model could be described as an image-generative AI.
[0043] The third AI model applies motion files to the generated emoticons, transforming them into dynamic emoticons. The third AI model can also be described as animating AI.
[0044] The processor (120) can input an image of an emoticon to be created into the first artificial intelligence model, thereby obtaining joint location information for the image. For example, if the emoticon to be created is a character of a player AAA, the processor (120) can input a photographic image of the player AAA into the first artificial intelligence model.
[0045] Joint position information refers to information about points that enable an image to move in a specific pose or motion. In other words, joint position information can be information indicating the location of a static emoticon where movement occurs, i.e., the joint position. Joints can be understood as a concept corresponding to human joints. Furthermore, joint position information includes style information to be applied to the emoticon, feature information related to content processed by at least one external device, and a request for the generation of a mask image related to the application of a motion file.
[0046] The processor (120) generates an input prompt including joint position information obtained using the first artificial intelligence model.
[0047] The processor (120) may also include an image corresponding to the emoticon to be created (e.g., a photo image of an AAA player, etc.) and a request for creating a mask image for the image in the input prompt.
[0048] According to another embodiment, the processor (120) may generate an input prompt that includes various pieces of information, such as style information to be generated, feature information obtained by analyzing the content being viewed by the user, etc.
[0049] The processor (120) inputs the generated input prompt into the second artificial intelligence model to obtain an emoticon with joint position information applied. The processor (120) inputs the generated emoticon into the third artificial intelligence model to convert it into a dynamic emoticon.
[0050] Specifically, the third artificial intelligence model can apply a motion file to an emoticon to generate a dynamic emoticon that moves based on a rigging point corresponding to joint position information.
[0051] The operations of the second artificial intelligence model and the third artificial intelligence model can be performed sequentially. That is, the output of the second artificial intelligence model can be used directly as the input of the third artificial intelligence model. Accordingly, the processor (120) can automatically and sequentially generate emoticons and dynamic emoticons by generating a single input prompt.
[0052] Among the information that can be included in the above-described input prompt, style information may include information indicating the style in which the emoticon will be displayed. Specifically, the style information may include various types of information, such as first style information for displaying the emoticon in a realistic style, second style information for displaying it in an animated style, and third style information for displaying it in a text style.
[0053] When creators want to create emoticons, they can also choose their own style.
[0054] For example, if you select the realistic style, the emoticon will be created as an actual photo, while if you select the animated style, the emoticon will be created as an illustration or cartoon. Additionally, various detailed styles, such as the size, color, and background type, can be optionally added.
[0055] The second artificial intelligence model can generate emoticons in a style corresponding to the style information contained in the input prompt.
[0056] Among the information that can be included in the above-described input prompt, the feature information acquired by analyzing the content may be information on various features of the content being displayed or played on the electronic device (100) or designated for reference in creating emoticons. Specifically, it includes at least one of the type of content, information on characters in the content, information on clothing worn by characters in the content, information on the appearance of the characters, and color information included in the content. The content analysis may be performed by another separately provided software module. The second artificial intelligence model may generate emoticons based on the acquired feature information.
[0057] For example, when a user is watching a baseball game, the processor (120) analyzes the content to obtain characteristic information such as information that the content the user is watching is a “baseball game,” information on the players who participated in the game, and the type of team they participated in. The processor (120) can include the obtained characteristic information in the input prompt. Accordingly, the second artificial intelligence model can generate an emoticon reflecting the characteristics of the content. In the above example, assuming that player AAA is a soccer player and a picture of player AAA is input while watching baseball content, the second artificial intelligence model can generate an emoticon of the AAA character wearing a baseball uniform or holding baseball equipment.
[0058] More details about the artificial intelligence model will be described later. The electronic device (100) may automatically generate emoticons when the user watches content, as described above, but may also generate emoticons according to an emoticon generation command directly input by the user. The method of inputting the emoticon generation command may be implemented in various ways depending on the characteristics of the electronic device (100). For example, when the electronic device (100) is implemented as a TV or a set-top box, a UI including an emoticon generation request menu is displayed on the screen according to the operation of the remote control of the TV or set-top box, and when the user selects the menu, the electronic device (100) performs the task of generating a dynamic emoticon as described above. The menu selection is typically made by the user's operation of the remote control, but is not limited thereto, and the menu selection may be received through voice recognition technology that recognizes the voice spoken by the user or motion recognition technology that recognizes the user's motion.
[0059] Data for the emoticon UI may be stored in memory (110) or received from an external device.
[0060] When the electronic device (100) is implemented as a server device, when an emoticon creation request menu is selected in the above-described manner from an external device such as a TV or set-top box connected to the electronic device (100), an emoticon creation request can be transmitted from the external device to the electronic device (100).
[0061] FIG. 2 is a drawing for explaining joint position information of a dynamic emoticon according to at least one embodiment of the present disclosure.
[0062] According to FIG. 2, rigging points corresponding to joint position information are indicated on the arms and legs of a human-shaped emoticon (10) standing in an upright posture (a, b, c, d). When a motion file is applied to the emoticon, the emoticon can move based on the rigging points (a, b, c, d). For example, when a motion file is applied to a human-shaped emoticon (10) standing in an upright posture as illustrated in FIG. 2, the emoticon can move based on points a and b, thereby creating a dynamic emoticon (20) performing a cheering gesture.
[0063] As described above, joint position information is information about points that enable an image to move in a specific pose or motion. The process of applying a motion file to an emoticon image is called "rigging" the joint information to the character image. Rigging is the process of creating a skeleton-like connection structure to give the emoticon movement. If each part of the connection structure is considered to be connected to each other by the joints described above, the part that controls the movement of the entire emoticon is located at the top.
[0064] Based on this connection structure, the processor (120) can calculate relative position and rotation information between each part. For example, if an upper part rotates around a single point, the rotation degree and direction of the lower part connected to it are calculated based on the relative position and rotation information with respect to the upper part. Accordingly, the lower part can be perceived as moving naturally based on the joint, thereby making the overall movement of the model more natural.
[0065] Once the overall framework of an emoticon is established, a mask image can be created to limit the range of motion, resulting in natural animation. A mask image can be used to specify the area within an emoticon where motion will occur. Mask images can be used to control the shape or visibility of specific areas. Mask images are typically black and white. The white portion represents the actual area of the emoticon visible to the user, while the black portion represents the area to be hidden. For example, when creating a dynamic emoticon that blinks, a mask image for the eyes is created. The mask image should be black and white, with the pupils appearing in white. By applying the mask image to the emoticon and adding keyframes that control the visibility of the pupils, the effect of the pupils appearing and disappearing can be created. Once the mask image is created, a motion file (BVH) can be applied to define the overall motion of the emoticon. Motion files will be discussed later.
[0066] Figure 3 is a diagram illustrating the operation of an electronic device that performs the various operations described above when implemented as a server device. As described as a server device, the reference numeral 100 is also changed to 300.
[0067] According to FIG. 3, the server device (300) can communicate with a plurality of external devices (200-1 to 200-n). In addition, the server device (300) can receive various types of information from the external devices (200-1 to 200-n).
[0068] The plurality of external devices (200-1 to 200-n) may be electronic devices for providing content. Specifically, the plurality of external devices (200-1 to 200-n) may be composed of a TV, a set-top box, a one-connect box, a PC, a laptop PC, a smart monitor, a mobile phone, a tablet PC, a kiosk, an electronic whiteboard, and other home appliances. The plurality of external devices (200-1 to 200-n) may be display devices directly equipped with a display, or electronic devices connected to an external display device. Hereinafter, the expression "providing content" may be understood to mean directly displaying content when implemented as a display device, and when implemented as an electronic device that does not include a display device and is connected to an external display device, it may be understood as an operation of transmitting content and a control signal for controlling the output of the content to the external display device.
[0069] Content can be provided from various content sources. For example, the server device (300) may directly provide content to a plurality of external devices (200-1 to 200-n), but the present invention is not limited thereto. The plurality of external devices (200-1 to 200-n) may receive content as external inputs such as an antenna port, a broadcast cable, the Internet, etc., or may receive content from various sources such as a multimedia player, PC, laptop PC, or console game machine connected to the external devices (200-1 to 200-n).
[0070] External devices (200-1 to 200-n) can download content in advance and play it for output, or stream it in real time and output it.
[0071] The server device (300) can receive input data when a user of each external device (200-1 to 200-n) inputs a chat message while outputting content from each external device (200-1 to 200-n).
[0072] When the same content is output from multiple external devices (200-1 to 200-n), the server device (300) can receive content information viewed by users of the external devices (200-1 to 200-n), i.e., viewers, text information entered while performing a chat function, etc. Specifically, users of each external device (200-1 to 200-n) can enter chat messages expressing their emotions while watching the content. Chat messages entered from each external device (200-1 to 200-n) can be displayed through the screen of each external device (200-1 to 200-n) or its connected device. In addition, information about chat messages entered from each external device (200-1 to 200-n) is transmitted to the server device (300).
[0073] When a request for emoticon creation is received from at least one of a plurality of external devices (200-1 to 200-n), the server device (300) can create an emoticon as described above. In this case, when the image of the emoticon to be created, style information, characteristic information of the content being output, etc. are received from the device, the received information can be referenced for emoticon creation.
[0074] Additionally, the server device (300) may generate dynamic emoticons based on text information received from multiple external devices (200-1 to 200-n). The process of generating dynamic emoticons based on text information in a chat message will be described later.
[0075] Fig. 4 is a block diagram showing an example of the configuration of the server device of Fig. 3.
[0076] According to FIG. 4, the server device (300) includes a memory (310), a processor (320), and a communication unit (330).
[0077] The memory (310) is configured to store various software, commands, control codes, and data required for the operation of the server device (300). In the present disclosure, the memory (310) stores an artificial intelligence model for generating dynamic emoticons. In addition, various data (e.g., layout data, graphic data, emoticon information, etc.) for configuring an emoticon UI or chat window can be stored. In addition, the memory (310) can also store data from each external device received through the communication unit (330). The memory (310) may also store information related to the user account viewing the corresponding content. To register a user account, the user can input various identification information, such as his or her name, age, address, phone number, email address, ID, and password, through the website. The user account information stored in this way is used when using the chat service while viewing the content. Details regarding the memory (310) are as described above with reference to FIG. 1.
[0078] The processor (320) is configured to control the overall operation of the server device (300). In the present disclosure, the processor (320) generates an input prompt that includes style information to be applied to an emoticon, feature information acquired by analyzing content to be processed in at least one external device (200-1 to 200-n), a request for generating a mask image that specifies a portion of the emoticon to which a motion file is to be applied, and joint position information for the image.
[0079] In addition, when the server device (300) provides an emoticon provision service or a chat service, the processor (320) can provide a chat window that can be used by multiple external devices (200-1 to 200-n) viewing the same content. The chat window can be a display area for displaying chat messages and emoticons. The chat window can be described in various ways, such as a chat area, a viewer area, a participation space, an additional display area, a related information display area, etc., but is described as a chat window in the present disclosure.
[0080] The server device (300) does not necessarily provide only one chat window, and may additionally provide a separate display area capable of displaying only emoticons. However, for convenience of explanation, the following description will assume that both emoticons and chat messages are displayed in a single display area, i.e., a chat window.
[0081] The processor (320) can transmit data regarding the chat window and a control signal for commanding the display thereof to each external device (200-1 to 200-n) so that a chat window including various emoticons or chat messages input from each external device (200-1 to 200-n) can be commonly displayed on multiple external devices (200-1 to 200-n). Accordingly, users using multiple external devices (200-1 to 200-n) can share their emotions with others through emoticons or chat messages. The processor (320) can also receive a request from the external devices (200-1 to 200-n) to display an emoticon.
[0082] When a request for displaying an emoticon is received from at least one of the external devices (200-1 to 200-n) connected via the communication unit (330), the processor (320) transmits emoticon information including a dynamic emoticon stored in the memory (310) to the external device (200-1 to 200-n). When a dynamic emoticon is selected from the external device (200-1 to 200-n), the selected dynamic emoticon can be provided to other external devices that are outputting the same content as the external device (200-1 to 200-n). Accordingly, the user can use the dynamic emoticon through the chat service provided by the server device (300).
[0083] When the processor (320) receives an emoticon creation request from at least one of the external devices (200-1 to 200-n) connected via the communication unit (330), the processor (320) can create an emoticon in the manner described in the above-described section. The processor (320) can store the created emoticon in the memory (310) and provide it to all the external devices (200-1 to 200-n). However, the present invention is not limited thereto, and the processor (320) can also provide the created emoticon only to the external device that transmitted the emoticon creation request.
[0084] As another example, the processor (320) may generate a dynamic emoticon based on an image uploaded to a chat window. When a request for dynamic emoticon generation for an image uploaded to the chat window is received from one of a plurality of external devices (200-1 to 200-n) through the communication unit (330), the processor (320) inputs the uploaded image into a first artificial intelligence model to generate joint position information for the image. The first artificial intelligence model generates an input prompt including the generated joint position information and feature information acquired by identifying the atmosphere of the chat window, and inputs the input prompt to a second artificial intelligence model to obtain an emoticon to which the joint position information has been applied by the second artificial intelligence model. The emoticon generated by the second artificial intelligence model is input to a third artificial intelligence model. The third artificial intelligence model converts the emoticon into a dynamic emoticon that moves based on a rigging point corresponding to the joint position information. The processor (320) provides the generated dynamic emoticon to the external device that transmitted the dynamic emoticon generation request or to all external devices through the communication unit (330).
[0085] The processor (320) can generate dynamic emoticons using not only images uploaded during the chat function but also text as a script. In this case, the processor (320) generates an input prompt including information about the text entered by the user and feature information analyzed from the content. The following process for generating dynamic emoticons is as described above. For example, if a user uploads an image of a puppy during the chat function and requests the generation of a dynamic emoticon for the image, the first artificial intelligence model can extract joint position information for generating a dynamic emoticon for the puppy image and analyze the content and the chat window to generate a command prompt including information about the chat message uploaded by the user, the mood of the chat window, and feature information of the content. If the chat message or the mood of the chat window is cheerful, a dynamic emoticon in the shape of a smiling and jumping puppy can be automatically generated. As a result, it is possible to generate dynamic emoticons tailored to individual users.
[0086] As another example, the server device (300) may provide a metaverse service.
[0087] The Metaverse is an extensible digital space that merges the real and virtual worlds. It refers to a space where 3D virtual environments, virtual reality (VR), augmented reality (AR), online games, social media, the internet, and the digital economy interact. In the Metaverse, users can create their own avatars and interact with them in real life. Users can input their desired avatar image to create a dynamic avatar of their choice, and can also customize it. This allows users to freely move their avatars within the virtual space, chat with other users, and interact with them through games offered by the Metaverse.
[0088] The communication unit (330) is configured to communicate with multiple external devices. In an environment such as FIG. 3, the communication unit (330) can communicate with at least one external device (200-1 to 200-n). Specifically, the communication unit (330) can receive information about emoticons entered by multiple users viewing the same content through their respective external devices, information about the content being viewed, and information about text entered while performing a chat function.
[0089] The communication unit (330) can transmit and receive various signals and data with external devices through various wired and wireless communication methods such as Bluetooth, AP-based Wi-Fi (Wi-Fi, Wireless LAN network), Zigbee, wired / wireless LAN (Local Area Network), WAN (Wide Area Network), Ethernet, IEEE 1394, HDMI (High-Definition Multimedia Interface), USB (Universal Serial Bus), MHL (Mobile High-Definition Link), AES / EBU (Audio Engineering Society / European Broadcasting Union), optical, coaxial, etc.
[0090] FIG. 5 is a drawing for explaining each module of a program for generating emoticons according to at least one embodiment of the present disclosure.
[0091] According to FIG. 5, various software and data for creating emoticons can be stored in memory (110, 310). Specifically, a first artificial intelligence model (311), a content analysis module (312), an input prompt creation module (313), a second artificial intelligence model (314), a third artificial intelligence model (315), etc. can be stored.
[0092] The artificial intelligence model is composed of a first artificial intelligence model (311) that obtains joint position information for an input image as described above and generates an input prompt that includes joint position information, style information to be generated, a request to generate an image and a mask image, and feature information obtained by analyzing the content being viewed by the user, a second artificial intelligence model (314) that generates an emoticon according to the generated input prompt, and a third artificial intelligence model (315) that applies a motion file to the generated emoticon.
[0093] An artificial intelligence model is a computer system or software module that implements human-level intelligence. It has the characteristics of a machine learning and making judgments on its own, and its recognition rate improves with use.
[0094] Artificial intelligence models are composed of machine learning (deep learning) technology that uses algorithms that classify / learn the characteristics of input data on their own, and element technologies that use machine learning algorithms to simulate the cognitive and judgment functions of the human brain.
[0095] The element technologies may include, for example, at least one of a linguistic understanding technology that recognizes human language / characters, a visual understanding technology that recognizes objects as if they were human vision, an inference / prediction technology that judges information and logically infers and predicts, and a knowledge representation technology that processes human experience information into knowledge data.
[0096] An artificial intelligence model can be created through learning. Creating a model through learning means that a basic artificial intelligence model is trained using a learning algorithm using a plurality of learning data, thereby creating a predefined set of operating rules or an artificial intelligence model configured to perform a desired characteristic (or purpose). This learning may be performed in the electronic device (100) or server device (300) itself, in which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0097] At least one of the first to third artificial intelligence models may be comprised of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the computational results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated during the learning process so that the loss value or cost value acquired by the artificial intelligence model is reduced or minimized.
[0098] Artificial neural networks may include deep neural networks (DNNs), such as, but not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), or deep Q-networks.
[0099] Additionally, in various embodiments of the present disclosure, the processor (120, 320) must generate appropriate emoticons for the user, which may be provided based on the results of an inference and prediction process utilizing an artificial intelligence model. Inference and prediction are technologies for logically inferring and predicting information by judging it, and include knowledge and probability-based inference, optimized prediction, preference-based planning, and recommendations.
[0100] In this disclosure, a Generative Adversarial Network (GAN) can also be used as a second artificial intelligence model to generate virtual images. A Generative Adversarial Network (GAN) is a deep learning model used to generate new data by mimicking the distribution of given data. A GAN is an artificial intelligence model that initially starts with random noise and, through repeated training, learns to generate data with a distribution similar to real data.
[0101] The first AI model checks for errors that may occur when applying motion files to ensure that the motion files are properly applied to the input character image. It then repeatedly learns joint rigging and generates optimal joint position values for the input image. The specifics of joint positioning and rigging are described above. Even if the AI model does not actually apply the motion files to the image to generate results, if deep learning is performed in advance using a large amount of diverse data, errors that may occur before the motion files are actually applied can be predicted. This process is called joint rigging reinforcement learning.
[0102] The content analysis module (312) analyzes the content as described above to obtain feature information, and the first artificial intelligence model (311) transmits information including a request for generating a character image and a corresponding mask image based on the emoticon style information to be generated and the input image to the input prompt generation module (313). At this time, the input prompt is generated in a format that the second artificial intelligence model (314) can understand. Instead of generating the user's text input value as a prompt, one artificial intelligence model generates a prompt so that another artificial intelligence model can understand it, which is called an artificial intelligence chain (AI chain) or artificial intelligence pipeline. In other words, it refers to a structure in which artificial intelligence models are connected so that they can perform tasks continuously.
[0103] As another example, when generating dynamic emoticons based on text and images entered by a user while performing a chat function, the processor (320) can use the content analysis module (312) to analyze the images or text entered by the user while performing the chat function to identify the atmosphere of the chat window. In this case, the atmosphere of the chat window can be information about the viewing atmosphere or the viewer's emotions.
[0104] Specifically, the processor (320) inputs texts used during the execution of the chat function into the content analysis module (312). The content analysis module (312) can output probability values of multiple moods inferred based on the text. The processor (120) can identify the mood with the maximum probability value among the output probability values as the mood of the chat window. To increase accuracy, the processor (120) can use multiple texts as input values, and can also determine the chat window mood when the difference between the maximum probability value and other probability values exceeds a threshold value. Through this process, an input prompt including the acquired chat window mood, the chat message and image input by the user, and joint information about the image is generated.
[0105] When the generated input prompt is input into the second artificial intelligence model (314), an emoticon is generated based on the transmitted information. The second artificial intelligence model (314) is a GEN-Image AI module, which is an artificial intelligence model that generates images.
[0106] The emoticon generated at this time is generated in the form of a static emoticon because it is a step before applying the motion file (BVH) file. The generated emoticon is input into the third artificial intelligence model (315), and the motion file (BVH) file is applied to it to have movement. BVH file is an abbreviation for BioVision Hierarchy and is a file format that describes the skeleton or bone structure of a 3D model. It is mainly used for motion capture, animation, and 3D computer graphics. The BVH file contains information about the hierarchical structure between each joint and bone, and includes information indicating the movement of a specific model or the rotation and movement of each joint.
[0107] FIG. 6 is a diagram illustrating an example of how an emoticon is displayed in a chat window when a user requests to display an emoticon according to at least one embodiment of the present disclosure. FIG. 7 is a diagram illustrating a chat window UI when a user requests to create an emoticon according to at least one embodiment of the present disclosure.
[0108] According to FIG. 6, the external device (200) can display an emoticon UI (404) on an area of the screen while outputting content (400) on the screen. Information about the emoticon UI (404) can be provided in real time from the server device (300), and can also be generated based on data that the external device (200) stores on its own. The external device (200) can automatically generate and display the emoticon UI (404) according to the type of content selected by the user. For example, if the server device (300) has determined content that provides an emoticon service or a chat service, the emoticon UI (404) can be displayed when the content is output. However, the present invention is not limited thereto, and the external device (200) can also display the emoticon UI (404) according to a user command. When the emoticon display request button (403) located in the middle of the chat window is clicked, emoticons available to the user are displayed on the emoticon UI. On the emoticon UI (404) of FIG. 6, emoticons selected based on various criteria, such as emoticons frequently used by the user, emoticons that match the content currently being displayed, and emoticons available to the user, may be displayed. Alternatively, all emoticons may be displayed. When the user selects an emoticon, the selected emoticon is displayed in the chat window (401). In addition, the external device (200) may display emoticons (402) entered by other users in the chat window, as shown in FIG. 6, based on emoticon information provided from the server device (300).
[0109] According to FIG. 7, when a user selects the emoticon creation button (405), a dynamic emoticon is created based on an image and chat message uploaded by the user. As illustrated in FIG. 7, when a user inputs the text "I'm bored" into a chat window (402-2), the processor (320) can analyze the context of the content and the meaning of the text to create an emoticon using the chat message as a script. "I'm bored" is a text with negative content, and the processor (320) can create an emoticon with an angry face (404-1) and a sleepy emoticon (404-2) based on this.
[0110] Additionally, since users use the chat service with their own identifiable IDs, the processor (320) can select only the chat messages directly entered by the user and generate dynamic emoticons that meet the user's needs.
[0111] FIG. 8 is a diagram illustrating another configuration of an electronic device according to at least one embodiment of the present disclosure.
[0112] According to FIG. 8, the electronic device (500) may include a memory (510), a processor (520), a communication unit (530), and a display (540).
[0113] The processor (520) can obtain dynamic emoticons using the first to third artificial intelligence models. At least some of the first to third artificial intelligence models may be stored in the memory (510), but this is not limited to the first to third artificial intelligence models, and all of the first to third artificial intelligence models may be stored in an external server. In this case, the processor (520) can communicate with the server via the communication unit (530) to utilize the first to third artificial intelligence models.
[0114] Specifically, the processor (520) can obtain joint position information for an image to be made into an emoticon using the first artificial intelligence model. For example, the processor (520) can transmit information including an image to a server via the communication unit (530), and the server can output joint position information using the image included in the received information as an input value for the first artificial intelligence model.
[0115] The processor (520) may obtain an emoticon with applied joint position information from the second artificial intelligence model based on the joint position information of the image corresponding to the emoticon obtained from the first artificial intelligence model. In addition, the processor (520) may obtain a dynamic emoticon converted from the emoticon based on the rigging point corresponding to the joint position information from the third artificial intelligence model for motion file application. The server may transmit the output value of the third artificial intelligence model, i.e., the dynamic emoticon, to the electronic device (500), and the processor (520) may receive it through the communication unit (530).
[0116] The processor (520) can store the received dynamic emoticon in the memory (510) or output it through the display (540).
[0117] The joint position information obtained from the server using the first artificial intelligence model may include various information, such as style information to be applied to the emoticon, feature information related to content processed by at least one external device, and a request for generating a mask image related to the application of the motion file. However, the joint position information is not limited thereto and may be varied in various ways. For example, the output value provided from the first artificial intelligence model may be in the form of an input prompt including joint position information, feature information of the content, style information to be applied to the emoticon, and a request for generating a mask image.
[0118] The content characteristic information may be information about various characteristics related to the content to be processed by the electronic device (500). For example, when the electronic device (500) is implemented as a set-top box, the electronic device (500) may be connected to a TV receiver or other external display devices to provide various screens. The electronic device (500) may further include an input / output interface to connect to these various external devices. In addition, the electronic device (500) may further include a signal processing unit (not shown) for processing content provided from various sources such as cable, satellite, or terrestrial broadcasting stations, and web servers.
[0119] As described above, when a server equipped with the first to third artificial intelligence models also serves as a content source that provides content, the server can directly obtain characteristic information of the content by analyzing the content to be output from the electronic device (500).
[0120] Alternatively, if content is provided from a content source other than the server, the processor (520) may control the signal processing unit to process the content, acquire characteristic information, and then provide the acquired characteristic information to the server through the communication unit (530).
[0121] The signal processing unit (not shown) decodes signals received from a TV receiver, processes audio and video signals, and displays them on an external device. Decoding refers to the process of decoding digital broadcast signals and converting them into a format that can be displayed on a screen. When the image frames constituting the content are acquired by the signal processing unit, the processor (520) analyzes feature points to generate dynamic emoticons from the image frames.
[0122] Specifically, the processor (520) divides an image frame into pixel units or pixel groups, and identifies the edges of objects included in the image frame based on the feature values of each pixel or pixel group. The processor (520) can identify what the object is (person, object, text, etc.) based on the shape and size of the identified boundary, the color value inside the boundary, etc. The processor (520) can obtain information corresponding to the identified feature from a pre-stored feature database. For example, in the case of sports content in which an AAA player plays while wearing the uniform of the team to which he belongs, the processor (520) can identify the uniform logo, number, facial features of the AAA player, etc., and estimate that the content is sports content in which the AAA player appears.
[0123] However, the present invention is not limited thereto, and according to another embodiment, the processor (520) may analyze the characteristics of the content based on metadata about the content, EPG (Electronic Program Guide) information, etc.
[0124] Alternatively, the processor (520) may perform an Internet search using content titles, etc. as keywords and use the searched information as feature information.
[0125] Meanwhile, the signal processing unit may include a digital signal processing unit, an analog-to-digital converter, a graphics processing unit, and a video output interface. The digital signal processing unit includes an MPEG video codec and an AC-3 audio codec and is primarily responsible for processing and decoding digital broadcast signals. The analog-to-digital converter converts analog signals, such as terrestrial TV signals, into digital signals for processing by the processor. The graphics processing unit generates and manages graphic elements displayed on the screen. In addition, the video output interface transmits the processed video signal to the television through a video output interface such as HDMI, component video, or SCART.
[0126] In addition, the program and data stored in the memory (510) may vary depending on the type of electronic device (500). For example, when implemented as a set-top box,
[0127] The memory (510) can store various software, commands, control codes, and data necessary for the operation of the set-top box (500). For example, the memory (830) can store various data (e.g., layout data, graphic data, emoticon information, etc.) and artificial intelligence models for configuring an emoticon UI or chat window. Since an example of the memory (510) has been described in FIG. 2, a detailed description thereof will be omitted.
[0128] The processor (520) can control the overall operation of the set-top box.
[0129] The processor (520) can process data, media content, etc. received from a network and transmit them to external devices (200-1 to 200-n). When a user command to view specific content is input, the processor (520) receives the content from the content source, generates a control signal to display the received content, and transmits the received content to the external device. In addition, as described above, the processor (520) can also perform an operation of receiving information about a chat message or emoticon input by the user, analyzing the same, and generating a dynamic emoticon.
[0130] The processor (520) can generate an emoticon according to a user's request to generate an emoticon. In this case, the user's request can be made using a method of voice recognition in addition to a method in which the user selects a UI displayed on an external device (200-1 to 200-n). If the remote control has a built-in microphone, the remote control can receive an analog voice signal, digitize it, and transmit it to the set-top box. The processor (520) can receive the voice signal through a Wi-Fi module, a Bluetooth module, or the like of the communication unit (530). Alternatively, the processor (520) can also receive the user's voice signal through a mobile phone with a remote control application installed, other terminal devices, or an AI speaker.
[0131] The communication unit (530) is a component for performing communication with an external device. In FIG. 8, only one communication unit (530) is illustrated, but the number and type of communication units are not limited. For example, the set-top box may separately include a communication unit for communicating with a server device and a communication unit for communicating with other external devices, such as a remote control. The communication unit (530) may include an IR signal receiving module capable of receiving a remote control signal. However, if the set-top box is implemented to be connected to a remote control according to a wireless communication standard, such as Bluetooth or Wi-Fi, the IR signal receiving module may be omitted. Specifically, communication may be performed with the server device via an Ethernet modem or Wi-Fi module, and communication may be performed with the remote control or the like via a Bluetooth module. However, the present invention is not limited thereto, and the set-top box may include an integrated communication unit equipped with multiple communication modules.
[0132] Also, although only the communication unit (530) is illustrated in FIG. 8, the set-top box may further include various input / output interfaces such as an HDMI port, DP, RGB, DVI, USB, Thunderbolt, etc. for receiving video / audio signals by being connected to external content sources. HMDI, DP, and Thunderbolt are ports that can transmit video and audio signals simultaneously. The processor (520) can receive images or contents to be converted into emoticons through the communication unit (530) and these input / output interfaces. In addition, dynamic emoticons obtained in the above-described manner can be transmitted to various external devices.
[0133] Meanwhile, the functions described in the other embodiments described above can be performed in the same manner in the embodiment of FIG. 8. For example, when a request for providing a chat function is received through the communication unit (530) from an external device outputting the same content, the processor (520) can control the communication unit (530) to transmit data for providing a chat function corresponding to the same content to the external device.
[0134] Additionally, when a request for dynamic emoticon generation for an image uploaded by an external device is received while executing a chat function based on transmitted data, the processor (520) may transmit a dynamic emoticon converted from the image uploaded by the external device to the external device.
[0135] Alternatively, the processor (520) may control the display (540) to display a chat window in an area of the screen where content is displayed when the chat function is executed, and to display text entered by the user or another user viewing the same content in the chat window. The processor (520) may identify not only feature information related to the content, joint position information for the image, but also the atmosphere of the chat window using the first artificial intelligence model, and then generate dynamic emoticons based on these various pieces of information as described above. This has been specifically described in the other embodiments described above, and thus a duplicate description will be omitted.
[0136] FIG. 9 is a diagram illustrating a configuration when an electronic device according to at least one embodiment of the present disclosure is implemented as a display device.
[0137] According to FIG. 9, the display device (600) is composed of a memory (610), a processor (620), and a display (630).
[0138] The memory (610) is configured to store various software, commands, control codes, and data required for the operation of the display device (600). For example, the memory (610) can store various data (e.g., layout data, graphic data, emoticon information, etc.) for configuring an emoticon UI or chat window, as well as a program and artificial intelligence model for dynamic emoticon generation. Since an example of the memory (610) has been described in FIG. 5, a detailed description thereof will be omitted.
[0139] The processor (620) is a component for controlling the overall operation of the display device (600). Specifically, the processor (620) can control the display (630) to display content. When a menu for inputting emoticons is selected, the processor (620) can control the display (630) to display at least one of various types of emoticon UIs and chat windows as illustrated in FIGS. 6 and 7 . When various user commands are input, the processor (620) can perform operations corresponding to the user commands. The user commands may be received through a remote control corresponding to the display device (600), input through buttons or other operation panels provided on the main body of the display device (600), or received through an external electronic device connected to the display device (600).
[0140] Another example of generating dynamic emoticons on a display device (600) may include a method of generating emoticons by recognizing the user's voice utterances and a method of generating emoticons by recognizing the user's motions. The part related to user voice recognition has been described above in FIG. 8.
[0141] Motion recognition is a technology that detects a user's actions or movements, collects and interprets this data, and utilizes it in various applications. In this disclosure, motion recognition technology is utilized to detect a user's actions in real time and generate or manipulate emoticons or avatars that match the user's movements.
[0142] To obtain image and joint information for creating emoticons or avatars using motion recognition technology, a method of additionally using a camera to capture the subject to detect the user's movements may be used. However, this is not limited to this method, and methods of detecting the user's movements using radar or ultrasonic sensors may also be used. Alternatively, information on the change patterns of measurements from the gyroscope and accelerometer built into the user's mobile phone or remote control may be received, and the user's movements may be estimated based on this information.
[0143] The processor (620) can determine a joint position based on a motion recognized by motion recognition technology, and generate an input prompt including the determined joint position, thereby generating a dynamic emoticon through the above-described process.
[0144] The display (630) is a configuration for displaying various screens under the control of the processor (620). The display (630) may be implemented as a display including a self-luminous element or a display including a non-luminous element and a backlight. For example, the display may be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an LED (Light Emitting Diodes), a micro LED, a Mini LED, a PDP (Plasma Display Panel), a QD (Quantum dot) display, a QLED (Quantum dot light-emitting diodes), etc. The display (630) may also include a driving circuit, a backlight unit, etc., which may be implemented in a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. The user may select an emoticon, request an emoticon display, or request an emoticon creation through the emoticon UI. Although FIG. 9 describes a method of processing information in a display device including a display (630), as described above, these operations may also be performed in an electronic device that does not include a display.
[0145] In this case, the processor (620) may generate screen data for a content screen, an emoticon UI, a chat window, etc., and then transmit the screen data and a control signal for controlling the display of the screen based on the screen data to a connected external display device (e.g., a TV, a monitor, a projector, etc.). Such data and control signals may be transmitted through an input / output interface connected between the electronic device and the external display device, or may be transmitted through a communication session when a communication session is connected between the electronic device and the external display device.
[0146] FIG. 10 is a flowchart illustrating a method for generating an emoticon of an electronic device according to at least one embodiment of the present disclosure.
[0147] According to FIG. 10, in order to create an emoticon using an electronic device, when an image for an emoticon to be created is input into the first artificial intelligence model, the method includes the steps of: obtaining joint position information for the image (S1010); generating an input prompt including the obtained joint position information by the electronic device (S1020); inputting the generated input prompt into the second artificial intelligence model (S1030); and obtaining the emoticon with the joint position information applied (S1040). Specific details for each artificial intelligence model and specific details for joint positions, input prompt generation, etc. have been described above.
[0148] The control method of FIG. 10 can be performed by an electronic device (100) or a server device (300) having the configuration described in FIGS. 1 and 3, but is not necessarily limited thereto, and can also be performed by a device having a different configuration from that of FIG. 3.
[0149] FIG. 11 is a flowchart illustrating a method for generating emoticons in an electronic device according to another embodiment of the present disclosure.
[0150] According to FIG. 11, the electronic device can generate dynamic emoticons based on images and texts uploaded by a user while providing a chat service.
[0151] As described above, when a request for providing a chat service is received from multiple external devices among external devices outputting the same content (S1110), a chat window to be shared by the multiple external devices is transmitted to each of the multiple external devices (S1120), and when a request for generating a dynamic emoticon for an image uploaded to a chat window is received from one of the multiple external devices (S1130), the input image is input into a first artificial intelligence model to generate joint position information for the image, an input prompt including the generated joint position information and feature information obtained by identifying the atmosphere of the chat window is generated and inputted into a second artificial intelligence model to obtain an emoticon to which the joint position information is applied, and the generated emoticon is input into a third artificial intelligence model to convert the emoticon into a dynamic emoticon that moves based on a rigging point corresponding to the joint position information (S1140). The dynamic emoticon thus generated is provided to the external device that transmitted the dynamic emoticon generation request (S1150). Specific examples thereof have been described above, so redundant description thereof will be omitted.
[0152] The control method described in Fig. 11 can be executed by electronic devices having various configurations as shown in Figs. 2, 4, 8, 9, etc., but is not necessarily limited thereto, and can also be executed by electronic devices having different configurations.
[0153] The programs or instructions for performing the various information processing methods described above may be provided stored on a non-transitory, readable medium. The non-transitory, readable medium may be loaded and used in a device capable of recalling the instructions stored in the storage medium and performing operations according to the recalled instructions. Accordingly, when the program or instructions stored on the non-transitory, readable medium are executed by a processor, the processor may directly, or under the control of the processor, utilize other components to perform the operations described in the various embodiments described above.
[0154] A non-transitory computer-readable medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of non-transitory computer-readable media include CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.
[0155] Instructions may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" means that the storage medium does not contain signals and is tangible, but does not distinguish between whether data is stored semi-permanently or temporarily on the storage medium.
[0156] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be implemented as a product distributed online through an application store, in addition to the non-transitory readable recording medium described above. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0157] Accordingly, in a non-transitory computer-readable medium storing one or more instructions executed by a control unit of an electronic device to cause the electronic device to perform an operation,
[0158] The operation of the recording medium is a step of obtaining joint position information for the image using a first artificial intelligence model when an image for an emoticon to be created is input, a step of obtaining an emoticon with the joint position information applied by inputting an input prompt including the obtained joint position information into a second artificial intelligence model, and
[0159] It includes a step of inputting an emoticon into a third artificial intelligence model and converting the emoticon into a dynamic emoticon that moves based on a rigging point corresponding to joint position information.
[0160] While the present invention has been described with reference to the attached drawings, the scope of the present invention is determined by the claims described below and should not be construed as being limited to the aforementioned embodiments and / or drawings. Furthermore, it should be clearly understood that improvements, modifications, and variations apparent to those skilled in the art, as defined in the claims, are also included within the scope of the present invention.
Claims
1. In electronic devices, memory; Department of Communications; display; and Processor; including; The above processor, Based on the joint position information of the image corresponding to the emoticon obtained from the first artificial intelligence model, the emoticon to which the joint position information is applied is obtained from the second artificial intelligence model, An electronic device that obtains a dynamic emoticon converted from the emoticon based on the rigging point corresponding to the joint position information from the third artificial intelligence model for motion file application, and controls the display to output the obtained dynamic emoticon.
2. In paragraph 1, The joint position information of the image corresponding to the above emoticon is An electronic device comprising style information to be applied to the emoticon, feature information related to content processed by at least one external device, and a request for generating a mask image related to application of the motion file.
3. In paragraph 1, It includes more input / output interfaces, The above processor, An electronic device that, when content to be provided to an external device connected through the input / output interface is received from a server through the communication unit, processes the received content based on the characteristics of the external device and controls the transmission to the external device through the input / output interface.
4. In paragraph 2, An electronic device, wherein the style information includes at least one of first style information for displaying the emoticon in a real-life style, second style information for displaying the emoticon in an animation style, and third style information for displaying the emoticon in a text style.
5. In paragraph 2, The above characteristic information is, An electronic device including at least one of the type of the content, information on characters in the content, information on clothing worn by characters in the content, information on the appearance of the characters, and color information included in the content.
6. In paragraph 1, The above processor, An electronic device that, when an emoticon display request is received from at least one of the external devices connected through the communication unit, controls to transmit emoticon information including the dynamic emoticon to the external device.
7. In paragraph 6, The above processor, An electronic device that, when the dynamic emoticon is selected from the external device, controls the communication unit to provide the selected dynamic emoticon to another external device that is outputting the same content as the external device.
8. In paragraph 1, The above processor, When a request for providing a chat function is received through the communication unit from an external device outputting the same content, the communication unit is controlled to transmit data for providing a chat function corresponding to the same content to the external device, An electronic device that controls the communication unit to transmit a dynamic emoticon converted from the image uploaded by the external device to the external device when a request for dynamic emoticon generation for the image uploaded by the external device is received while executing a chat function based on the transmitted data.
9. In paragraph 3, The above processor, An electronic device that, when receiving the image through the communication unit or the input / output interface, obtains feature information of the content and joint position information for the image using the first artificial intelligence model, and generates an input prompt including feature information of the content, the joint position information, style information to be applied to the emoticon, and a request for creating a mask image.
10. In paragraph 8, The above processor, An electronic device that controls the display to output text input by a user or another user viewing the same content as the content to a chat window displayed in an area of the screen on which the content is output based on the execution of the chat function, and generates the atmosphere of the chat window, feature information related to the content, and joint position information for the image based on the displayed text using the first artificial intelligence model.
11. In a method for creating an emoticon for an electronic device, A step of obtaining an emoticon to which the joint position information has been applied from a second artificial intelligence model based on joint position information of an image corresponding to the emoticon obtained from a first artificial intelligence model; A step of obtaining a dynamic emoticon converted from the emoticon based on the rigging point corresponding to the joint position information from the third artificial intelligence model for motion file application; and An emoticon generation method, comprising a step of outputting the dynamic emoticon obtained above.
12. In paragraph 11, The joint position information of the image corresponding to the above emoticon is An emoticon generation method comprising: style information to be applied to the emoticon, feature information related to content processed by at least one external device, and a request for generation of a mask image related to application of the motion file.
13. In paragraph 11, An emoticon generation method further comprising: when content to be provided to an external device connected to the electronic device is received, a step of processing the received content based on characteristic information of the external device and transmitting the received content to the external device; 14. In paragraph 12, A method for generating an emoticon, wherein the style information includes at least one of first style information for displaying the emoticon in a photorealistic style, second style information for displaying the emoticon in an animation style, and third style information for displaying the emoticon in a text style.
15. In paragraph 12, The above characteristic information is, An emoticon generation method comprising at least one of the type of the content, information on characters in the content, information on clothing worn by characters in the content, information on the appearance of the characters, and color information included in the content.
Citation Information
Patent Citations
Method and apparatus for mobile messenger service by using avatar
KR101719742B1
Deep learning based method and apparatus for the auto generation of character rigging
KR102437212B1
Video generating device and method therfor
KR102621814B1
Personalized avatar real-time motion capture
US20220383577A1
KR20210147654A