Electronic device, method, and non-transitory computer-readable storage medium for providing animation for character
The electronic device generates animation characters that speak and change facial expressions to match user emotions and context, addressing the lack of emotional expression in existing communication technologies, enhancing the intuitive and engaging nature of text-based interactions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing communication technologies lack the ability to intuitively express user emotions through dynamic character animations that correspond to the context and emotional content of user-generated text.
An electronic device and method that acquires context and emotion information from user input, generates an animation character that speaks and changes facial expressions to match the emotional content of the text, and inputs this character into the user interface for communication.
Enhances emotional expression in communication by providing dynamic character animations that reflect the user's emotions and context, improving the intuitive and engaging nature of text-based interactions.
Smart Images

Figure KR2025015989_15052026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transient computer-readable storage medium for providing animation of a character
[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for providing animation of a character.
[0002] Recently, the distribution of various types of portable electronic devices, such as smartphones, tablet PCs, wireless earphones, and smartwatches, is expanding.
[0003] The use of emojis, which can intuitively express users' various emotions, is increasing significantly in services such as text messaging and social network services (SNS) for communication with other users.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] According to one embodiment, the electronic device comprises: a display; a memory for storing instructions; and at least one processor including a processing circuitry; wherein, when the instructions are executed individually or collectively by the at least one processor, the electronic device acquires context information corresponding to the acquired text through an input UI (user interface) provided on the display, acquires user emotion information corresponding to each sentence unit of the text based on the context information, acquires an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes corresponding to each sentence unit of the text based on the emotion information, and allows the animation character to be input into the input UI.
[0006] A control method for an electronic device according to one embodiment may include: an operation of acquiring context information corresponding to the acquired text when text is acquired through an input UI; an operation of acquiring user emotion information corresponding to each sentence unit of the text based on the context information; an operation of acquiring an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes in correspondence with each sentence unit of the text based on the emotion information; and an operation of inputting the animation character into the input UI.
[0007] In a non-transient computer-readable medium storing instructions that cause the electronic device to perform an operation when executed by a processor of an electronic device according to one embodiment, the operation may include: an operation of obtaining context information corresponding to the obtained text when text is obtained through an input UI; an operation of obtaining user emotion information corresponding to each sentence unit of the text based on the context information; an operation of obtaining an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes corresponding to each sentence unit of the text based on the emotion information; and an operation of inputting the animation character into the input UI.
[0008] The above and other aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken together with the accompanying drawings.
[0009] FIG. 1 is a diagram illustrating the schematic operation of an electronic device according to one embodiment.
[0010] FIG. 2 illustrates an example of a block diagram of an electronic device according to one embodiment.
[0011] FIG. 3 is a flowchart illustrating the operation of an electronic device according to one embodiment.
[0012] FIG. 4 is a drawing for illustrating a generative artificial intelligence model according to an embodiment of the present disclosure.
[0013] FIGS. 5a to 5e are drawings for explaining a method of providing a UI screen according to one embodiment.
[0014] FIGS. 6a and 6b are drawings for explaining a method of providing a UI screen according to one embodiment.
[0015] FIGS. 7a to 7c are drawings for explaining a method of switching between an animation character and text according to one embodiment.
[0016] FIGS. 8a to 8c are drawings for illustrating a rapid response method according to one embodiment.
[0017] FIGS. 9a to 9c are drawings for explaining an audio response method according to one embodiment.
[0018] FIGS. 10a to 10c are drawings for illustrating a customized rapid response method according to one embodiment.
[0019] FIGS. 11 and FIGS. 12 are drawings for explaining a method for generating an animated character according to one embodiment.
[0020] FIGS. 13a to 13f are drawings for explaining a character selection method according to one embodiment.
[0021] FIGS. 14a and FIGS. 14b are drawings for explaining a method of providing a UI screen according to one embodiment.
[0022] FIGS. 15a and FIGS. 15b are drawings for explaining a voice personalization method for an animation character according to one embodiment.
[0023] FIGS. 16a and FIGS. 16b are drawings for illustrating a fast response method in a specific type of device according to one embodiment.
[0024] FIGS. 17a to 17c are drawings for illustrating an audio response method in a specific type of device according to one embodiment.
[0025] FIGS. 18a to 18e are drawings for explaining a method for creating an animation widget in a specific type of device according to one embodiment.
[0026] FIGS. 19a to 19c are drawings for explaining a method of providing interaction for a widget in a specific type of device according to one embodiment.
[0027] FIGS. 20a to 20c are drawings for explaining a method of providing an animation widget when a notification is received in a specific type device according to one embodiment.
[0028] FIGS. 21a and FIGS. 21b are drawings for explaining a method of generating an animated character in a specific type device according to one embodiment.
[0029] FIG. 22 is a drawing for explaining an application capable of supporting animation characters according to one embodiment.
[0030] FIG. 23 is a flowchart illustrating a method for generating an animated character based on received text according to one embodiment.
[0031] FIG. 24 is a block diagram of an electronic device in a network environment according to various embodiments.
[0032] The present disclosure will be described in detail below with reference to the attached drawings.
[0033] The terms used in the embodiments of this disclosure have been selected to be as widely used as possible, taking into account their functions within this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, or the emergence of new technologies. Additionally, in specific cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the description section of the disclosure. Therefore, terms used in this disclosure should be defined based on their meanings and the overall content of this disclosure, rather than merely their names (e.g., call, message, analyzed schedule).
[0034] In this specification, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of the above features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.
[0035] The expression "at least one of A and / or B" should be understood as representing either "A" or "B" or "A and B".
[0036] Expressions such as "first," "second," "first," or "second" used in this specification may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.
[0037] Where it is stated that a component (e.g., a first component) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., a second component), it should be understood that the component may be directly connected to the other component or connected through the other component (e.g., a third component).
[0038] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “consisting” are intended to specify the existence of the features, numbers, actions, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, actions, actions, components, parts, or combinations thereof.
[0039] In the embodiments, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of "modules" or a plurality of "parts" may be integrated into at least one module and implemented by at least one processor, except for a "module" or "part" that needs to be implemented in specific hardware.
[0040] In the present disclosure, the term "user" may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence electronic device).
[0041] The various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0042] Embodiments of the present disclosure will be described in more detail below with reference to the attached drawings.
[0043] FIG. 1 is a diagram for explaining the schematic operation of an electronic device (100) according to one embodiment.
[0044] According to one embodiment, the electronic device (100) may provide an execution screen of a preset application. For example, as illustrated in FIG. 1, the electronic device (100) may provide an execution screen of a messenger application. For example, the messenger application may be a platform application that allows users to send and receive messages.
[0045] According to one example, an electronic device (100) may provide an execution screen (10) of a messenger application as illustrated in FIG. 1. For example, the execution screen (10) of a messenger application may include a message display area (11), a text input window (12), a message sending button (13), a function button (14), and a virtual keyboard area (15). According to one example, the virtual keyboard area (15) may be provided by a platform separate from the third-party platform providing the messenger application (e.g., the platform (or operating system) of the electronic device (100)), but for convenience of explanation, it will be described as being included in the execution screen (10) of the messenger application.
[0046] According to one example, when a message is received through a messenger application, the received message (16) can be displayed in the message display area (11) as shown in FIG. 1.
[0047] According to one example, when a user of an electronic device (100) inputs a user's response (e.g., text, voice) to a received message (16), the user can generate an animated character (17) based on the input response and send it as a response to the received message (16). For example, the animated character (17) can express the emotions contained in the sentence through facial expressions while reproducing the pronunciation that matches the phonetic value of the sentence corresponding to the user's response and reading the sentence aloud.
[0048] Below, various embodiments for creating animated characters will be described.
[0049] FIG. 2 illustrates an example of a block diagram of an electronic device according to one embodiment.
[0050] According to various embodiments, the electronic device (100) of FIG. 2 may be at least partially similar to the electronic device (2401) of FIG. 24, or may include other embodiments of the electronic device.
[0051] In one embodiment, in terms of being owned by a user, the electronic device (100) may be referred to as a terminal (or user terminal). The terminal may include, for example, a personal computer (PC) such as a laptop and a desktop. The terminal may include, for example, a smartphone, a smartpad, and / or a tablet PC. The terminal may include smart accessories such as a smartwatch and / or a head-mounted device (HMD). According to one embodiment, the electronic device (100) may include a deformable housing. Based on the deformability, the housing of the electronic device (100) may be divided into a plurality of parts. According to one example, the electronic device (100) may be implemented as a user terminal (40) illustrated in FIG. 1.
[0052] According to one embodiment, the electronic device (100) may include at least one of a processor (110), a memory (120), a display (130), a communication circuit (140), a camera (150), a sensor (160), or a microphone (170). The processor (110), memory (120), display (130), communication circuit (140), camera (150), sensor (160), or microphone (170) may be electrically and / or operably coupled with each other by an electronic component such as a communication bus.
[0053] In one embodiment, the hardware of the electronic device (100) being operatively coupled may mean that a direct or indirect connection between the hardware is established via wired or wireless means so that the second hardware is controlled by the first hardware among the hardware. Although illustrated based on different blocks, the embodiment is not limited thereto, and some of the hardware of FIG. 2 (e.g., at least some of the processor (110), memory (120), and communication circuit (140)) may be included in a single integrated circuit, such as a system on a chip (SoC). The type and / or number of hardware included in the electronic device (100) is not limited to that shown in FIG. 2. For example, the electronic device (100) may include only some of the hardware components shown in FIG. 2.
[0054] According to one embodiment, the processor (110) of the electronic device (100) may include hardware for processing data based on one or more instructions. The hardware for processing data may include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), a field programmable gate array (FPGA), a central processing unit (CPU), a graphic processing unit (GPU), a neural processing unit (NPU), and / or an application processor (AP). The number of processors (110) may be one or more. For example, the processor (110) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.
[0055] The processor (110) can control the operations of the electronic device (100) by executing instructions stored in the memory (120). For example, the processor (110) may correspond to a plurality of processors that divide and collectively perform a plurality of operations among the processors.
[0056] A CPU (central processing unit) is a general-purpose processor capable of performing not only general operations but also artificial intelligence operations, and it can efficiently execute complex programs through a multi-layered cache structure. The CPU is advantageous for serial processing methods, which enable the organic linkage between previous and next calculation results through sequential computation. General-purpose processors are not limited to the examples mentioned above, except for cases specified as the aforementioned CPU.
[0057] A GPU (graphic processing unit) is a processor designed for massive computations, such as floating-point operations used in graphics processing, and can perform large-scale computations in parallel by integrating a large number of cores. In particular, GPUs may be advantageous over CPUs for parallel processing methods such as convolution operations. Additionally, GPUs can be used as co-processors to complement the functions of CPUs. Processors for massive computation are not limited to the examples mentioned above, except for cases specified as GPUs.
[0058] A Neural Processing Unit (NPU) is a processor specialized for artificial intelligence computations using artificial neural networks, and each layer constituting the neural network can be implemented in hardware (e.g., silicon). In this case, since the NPU is designed specifically according to the specifications required by the vendor, it has a lower degree of flexibility compared to CPUs or GPUs, but it can efficiently process the artificial intelligence computations required by the vendor. Meanwhile, as a processor specialized for artificial intelligence computations, the NPU can be implemented in various forms such as Tensor Processing Units (TPUs), Intelligence Processing Units (IPUs), and Vision Processing Units (VPUs). Artificial intelligence processors are not limited to the examples mentioned above, except for cases specified as the aforementioned NPU.
[0059] According to one embodiment, the memory (120) of the electronic device (100) may include a hardware component for storing data and / or instructions that are input and / or output to the processor (110). The memory (120) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). Volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disk, solid status drive (SSD), or embedded multimedia card (eMMC).
[0060] According to one embodiment, within the memory (120) of the electronic device (100), one or more instructions (or commands) representing operations and / or operations to be performed on data by the processor (110) may be stored. A set of one or more instructions may be referred to as firmware, an operating system, a process, a routine, a sub-routine, and / or an application. For example, the electronic device (100) and / or the processor (110) may perform various operations when a set of a plurality of instructions distributed in the form of an operating system, firmware, a driver, and / or an application is executed. In the following, the statement that an application is installed on an electronic device (100) means that one or more instructions provided in the form of an application are stored in the memory (120) of the electronic device (100), and that the one or more applications are stored in an executable format (e.g., a file having an extension specified by the operating system of the electronic device (100)) that is executable by the processor (110) of the electronic device (100).
[0061] One or more processors (110) control input data to be processed according to a predefined operation rule or AI model (artificial-intelligence model) stored in memory (120). The predefined operation rule or AI model is characterized by being created through learning. Being created through learning means that a predefined operation rule or AI model with desired characteristics is created by applying a learning algorithm to a number of learning data. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is performed, or it may be performed through a separate server / system.
[0062] An AI model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs the layer's operation through the result of the operation of the previous layer and at least one defined operation. Examples of neural networks include convolutional neural networks (CNN), recurrent neural networks (RNN), deep neural networks (DNN), restricted Boltzmann machines (RBM), deep belief networks (DBN), bidirectional recurrent deep neural networks (BRDNN), deep Q-networks, and Transformers; however, the neural networks in this disclosure are not limited to the aforementioned examples except where specified.
[0063] A learning algorithm is a method of training a specific target device (e.g., a robot) using a number of learning data to enable the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, and the learning algorithms in this disclosure are not limited to the aforementioned examples except where specified.
[0064] According to one embodiment, a display (130) of an electronic device (100) can output visualized information to a user. For example, the display (130) can be controlled by a controller, such as a GPU (graphic processing unit), to output visualized information to a user. The display (130) may include an OLED (organic light emitting diodes) display, an LED (light emitting diodes), a micro LED, a mini LED, a PDP (plasma display panel), a QD (quantum dot) display, a QLED (quantum dot light-emitting diodes) and / or an e-ink display and / or an e-paper display. According to one example, the display (130) may be implemented as a flat display, a curved display, a folding and / or rolling flexible display.
[0065] A communication circuit (140) of an electronic device (100) according to one embodiment may include hardware for supporting the transmission and / or reception of electrical signals between the electronic device (100) and an external device (e.g., a server). The communication circuit (140) may include, for example, at least one of a modem (modulator and demodulator), an antenna, and an optic / electronic converter. The communication circuit (140) may support the transmission and / or reception of electrical signals based on various types of protocols such as Ethernet, LAN (local area network), WAN (wide area network), WiFi (wireless fidelity), NFC (near field communication), Bluetooth, BLE (bluetooth low energy), ZigBee, LTE (long term evolution), 5G NR (new radio), and / or 6G.
[0066] According to one example, the electronic device (100) may be connected to a server and to each other based on a wired network and / or a wireless network. The wired network may include a network such as the Internet, a LAN (local area network), a WAN (wide area network), Ethernet, or a combination thereof. The wireless network may include a network such as LTE (long term evolution), 5G NR (new radio), WiFi (wireless fidelity), Zigbee, NFC (near field communication), Bluetooth, BLE (bluetooth low-energy), or a combination thereof. According to one example, the electronic device (100) and the server may be connected indirectly through an intermediate node within the network.
[0067] A camera (150) of an electronic device (100) according to one embodiment can convert a captured image into an electrical signal and generate image data based on the converted signal. For example, the camera (150) may include at least one of a general (or basic) camera, a depth camera, and an ultra-wide angle camera.
[0068] A sensor (160) of an electronic device (100) according to one embodiment can sense various user information. The sensor (160) can be implemented as various types of sensors capable of user sensing. For example, the sensor (160) may include at least one sensor among a time of flight (ToF) sensor, an ultrasonic sensor, a radio detection and ranging (RADAR) sensor, a photodiode sensor, a proximity sensor, a passive infrared (PIR) sensor, a pinhole sensor, a pinhole camera, an infrared human body detection sensor, a complementary metal oxide semiconductor (CMOS) image sensor, a thermal sensor, a light sensor, and a motion sensor.
[0069] The sensor (160) may include a touch sensor that detects touch actions, having a form such as a touch film, a touch sheet, or a touch pad.
[0070] The sensor (160) may include at least one of a CO2 sensor and an atmospheric pressure sensor. The CO2 sensor is a sensor for measuring carbon dioxide concentration. The atmospheric pressure sensor is a sensor for sensing ambient pressure.
[0071] The sensor (160) may further include at least one sensor capable of sensing ambient illuminance, ambient temperature, and the direction of incidence of light. In this case, the sensor (160) may be implemented as an illuminance sensor, a temperature sensing sensor, a light intensity sensing layer, and a camera.
[0072] The sensor (160) may further include at least one of an acceleration sensor (or gravity sensor), a geomagnetic sensor, and a gyro sensor. For example, the acceleration sensor may be a 3-axis acceleration sensor. The 3-axis acceleration sensor may measure gravitational acceleration by axis and provide raw data to the processor (140). The geomagnetic sensor or the gyro sensor may be used to obtain attitude information. Here, the attitude information may include at least one of roll information, pitch information, or yaw information.
[0073] A microphone (170) of an electronic device (100) according to one embodiment is configured to receive user voice or other sounds and convert them into audio data.
[0074] FIG. 3 is a flowchart illustrating the operation of an electronic device according to one embodiment.
[0075] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0076] According to one embodiment, operations 310 to 340 can be understood as being performed in the processor (110) of the electronic device (100).
[0077] According to FIG. 3, in operation 310, an electronic device (100) according to one embodiment can identify whether text is obtained through an input UI (or text input field).
[0078] According to one example, the input UI may be provided on the execution screen of an application. However, it is not limited thereto and may also be provided on inactive screens such as a lock screen or a background screen. For example, the input UI may be a UI provided by the platform (or operating system) of the electronic device (100) to receive user input.
[0079] According to one example, the input UI may provide at least one of a text input UI for inputting text or a voice input widget UI for inputting user voice.
[0080] According to one example, the electronic device (100) can acquire text entered through a text input UI or acquire text based on user voice entered through a voice input widget UI.
[0081] A text input UI (user interface) may be an interface element that allows a user to input text. In one example, the text input UI may include at least one of an input window, an input menu, or a keyboard. In one example, when a user clicks or touches the text input UI, a keyboard (e.g., a virtual keyboard) may be activated.
[0082] The voice input widget UI may be an interface element that allows a user to input voice. According to one example, the voice input widget UI may include at least one of a voice input button, a text conversion window, a cancel / confirm button, a language selection option, a volume control option, or a sensitivity control option.
[0083] When text is obtained through the input UI (310:Y), in operation 320, the electronic device (100) according to one embodiment can obtain context information corresponding to the entered text. When text is not obtained through the input UI (310:N), the electronic device (100) according to one embodiment may not perform subsequent operations.
[0084] Contextual information corresponding to the text may include background information, environmental information, and / or situational information necessary to more clearly understand the meaning the text intends to convey.
[0085] According to one example, the electronic device (100) can identify context information corresponding to the text based on information related to the text. For example, the electronic device (100) can identify context information corresponding to the input text based on at least one of the content included in the text, keywords, user profile information, the relationship between the user and the recipient to whom the text is sent (e.g., the recipient of the text message), or the context of the conversation with the recipient. For example, keywords may include words representing the core content of the text. For example, keywords may include at least one of words representing an object (e.g., subject), words representing the action of an object, words for describing an object or the action of an object, words representing a background or place, words representing time, words representing emotions or tastes, or words representing importance. For example, profile information may include at least one of the user's gender, age, address, or occupation.
[0086] According to one example, the electronic device (100) may identify context information corresponding to the text based on additional information other than information related to the text. For example, the electronic device (100) may obtain context information corresponding to the text based on at least one of the user's facial expression information obtained through a camera (150), the user's gesture information obtained through a sensor (160), the user's voice information obtained through a microphone (170), time information, location information, or biometric information obtained through a wearable device (e.g., watch, ring, earbuds). The gesture information may include at least one of the user's body gesture (e.g., movement of body structures such as the body, hand, arm, and leg), touch gesture, air gesture, and gesture on the electronic device (100) itself. For example, an air gesture may be a gesture of operating by waving or moving the hand without touching the screen. For example, a gesture on the electronic device (100) itself may be a gesture of moving the device itself. For example, gestures on the electronic device (100) itself may include at least one of tilting the device, shaking the device, flipping the device, or changing the orientation of the device. Additionally, further information may include at least one of habit information, app execution information, or work information. For example, habit information may include information on sleep start times, work times (e.g., reading), meal times, and exercise times based on usual habits. For example, app execution information may include information on the type of app used, execution time, and operating system. For example, ambient environment information may include information about the surrounding environment, such as weather information, temperature information, and light intensity information. For example, time information may include time-related information such as the current time, day of the week, and season. For example, location information may include at least one of the current location, a previously visited location, or a future planned location.For example, work information can be analyzed for calls, messages, and schedules, and may include workload.
[0087] According to one embodiment, the electronic device (100) can obtain context information by pre-processing data collected through a sensor (140) or a communication circuit (150). For example, pre-processing may be a process of cleaning and transforming raw data to analyze data using an artificial intelligence model. For example, the electronic device (100) may pre-process the collected data through at least one of data cleaning, data transformation, feature selection, data splitting, data augmentation, and data integration. For example, in the case of a face image, pre-processing such as image cropping centered on facial features, histogram equalization, gGamma correction, and Gaussian filtering may be performed. However, this is merely an example, and various pre-processing may be performed depending on the type of sensing data.
[0088] In operation 330, the electronic device (100) according to one embodiment may acquire user emotion information based on context information corresponding to text. This is not limited thereto, and if the text contains multiple sentences, user emotion information may be acquired on a sentence-by-sentence basis.
[0089] According to one embodiment, the electronic device (100) can obtain user emotion information corresponding to each sentence unit of the text based on context information corresponding to the text. The user emotion information may include at least one of happy, scared, sad, suprised, angry, disgusted, or neutral.
[0090] According to one example, the electronic device (100) can obtain user emotion information by inputting context information corresponding to text into an artificial intelligence model. According to one example, the electronic device (100) can obtain a first prompt based on context information corresponding to text and input the obtained first prompt into a first artificial intelligence model.
[0091] A prompt may refer to an input for initiating interaction with an artificial intelligence model that generates an image. The prompt may be a text input containing at least one word and / or at least one sentence. For example, the prompt may contain natural language text. The natural language text may contain various information that the artificial intelligence model can use to generate an image, such as an intent, a task, and / or constraints. The electronic device (100) may process the natural language text using a natural language processing (NLP) model.
[0092] According to one example, an electronic device (100) can generate a prompt requesting the creation of an image related to a notification by inserting a keyword into a placeholder of a prompt template. For example, a placeholder may be a location in the prompt template where the keyword is inserted. A placeholder to which the keyword is mapped may be pre-specified according to the attributes of the keyword. For example, the electronic device (100) can generate a prompt by inserting a keyword into a placeholder of a prompt template. For example, the electronic device (100) can generate a prompt by inserting a keyword into a placeholder of a prompt template and specifying each element of the prompt template into which the keyword is inserted.
[0093] According to one example, the electronic device (100) can obtain the user's emotional information by applying at least one of a pre-set algorithm or a pre-set formula to context information corresponding to text.
[0094] According to one embodiment, the electronic device (100) can acquire user emotion information that changes by sentence unit of text based on context information corresponding to text.
[0095] A sentence unit may include at least one of a word, phrase, clause, or word unit. A phrase may be a sentence component in which two or more words are combined to express a single meaning. A clause is a syntactic unit that possesses both a subject and a predicate, and may be a sentence or a part of another sentence. A word unit is a unit separated by spaces, and may be a unit that represents one unit of meaning at a time.
[0096] According to one example, an electronic device (100) can obtain user sentiment information for the same or different sentence composition units in a sentence. For example, the electronic device (100) can obtain user sentiment information based on context information corresponding to a first text corresponding to a first sentence composition unit included in the text, and can obtain user sentiment information based on context information corresponding to a second text corresponding to a second sentence composition unit entered after the first text. For example, the first sentence composition unit may be the same or different. For example, the electronic device (100) can obtain user sentiment information corresponding to each "phrase" unit in a sentence. For example, the electronic device (100) can obtain user sentiment information in "phrase" units and "clause" units in a sentence. For example, the electronic device (100) can obtain user sentiment information in "phrase" units and "clause" units respectively in the same text included in a sentence, and then select one of the obtained sentiment information. For example, the electronic device (100) can obtain user information in "phrase" units from a first text included in a sentence, and obtain user emotion information in "clause" units from a second text following the first text.
[0097] In operation 340, an electronic device (100) according to one embodiment may acquire animation elements based on emotional information. According to one example, the electronic device (100) may acquire animation elements that provide an animation effect of speaking text while changing facial expressions based on emotional information. For example, the electronic device (100) may acquire animation elements that provide a dynamic animation effect of speaking text while changing facial expressions in correspondence with each sentence composition unit. Animation elements may include various images including a mouth capable of making facial expressions and speaking, but for convenience of explanation, they will be collectively referred to as animation characters below.
[0098] A character may be an image designed to resemble at least one of a person, an animal, or a plant. However, it is not limited thereto, and various images including a mouth capable of making facial expressions and speaking may be used as the character of the present disclosure. For example, the character may include images such as emojis or emoticons.
[0099] Animation effects are techniques that make a character move. Accordingly, an animated character may be a character that moves. For example, an animated character may be a character in which at least one of the eyes, eyebrows, nose, mouth, lips, forehead, wrinkles, cheeks (or cheeks), or chin moves. However, it is not limited to this, and an animated character may also have at least one of the gestures or poses that move.
[0100] An animated character may be a character in which at least one of the eyes, nose, mouth, or facial expressions moves. However, it is not limited thereto, and an animated character may also have at least one of the gestures or poses moved.
[0101] According to one example, the electronic device (100) may acquire first emotional information based on a first text included in a sentence according to one example, and second emotional information based on a second text, and may acquire an animated character in which the character sequentially speaks the first text and the second text while the character's facial expression changes sequentially based on the first emotional information and the second emotional information. For example, the electronic device (100) may change at least one of the character's eyes, eyebrows, nose, mouth, lips, forehead, wrinkles, cheeks (or cheeks), or chin.
[0102] According to one embodiment, the electronic device (100) converts text into speech through TTS (text-to-speech), obtains mouth shape information in phoneme units according to the pronunciation of the speech, and obtains an animated character whose mouth shape changes based on the mouth shape information.
[0103] According to one example, the electronic device (100) can obtain an animated character by inputting text, user's changing emotional information, and character information into a second artificial intelligence model. For example, the electronic device (100) can obtain a second prompt based on text, user's changing emotional information, and character information, and input the obtained second prompt into the second artificial intelligence model.
[0104] According to one example, the electronic device (100) can obtain an animated character by inputting voice corresponding to text, user's changing emotional information, and character information into a second artificial intelligence model. For example, the electronic device (100) can obtain voice corresponding to text through TTS and input the voice corresponding to text, user's changing emotional information, and character information into the second artificial intelligence model. For example, the electronic device (100) can obtain a second prompt based on user's changing emotional information and character information and input the obtained second prompt and voice corresponding to text into the second artificial intelligence model. According to one example, the second artificial intelligence model can be implemented as a multimodal AI model. For example, the multimodal AI model can be implemented to simultaneously understand and analyze data of different formats, such as text, images, voice, and video.
[0105] According to one example, an electronic device (100) can obtain an animated character by applying at least one of a predefined rule, pattern, format, and template to text (or voice corresponding to text), user's changing emotion information, and character information.
[0106] According to one embodiment, the electronic device (100) may acquire an animated character in which the character speaks text based on the user's pronunciation feature information. The pronunciation feature information may be information regarding how an individual pronounces a particular language. Pronunciation features are influenced by various factors, and a unique pronunciation style may be formed for each individual. For example, the user's pronunciation feature information may include at least one of speed, intonation, tone, stress, phoneme pronunciation, linking, or assimilation.
[0107] According to one example, when using a second artificial intelligence model, the electronic device (100) can obtain an animated character by inputting the user's pronunciation feature information, voice (or text) corresponding to the text, the user's changing emotion information, and character information into the artificial intelligence model.
[0108] In operation 350, the electronic device (100) according to one embodiment can input an animated character.
[0109] According to one example, the electronic device (100) can input an animated character into the input UI. After the animated character is input into the input UI, if the send button is selected, the animated character can be transmitted as multi-video content to an application on the execution screen (e.g., a messenger application). For example, the input UI may be the same as and / or different from the input UI that receives text or voice input in operation 310, but it is collectively referred to as an input UI in that it is a UI provided by the platform (or operating system) of the electronic device (100) to receive user input.
[0110] According to one embodiment, the electronic device (100) can identify font information for font visualization based on emotional information.
[0111] According to one example, the electronic device (100) can identify at least one font information for font visualization for each sentence unit of text based on user sentiment information corresponding to each sentence unit of text.
[0112] According to one example, the electronic device (100) can provide visualization effects to the fonts for each sentence unit of text based on at least one font information. For example, the at least one font information may include at least one of font type information, font size information, font color information, font in / out information, font rotation information, or font emphasis information. For example, the electronic device (100) can display specific text with emphasis or display it with the size enlarged or reduced.
[0113] According to one embodiment, the electronic device (100) may provide a UI including a plurality of characters based on context information corresponding to text. According to one example, when one of the plurality of characters is selected, the electronic device (100) may obtain an animated character based on the selected character.
[0114] According to one embodiment, the electronic device (100) may acquire emotional information of a user based on at least one of the user's voice information acquired through a microphone (170), the user's facial expression information acquired through a camera (150), or the user's gesture information acquired through a sensor (160) while a widget is displayed on a display (130). According to one example, the electronic device (100) may acquire an animation widget that provides an animation effect in which the facial expression of a character included in the widget changes based on the user's emotional information. According to one example, the electronic device (100) may store the animation widget in memory (120) according to a user command. According to one example, the electronic device (100) may transmit the animation widget to an external source (e.g., a server, another electronic device).
[0115] According to one embodiment, when the electronic device (100) receives text input by a user of an external device into the external device, it can acquire emotional information of the user that changes for each sentence unit of the received text based on context information corresponding to the received text. According to one example, the electronic device (100) can acquire an animated character that provides an animation effect in which the character speaks text while the character's facial expression changes based on the changing emotional information. According to one example, the electronic device (100) can display the received text and the animated character on a display (130).
[0116] FIG. 4 is a drawing for illustrating a generative artificial intelligence model according to an embodiment of the present disclosure.
[0117] According to one embodiment, at least one of the first artificial intelligence model for acquiring emotional information or the second artificial intelligence model for generating an animated character may be implemented as a Generative AI Model (405) as shown in FIG. 4.
[0118] According to FIG. 4, the User Query / Response Interface (401) can receive user input. The user input may be in the form of natural language, images and / or videos.
[0119] For example, user input may include user voice received through a microphone. However, it is not limited thereto, and in addition to voice, user input may include text corresponding to the voice generated by a speech-to-text (STT) model. Furthermore, context information may also be transmitted along with the user input. Context information may include various additional information at the time of user input. For example, this may include information about the application currently being used by the user or the user's location information. Additionally, user input may take the form of a mixture of the aforementioned natural language, images, sounds, and context information. Furthermore, user input may also take the form of non-natural language input, such as menu selection.
[0120] The User Query / Response Interface (401) can output results from the Generative AI system to the user. The output can be in the form of natural language or specific content, and it can also be provided in the form of an action requested by the user. The User query interface (401) can output results from the Generative AI system to the user. The output can be in the form of natural language or specific content, and it can also be provided in the form of an action requested by the user. For example, the User Query / Response Interface (401) can output content generated by the Generative AI Model (405) based on voice received from the user.
[0121] The AI framework (402) can receive user input and coordinate and control each component necessary to perform user intent based on user query.
[0122] User input received from the User Query / Response Interface (401) can be sent to the Prompt design component (402-1). The Prompt design component (402-1) can be used to generate a prompt suitable for inputting user input into a large language model (LLM) or a large multimodal model (LMM). The Prompt design component (402-1) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. Based on user input, the Prompt design component (402-1) can generate a prompt by accessing knowledge repositories (1003) containing user preference data, a prompt library, and prompt examples, and can send the generated prompt to the LLM or LMM.
[0123] The API / Plug-in management component (402-2) can perform the role of communicating with external information when there is a request for additional information when passing user input as input to a generative model. The API / Plug-in management component (402-2) establishes a channel to communicate with the outside of the AI Interface via the API, and can enable access to various data sources through the established channel. Additionally, the API / Plug-in management component (402-2) can request an action via the API if the application or service needs to perform an action that executes the user input as a final step, rather than an intermediate result. Information obtained from the outside (e.g., the Applications / service component (404)) may be used to generate a prompt in the Prompt design component (402-1) along with the user input, or it may be passed as input to the generative model.
[0124] The Output modification Component (402-3) (or Refiner component) can fine-tune the output of the generative model. For example, the Output modification Component (402-3) can verify whether the content generated through the LLM and / or LMM is irrelevant, contains biased content, or contains harmful content. Additionally, the Output modification Component (402-3) can determine the extent to which the output matches the desired result and, if additional processing is required, proceed with that process. Furthermore, the Output modification Component (402-3) can configure and provide hints to the user to avoid unwanted output.
[0125] A Generative AI Model (1005) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A Generative AI Model (405) may include models that generate images and / or models that generate language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational autoencoders), and examples include Diffusion-based generative models that use VAEs and Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4 (e.g., CHAT-GPT 4o). There are also LMMs (large multimodal models) that can recognize various forms of data input, such as text, images, and voice, and generate new data corresponding to them.
[0126] According to an embodiment, when a prompt is input from the Prompt design component (402-1), the Generative AI Model (405) generates content corresponding to the prompt based on the instructions and can output the content through the User Query / Response Interface (401).
[0127] FIGS. 5a to 5e are drawings for explaining a method of providing a UI screen according to one embodiment.
[0128] According to one embodiment, the electronic device (100) may provide a screen for executing a pre-configured application. The pre-configured application may include at least one of a messenger application, a social network service (SNS) application, an email application, a calendar application, a gallery application, or a chat application. According to one example, at least one of a messenger application, a social network service (SNS) application, an email application, a calendar application, a gallery application, or a chat application may be provided by a platform (or operating system) of the electronic device (100) and a platform separate from the platform.
[0129] For the sake of convenience, the following explanation assumes that the execution screen of a pre-configured application is the execution screen of a messenger application. For example, a messenger application may be a platform application that allows users to send and receive messages.
[0130] According to one example, an electronic device (100) may provide an execution screen (510) of a messenger application as illustrated in FIG. 5a. For example, the execution screen (510) of the messenger application may include a message display area (511), a text input window (512), a message send button (513), a function button (514), and a virtual keyboard area (515). According to one example, the virtual keyboard area (515) may be provided by a platform separate from the platform of a third party providing the messenger application (e.g., the platform of the electronic device (100)). However, for the sake of convenience of explanation, the message display area (511), the text input window (512), the message send button (513), the function button (514), and the virtual keyboard area (515) will be described as being included in the execution screen (510) of the messenger application.
[0131] According to one example, when a message is received through a messenger application, the received message (511-1) can be displayed in the message display area (511) as shown in FIG. 5a.
[0132] According to one example, when the first button (514-1) for transmitting character animation among the function buttons (514) in FIG. 5a is selected (e.g., touch input), an input window (516) for character input, hereinafter, a character input window (516) may be provided as shown in FIG. 5b.
[0133] According to one example, when a text is entered into the text input window (512) in FIG. 5c and a specific type of character is selected, an animated character can be created based on the text and the selected character. For example, the animated character (518) may be in the form of a video that provides an animation effect of speaking text while changing facial expressions.
[0134] According to one example, when an animated character is created, a character transmission window (519) for transmitting the created animated character (518) may be provided as illustrated in FIG. 5d. For example, the character transmission window (519) may provide a preview of the animated character (518).
[0135] According to one example, when the message sending button (513) in FIG. 5d is selected, an animated character (518) can be displayed in the message display area (511) as it is delivered to the conversation partner, as shown in FIG. 5e. For example, the animated character (518) can be delivered from a keyboard application to a messenger application and then transmitted to the conversation partner's device through the messenger application.
[0136] FIGS. 6a and 6b are drawings for explaining a method of providing a UI screen according to one embodiment.
[0137] According to one embodiment, a menu for creating an animated character may be provided together with a specific UI menu. For example, the specific UI menu may include an example of an AI assistant menu.
[0138] According to one example, as illustrated on the left side of FIG. 6a, a menu UI (611) including a menu (611-1) for creating an animation character and AI assistant menus (611-2 to 611-4) may be provided in an application execution screen (610).
[0139] According to one example, when a menu (611-1) for creating a character animation is selected and text is entered into a text input window (612) as shown on the right side of FIG. 6a, a preview image of an animation character (614) corresponding to the text entered into a character input window (613) may be provided. For example, text input and the provision of a preview image may be implemented so that they cannot be done in parallel.
[0140] According to one embodiment, a UI menu may be deactivated while an animated character is being generated. For example, the UI menu that is deactivated may include a text input field.
[0141] According to one example, as shown on the left side of FIG. 6b, the text input window (621) and the completion button (622) provided on the screen (620) while the animated character is being created may be disabled.
[0142] According to one example, when an animated character is created as shown on the right side of FIG. 6b, a text input window (621) and a completion button (622) are activated and a preview image of the animated character (623) may be provided. For example, text input and the provision of the preview image may be implemented to be done in parallel (or simultaneously).
[0143] FIGS. 7a to 7c are drawings for explaining a method of switching between an animation character and text according to one embodiment.
[0144] According to one embodiment, when an electronic device (100) receives a message containing an animated character through a messenger application, it can play the animated character or display text corresponding to the animated character on an application execution screen (hereinafter, screen) according to user selection.
[0145] According to FIG. 7a, when an animated character is received, an animated character (711) and a text switching button (712) may be displayed on the screen (710). The text switching button (712) may be a menu button for deactivating the animated character (711) and displaying text corresponding to the animated character (711).
[0146] According to FIG. 7b, when a play button (713) for playing an animated character (711) is selected, the animated character (711) is played and text corresponding to the animated character (711) can be displayed. For example, the animated character (711) may be a video that speaks text while changing facial expressions.
[0147] According to FIG. 7c, when the button (712) for text switching is selected, the animated character (711) can be switched to display the disabled image (715) and text (716) of the animated character (711).
[0148] FIGS. 8a to 8c are drawings for illustrating a quick response method according to one embodiment.
[0149] According to one embodiment, the electronic device (100) can use a quick response function for a message received through a messenger application to generate an animated character as a response to the message and send it to the messenger application.
[0150] According to FIG. 8a, when a message is received through a messenger application and the received message (811) is displayed on the screen (810) according to one example, the user can execute a quick response function (or quick reply function) for the message (811) by inputting a specific touch operation (e.g., touch and drag).
[0151] According to FIG. 8b, in one example, when a quick response function is executed, the electronic device (100) may provide an area (814) containing multiple characters (814-1, 814-2, 813-3) to be recommended as an answer by understanding and analyzing the context of a received message (811) even if no text is entered in the text input window (812). In one example, the electronic device (100) may obtain multiple recommended characters (814-1, 814-2, 813-3) using an artificial intelligence model. For example, the electronic device (100) may display a guide (813) indicating that it is in a state of entering an answer to a received message (813).
[0152] According to FIG. 8c, when one of a plurality of characters (814-1, 814-2, 813-3) is selected, the electronic device (100) can transmit an animated character (816) generated based on the selected character (814) to a messenger application. For example, the electronic device (100) can display a received message (815) in which the animated character (816) is provided as a response for user confirmation.
[0153] FIGS. 9a to 9c are drawings for explaining an audio response method according to one embodiment.
[0154] According to one embodiment, the electronic device (100) can use a voice input function to generate an animated character in response to a message received through a messenger application and send it to the messenger application.
[0155] According to FIG. 9a, in one example, when an electronic device (100) receives a message through a messenger application and the received message (911) is displayed on a screen (910), it may provide a guide (912) indicating that the user is in a state to input a response to the received message (911). In one example, the electronic device (100) may receive voice input from a user using a voice input button (914).
[0156] According to FIG. 9b, in one example, the electronic device (100) can receive the voice “Thanks! That’s exactly what I needed today!” while the voice input button (914) is pressed, and display text corresponding to the received voice.
[0157] According to FIG. 9c, in one example, when voice input is received, the electronic device (100) can input text corresponding to the input voice into a text input window (913) and input a character (916) into a character input window (915). For example, a recommended character based on the input voice or a default character may be displayed. Afterward, the electronic device (100) can generate an animated character based on the text entered into the text input window (913) and send the generated animated character to a messenger application as a response to a received message (911).
[0158] FIGS. 10a to 10c are drawings for illustrating a quick response customization method according to one embodiment.
[0159] According to one embodiment, when one of a plurality of characters is selected for a message received through a messenger application, the electronic device (100) can provide a user-customized response based on the selected character.
[0160] According to FIG. 10a, in one example, when an electronic device (100) receives a message through a messenger application and the received message (1011) is displayed on a screen (1010), it may provide a guide (1012) indicating that a response to the received message (1011) is being entered. In one example, the electronic device (100) may provide a plurality of characters to be entered as a response to the received message (1011).
[0161] According to FIG. 10b, in one example, when one of a plurality of characters (1014) is selected, the electronic device (100) can input the selected character (1014) into a character input window (1015) and input default text corresponding to the selected character (105) into a text input window (1013). The default text may include at least one sentence that is mapped to each character.
[0162] According to FIG. 10c, according to one example, the electronic device (100) can generate and display an animated character (1016) based on a character (105) entered in a character input window (1015) and text entered in a text input window (1013).
[0163] FIGS. 11 and FIGS. 12 are drawings for explaining a method for generating an animated character according to one embodiment.
[0164] According to one example illustrated in FIG. 11, the electronic device (100) can identify a text entry. A text entry may refer to the process of inputting text into the electronic device (100). For example, the input text may include the sentence “Oh Wow! That's amazing!”.
[0165] According to one example, as illustrated in FIG. 11, an electronic device (100) can acquire emotional information based on context information of input text. For example, the emotional information may include at least one of happy, scared, sad, suprised, angry, disgusted, or neutral. For example, the electronic device (100) can acquire user emotional information that changes by sentence unit of text based on context information corresponding to the input text. For example, a sentence unit may include at least one of a word, a phrase, a clause, or a word element. It is not limited thereto, and if the text contains multiple sentences, user emotional information may be acquired by sentence unit.
[0166] According to one example, as illustrated in FIG. 11, the electronic device (100) can change the facial expression of a character based on emotional information. For example, the electronic device (100) can change at least one of the character's eyes, eyebrows, nose, mouth, lips, forehead, wrinkles, cheeks (or cheeks), or chin.
[0167] According to one example, as illustrated in FIG. 11, an electronic device (100) can convert text input via TTS (text-to-speech) into speech and obtain mouth shapes in phoneme units according to the pronunciation of the speech. A phoneme may be the smallest unit of sound that distinguishes the meaning of a word in the speech system of a language.
[0168] According to one example, the electronic device (100) can convert input text into speech based on the user's pronunciation feature information. The pronunciation feature information may be information regarding how an individual pronounces a particular language when speaking it. For example, the user's pronunciation feature information may include at least one of speed, intonation, tone, stress, phoneme pronunciation, linking, or assimilation.
[0169] According to one example, the electronic device (100) can acquire mouth shapes in phoneme units based on the user's pronunciation feature information. For example, corresponding mouth shape information for each phoneme unit may be stored in the electronic device (100) (e.g., lookup table) or acquired through a pre-set rule (e.g., mouth shape conversion rule). For example, different mouth shapes may be stored for the same phoneme according to the user's pronunciation feature information.
[0170] According to one example, the electronic device (100) can acquire a mouth shape in phoneme units based on the user's mouth shape feature information. For example, the electronic device (100) can acquire different mouth shapes for the same phoneme according to the mouth shape feature information. The mouth shape feature information may include at least one of the position of the corners of the mouth, the shape of the lips, or the size of the mouth opening.
[0171] According to one example, as illustrated in FIG. 11, an electronic device (100) can obtain an animated character based on facial expressions and mouth shapes corresponding to each sentence unit. For example, as illustrated in FIG. 12, if the input text contains multiple sentences (or paragraphs) (e.g., Sentence A, Sentence B, Sentence C), the animated character can sequentially play animations of emotions corresponding to each sentence.
[0172] FIGS. 13a to 13f are drawings for explaining a character selection method according to one embodiment.
[0173] According to one embodiment, the electronic device (100) may provide a UI screen that allows selecting a basic character used to create an animation character.
[0174] According to one example, when a messenger application is launched, a UI screen (1310) including a text input window (1311), a character input window (1312), a character switching button (1314), a preview button (1315), and a voice input button (1316) may be provided as illustrated in FIG. 13a. The character switching button (1314) may be a button for changing the type of character. According to one example, a character (1313-1) may be displayed in the character input window (1312). In this case, the preview button (1315) may be disabled because the animated character has not yet been created.
[0175] According to one example, the electronic device (100) can generate an animated character based on text when text is entered into a text input window (1311) as shown in FIG. 13b. In this case, the preview button (1315) may be displayed in a loading state indicating that the animated character (1313-2) is being generated. For example, the voice input button (1316) may be changed to a character animation transmission button (1317), but the animation character transmission button (1317) may be disabled while the animated character is being generated.
[0176] According to one example, when the creation of the animated character (1313-3) is completed as shown in FIG. 13c, the preview button (1315) and the character transfer button (1317) can be activated.
[0177] According to one example, when the character transmission button (1317) is selected as shown in FIG. 13d, the generated animated character (1318) can be transmitted. When the animated character (1318) is transmitted, a voice input button (1316) may be displayed instead of the character transmission button (1317).
[0178] According to one example, when a user's voice is input using the voice input button (1316) as shown in FIG. 13e, an animated character (1318) can be generated based on the user's voice.
[0179] According to one example, when a character type is selected via the character switching button (1314) as shown in FIG. 13f, the type of the basic character for creating an animated character can be changed to the character of the selected type (1319).
[0180] FIGS. 14a and FIGS. 14b are drawings for explaining a method of providing a UI screen according to one embodiment.
[0181] According to one embodiment, the electronic device (100) can provide a UI screen that allows a user to select a desired character among various basic characters for creating an animation character.
[0182] According to one example, an electronic device (100) may provide a UI screen (1410) including a text input window (1411) and a character input window (1412) as shown in FIG. 14a. For example, when text is entered into the text input window (411), the electronic device (100) may provide a plurality of characters in the character input window (1412). For example, as shown in FIG. 14a, the electronic device (100) may provide category information (e.g., Excited, Formal, Neutral) corresponding to emotional information and provide a character corresponding to the corresponding emotional category according to the user's selection.
[0183] According to one example, the electronic device (100) can transmit an animated character (1413) generated based on a character selected in a character input window (1412) and text entered in a text input window (1411), as shown in FIG. 14b.
[0184] FIGS. 15a and FIGS. 15b are drawings for explaining a voice personalization method for an animation character according to one embodiment.
[0185] According to one embodiment, the electronic device (100) can generate the spoken voice of an animated character based on the user's personalized voice. For example, the spoken voice of an animated character can be generated based on the user's voice characteristics. For example, voice characteristics may include at least one of speed, intonation, stress, phonemic pronunciation, linking, dialect, or assimilation. For example, the electronic device (100) can obtain the user's voice characteristics by using a previously stored user voice recording or by recording the user's voice through a separate function.
[0186] According to one example, the electronic device (100) may provide a UI screen (1510) for recording a user's voice as illustrated in FIG. 15a. For example, the UI screen (1510) may include a user avatar (1511), a guide for voice recording (1512), sample text (1513), a record button (1514), a retry button (1515-1), and a next button (1515-2).
[0187] According to one example, the electronic device (100) may provide a UI screen (1520) for setting personalized voice characteristics as illustrated in FIG. 15b. The UI screen (1520) may include a user avatar (1512), a text input window (1522), an expression characteristic control item (1523), a Skip button (1524-1), and a Done button (1524-2). For example, the expression characteristic control item (1523) may include control options for adjusting expressive traits such as Speed, Exictiement, and Pitch, respectively.
[0188] FIGS. 16a and FIGS. 16b are drawings for illustrating a quick response method in a specific type of device according to one embodiment.
[0189] According to one embodiment, the electronic device (100) may be implemented as a flip-type device. According to one example, the electronic device (100) may generate and transmit an animated character in response to a received message using a quick response function on a flip cover screen. The flip cover screen may be a screen provided through the front display of a flip cover that covers the main display of the flip-type electronic device (100).
[0190] According to FIG. 16a, when a notification (1611) corresponding to a received message is displayed on a flip cover screen (1610) according to one example, a quick response function (or quick reply function) can be executed through interaction with the notification (1611).
[0191] According to FIG. 16b, in one example, when a quick response function is executed, the electronic device (100) may provide a UI (1612) that includes a plurality of recommended characters to be sent as a response, and displays full text (1611-1) corresponding to the notification (1611).
[0192] FIGS. 17a and 17b are drawings for explaining an audio response method in a specific type of device according to one embodiment.
[0193] According to one embodiment, the electronic device (100) may be implemented as a flip-type device. According to one example, the electronic device (100) may generate and transmit an animated character in response to a received message using an audio response function on a flip cover screen.
[0194] According to FIG. 17a, according to one example, the electronic device (100) can display a received message (1711) on a flip cover screen (1710).
[0195] According to FIG. 17b, in one example, when a user's voice input is received by using a voice input button (1713) on a flip cover screen (1710), an electronic device (100) can generate an animated character using the input voice and a basic character (1712).
[0196] According to FIG. 17c, according to one example, the electronic device (100) can transmit an animated character (1714) in response to a message (1711) received on a flip cover screen (1710).
[0197] FIGS. 18a to 18e are drawings for explaining a method for creating an animation widget in a specific type of device according to one embodiment.
[0198] According to one embodiment, the electronic device (100) may be implemented as a flip-type device. According to one example, the electronic device (100) may provide a widget on a flip cover screen and generate an animated widget based on user input. The widget may be a small application that provides specific information or functions. For example, the user can use the widget to quickly access real-time data or frequently used functions, such as weather updates, a calendar, and a news feed.
[0199] According to FIG. 18a, according to one example, the electronic device (100) can display a widget (1811), a voice input button (1812), and a camera button (1813) on a flip cover screen (1810).
[0200] According to FIG. 18b, in one example, the electronic device (100) may provide a preset animation for the widget (1811) based on a preset event while the widget is in an idle state. An idle state may be a state in which the widget is active but not interacting with the user. For example, the electronic device (100) may provide a preset animation for the widget (1811) based on an alarm and / or reminder.
[0201] According to FIG. 18c, according to one example, the electronic device (100) can receive a user's voice using a voice input button (1812) while a widget (1811) is provided on a flip cover screen (1810).
[0202] According to FIG. 18d, in one example, the electronic device (100) can generate an animation widget based on a widget (1811) based on received user voice. For example, the electronic device (100) can provide a close button (1414) and a next button (1415). For example, the electronic device (100) can use the camera button (1813) to additionally receive information such as facial expression information and / or gesture information to generate the animation widget.
[0203] According to FIG. 18e, in one example, the electronic device (100) may provide a UI (1816) including applications using the generated animation widget. For example, a user may use the animation widget by selecting a desired application.
[0204] FIGS. 19a to 19c are drawings for explaining a method of providing interaction for a widget in a specific type of device according to one embodiment.
[0205] According to one embodiment, the electronic device (100) can provide a widget on a flip cover screen and provide playful interactions to the widget based on the type of user input.
[0206] According to FIG. 19a, in one example, when a first type input is received on a flip cover screen (1910), the electronic device (100) may provide a first interaction (1911) for a widget. For example, when a swipe input is received, the electronic device (100) may provide a widget based on a first expression.
[0207] According to FIG. 19b, in one example, when a second type input is received on a flip cover screen (1910), the electronic device (100) may provide a second interaction (1912) for a widget. For example, when a swipe input is received, the electronic device (100) may provide a widget based on a second expression.
[0208] According to FIG. 19c, in one example, when a third type input is received on a flip cover screen (1910), the electronic device (100) may provide a third interaction (1913) for a widget. For example, when a swipe input is received, the electronic device (100) may provide a widget based on a third expression.
[0209] FIGS. 20a to 20c are drawings for explaining a method of providing an animation widget when a notification is received in a specific type device according to one embodiment.
[0210] According to one embodiment, the electronic device (100) can provide a widget on a flip cover screen and, when a notification is received, provide an animation widget corresponding to the received notification.
[0211] According to FIG. 20a, according to one example, the electronic device (100) can provide a widget (2011) on a flip cover screen (2010).
[0212] According to FIG. 20b, when a portion of a received notification (2012) is displayed on a flip cover screen (2010) according to one example, the widget (2011) may provide an animation corresponding to the notification (2012). For example, the widget (2011) may provide an animation of the notification (2012) being spoken while changing its expression to correspond to the context of the notification (2012).
[0213] According to FIG. 20c, when all received notifications (2012) are displayed on the flip cover screen (2010) according to one example, a guide (2013) can be provided to inquire whether to run an application corresponding to the notification (2012).
[0214] FIGS. 21a and FIGS. 21b are drawings for explaining a method of generating an animated character in a specific type device according to one embodiment.
[0215] According to one embodiment, the electronic device (100) can generate an animated character by tracking a user's facial expression based on a captured image obtained through a camera (150).
[0216] According to FIGS. 21a and 21b, according to one example, an electronic device (100) can track a user's facial expression and generate an animation character corresponding to the user's facial expression (2111, 2112) that changes over time, and provide it to a flip cover screen (2110).
[0217] FIG. 22 is a drawing for explaining an application capable of supporting animation characters according to one embodiment.
[0218] According to one embodiment, animated characters can be provided in various types of applications. According to one example, various types of applications may include at least one of a messenger application, an SNS application, an email application, a calendar application, a gallery application, or a chat application. According to one example, various types of applications may include third-party applications.
[0219] According to FIG. 22, in one example, an electronic device (100) can create an animated character using a content creation tool provided by a third-party application (2200) that provides SNS services. The third-party application may be an application created by an external developer or company that is not the official developer of the platform or operating system of the electronic device (100).
[0220] FIG. 23 is a flowchart illustrating a method for generating an animated character based on received text according to one embodiment.
[0221] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0222] According to one embodiment, operations 2310 to 2340 can be understood as being performed in the processor (110) of the electronic device (100).
[0223] According to FIG. 23, in operation 2310, an electronic device (100) according to one embodiment can identify whether text entered by a user of an external device into the external device is received.
[0224] When it is identified that text has been received (2310:Y), in operation 2320, the electronic device (100) according to one embodiment can obtain user sentiment information corresponding to each sentence unit of the received text based on context information corresponding to the received text.
[0225] According to one example, the electronic device (100) can acquire the user's emotional information in the same / similar way as operation 320 of FIG. 3.
[0226] If it is identified that the text has not been received (2310:N), the electronic device (100) according to one embodiment may not perform any subsequent operations.
[0227] In operation 2330, the electronic device (100) according to one embodiment can obtain an animated character that provides an animation effect of speaking text while changing the character's facial expression based on emotional information.
[0228] According to one example, the electronic device (100) can obtain the user's emotional information in the same / similar way as operation 330.
[0229] In operation 2340, an electronic device (100) according to one embodiment may display received text and an animated character. According to one example, the electronic device (100) may display the received text as well as an animated character generated based on the received text in the message display area (511) shown in FIG. 5.
[0230] FIG. 24 is a block diagram of an electronic device (2401) in a network environment (2400) according to various embodiments. The electronic device (2401) may be implemented as the electronic device (100) shown in FIG. 2 according to one example.
[0231] Referring to FIG. 24, in a network environment (2400), an electronic device (2401) may communicate with an electronic device (2402) through a first network (2498) (e.g., a short-range wireless communication network) or with at least one of an electronic device (2404) or a server (2408) through a second network (2499) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (2401) may communicate with the electronic device (2404) through a server (2408). According to one embodiment, the electronic device (2401) may include a processor (2420), memory (2430), input module (2450), sound output module (2455), display module (2460), audio module (2470), sensor module (2476), interface (2477), connection terminal (2478), haptic module (2479), camera module (2480), power management module (2488), battery (2489), communication module (2490), subscriber identification module (2496), or antenna module (2497). In some embodiments, at least one of these components (e.g., connection terminal (2478)) may be omitted from the electronic device (2401), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (2476), camera module (2480), or antenna module (2497)) may be integrated into a single component (e.g., display module (2460)).
[0232] The processor (2420) can, for example, execute software (e.g., program (2440)) to control at least one other component (e.g., hardware or software component) of the electronic device (2401) connected to the processor (2420) and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (2420) can store commands or data received from other components (e.g., sensor module (2476) or communication module (2490)) in volatile memory (2432), process the commands or data stored in volatile memory (2432), and store the resulting data in non-volatile memory (2434). According to one embodiment, the processor (2420) may include a main processor (2421) (e.g., a central processing unit or an application processor) or an auxiliary processor (2423) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (2401) includes a main processor (2421) and an auxiliary processor (2423), the auxiliary processor (2423) may be configured to use lower power than the main processor (2421) or to be specialized for a specified function. The auxiliary processor (2423) may be implemented separately from the main processor (2421) or as part thereof.
[0233] The auxiliary processor (2423) may control at least some of the functions or states associated with at least one component of the electronic device (2401) (e.g., display module (2460), sensor module (2476), or communication module (2490)) on behalf of the main processor (2421) while the main processor (2421) is in an inactive (e.g., sleep) state, or together with the main processor (2421) while the main processor (2421) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (2423) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (2480) or communication module (2490)). According to one embodiment, the auxiliary processor (2423) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (2401) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (2408)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, an artificial intelligence model may include a software structure, either additionally or substantially.
[0234] The memory (2430) can store various data used by at least one component of the electronic device (2401) (e.g., processor (2420) or sensor module (2476)). The data may include, for example, input data or output data for software (e.g., program (2440)) and related commands. The memory (2430) may include volatile memory (2432) or non-volatile memory (2434).
[0235] The program (2440) may be stored as software in memory (2430) and may include, for example, an operating system (1442), middleware (1444), or an application (1446).
[0236] The input module (2450) can receive commands or data to be used for a component of the electronic device (2401) (e.g., processor (2420)) from outside the electronic device (2401) (e.g., user). The input module (2450) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0237] The sound output module (2455) can output a sound signal to the outside of the electronic device (2401). The sound output module (2455) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0238] The display module (2460) can visually provide information to an external (e.g., user) of the electronic device (2401). The display module (2460) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (2460) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0239] The audio module (2470) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (2470) can acquire sound through the input module (2450) or output sound through the sound output module (2455) or an external electronic device (e.g., electronic device (2402)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (2401).
[0240] The sensor module (2476) can detect the operating state of the electronic device (2401) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (2476) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0241] The interface (2477) may support one or more specified protocols that can be used for the electronic device (2401) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (2402)). According to one embodiment, the interface (2477) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0242] The connection terminal (2478) may include a connector through which the electronic device (2401) can be physically connected to an external electronic device (e.g., electronic device (2402)). According to one embodiment, the connection terminal (2478) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0243] The haptic module (2479) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (2479) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0244] The camera module (2480) can capture still images and video. According to one embodiment, the camera module (2480) may include one or more lenses, image sensors, image signal processors, or flashes.
[0245] The power management module (2488) can manage power supplied to the electronic device (2401). According to one embodiment, the power management module (2488) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0246] The battery (2489) can supply power to at least one component of the electronic device (2401). According to one embodiment, the battery (2489) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0247] The communication module (2490) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (2401) and an external electronic device (e.g., electronic device (2402), electronic device (2404), or server (2408)), and the performance of communication through the established communication channel. The communication module (2490) may include one or more communication processors that operate independently of the processor (2420) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (2490) may include a wireless communication module (2492) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (1494) (e.g., LAN (local area network) communication module, or power line communication module). Among these communication modules, the communication module described above can communicate with an external electronic device (2404) through a first network (2498) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (2499) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (2492) can identify or authenticate the electronic device (2401) within a communication network such as the first network (2498) or the second network (2499) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (2496).
[0248] The wireless communication module (2492) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (2492) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (2492) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (2492) can support various requirements specified in the electronic device (2401), external electronic device (e.g., electronic device (2404)), or network system (e.g., second network (2499)). According to one embodiment, the wireless communication module (2492) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0249] An antenna module (2497) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (2497) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (2497) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (2498) or a second network (2499), may be selected from the plurality of antennas, for example, by a communication module (2490). A signal or power may be transmitted or received between the communication module (2490) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (2497).
[0250] According to various embodiments, the antenna module (2497) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0251] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0252] According to one embodiment, commands or data may be transmitted or received between the electronic device (2401) and an external electronic device (2404) through a server (2408) connected to a second network (2499). Each of the external electronic devices (2402, or 2404) may be the same or a different type of device as the electronic device (2401). According to one embodiment, all or part of the operations performed on the electronic device (2401) may be performed on one or more of the external electronic devices (2402, 2404, or 2408). For example, if the electronic device (2401) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (2401) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (2401). The electronic device (2401) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (2401) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (2404) may include an Internet of Things (IoT) device. The server (2408) may be an intelligent server using machine learning and / or neural networks.According to one embodiment, an external electronic device (2404) or server (2408) may be included within the second network (2499). The electronic device (2401) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0253] According to one embodiment, the electronic device comprises: a display; a memory for storing instructions; and at least one processor including a processing circuitry; wherein, when the instructions are executed individually or collectively by the at least one processor, the electronic device acquires context information corresponding to the acquired text through an input UI (user interface) provided on the display, acquires user emotion information corresponding to each sentence unit of the text based on the context information, acquires an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes corresponding to each sentence unit of the text based on the emotion information, and allows the animation character to be input into the input UI.
[0254] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may acquire a user's first emotion information based on context information corresponding to a first text corresponding to a first sentence unit included in the text, acquire a user's second emotion information based on context information corresponding to a second text corresponding to a second sentence unit acquired after the first text, and acquire an animation character in which the character sequentially utters the first text and the second text while the character's facial expression changes sequentially based on the first emotion information and the second emotion information.
[0255] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may identify context information corresponding to the text based on at least one of a keyword included in the text, the user's profile information, the relationship between the user and the recipient to whom the text is transmitted, and the conversation context with the recipient.
[0256] According to one embodiment, the electronic device further comprises a microphone; a camera; and a sensor; and when the instructions are executed individually or collectively by the at least one processor, the electronic device may obtain context information corresponding to the text based on at least one of the user's voice information obtained through the microphone, the user's facial expression information obtained through the camera, or the user's gesture information obtained through the sensor.
[0257] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may input at least one of the user's voice information, the user's facial expression information, and the user's gesture information, and the text into a learned artificial intelligence model to obtain the animation character.
[0258] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device identifies at least one font information for font visualization of each sentence unit of the text based on user sentiment information corresponding to each sentence unit of the text, and provides a visualization effect on the font of each sentence unit of the text input to the input UI based on the at least one font information, and the at least one font information may include at least one of font type information, font size information, font color information, font in / out information, font rotation information, or font emphasis information.
[0259] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device obtains the animation character in which the character speaks the text based on the user's pronunciation feature information, and the user's pronunciation feature information may include at least one of speed, intonation, stress, phoneme pronunciation, linking, or assimilation.
[0260] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may convert the text into speech through TTS (text-to-speech), acquire mouth shape information in phoneme units according to the pronunciation of the speech, and acquire the animation character whose mouth shape changes based on the mouth shape information.
[0261] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may provide a UI including a plurality of characters based on the context information of the text, and when one of the plurality of characters is selected, the animation character may be obtained based on the selected character.
[0262] According to one embodiment, the input UI may provide at least one of a text input UI for inputting text or a voice input widget UI for inputting user voice. When the instructions are executed individually or collectively by the at least one processor, the electronic device may acquire the text entered through the text input UI or acquire the text based on the user voice entered through the voice input widget UI.
[0263] According to one embodiment, the electronic device further comprises a microphone; a camera; and a sensor; and when the instructions are executed individually or collectively by the at least one processor, the electronic device may acquire emotional information of the user based on at least one of the user’s voice information acquired through the microphone, the user’s facial expression information acquired through the camera, and the user’s gesture information acquired through the sensor while a widget is displayed on the display, acquire an animation widget that provides an animation effect in which the facial expression of a character included in the widget changes based on the emotional information of the user, and store the animation widget in the memory or transmit it outside the electronic device according to a user command.
[0264] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may, upon receiving text from an external device, acquire user emotion information corresponding to each sentence unit of the received text based on context information corresponding to the received text, acquire an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes corresponding to each sentence unit of the text based on the emotion information, and display the received text and the animation character on the display.
[0265] A control method for an electronic device according to one embodiment may include: an operation of acquiring context information corresponding to the acquired text when text is acquired through an input UI; an operation of acquiring user emotion information corresponding to each sentence unit of the text based on the context information; an operation of acquiring an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes in correspondence with each sentence unit of the text based on the emotion information; and an operation of inputting the animation character into the input UI.
[0266] According to one embodiment, the operation of acquiring the user's emotional information may include: acquiring the user's first emotional information based on context information corresponding to a first text corresponding to a first sentence unit included in the text, and acquiring the user's second emotional information based on context information corresponding to a second text corresponding to a second sentence unit acquired after the first text; and acquiring the animation character in which the character sequentially utters the first text and the second text while the character's facial expression changes sequentially based on the first emotional information and the second emotional information.
[0267] According to one embodiment, the operation of acquiring the user's emotional information may include an operation of identifying context information corresponding to the acquired text based on at least one of keywords included in the text, the user's profile information, the relationship between the user and the recipient to whom the text is transmitted, and the conversation context with the recipient.
[0268] According to one embodiment, the operation of acquiring the user's emotion information may include acquiring context information corresponding to the text based on at least one of the user's voice information acquired through a microphone, the user's facial expression information acquired through a camera, or the user's gesture information acquired through a sensor.
[0269] According to one embodiment, the operation of acquiring the animation character may include at least one of the user's voice information, the user's facial expression information, and the user's gesture information, and the operation of acquiring the animation character by inputting the text into a learned artificial intelligence model.
[0270] According to one embodiment, the control method further comprises: identifying at least one font information for font visualization for each sentence unit of the text based on user sentiment information corresponding to each sentence unit of the text; and providing a visualization effect on the font for each sentence unit of the text input to the input UI based on the at least one font information; wherein the at least one font information may include at least one of font type information, font size information, font color information, font in / out information, font rotation information, or font emphasis information.
[0271] According to one embodiment, the operation of acquiring the animation character includes the operation of acquiring the animation character in which the character utters the text based on the pronunciation feature information of the user; wherein the pronunciation feature information of the user may include at least one of speed, intonation, stress, phoneme pronunciation, linking, or assimilation.
[0272] In a non-transient computer-readable medium storing instructions that cause the electronic device to perform an operation when executed by a processor of an electronic device according to one embodiment, the operation may include: an operation of obtaining context information corresponding to the obtained text when text is obtained through an input UI; an operation of obtaining user emotion information corresponding to each sentence unit of the text based on the context information; an operation of obtaining an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes corresponding to each sentence unit of the text based on the emotion information; and an operation of inputting the animation character into the input UI.
[0273] According to the various embodiments described above, by delivering an animated character that reads a sentence while expressing the user's emotions, it becomes possible to express nuances that cannot be accurately conveyed through text messages alone. Accordingly, emotionally rich and accurate communication including active and passive input information becomes possible.
[0274] Although the various embodiments described above use multiple individual neural network models, the operation of at least two of the multiple neural network models may be implemented in a single neural network model.
[0275] Each operation according to the various embodiments described above may be performed by the processor (110), but if necessary, a module for each operation may be used. For example, each module may be implemented with at least one software, at least one hardware, and / or a combination thereof. Each module may be implemented to use a predefined algorithm, a predefined formula, and / or a learned artificial intelligence model to perform the operation. However, at least some modules may be distributed to external devices.
[0276] The methods according to the various embodiments of the present disclosure described above may be implemented in the form of an application that can be installed on an existing electronic device. Alternatively, the methods according to the various embodiments of the present disclosure described above may be performed using a deep learning-based artificial neural network (or deep artificial neural network), that is, a learning network model.
[0277] The methods according to the various embodiments of the present disclosure described above can be implemented by software upgrades or hardware upgrades alone for existing electronic devices.
[0278] The various embodiments of the present disclosure described above may also be performed through an embedded server equipped in an electronic device or an external server of the electronic device.
[0279] According to a specific example of the present disclosure, the various embodiments described above may be implemented as software comprising instructions stored on a machine-readable storage medium (e.g., a computer). The machine may include an electronic device (e.g., electronic device (A)) according to the disclosed embodiments, which is a device capable of calling instructions stored from the storage medium and operating according to the called instructions. When instructions are executed by a processor, the processor may perform a function corresponding to the instructions directly or by using other components under the control of the processor. Instructions may include code generated or executed by a compiler or an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" means only that the storage medium does not contain a signal and is tangible, and does not distinguish whether data is stored semi-permanently or temporarily in the storage medium.
[0280] Additionally, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a storage medium such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0281] Additionally, each component (e.g., module or program) according to the various embodiments described above may be composed of a single or multiple entities, and some of the aforementioned sub-components may be omitted, or other sub-components may be further included in the various embodiments. Generally or additionally, some components (e.g., module or program) may be integrated into a single entity to perform the functions performed by each of the respective components prior to integration in the same or similar manner. The operations performed by the module, program, or other components according to the various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations added.
[0282] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
Claims
1. In an electronic device (100), Display (130); Memory (120) for storing instructions; and It includes at least one processor (110) including a processing circuitry; and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, When text is obtained through the input UI (user interface) provided on the above display, context information corresponding to the obtained text is obtained, and Based on the above context information, user emotion information corresponding to each sentence unit of the above text is obtained, and Based on the above emotional information, an animation character is obtained that provides an animation effect in which the character speaks the text while the character's facial expression changes in correspondence with each sentence unit of the text, and An electronic device that inputs the above-mentioned animated character into the above-mentioned input UI.
2. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, A user's first emotion information is obtained based on context information corresponding to a first text corresponding to a first sentence unit included in the above text, and a user's second emotion information is obtained based on context information corresponding to a second text corresponding to a second sentence unit obtained after the first text. An electronic device that enables the acquisition of an animation character in which the character sequentially changes facial expressions based on the first emotional information and the second emotional information, and sequentially speaks the first text and the second text.
3. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device for identifying context information corresponding to the text based on at least one of keywords included in the text, profile information of the user, the relationship between the user and the recipient to whom the text is transmitted, and the context of a conversation with the recipient.
4. In Paragraph 1 or 2, The above electronic device is, mike; Camera; and It further includes a sensor; When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that obtains context information corresponding to the text based on at least one of the voice information of the user obtained through the microphone, facial expression information of the user obtained through the camera, or gesture information of the user obtained through the sensor.
5. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that inputs at least one of the voice information of the user, the facial expression information of the user, and the gesture information of the user, and the text into a learned artificial intelligence model to obtain the animated character.
6. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device identifies at least one font information for font visualization for each sentence unit of the text based on user sentiment information corresponding to each sentence unit of the text, and Based on the above at least one font information, visualization effects are provided for the fonts of each sentence composition unit of the text entered into the input UI, and The above at least one font information is, An electronic device comprising at least one of font type information, font size information, font color information, font in / out information, font rotation information, or font emphasis information.
7. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Based on the pronunciation feature information of the user, the animation character that speaks the text is obtained, and The above user's pronunciation characteristic information is, An electronic device comprising at least one of speed, intonation, tone, stress, phonemic pronunciation, linking, or assimilation.
8. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Convert the above text into speech using TTS (text-to-speech), and Based on the pronunciation of the above voice, mouth shape information is obtained in phoneme units, and An electronic device for obtaining the animation character whose mouth shape changes based on the above mouth shape information.
9. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Based on the context information of the above text, a UI including multiple characters is provided, and An electronic device that, when one of the above multiple characters is selected, obtains the animation character based on the selected character.
10. In Paragraph 1 or 2, The above input UI is, Provides at least one of a text input UI for inputting text or a voice input widget UI for inputting user voice, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that acquires the text entered through the text input UI or acquires the text based on the user's voice entered through the voice input widget UI.
11. In Paragraph 1 or 2, The above electronic device is, mike; Camera; and It further includes a sensor; When the above instructions are executed individually or collectively by the at least one processor, the electronic device, While a widget is displayed on the display, emotional information of the user is obtained based on at least one of the user's voice information obtained through the microphone, the user's facial expression information obtained through the camera, and the user's gesture information obtained through the sensor. Based on the emotional information of the user, an animation widget is obtained that provides an animation effect in which the facial expression of a character included in the widget changes. An electronic device that stores the animation widget in the memory or transmits it outside the electronic device according to a user command.
12. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, When text is received from an external device, user emotion information corresponding to each sentence unit of the received text is obtained based on context information corresponding to the received text, and Based on the above emotional information, an animation character is obtained that provides an animation effect in which the character speaks the text while the character's facial expression changes in correspondence with each sentence unit of the text, and An electronic device that displays the received text and the animated character on the display.
13. In a method for controlling an electronic device, When text is obtained through an input UI, an operation to obtain context information corresponding to the obtained text; An operation to obtain user emotion information corresponding to each sentence unit of the text based on the above context information; An action of acquiring an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes in correspondence with each sentence unit of the text based on the emotional information of the user; and A control method comprising the action of inputting the above-mentioned animation character into the above-mentioned input UI.
14. In Paragraph 13, The operation of acquiring the emotional information of the above-mentioned user is, The operation of obtaining a user's first emotion information based on context information corresponding to a first text corresponding to a first sentence unit included in the above text, and obtaining a user's second emotion information based on context information corresponding to a second text corresponding to a second sentence unit obtained after the first text; and A control method comprising: an action of acquiring an animation character in which the character sequentially changes its facial expression based on the first emotional information and the second emotional information while sequentially speaking the first text and the second text.
15. A non-transient computer-readable medium storing instructions that cause said electronic device to perform an operation when executed by a processor of said electronic device, The above operation is, When text is obtained through an input UI, an operation to obtain context information corresponding to the obtained text; An operation to obtain user emotion information corresponding to each sentence unit of the text based on the above context information; An action of acquiring an animation character that provides an animation effect in which the character speaks the text while the character's facial expression changes in correspondence with each sentence unit of the text based on the emotional information of the user; and A control method comprising the action of inputting the above-mentioned animation character into the above-mentioned input UI.