Electronic device, method, and non-transitory computer-readable storage medium for changing pose of avatar
By processing user input to dynamically change the pose of avatars using trained models, the device addresses the limitations of turn-based frameworks, offering a more immersive and realistic conversational experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NCSOFT CORP
- Filing Date
- 2024-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Existing technologies for interacting with users through avatars lack the ability to provide a realistic and immersive conversational experience, as they often rely on turn-based frameworks that do not accurately mimic human conversation and fail to integrate multiple modalities of interaction.
An electronic device is equipped with processors and models that process user input, including voice and image data, to dynamically change the pose of an avatar through a display and speaker, using trained models to generate continuous and natural animations based on user interactions.
The solution enhances user immersion by providing avatars that can realistically converse and gesture, mimicking human-like interactions through integrated modalities, thereby improving the overall user experience.
Smart Images

Figure KR2024016572_07052026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transient computer-readable storage medium for changing the pose of an avatar
[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for changing the pose of an avatar.
[0002] The electronic device may include a display. The electronic device may display an avatar for interacting with a user of the electronic device through the display. The electronic device may control the form of the avatar based on receiving user input. The electronic device may display an animation of the avatar's pose changing through the display. The electronic device may display the animation through the display based on receiving user input that causes the avatar's pose to change.
[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure.
[0004] No claim or determination is made as to whether any of the foregoing can be applied as prior art related to the present disclosure.
[0005] An electronic device is described. The electronic device may include a memory that stores instructions and includes one or more storage media. The electronic device may include a speaker. The electronic device may include a display. The electronic device may include at least one processor that includes a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to display, through the display, an avatar of a first pose for interacting with a user of the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to obtain text data including a natural language response expressed in text by providing information related to the user data to a trained first model based on identifying user data for interaction between the electronic device and the user. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to acquire first joint data representing a second pose continuous with the first pose by providing audio data acquired using the text data and first pose data representing the features of the first pose to a trained second model. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to display, through the display, an animation of the avatar changing from the first pose to the second pose using the first joint data when outputting voice corresponding to the text data through the speaker.
[0006] A method is provided. The method may be executed within an electronic device comprising a speaker and a display. The method may include an action of displaying an avatar of a first pose for interacting with a user of the electronic device through the display. The method may include an action of obtaining text data including a natural language response expressed in text by providing information related to the user data to a trained first model based on identifying user data for interaction between the electronic device and the user. The method may include an action of obtaining first joint data representing a second pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the characteristics of the first pose to a trained second model. The method may include an action of displaying an animation of the avatar changing from the first pose to the second pose using the first joint data through the display when voice corresponding to the text data is output through the speaker.
[0007] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to display, through the display, an avatar of a first pose for interacting with a user of the electronic device when executed by the electronic device having a speaker and a display. The one or more programs may include instructions that cause the electronic device to obtain text data including a natural language response expressed in text by providing information related to the user data to a trained first model, based on identifying user data for interaction between the electronic device and the user when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain first joint data representing a second pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the features of the first pose to a trained second model when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display, through the display, an animation of the avatar changing from the first pose to the second pose using the first joint data, when the electronic device is executed and outputs voice corresponding to the text data through the speaker.
[0008] An electronic device is described. The electronic device may include a memory that stores instructions and includes one or more storage media. The electronic device may include a speaker. The electronic device may include a display. The electronic device may include at least one processor that includes a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to display, through the display, an avatar of a first pose for interacting with a user of the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to obtain response data based on said information related to said user data by providing said information to a trained first model based on identifying said user data for interaction between the electronic device and said user. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to display, through the display, a first animation of the avatar changing from the first pose to the second pose through the third pose, by using first joint data representing a third pose between the first pose and the second pose, when outputting a voice corresponding to the text data through the speaker based on a determination that the response data includes text data including a natural language response expressed as text and includes a directive representing a second pose of the avatar.The above instructions may cause the electronic device to acquire second joint data representing a fourth pose continuous with the first pose by providing audio data acquired using the text data and first pose data representing the features of the first pose to a trained second model, based on a determination that the response data includes the text data and does not include the directive, when the above instructions are executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to display, through the display, a second animation of the avatar changing from the first pose to the fourth pose using the second joint data, when the voice corresponding to the text data is output through the speaker, based on a determination that the response data includes the text data and does not include the directive, when the above instructions are executed individually or collectively by the at least one processor.
[0009] A method is provided. The method may be executed within an electronic device having a speaker and a display. The method may include an action of displaying an avatar of a first pose for interacting with a user of the electronic device through the display. The method may include an action of obtaining response data based on information related to the user data by providing information related to the user data to a trained first model based on identifying user data for interaction between the electronic device and the user. The method may include an action of displaying a first animation of the avatar changing from the first pose to the second pose through the display, using first joint data representing a third pose between the first pose and the second pose, when a voice corresponding to the text data is output through the speaker based on a determination that the response data includes text data including a natural language response expressed as text and includes a directive representing a second pose of the avatar. The above method may include an operation of obtaining second joint data representing a fourth pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the characteristics of the first pose to a trained second model, based on a determination that the response data includes the text data and does not include the directive. The above method may include an operation of displaying a second animation of the avatar changing from the first pose to the fourth pose using the second joint data through the display, based on a determination that the response data includes the text data and does not include the directive, when the voice corresponding to the text data is output through the speaker.
[0010] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to display, through the display, an avatar of a first pose for interacting with a user of the electronic device when executed by the electronic device having a speaker and a display. The one or more programs may include instructions that cause the electronic device to obtain response data based on said information related to said user data by providing said information to a trained first model, based on identifying said user data for interaction between the electronic device and said user when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display, through the display, a first animation of the avatar changing from the first pose to the second pose through the third pose, using first joint data representing a third pose between the first pose and the second pose, when outputting a voice corresponding to the text data through the speaker, based on a determination that the response data includes text data including a natural language response expressed as text when executed by the electronic device and includes a directive representing a second pose of the avatar.The above one or more programs may include instructions that cause the electronic device to acquire second joint data representing a fourth pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the features of the first pose to a trained second model, based on a determination that when executed by the electronic device, the response data includes the text data and does not include the directive. The above one or more programs may include instructions that cause the electronic device to display, through the display, a second animation for the avatar changing from the first pose to the fourth pose using the second joint data, based on a determination that when executed by the electronic device, the response data includes the text data and does not include the directive, the voice corresponding to the text data is output through the speaker.
[0011] Figure 1 illustrates an example of an environment including an electronic device that displays an avatar.
[0012] Figure 2 is a simplified block diagram of an exemplary electronic device.
[0013] Figure 3 is a flowchart illustrating the operation of an electronic device that displays animation using a model.
[0014] FIGS. 4a and 4b illustrate an exemplary operation of an electronic device that acquires joint data of an avatar using a model.
[0015] Figure 5 is a flowchart illustrating the operation of an electronic device that acquires joint data using pose data and potential data.
[0016] FIG. 6 illustrates an exemplary operation of an electronic device that generates pose data using audio data.
[0017] Figure 7 is a flowchart illustrating the operation of an electronic device that displays an animation using response data containing a directive.
[0018] FIGS. 8A and 8B illustrate exemplary operation of an electronic device that displays animation using a database of poses.
[0019] Figure 1 illustrates an example of an environment including an electronic device that displays an avatar.
[0020] Referring to FIG. 1, the environment (150) may include an electronic device (100) and a user (120). The electronic device (100) may be used to display an avatar (160). For example, the avatar (160) may be described as an interface for interacting with the user (120). For example, the electronic device (100) may display the avatar (160) differently as it receives user input. For example, the electronic device (100) may display the avatar (160) interacting with the user (120) through a display (e.g., the display (208) in FIG. 2) based on identifying information of the user (120).
[0021] An avatar (160) may be used to interact with a user (120). For example, the avatar (160) may be used to converse with the user (120). For example, the conversation may include verbal and non-verbal elements (e.g., gestures, poses). For example, the avatar (160) may mimic a person. For example, the avatar (160) may be used to provide a realistic conversational experience to the user (120). For example, an electronic device (100) capable of displaying the avatar (160) may be required to interact with the user (120) in various ways. For example, the electronic device (100) may provide an enhanced user experience to the user (120) based on providing an avatar (160) capable of interacting in various ways. For example, the electronic device (100) can cause the user (120) to immerse himself in the avatar (160) by displaying the avatar (160) having various gestures for communication on the display.
[0022] For example, an electronic device (100) may provide an avatar (160) capable of interacting with a user (120) based on a turn-based framework. For example, the turn-based framework may be performed based on an idle state, a listening state, and a speaking state. For example, the turn-based framework may not provide the same sensation as human conversation. For example, a person may communicate using facial expressions, gestures, and voice effects. For example, a person may interrupt another person while they are speaking. For example, a person may pass their turn to speak, nod, or shake their head. For example, the electronic device (100) may be required to acquire images of the user (120). For example, the electronic device (100) may be required to acquire information about the user's (120) voice. For example, the electronic device (100) may be required to provide an avatar (160) capable of communicating with the user (120) through multiple modalities. For example, the electronic device (100) may be required to provide an avatar (160) that is not limited to turn-based logic.
[0023] For example, the electronic device (100) may provide an avatar (160) for interacting with a user (120) by using trained model(s). For example, the electronic device (100) may use a model trained to acquire joint data representing the pose of the avatar (160) using an audio signal. For example, the electronic device (100) may acquire the audio signal by performing text-to-speech (TTS) on text data based on information of the user (120). For example, the electronic device (100) may acquire joint data representing the pose of the avatar (160) by providing the audio signal to the model. For example, the electronic device (100) may use the acquired joint data to display an animation of the pose of the avatar (160) changing through the display.
[0024] Figure 2 is a simplified block diagram of an exemplary electronic device.
[0025] Referring to FIG. 2, the electronic device (100) may include a display (208), a speaker (210), at least one processor (207) and a memory (206).
[0026] At least one processor (207) may include a hardware component for processing data using instructions stored in memory (206). The hardware component for processing data may include a CPU (central processing unit) (e.g., including processing circuits). The hardware component for processing data may include a GPU (graphic processing unit) (e.g., including processing circuits). The hardware component for processing data may include a DPU (display processing unit) (e.g., including processing circuits). The hardware component for processing data may include a NPU (neural processing unit) (e.g., including processing circuits).
[0027] At least one processor (207) may include one or more cores. For example, at least one processor (207) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.
[0028] Memory (206) may include a hardware component for storing data and / or instructions that are input to and / or output from at least one processor (207). Memory (206) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). Volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disk, and embedded multimedia card (EMMC).
[0029] The display (208) can output visualized information. For example, the display (208) can output visualized information to a user under the control of at least one processor (207). The display (208) may include hardware components of an electronic device (100) used to display a screen. For example, the display (208) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements may include an organic light-emitting diode (OLED) or a micro LED. However, it is not limited thereto. For example, the display (208) may include a liquid crystal display (LCD).
[0030] The speaker (210) may include a hardware component for supporting the output of audio or an audio signal. The speaker (210) may be used to output verbal audio. The speaker (210) may be used to output non-verbal audio. The speaker (210) may be used to provide auditory feedback to the user (120).
[0031] At least one processor (207) may display an avatar (160) of a first pose for interacting with a user (120) of an electronic device (100) through a display (208). For example, the display (208) may be used to display the avatar (160). At least one processor (207) may obtain text data including a natural language response based on (or expressed in text) text by providing information related to the user input to a trained first model, based on identifying user input for interaction between the electronic device (100) and the user (120). For example, at least one processor (207) may obtain first joint data representing a second pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the features of the first pose to a trained second model. For example, when at least one processor (207) outputs voice corresponding to text data through a speaker (210), it can display an animation of an avatar (160) changing from a first pose to a second pose using first joint data through a display (208). For example, the speaker (210) can be used to output voice corresponding to text data. For example, the display (208) can be used to display an animation of an avatar (160).
[0032] FIG. 3 is a flowchart illustrating the operation of an electronic device that displays an animation using a model. This method may be executed by the electronic device (100) illustrated in FIG. 2 or by at least one processor (207) of the electronic device (100).
[0033] Referring to FIG. 3, in operation 310, at least one processor (207) may display an avatar (160) in a first pose for interacting with a user (120) of the electronic device (100) through a display (208). For example, the first pose may include a basic pose of the avatar (160). For example, the first pose may include a previous pose of the avatar (160). For example, the first pose may include a pose prior to the avatar (160) newly interacting with the user (120).
[0034] In operation 320, at least one processor (207) can obtain text data including a natural language response based on (or expressed in text) text by providing information related to said user input (e.g., user information (422) in FIG. 4a) to a trained first model (e.g., first model (424) in FIG. 4a) based on identifying user input for interaction between the electronic device (100) and the user (120). For example, said user input may be referred to as user data. For example, said user input may include the voice of the user (120) obtained through a microphone. For example, at least one processor (207) may perform voice recognition through said microphone. For example, said user input may include images of the user (120) obtained through a camera. For example, at least one processor (207) can obtain the text data by providing the information to the first model based on identifying an event that obtains information about the user (120) (e.g., user information (422) in FIG. 4a). For example, response data (e.g., response data (426) in FIG. 4a) may include the text data.
[0035] For example, at least one processor (207) can perform voice recognition using the voice of the user (120) obtained through the microphone. For example, at least one processor (207) can perform text-to-speech (TTS) using the voice of the user (120) obtained through the microphone. For example, at least one processor (207) can recognize or generate full-body gestures using the voice of the user (120) obtained through the microphone. For example, at least one processor (207) can recognize a person using an image of the user (120) obtained through the camera. For example, at least one processor (207) can perform recognition and / or synthesis of facial expressions using an image of the user (120) obtained through the camera. For example, at least one processor (207) can detect the emotions of the user (120) using an image of the user (120) obtained through the camera. For example, at least one processor (207) can track the gaze of the user (120) using an image of the user (120) obtained through the camera. For example, at least one processor (207) can perform lip sync using an image of the user (120) obtained through the camera. For example, at least one processor (207) can recognize the props and / or accessories of the user (120) using an image of the user (120) obtained through the camera. For example, at least one processor (207) can detect hand-object interactions using an image of the user (120) obtained through the camera.
[0036] In operation 330, at least one processor (207) can obtain audio data obtained using the text data (e.g., audio data (610) of FIG. 6) and first pose data (e.g., first pose data (620) of FIG. 6) representing the features of the first pose, thereby providing the first joint data (e.g., first joint data (670) of FIG. 6) representing the first pose and a second pose consecutive to the first pose to a trained second model (e.g., second model (630) of FIG. 6).
[0037] In operation 340, when at least one processor (207) outputs voice corresponding to the text data through the speaker (210), it can display an animation of an avatar (160) changing from a first pose to a second pose using the first joint data (e.g., the first joint data (670) of FIG. 6) through the display (208).
[0038] For example, the electronic device (100) may provide an avatar (160) for interacting with a user (120). For example, the electronic device (100) may enhance the interaction between the avatar (160) and the user (120) by using a first model (424), a second model (630), a third model (650), and a fourth model (660). For example, the user (120) may interact with the avatar (160) through an interface (e.g., a computer, a telephone, or a kiosk). However, it is not limited thereto. For example, the electronic device (100) may generate a cut-scene that encompasses movements included in a conversation scenario between the avatars. For example, the electronic device (100) may use recorded voice to display animations of the avatar (160). For example, the models included in the electronic device (100) may generate directives for the recorded voice as annotations.
[0039] At least one processor (207) can obtain response data (e.g., response data (426) of FIG. 4a) by providing information about the user (120) (e.g., user information (422) of FIG. 4a) to the first model (424). For example, at least one processor (207) can obtain joint data representing the pose of the avatar (160) based on obtaining the response data. For example, at least one processor (207) can display the avatar (160) through the display (208) using the joint data. The acquisition of the response data and the display of the avatar (160) are described and illustrated in more detail with reference to FIG. 4a and FIG. 4b.
[0040] FIGS. 4a and 4b illustrate an exemplary operation of an electronic device that acquires joint data of an avatar using a model.
[0041] Referring to FIG. 4a, at least one processor (207) can acquire a user voice (402) and / or a user image (404). For example, at least one processor (207) can acquire the voice of the user (120) through a microphone (not shown). For example, at least one processor (207) can acquire image(s) of the user (120) through a camera (not shown). For example, at least one processor (207) can identify an event for acquiring the user voice (402) and / or a user image (404). For example, at least one processor (207) can identify a user input for receiving the user voice (402). For example, at least one processor (207) can identify a user input for acquiring the user image (404). For example, the user input may be referred to as user data.
[0042] At least one processor (207) can obtain information for each of the following using the user voice (402) and / or user image (404): speech (406), co-speech gesture (408), body language (410), facial expression (412), identification (414), accessory (416), and eye contact (418). For example, information for speech (406) may include information indicating "I had a hard day today." For example, information for the co-speech gesture (408) may include information for a baton gesture. For example, information for body language (410) may include information for a crossing arms gesture. For example, information for facial expression (412) may include information indicating a frown. For example, information regarding ID (414) may include information indicating "John Doe". For example, information regarding accessory (416) may include information regarding glasses. For example, information regarding eye contact (418) may include information indicating engaging eye contact.
[0043] At least one processor (207) can obtain user information (422) by providing information for each of the following to the input converter (420): speech (406), co-speech gesture (408), body language (410), facial expression (412), identification (414), accessory (416), and eye contact (418). For example, the user information (422) may be referenced by a prompt. For example, the user information (422) may be information related to user input for interaction between the electronic device (100) and the user (120).
[0044] For example, at least one processor (207) can obtain response data (426) by providing user information (422) to the first model (424). For example, the first model (424) may include a model trained with machine learning techniques. For example, the first model (424) may include a model trained with deep learning techniques. For example, the first model (424) may include a large language model (LLM). For example, the first model (424) may include a large multimodal model (LMM). For example, the first model (424) may be described as a model trained to output response data (426) based on receiving user information (422). For example, the first model (424) may be trained to interpret text enclosed in various forms of brackets. For example, the response data (426) may include text data related to the utterance of the avatar (160).
[0045] For example, at least one processor (207) may output a conditional directive and / or a triggering directive using the first model (424). For example, the conditional directive may be used to provide situational context, such as a consistently maintained ID (414). For example, the conditional directive may correspond to a conditional gesture. For example, at least one processor (207) may display the avatar (160) of the conditional gesture until the triggering directive is identified. For example, the conditional gesture may be associated with the actions of the avatar (160) listening, resting, thinking, or speaking. For example, the triggering directive may be used to indicate a spontaneous event within a conversation, to coordinate the actions of the avatar (160), or to immediately execute a specific action. For example, the triggering directive may correspond to a triggering gesture. For example, the triggering gesture may include an immediate movement to convey semantic body language. For example, at least one processor (207) may be required to rapidly express the triggering gesture.
[0046] For example, the response data (426) may include a triggering directive related to the pose of the avatar (160). For example, the triggering directive may be referred to as a directive. For example, the response data (426) may include a triggering directive related to the pose of the avatar (160). For example, the triggering directive may be described as text indicating the pose of the avatar (160). For example, the triggering directive may be described as text indicating the pose of the avatar (160).
[0047] Referring to FIG. 4b, at least one processor (207) may provide response data (426) to an output converter (428). For example, by providing the response data (426) to the output converter (428), at least one processor (207) may obtain information regarding each of the persona (430), eye contact (432), gesture (434), facial expression (436), audio (438), and silence (440). For example, the output converter (428) may include a second model (e.g., the second model (630) of FIG. 6), a third model (e.g., the third model (650) of FIG. 6), and a fourth model (e.g., the fourth model (660) of FIG. 6). For example, the output converter (428) may be used to render an avatar (160). For example, information about the persona (430) may include information indicating kindness. For example, information about eye contact (432) may include information indicating engaging eye contact. For example, information about the gesture (434) may include information indicating a shrug and rhythmic gesture. For example, information about the facial expression (436) may include information about an expression expressing concern. For example, information about the audio (438) may include information indicating "I'm sorry." For example, information about silence (440) may include information indicating a time to be silent during speech.
[0048] For example, at least one processor (207) can display an avatar (160) through a display (208) using information regarding each of the persona (430), eye contact (432), gesture (434), facial expression (436), audio (438), and silence (440). For example, at least one processor (207) can enhance the user's (120) immersion using information regarding each of the persona (430), eye contact (432), gesture (434), facial expression (436), audio (438), and silence (440).
[0049] For example, at least one processor (207) may use a second model (e.g., the second model (630) of FIG. 6), a third model (e.g., the third model (650) of FIG. 6), and a fourth model (e.g., the fourth model (660) of FIG. 6) to obtain information about a gesture (434). For example, at least one processor (207) may render an avatar (160) using response data (426) obtained by providing information about the user (120) to the first model (424). For example, at least one processor (207) may obtain an audio signal corresponding to said text data by performing text-to-speech (TTS) on text data related to the utterance of the avatar (160) included in said response data (426). For example, at least one processor (207) may generate or obtain joint data representing the pose of the avatar (160) using said audio signal. The acquisition of the above joint data is explained and illustrated in more detail with reference to FIG. 5.
[0050] FIG. 5 is a flowchart illustrating the operation of an electronic device for acquiring joint data using pose data and potential data. This method may be executed by the electronic device (100) illustrated in FIG. 2 or by at least one processor (207) of the electronic device (100).
[0051] Referring to FIG. 5, in operation 510, at least one processor (207) can obtain second pose data of the avatar (160) (e.g., second pose data (640) of FIG. 6) and first potential data for the second pose data (e.g., first potential data (645) of FIG. 6) by providing audio data (610) and first pose data representing the characteristics of the first pose (e.g., first pose data (620) of FIG. 6) to a second model (e.g., second model (630) of FIG. 6). For example, the pose data may be described as data regarding the characteristics of a (specific) pose of the avatar (160). For example, the pose data may include data representing the position and rotation of a part of the joint of the avatar (160). For example, the pose data may include data regarding the trajectory of the part of the joint.
[0052] In operation 520, at least one processor (207) can obtain first joint data (e.g., first joint data (670) of FIG. 6) representing the second pose by providing second pose data (e.g., second pose data (640) of FIG. 6) and first potential data (e.g., first potential data (645) of FIG. 6) to a fourth model (e.g., fourth model (660) of FIG. 6). For example, the joint data can be described as data regarding the positions of all joints of the avatar (160). For example, the joint data can be used to render the avatar (160) in a specific pose. For example, at least one processor (207) can directly compute pose data from the joint data. For example, since it is difficult to compute pose data directly from the joint data, at least one processor (207) can use a model (or network) to obtain pose data from the joint data.
[0053] In operation 530, at least one processor (207) can obtain third pose data (e.g., third pose data (680) of FIG. 6) and second potential data (e.g., second potential data (685) of FIG. 2) for generating second joint data (e.g., second joint data (690) of FIG. 6) representing a third pose and a third pose consecutive to the second pose by providing second pose data (e.g., second pose data (640) of FIG. 6) and first potential data (e.g., second potential data (685) of FIG. 2). The operation of generating joint data using the pose data and potential data is described and illustrated in more detail with reference to FIG. 6.
[0054] FIG. 6 illustrates an exemplary operation of an electronic device that generates pose data using audio data.
[0055] Referring to FIG. 6, at least one processor (207) can obtain second pose data (640) and first potential data (645) by providing audio data (610) and first pose data (620) to a second model (630). For example, the first pose data (620) may be described as pose data representing the characteristics of a first pose of an avatar (160) before receiving user input. For example, the first pose data (620) may be described as data for indirectly rendering the avatar (160) of the first pose.
[0056] The second model (630) may be described as a model trained to output pose data representing the joints (or features of the pose) of the avatar (160) for the pose and latent data for the pose data, in order to generate joint data representing the pose for the audio data (610) based on receiving the audio data (610). For example, the latent data may be referred to as a latent vector. For example, the second model (630) may include a model trained using a generative adversarial network technique. For example, the second model (630) may include a generative model and a discriminative model. For example, the second model (630) may be trained using noise sampled from a normal sequence. For example, the second model (630) may be trained based on the provision of pose data and feature data to enhance generative characteristics.
[0057] For example, the second model (630) may include multiple discrimination models. For example, in the case of the second model (630) having only one discrimination model, mode collapse or learning failure may occur. For example, in the case of the second model (630) having only one discrimination model, the same gesture or the same pose may be repeated independently of the input value. For example, at least one processor (207) may be required to use multiple discrimination models to avoid avatars (160) that repeat the same gesture or the same pose. For example, at least one processor (207) may assign different goals to each of the discrimination models. For example, at least one processor (207) may generate or acquire pose data and latent data of the avatar (160) using the second model (630) which includes discrimination models with different goals.
[0058] For example, at least one processor (207) may include a second model (630) comprising a pose transition determination model, an audio-pose determination model, and a long-term determination model. For example, the second model (630) may include a pose transition determination model, an audio-pose determination model, and a long-term determination model. For example, the pose transition determination model may be used to evaluate the validity of a pose transition by analyzing the first pose data (620) and the first joint data (670) of consecutive frames. For example, at least one processor (207) may evaluate the validity of a pose transition for each of consecutive poses using the pose transition determination model. For example, the determination method of the pose transition determination model may be referenced by the following mathematical formula.
[0059]
[0060] The above represents a pose transition determination model. The above represents a decompressor network. The above represents the pose data at t-1. The above represents the latent data at t-1. The above represents pose data at t. The above represents latent data at t. For example, the pose transition model has a value of 1 if it is determined that the pose transition is valid, and a value of 0 if it is determined that the pose transition is invalid.
[0061] For example, the audio-pose discrimination model may be used to evaluate consistency and synchronization between audio and pose features. For example, at least one processor (207) may determine the authenticity of the data by using the features of the current pose, latent data, feature data, and bit data through the audio-pose discrimination model. For example, the ability of the audio-pose discrimination model to identify discrepancies by changing audio or bit features may be trained. For example, the ability of the audio-pose discrimination model may be trained by using non-human sounds paired with the pose of the user (120). For example, the sensitivity of the audio-pose discrimination model to natural voice patterns may be trained by generating atypical human voices converted through text-to-speech technology. For example, the ability of the audio-pose discrimination model to distinguish between natural and artificial inputs may be trained by using white noise.
[0062]
[0063] The above represents an audio-pause discrimination model. The above represents pose data at t. The above represents latent data at t. The above represents feature data at t. The above represents bit data at t. For example, the audio-pose determination model has a value of 1 if it determines that the generated pose for the audio signal is valid, and a value of 0 if it determines that the generation of the pose is invalid.
[0064] For example, a long-term discrimination model can be used to evaluate the consistency of a gesture sequence over a long period. For example, a long-term discrimination model can be used to maintain the naturalness of the gesture over a long period. For example, the electronic device (100) can improve the authenticity of the animation for the avatar (160) by evaluating the continuity and realism of the gesture sequence over a long period using a long-term discrimination model.
[0065]
[0066] The above represents a long-term discriminant model. The above represents the pose data at t-1. The above represents the latent data at t-1. The above represents pose data at t. The above represents latent data at t. For example, the above long-term discrimination model has a value of 1 if it determines that the pose generated for the audio signal maintains naturalness over the long term, and a value of 0 if it determines that the pose generated for the audio signal does not maintain naturalness over the long term. For example, the above It can be referenced by the following mathematical formula.
[0067]
[0068] The above S represents a step model. The above represents the M-times configuration of the above S. For example, the above mathematical formula 3 may use the above mathematical formula 4.
[0069] For example, at least one processor (207) can obtain second pose data (640) and first potential data (645) by providing audio data (610) to the second model (630). For example, audio data (610) may be obtained or generated using text data included in response data (426). For example, the text data may be the subject of TTS. For example, the avatar (160) may be linked to an audio signal on which TTS is performed on the text data. For example, at least one processor (207) may display the avatar (160) when outputting the audio signal through the speaker (210). For example, at least one processor (207) may display an animation of the avatar (160) changing its pose through the display (208) when outputting the audio signal through the speaker (210).
[0070] For example, at least one processor (207) can acquire audio data (610) using an audio signal corresponding to text data. For example, at least one processor (207) can acquire data regarding the characteristics of the audio signal from the audio signal through a mel-frequency cepstral coefficient (MFCC) technique. For example, at least one processor (207) can acquire audio data (610) using the audio signal. For example, at least one processor (207) can acquire audio data (610) representing the characteristics of the audio signal using the audio signal. For example, the audio data (610) may include feature data representing the characteristics of the audio signal from the audio signal corresponding to the text data through an MFCC technique. For example, the feature data may represent the overall characteristics of the audio signal.
[0071] For example, at least one processor (207) can acquire audio data (610) using an audio signal corresponding to text data. For example, at least one processor (207) can divide the audio signal into predefined time intervals. For example, at least one processor (207) can divide the audio signal according to a time interval (e.g., 1 second). For example, the time interval may be referred to as a beat. For example, the audio signal may be divided based on syllables. For example, the audio signal may be divided based on morphemes. For example, the audio signal may be divided based on phonemes. For example, the audio data (610) may include bit data indicating the start time of the voice within the bits divided according to the time interval of the audio signal corresponding to the text data. For example, the audio data (610) may include bit data indicating the start time of the voice corresponding to the text data within the bits divided according to the time interval of the audio signal. For example, beat data may represent data regarding the timing at which a voice or audio signal begins within a beat. For example, beat data may be used to reinforce the association between rhythmic elements and gestures. For example, audio data (610) may include the feature data and the beat data.
[0072] For example, the above time intervals may be changed. For example, at least one processor (207) may determine the time intervals differently in rhythmic chunks. For example, the length of the first time interval of the audio signal may differ from the length of the second time interval following the first time interval. For example, the length of the time intervals may not be uniform. For example, the electronic device (100) may improve synchronization between gestures and voice rhythms by dividing the audio signal into non-uniform time intervals. For example, the electronic device (100) may generate or acquire joint data of the avatar (160) from the audio data (610) by utilizing periods of silence during which voice activity is relatively low.
[0073] At least one processor (207) can obtain first joint data (670) representing a first pose and a second pose consecutive to the first pose by providing the second pose data (640) and the first potential data (645) to the fourth model (660). For example, at least one processor (207) can render an avatar (160) using the first joint data (670). For example, at least one processor (207) can display the avatar (160) in the second pose through a display (208) using the first joint data (670). For example, at least one processor (207) can display an animation through the display (208) in which the pose of the avatar (160) changes from the first pose to the second pose using the first joint data (670). For example, the first joint data (670) can be used to render the avatar (160) in the second pose.
[0074] For example, the fourth model (660) may be referred to as a decompression model or a decompressor. For example, the fourth model (660) may be trained to output joint data corresponding to the pose data by receiving pose data for the joints of the avatar (160) and latent data for said pose data. For example, the fourth model (660) may be used to obtain joint data representing the position of a complete joint from the pose data. For example, said latent data may be used for the fourth model (660) to output said joint data corresponding to said pose data. For example, the fourth model (660) may be used to represent a fully realized jointed body movement by decoding the pose data representing compressed features into joint data.
[0075] For example, the compression model can convert joint data into latent data of a low-dimensional representation. For example, at least one processor (207) can provide the fourth model (660) based on concatenating the latent data and pose data converted by the compression model. For example, the compression model can be used to train the fourth model (660).
[0076] For example, at least one processor (207) can obtain third pose data (680) and second potential data (685) by providing second pose data (640) and first potential data (645) to a third model (650). For example, the third model (650) can be described as a model trained to output different pose data and different potential data by receiving pose data and potential data. For example, the third pose data (680) and second potential data (685) may be pose data and potential data for the third pose. For example, the third pose may be described as a pose continuous with the second pose. For example, the third pose may be described as a pose naturally connected from the second pose. For example, the third pose may be described as a pose after the second pose. For example, at least one processor (207) can use audio data (610) to display, through a display (208), an animation changing from an avatar (160) of a first pose followed by a second pose followed by a third pose followed by a second pose followed by a third pose. For example, the third model (650) may be referred to as a step model or a stepper.
[0077] For example, at least one processor (207) may acquire or generate second joint data (690) by providing third pose data (680) and second potential data (685) to the fourth model (660). For example, at least one processor (207) may use the second joint data (690) to display the avatar (160) of the third pose through the display (208). For example, the second joint data (690) may be used to render the avatar (160) of the third pose.
[0078] At least one processor (207) can display the avatar (160) through the display (208) using a directive that is included in the response data (426) and indicates the pose of the avatar (160). For example, at least one processor (207) can determine the pose of the avatar (160) using the directive. The operation of determining the pose of the avatar (160) based on the directive is described and illustrated in more detail with reference to FIG. 7.
[0079] FIG. 7 is a flowchart illustrating the operation of an electronic device that displays an animation using response data containing a directive. This method may be executed by the electronic device (100) illustrated in FIG. 2 or by at least one processor (207) of the electronic device (100).
[0080] Referring to FIG. 7, in operation 710, at least one processor (207) may display an avatar (160) in a first pose for interacting with a user (120) of the electronic device (100) through a display (208). For example, operation 710 may correspond to operation 310 of FIG. 3.
[0081] In operation 720, at least one processor (207) can obtain response data (426) based on the information related to the user input by providing information related to the user input to a trained first model (424) based on identifying user input for interaction between the electronic device (100) and the user (120). For example, operation 720 may correspond to operation 320 of FIG. 3.
[0082] In operation 730, at least one processor (207) can identify whether the response data (426) includes text data and a directive indicating a second pose of the avatar (160). For example, at least one processor (207) can execute operation 740 under the condition that the response data (426) includes text data containing a text-based (or text-expressed) natural language response and a directive indicating a second pose of the avatar (160), and execute operation 750 under the condition that the response data (426) does not include said directive.
[0083] In operation 740, at least one processor (207) may, based on the determination that response data (426) includes text data containing a text-based natural language response and includes a directive indicating a second pose of the avatar (160), output a voice corresponding to said text data through a speaker (210), and use first joint data indicating a third pose between the first pose and the second pose, display a first animation of the avatar (160) changing from the first pose to the second pose through the third pose through a display (208). For example, said directive may be referred to as a triggering directive. For example, said directive may be described as text indicating a pose of the avatar (160). For example, said third pose may be described as a pose within the process of transitioning from the first pose to the second pose. For example, the acquisition of the first joint data representing the third pose is described and illustrated in more detail with reference to FIGS. 8a and FIGS. 8b.
[0084] FIGS. 8A and 8B illustrate exemplary operation of an electronic device that displays animation using a database of poses.
[0085] Referring to FIG. 8a, the state (810) can be described as a state in which an avatar (160) in a first pose is displayed through a display (208). Referring to FIG. 8b, the state (830) can be described as a state in which an avatar (160) in a second pose indicated by a directive included in response data (426) is displayed through a display (208). For example, at least one processor (207) may determine the pose of the avatar (160) based on acquiring the directive. For example, at least one processor (207) may display the avatar (160) in a second pose through a display (208) based on acquiring a directive indicating or indicating the second pose of the avatar (160). For example, the electronic device (100) may store joint data for the pose indicated by the directive in memory (206). For example, the electronic device (100) may store a first database for a first pose in memory (206). For example, the electronic device (100) may store a second database for a second pose in memory (206). For example, the electronic device (100) may store a third database for a third pose between the first pose and the second pose in memory (206). For example, the electronic device (100) may store a fourth database for a fourth pose between the second pose and the first pose following the second pose in memory (206). For example, the first pose may be referred to as a base pose. For example, the third pose may be referred to as a stroke pose. For example, the stroke pose may include an initial movement from the base pose to a key component. For example, the second pose may be referred to as a hold pose. For example, the hold pose may be described as an essential movement to emphasize the unique characteristics of the gesture.
[0086] For example, at least one processor (207) may retrieve first joint data for a first pose from a first database to display an avatar (160) of a first pose. For example, at least one processor (207) may retrieve second joint data for a second pose from a second database to display an avatar (160) of a second pose. For example, at least one processor (207) may retrieve third joint data for a third pose from a third database to display an avatar (160) of a third pose. For example, at least one processor (207) may retrieve fourth joint data for a fourth pose from a fourth database to display an avatar (160) of a fourth pose. For example, the first joint data may be described as data for rendering the avatar (160) of the first pose. For example, the first joint data may include data for the joints of the avatar (160) of the first pose. For example, the second joint data may be described as data for rendering the avatar (160) of the second pose. For example, the second joint data may include data regarding the joints of the avatar (160) of the second pose. For example, the third joint data may be described as data for rendering the avatar (160) of the third pose. For example, the third joint data may include data regarding the joints of the avatar (160) of the third pose. For example, the fourth joint data may be described as data for rendering the avatar (160) of the fourth pose. For example, the fourth joint data may include data regarding the joints of the avatar (160) of the fourth pose.
[0087] For example, at least one processor (207) can retrieve third joint data representing the third pose from a third database for poses between the first pose and the second pose, based on a determination that the response data (426) includes the text data and includes the directive representing the second pose.
[0088] For example, at least one processor (207) may retrieve second joint data for a second pose corresponding to a directive from a second database based on identifying the pose represented by a directive included in the response data (426). For example, at least one processor (207) may display an avatar (160) of the second pose through a display (208) based on retrieving second joint data for the second pose from the second database. For example, at least one processor (207) may display an avatar (160) of the second pose through a display (208) using second joint data for the second pose stored in memory (206).
[0089] For example, the state (820) may be described as a state that displays an animation transitioning from the avatar (160) of the first pose to the avatar (160) of the second pose. For example, the third pose between the first pose and the second pose may be a pose continuous with the first pose. For example, at least one processor (207) may retrieve third joint data for the third pose from a third database to display the avatar (160) of the third pose. For example, at least one processor (207) may display the avatar (160) of the third pose using the third joint data for the third pose retrieved from the third database. For example, at least one processor (207) may display the first animation changing from the avatar (160) of the first pose to the avatar (160) of the second pose through the display (208) using the third joint data for the third pose.
[0090] The state (840) may be described as a state in which a second animation is displayed transitioning from the avatar (160) of the second pose to the avatar (160) of the first pose. For example, at least one processor (207) may display the avatar (160) of the first pose based on the completion of displaying the avatar (160) of the second pose. For example, at least one processor (207) may display the avatar (160) of the fourth pose between the second pose and the first pose through the display (208) based on the completion of displaying the avatar (160) of the second pose. For example, the fourth pose may be referred to as a return pose. For example, the return pose may be described as a pose for returning the movement of the avatar (160) to the base pose. For example, at least one processor (207) may display a second animation changing from the avatar (160) of the second pose to the avatar (160) of the first pose through the display (208) based on the completion of displaying the avatar (160) of the second pose. For example, at least one processor (207) may retrieve fourth joint data of the fourth pose from the fourth database to display the avatar (160) of the fourth pose. For example, at least one processor (207) may display the second animation through the display (208) using the fourth joint data of the fourth pose. For example, at least one processor (207) may retrieve or obtain fourth joint data representing the fourth pose from the fourth database for poses between the second pose and the first pose based on the completion of displaying the first animation. For example, at least one processor (207) can use the fourth joint data to display a second animation of an avatar (160) changing from the second pose to the first pose through the fourth pose via a display (208).
[0091] According to one embodiment, at least one processor (207) can switch the pose of the avatar (160) while the avatar (160) in a third pose is displayed. For example, at least one processor (207) can limit the interval during which the pose of the avatar (160) is switched so that the second pose of the avatar (160) can be maintained. For example, at least one processor (207) can determine the pose of the avatar (160) while the avatar (160) in a third pose is displayed. For example, while the avatar (160) in a second pose is displayed, changing the pose of the avatar (160) may not be allowed.
[0092] According to one embodiment, at least one processor (207) may loop the trigger gesture based on a decision that the trigger gesture should be sustained. For example, at least one processor (207) may maintain the avatar (160) of the trigger gesture based on a decision that the trigger gesture should be sustained.
[0093] According to one embodiment, at least one processor (207) can control a display (208) so that the next pose of the avatar (160) is changed in a return step corresponding to the return pose. For example, the electronic device (100) may be required to control the display (208) so that the change in the pose of the avatar (160) appears fast and smooth.
[0094] According to one embodiment, at least one processor (207) can display the avatar (160) of a conditional gesture through a display (208) after the display of the avatar (160) of a trigger gesture is completed.
[0095] According to one embodiment, at least one processor (207) can switch the pose of the avatar (160). For example, at least one processor (207) can switch the pose of the avatar (160) using a static inertialization technique or a dynamic inertialization technique. For example, the static inertialization technique may be described as a technique of inserting blending frames to fill the gap between source and target motion clips. For example, the dynamic inertialization technique may be described as a technique of generating a transition of pose by warping the target motion clip without inserting auxiliary frames. For example, at least one processor (207) can search for a matching frame among the stroke steps corresponding to the second database. For example, if the transition period exceeds the period corresponding to the stroke step, at least one processor (207) can switch the motion of the avatar (160) through a static inertialization technique. For example, at least one processor (207) can switch the operation of the avatar (160) through a dynamic inertia technique when the switching period is less than or equal to the period corresponding to the stroke phase.
[0096] Referring again to FIG. 7, in operation 750, at least one processor (207) can obtain second joint data representing a fourth pose continuous with the first pose by providing audio data (610) obtained using the text data and first pose data representing the features of the first pose to a trained second model (630), based on a determination that the response data (426) includes the text data and does not include the directive. For example, operation 750 may correspond to operation 330 of FIG. 3.
[0097] According to one embodiment, at least one processor (207) can obtain an audio signal corresponding to the text data by performing TTS on the text data. For example, at least one processor (207) can obtain feature data representing the features of the audio signal by performing an MFCC technique on the audio signal. For example, at least one processor (207) can obtain bit data representing the start time of the audio signal for the divided audio signal or the voice corresponding to the text data after dividing the audio signal into time intervals. For example, at least one processor (207) can obtain potential data for a second pose following (or after) the first pose and pose data for the second pose by providing the feature data and bit data included in the audio data (610) to a second model (630). For example, the second model (630) may include a plurality of discrimination models for determining the authenticity of pose data generated based on learned data. For example, the audio data (610) may include feature data representing the characteristics of the audio signal through the MFCC technique from the audio signal corresponding to the text data. For example, the audio data (610) may include bit data representing the start time of the voice within bits divided according to the time interval of the audio signal corresponding to the text data. For example, at least one processor (207) may obtain first joint data (e.g., first joint data (670)) representing the second pose by providing the potential data and the pose data to the fourth model (660). For example, at least one processor (207) may use the first joint data to display an animation of the avatar (160) changing from the first pose to the second pose through the display (208).For example, at least one processor (207) can obtain other potential data and other pose data for the second pose and a third pose consecutive to the second pose by providing the potential data and the pose data to the third model (650). For example, at least one processor (207) can obtain second joint data (e.g., second joint data (690)) representing the third pose by providing the other potential data and the other pose data to the fourth model (660). For example, at least one processor (207) can use the second joint data to display, through the display (208), another animation in which the pose of the avatar (160) changes from the second pose to the third pose.
[0098] In operation 760, when at least one processor (207) outputs a voice corresponding to the text data through a speaker (210) based on a determination that the response data (426) includes text data and does not include the directive, a second animation of an avatar (160) changing from the first pose to the fourth pose using the second joint data can be displayed through a display (208). For example, operation 760 may correspond to operation 340 of FIG. 3.
[0099] An electronic device as described above may include a memory for storing instructions. The electronic device may include at least one processor. The electronic device may include a speaker. The electronic device may include a display. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause the electronic device to display, through the display, an avatar of a first pose for interacting with a user of the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause the electronic device to obtain text data including a natural language response expressed in text by providing information related to the user data to a trained first model based on identifying user data for interaction between the electronic device and the user. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to acquire first joint data representing a second pose continuous with the first pose by providing audio data acquired using the text data and first pose data representing the features of the first pose to a trained second model. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to display, through the display, an animation of the avatar changing from the first pose to the second pose using the first joint data when outputting voice corresponding to the text data through the speaker.
[0100] According to one embodiment, the audio data may include feature data representing the characteristics of the audio signal through the Mel-frequency cepstral coefficients (MFCC) technique from the audio signal corresponding to the text data.
[0101] According to one embodiment, the audio data may include bit data indicating the start time of the voice within a bit divided according to the time interval of the audio signal corresponding to the text data.
[0102] According to one embodiment, the instructions may cause the electronic device to obtain second pose data of the avatar and first latent data for the second pose data by providing the audio data and the first pose data to the second model when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain third pose data and second latent data for generating second joint data representing a third pose continuous with the second pose by providing the second pose data and the first latent data to the third model when executed individually or collectively by the at least one processor.
[0103] According to one embodiment, the instructions may cause the electronic device to obtain the first joint data representing the second pose by providing the second pose data and the first potential data to the fourth model when executed individually or collectively by the at least one processor.
[0104] According to one embodiment, the second model may include a plurality of discrimination models for identifying the authenticity of pose data generated based on data or joint data generated based on said pose data.
[0105] According to one embodiment, the user data may include images of the user obtained through a camera and another voice of the user obtained through a microphone.
[0106] A method performed by an electronic device having a speaker and a display as described above may include an action of displaying an avatar of a first pose for interacting with a user of the electronic device through the display. The method may include an action of obtaining text data including a natural language response expressed in text by providing information related to the user data to a trained first model based on identifying user data for interaction between the electronic device and the user. The method may include an action of obtaining first joint data representing a second pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the characteristics of the first pose to a trained second model. The method may include an action of displaying an animation of the avatar changing from the first pose to the second pose using the first joint data through the display when a voice corresponding to the text data is output through the speaker.
[0107] According to one embodiment, the audio data may include feature data representing the characteristics of the audio signal through the Mel-frequency cepstral coefficients (MFCC) technique from the audio signal corresponding to the text data.
[0108] According to one embodiment, the audio data may include bit data indicating the start time of the voice within a bit divided according to the time interval of the audio signal corresponding to the text data.
[0109] According to one embodiment, the method may include an operation of obtaining second pose data of the avatar and first latent data for the second pose data by providing the audio data and the first pose data to the second model. The method may include an operation of obtaining third pose data and second latent data for generating second joint data representing a third pose continuous with the second pose by providing the second pose data and the first latent data to a third model.
[0110] According to one embodiment, the method may include the operation of obtaining the first joint data representing the second pose by providing the second pose data and the first potential data to the fourth model.
[0111] According to one embodiment, the second model may include a plurality of discrimination models for identifying the authenticity of pose data generated based on data or joint data generated based on said pose data.
[0112] According to one embodiment, the user data may include images of the user obtained through a camera and another voice of the user obtained through a microphone.
[0113] In a computer-readable storage medium storing one or more programs as described above, the one or more programs may include instructions that cause the electronic device to display, through the display, an avatar of a first pose for interacting with a user of the electronic device when executed by the electronic device having a speaker and a display. The one or more programs may include instructions that cause the electronic device to obtain text data including a natural language response expressed in text by providing information related to the user data to a trained first model, based on identifying user data for interaction between the electronic device and the user when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain first joint data representing a second pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the features of the first pose to a trained second model when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display, through the display, an animation of the avatar changing from the first pose to the second pose using the first joint data, when the electronic device is executed and outputs voice corresponding to the text data through the speaker.
[0114] According to one embodiment, the audio data may include feature data representing the characteristics of the audio signal through the Mel-frequency cepstral coefficients (MFCC) technique from the audio signal corresponding to the text data.
[0115] According to one embodiment, the audio data may include bit data indicating the start time of the voice within a bit divided according to the time interval of the audio signal corresponding to the text data.
[0116] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain second pose data of the avatar and first latent data for the second pose data by providing the audio data and the first pose data to the second model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain third pose data and second latent data for generating second joint data representing a third pose continuous with the second pose by providing the second pose data and the first latent data to a third model when executed by the electronic device.
[0117] According to one embodiment, the one or more programs may include instructions that cause the electronic device to acquire the first joint data representing the second pose by providing the second pose data and the first potential data to the fourth model when executed by the electronic device.
[0118] According to one embodiment, the second model may include a plurality of discrimination models for identifying the authenticity of pose data generated based on data or joint data generated based on said pose data.
[0119] According to one embodiment, the user data may include images of the user obtained through a camera and another voice of the user obtained through a microphone.
[0120] An electronic device as described above may include a memory for storing instructions. The electronic device may include a speaker. The electronic device may include a display. The electronic device may include at least one processor. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to display, through the display, an avatar of a first pose for interacting with a user of the electronic device. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to obtain response data based on said information related to said user data by providing said information to a trained first model based on identifying said user data for interaction between the electronic device and said user. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to display, through the display, a first animation of the avatar changing from the first pose to the second pose through the third pose, by using first joint data representing a third pose between the first pose and the second pose, when outputting a voice corresponding to the text data through the speaker based on a determination that the response data includes text data including a natural language response expressed as text and includes a directive representing a second pose of the avatar.The above instructions may cause the electronic device to acquire second joint data representing a fourth pose continuous with the first pose by providing audio data acquired using the text data and first pose data representing the features of the first pose to a trained second model, based on a determination that the response data includes the text data and does not include the directive, when the above instructions are executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to display, through the display, a second animation of the avatar changing from the first pose to the fourth pose using the second joint data, when the voice corresponding to the text data is output through the speaker, based on a determination that the response data includes the text data and does not include the directive, when the above instructions are executed individually or collectively by the at least one processor.
[0121] According to one embodiment, the instructions may cause the electronic device to search for the first joint data representing the third pose from a database of poses between the first pose and the second pose, based on a determination that the response data includes the text data and the directive representing the second pose, when executed individually or collectively by the at least one processor.
[0122] According to one embodiment, the instructions may cause the electronic device to retrieve third joint data representing a fifth pose between the second pose and the first pose from another database for poses between the second pose and the first pose, based on completing the display of the first animation when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to display, through the display, a third animation for the avatar changing from the second pose to the first pose through the fifth pose, using the third joint data when executed individually or collectively by the at least one processor.
[0123] According to one embodiment, the audio data may include feature data representing the characteristics of the audio signal through the Mel-frequency cepstral coefficients (MFCC) technique from the audio signal corresponding to the text data.
[0124] According to one embodiment, the audio data may include bit data indicating the start time of the voice within a bit divided according to the time interval of the audio signal corresponding to the text data.
[0125] According to one embodiment, the instructions may cause the electronic device to obtain second pose data of the avatar and first latent data for the second pose data by providing the audio data and the first pose data to the second model when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain third pose data and second latent data for generating third joint data representing the fourth pose and a fifth pose consecutive to the third model by providing the second pose data and the first latent data to the third model when executed individually or collectively by the at least one processor.
[0126] According to one embodiment, the instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to obtain the second joint data representing the fourth pose by providing the second pose data and the first potential data to the fourth model.
[0127] According to one embodiment, the second model may include a plurality of discrimination models for identifying the authenticity of pose data generated based on data or joint data generated based on said pose data.
[0128] According to one embodiment, the user data may include images of the user obtained through a camera and another voice of the user obtained through a microphone.
[0129] A method performed by an electronic device having a speaker and a display as described above may include an operation of obtaining a first text for a first prompt using a language model. The method may include an operation of displaying an avatar of a first pose for interacting with a user of the electronic device through the display. The method may include an operation of obtaining response data based on information related to the user data by providing information related to the user data to a trained first model based on identifying user data for interaction between the electronic device and the user. The method may include an operation of displaying a first animation of the avatar changing from the first pose to the second pose through the display using first joint data representing a third pose between the first pose and the second pose, based on a determination that the response data includes text data including a natural language response expressed as text and includes a directive representing a second pose of the avatar, when outputting a voice corresponding to the text data through the speaker. The above method may include the operation of obtaining second joint data representing a fourth pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the features of the first pose to a trained second model, based on a determination that the response data includes the text data and does not include the directive.The above method may include an operation of displaying, through the display, a second animation for the avatar changing from the first pose to the fourth pose using the second joint data when the voice corresponding to the text data is output through the speaker based on a determination that the response data includes the text data and does not include the directive.
[0130] According to one embodiment, the method may include the operation of searching for the first joint data representing the third pose from a database of poses between the first pose and the second pose, based on a determination that the response data includes the text data and includes the directive representing the second pose.
[0131] According to one embodiment, the method may include an operation of retrieving third joint data representing a fifth pose between the second pose and the first pose from another database for poses between the second pose and the first pose, based on the completion of displaying the first animation. The method may include an operation of displaying, through the display, a third animation for the avatar changing from the second pose to the first pose through the fifth pose using the third joint data.
[0132] According to one embodiment, the audio data may include feature data representing the characteristics of the audio signal through the Mel-frequency cepstral coefficients (MFCC) technique from the audio signal corresponding to the text data.
[0133] According to one embodiment, the audio data may include bit data indicating the start time of the voice within a bit divided according to the time interval of the audio signal corresponding to the text data.
[0134] According to one embodiment, the method may include an operation of obtaining second pose data of the avatar and first latent data for the second pose data by providing the audio data and the first pose data to the second model. The method may include an operation of obtaining third pose data and second latent data for generating third joint data representing a fifth pose continuous with the fourth pose by providing the second pose data and the first latent data to a third model.
[0135] According to one embodiment, the method may include the operation of obtaining the second joint data representing the fourth pose by providing the second pose data and the first potential data to the fourth model.
[0136] According to one embodiment, the method may include an operation of obtaining second pose data of the avatar and first latent data for the second pose data by providing the audio data and the first pose data to the second model. The method may include an operation of obtaining second joint data representing the fourth pose by providing the second pose data and the first latent data to the fourth model.
[0137] According to one embodiment, the method may include the operation of obtaining third pose data and second potential data for generating third joint data representing a fifth pose continuous with the fourth pose by providing the second pose data and the first potential data to a third model.
[0138] According to one embodiment, the second model may include a plurality of discrimination models for identifying the authenticity of pose data generated based on data or joint data generated based on said pose data.
[0139] According to one embodiment, the user data may include images of the user obtained through a camera and another voice of the user obtained through a microphone.
[0140] In a computer-readable storage medium in which one or more programs as described above are stored, said one or more programs may include instructions that cause said electronic device to display, through said display, an avatar of a first pose for interacting with a user of said electronic device when executed by said electronic device having a speaker and a display. said one or more programs may include instructions that cause said electronic device to obtain response data based on said information related to said user data by providing said information to a trained first model, based on identifying said user data for interaction between said electronic device and said user when executed by said electronic device. The above one or more programs may include instructions that cause the electronic device to display, through the display, a first animation of the avatar changing from the first pose to the second pose through the third pose, by using first joint data representing a third pose between the first pose and the second pose, when outputting voice corresponding to the text data through the speaker, based on a determination that when executed by the electronic device, the response data includes text data comprising a natural language response expressed as text and includes a directive representing a second pose of the avatar. The above one or more programs may include instructions that cause the electronic device to obtain second joint data representing a fourth pose continuous with the first pose by providing audio data obtained using the text data and first pose data representing the features of the first pose to a trained second model, based on a determination that when executed by the electronic device, the response data includes the text data and does not include the directive.The above one or more programs may include instructions that cause the electronic device to display, through the display, a second animation for the avatar changing from the first pose to the fourth pose using the second joint data, based on a determination that when executed by the electronic device, the response data includes the text data and does not include the directive, and when the voice corresponding to the text data is output through the speaker.
[0141] According to one embodiment, the one or more programs may include instructions that cause the electronic device to search for the first joint data representing the third pose from a database of poses between the first pose and the second pose, based on a determination that when executed by the electronic device, the response data includes the text data and includes the directive representing the second pose.
[0142] According to one embodiment, the one or more programs may include instructions that cause the electronic device to retrieve third joint data representing a fifth pose between the second pose and the first pose from another database for poses between the second pose and the first pose, based on completing the display of the first animation when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display, through the display, a third animation for the avatar changing from the second pose to the first pose through the fifth pose, using the third joint data when executed by the electronic device.
[0143] According to one embodiment, the audio data may include feature data representing the characteristics of the audio signal through the Mel-frequency cepstral coefficients (MFCC) technique from the audio signal corresponding to the text data.
[0144] According to one embodiment, the audio data may include bit data indicating the start time of the voice within a bit divided according to the time interval of the audio signal corresponding to the text data.
[0145] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain second pose data of the avatar and first latent data for the second pose data by providing the audio data and the first pose data to the second model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain third pose data and second latent data for generating third joint data representing a fifth pose continuous with the fourth pose by providing the second pose data and the first latent data to the third model when executed by the electronic device.
[0146] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain the second pose data of the avatar and the first latent data for the second pose data by providing the audio data and the first pose data to the second model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the second joint data representing the fourth pose by providing the second pose data and the first latent data to the fourth model when executed by the electronic device.
[0147] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain third pose data and second potential data for generating third joint data representing a fifth pose consecutive to the fourth pose by providing the second pose data and the first potential data to a third model when executed by the electronic device.
[0148] According to one embodiment, the second model may include a plurality of discrimination models for identifying the authenticity of pose data generated based on data or joint data generated based on said pose data. The first model may be trained to output said text data corresponding to said information using information about an image of said user and / or the voice of said user.
[0149] According to one embodiment, the user data may include images of the user obtained through a camera and another voice of the user obtained through a microphone.
[0150] According to one embodiment, the third model can be trained to output joint data regarding the position of all joints of the avatar using pose data representing the position and rotation of the joints of the avatar and latent data regarding the pose of the avatar. The fourth model can be trained to output other pose data and other latent data regarding other poses consecutive to the pose of the avatar using the pose data and the latent data.
[0151] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0152] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0153] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may continuously store a computer-executable program, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or several combined hardware, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Additionally, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.
[0154] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0155] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
1. In an electronic device, Memory comprising one or more storage media and storing instructions; speaker; Display; and It includes at least one processor comprising processing circuitry, and When the above instructions are executed individually or collectively by the at least one processor, Displaying an avatar of a first pose for interacting with a user of the electronic device through the display, and Based on identifying user data for interaction between the electronic device and the user, information related to the user data is provided to a trained first model to obtain text data including a natural language response expressed as text, and By providing audio data obtained using the text data and first pose data representing the characteristics of the first pose to a trained second model, first joint data representing a second pose continuous with the first pose is obtained, and When outputting voice corresponding to the above text data through the speaker, an animation of the avatar changing from the first pose to the second pose using the first joint data is displayed through the display. causing the above electronic device, Electronic device.
2. In claim 1, the audio data is, including feature data representing the characteristics of the audio signal through the MFCC (Mel-frequency cepstral coefficients) technique from the audio signal corresponding to the text data, Electronic device.
3. In claim 1, the audio data is, A bit data including bit data indicating the start time of the voice within a bit divided according to the time interval of the audio signal corresponding to the text data, Electronic device.
4. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, By providing the above audio data and the above first pose data to the above second model, the second pose data of the avatar and the first latent data for the second pose data are obtained, and By providing the second pose data and the first potential data to the third model, the first joint data representing the second pose is obtained. causing the above electronic device, Electronic device.
5. In Claim 4, When the above instructions are executed individually or collectively by the at least one processor, By providing the second pose data and the first potential data to the fourth model, in order to obtain third pose data and second potential data for generating second joint data representing a third pose continuous with the second pose, causing the above electronic device, Electronic device.
6. In claim 5, the third model is, Using pose data representing the position and rotation of the joints of the avatar and latent data regarding the pose of the avatar, the avatar is trained to output joint data regarding the position of all joints of the avatar, and The above fourth model is, Trained to output other pose data and other potential data for a pose consecutive to another pose of the avatar using the above pose data and the above potential data, Electronic device.
7. In claim 1, the first model is, Using information regarding the image and / or voice of the user, the system is trained to output the text data corresponding to the information, and The above second model is, A plurality of discrimination models for identifying the authenticity of pose data generated based on data or joint data generated based on said pose data, Electronic device.
8. In Claim 1, the user data is, including images of the user obtained through a camera and another voice of the user obtained through a microphone, Electronic device.
9. In an electronic device, Memory comprising one or more storage media and storing instructions; speaker; Display; and It includes at least one processor comprising processing circuitry, and When the above instructions are executed individually or collectively by the at least one processor, Displaying an avatar of a first pose for interacting with a user of the electronic device through the display, and Based on identifying user data for interaction between the electronic device and the user, by providing information related to the user data to a trained first model, response data based on the information related to the user data is obtained, and Based on the determination that the above response data includes text data containing a natural language response expressed as text and includes a directive indicating a second pose of the avatar, when a voice corresponding to the text data is output through the speaker, a first animation of the avatar changing from the first pose to the second pose through the third pose is displayed through the display by using first joint data indicating a third pose between the first pose and the second pose. Based on the decision that the above response data includes the above text data and does not include the above directive: By providing audio data obtained using the text data and first pose data representing the characteristics of the first pose to a trained second model, second joint data representing a fourth pose continuous with the first pose is obtained, and When the voice corresponding to the above text data is output through the speaker, a second animation for the avatar changing from the first pose to the fourth pose using the second joint data is displayed through the display. causing the above electronic device, Electronic device.
10. In Claim 9, When the above instructions are executed individually or collectively by the at least one processor, Based on the determination that the above response data includes the text data and includes the directive representing the second pose, to retrieve the first joint data representing the third pose from a database for poses between the first pose and the second pose. causing the above electronic device, Electronic device.
11. In Claim 9, When the above instructions are executed individually or collectively by the at least one processor, Based on the completion of the display of the first animation above, retrieve third joint data representing a fifth pose between the second pose and the first pose from another database regarding poses between the second pose and the first pose, and Using the third joint data above, to display, through the display, a third animation for the avatar changing from the second pose to the first pose through the fifth pose, causing the above electronic device, Electronic device.
12. In claim 9, the audio data is, including feature data representing the characteristics of the audio signal through the MFCC (Mel-frequency cepstral coefficients) technique from the audio signal corresponding to the text data, Electronic device.
13. In claim 9, the audio data is, A bit data including bit data indicating the start time of the voice within a bit divided according to the time interval of the audio signal corresponding to the text data, Electronic device.
14. In Claim 9, When the above instructions are executed individually or collectively by the at least one processor, By providing the above audio data and the above first pose data to the above second model, the second pose data of the avatar and the first latent data for the second pose data are obtained, and By providing the second pose data and the first potential data to the third model, in order to obtain third pose data and second potential data for generating third joint data representing a fifth pose continuous with the fourth pose, causing the above electronic device, Electronic device.
15. In Claim 14, When the above instructions are executed individually or collectively by the at least one processor, By providing the second pose data and the first potential data to the fourth model, the second joint data representing the fourth pose is obtained. causing the above electronic device, Electronic device.
16. In claim 9, the second model is, A plurality of discrimination models for identifying the authenticity of pose data generated based on data or joint data generated based on said pose data, Electronic device.
17. In claim 9, the user data is, including images of the user obtained through a camera and another voice of the user obtained through a microphone, Electronic device.
18. In a non-transient computer-readable storage medium storing one or more programs, said one or more programs, said one or more programs, When executed by an electronic device having a speaker and a display, Displaying an avatar of a first pose for interacting with a user of the electronic device through the display, and Based on identifying user data for interaction between the electronic device and the user, information related to the user data is provided to a trained first model to obtain text data including a natural language response expressed as text, and By providing audio data obtained using the text data and first pose data representing the characteristics of the first pose to a trained second model, first joint data representing a second pose continuous with the first pose is obtained, and When outputting voice corresponding to the above text data through the speaker, an animation of the avatar changing from the first pose to the second pose using the first joint data is displayed through the display. Including instructions that cause the above electronic device, Non-transient computer-readable storage media.
19. In Claim 18, When the above one or more programs are executed by the electronic device, By providing the above audio data and the above first pose data to the above second model, the second pose data of the avatar and the first latent data for the second pose data are obtained, and By providing the second pose data and the first potential data to the third model, the first joint data representing the second pose is obtained. Including instructions that cause the above electronic device, Non-transient computer-readable storage media.
20. In Claim 19, When the above one or more programs are executed by the electronic device, By providing the second pose data and the first potential data to the fourth model, in order to obtain third pose data and second potential data for generating second joint data representing a third pose continuous with the second pose, Including instructions that cause the above electronic device, Non-transient computer-readable storage media.
Citation Information
Patent Citations
Avatar movement control device and avatar movement control method
JP2024102698A
Metal foil carrier for metal foil printing machine
KR102272674B1
Error identification and model update method through information matrix-based trial design
KR102433219B1
Avatar animation using markov decision process policies
US20220130094A1
KR20240052578A