Artificial intelligence device and personalized agent generation method thereof
Patent Information
- Application Number
- PCT/KR2023/016218
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-19
- Publication Date
- 2025-09-11
AI Technical Summary
Current artificial intelligence technologies struggle to generate personalized three-dimensional agencies from two-dimensional images, as they only output three-dimensional landmarks and joint positions, failing to fully implement target targets in virtual spaces.
An artificial intelligence device that includes a processor to store learning datasets for each character, generating personalized 3D agencies by using character-specific learning datasets. The device extracts learning datasets from two-dimensional images containing characters, including operation information corresponding to their appearance, and enters this data into a pre-learned agency generation model to create customized 3D agencies.
Enables the creation of personalized three-dimensional agencies with desired appearance and operation characteristics, providing fun and interest by expressing unique movements and voices, and offering various services by learning actual or applied voices.
Smart Images

Figure KR2023016218_12092025_PF_FP_ABST
Abstract
Description
Method for creating an artificial intelligence device and its personalized agency
[0001] The present disclosure relates to an artificial intelligence device capable of generating a personalized three-dimensional agency having a user-desired appearance and movement characteristics, and a method for generating the personalized agency.
[0002] In general, artificial intelligence is a field of computer engineering and information technology that studies ways to enable computers to think, learn, and develop themselves in ways that human intelligence can do. It means enabling computers to imitate human intelligent behavior.
[0003] Furthermore, artificial intelligence does not exist in isolation; rather, it is closely related, both directly and indirectly, to other fields of computer science. In particular, in modern times, there are active attempts to introduce AI elements into various fields of information technology and utilize them to solve problems in those fields.
[0004] Meanwhile, technologies that utilize artificial intelligence to recognize and learn the surrounding situation, provide the user with desired information in the desired format, or perform desired actions or functions are being actively researched.
[0005] And, electronic devices that provide these various operations and functions can be called artificial intelligence devices.
[0006] Recently, research is being conducted on artificial intelligence technology that can learn two-dimensional images using artificial intelligence and create three-dimensional images of target objects from the learned two-dimensional images.
[0007] However, these artificial intelligence technologies have limitations in that they cannot be used in virtual spaces such as the metaverse because they only output 3D landmarks and 3D joint positions rather than implementing all parts of the target object as a 3D image through learning of 2D images, making it impossible to implement the unique detailed movement expression of the target object in 3D.
[0008] Therefore, in the future, it is necessary to develop artificial intelligence technology that can create personalized 3D agencies with the appearance and movement characteristics desired by the user through learning from 2D images.
[0009] The present disclosure aims to solve the above-mentioned problems and other problems.
[0010] The present disclosure provides an artificial intelligence device capable of generating a personalized 3D agency having an appearance and motion characteristics desired by a user through learning of a 2D image by receiving user setting information including a personalized character from a user client, inputting the user setting information into a pre-learned agency generation model, and generating a 3D agency expressing motion characteristics corresponding to the appearance of the personalized character, and a personalized agency generation method thereof.
[0011] An artificial intelligence device according to one embodiment of the present disclosure includes a memory for storing a character-specific learning dataset, and a processor for generating a personalized three-dimensional agency using the character-specific learning dataset, wherein the processor preprocesses a two-dimensional image including a character to extract a learning dataset including motion information corresponding to the appearance of the character, inputs the learning dataset into an agency generation model to learn motion characteristics of the character, and when receiving user setting information including a personalized character from a user client, inputs the user setting information into a pre-trained agency generation model to generate a three-dimensional agency expressing motion characteristics corresponding to the appearance of the personalized character.
[0012] A method for generating a personalized agency of an artificial intelligence device according to one embodiment of the present disclosure may include a step of acquiring a two-dimensional image including a character, a step of preprocessing the two-dimensional image including the character, a step of extracting a learning dataset including motion information corresponding to the appearance of the character, a step of inputting the learning dataset into an agency generation model to learn motion characteristics of the character, a step of receiving user setting information including the personalized character, and a step of inputting the user setting information into a pre-trained agency generation model to generate a three-dimensional agency expressing motion characteristics corresponding to the appearance of the personalized character.
[0013] According to one embodiment of the present disclosure, when an artificial intelligence device receives user setting information including a personalized character from a user client, the artificial intelligence device inputs the user setting information into a pre-learned agency generation model to generate a three-dimensional agency that expresses motion characteristics corresponding to the appearance of the personalized character, thereby generating a personalized three-dimensional agency having an appearance and motion characteristics desired by the user through learning of two-dimensional images.
[0014] In addition, the present disclosure can provide fun and interest to customers by extracting unique and special movement characteristic information for each personalized character and learning the unique specific movements of the personalized character, thereby generating a three-dimensional agency that expresses the unique specific movements of the personalized character.
[0015] In addition, the present disclosure can provide various services to customers by generating a three-dimensional agency that expresses not only the unique voice of a personalized character but also the user's applied voice by learning the actual voice of a personalized character or the voice that the user wishes to apply.
[0016] FIG. 1 illustrates an artificial intelligence device according to one embodiment of the present disclosure.
[0017] FIG. 2 illustrates an artificial intelligence server according to one embodiment of the present disclosure.
[0018] FIG. 3 illustrates an artificial intelligence system according to one embodiment of the present disclosure.
[0019] FIG. 4 is a diagram for explaining a personalized agency creation operation of an artificial intelligence device according to an embodiment of the present disclosure.
[0020] FIG. 5 and FIG. 6 are diagrams for explaining a two-dimensional image acquisition process including a personalized character of an artificial intelligence device according to one embodiment of the present disclosure.
[0021] FIG. 7 is a diagram for explaining a processor of an artificial intelligence device according to one embodiment of the present disclosure.
[0022] FIG. 8 is a diagram for explaining a facial information extraction process of an artificial intelligence device according to an embodiment of the present disclosure.
[0023] FIG. 9 is a diagram for explaining a process of extracting operation information of an artificial intelligence device according to an embodiment of the present disclosure.
[0024] FIG. 10 is a diagram for explaining a voice information extraction process of an artificial intelligence device according to an embodiment of the present disclosure.
[0025] FIG. 11 is a diagram illustrating a processor of an artificial intelligence device according to another embodiment of the present disclosure.
[0026] FIG. 12 and FIG. 13 are diagrams for explaining a special movement feature information extraction process of an artificial intelligence device according to an embodiment of the present disclosure.
[0027] FIG. 14 is a diagram for explaining an agency generation model of an artificial intelligence device according to one embodiment of the present disclosure.
[0028] FIG. 15 is a diagram for explaining an agency generation model of an artificial intelligence device according to another embodiment of the present disclosure.
[0029] FIG. 16 is a diagram for explaining an agency generation model of an artificial intelligence device according to another embodiment of the present disclosure.
[0030] FIG. 17 is a diagram for explaining the learning process of an agency generation model of an artificial intelligence device according to one embodiment of the present disclosure.
[0031] FIG. 18 and FIG. 19 are drawings for explaining a three-dimensional agency generation process of an artificial intelligence device according to one embodiment of the present disclosure.
[0032] FIG. 20 and FIG. 21 are drawings for explaining a three-dimensional agency generation process of an artificial intelligence device according to another embodiment of the present disclosure.
[0033] FIG. 22 is a diagram for explaining the overall operation flow of an artificial intelligence device according to one embodiment of the present disclosure.
[0034] FIG. 23 is a diagram for explaining an operation flow for providing recommended content related to a specific celebrity by an artificial intelligence device according to one embodiment of the present disclosure.
[0035] FIG. 24 and FIG. 25 are diagrams showing the provision of recommended content related to a specific celebrity by an artificial intelligence device according to one embodiment of the present disclosure.
[0036] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably only for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. In addition, when describing the embodiments disclosed in this specification, if it is determined that a specific description of a related known technology may obscure the gist of the embodiments disclosed in this specification, a detailed description thereof will be omitted. In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, and substitutes included in the spirit and technical scope of the present disclosure.
[0037] Terms that include ordinal numbers, such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used solely to distinguish one component from another.
[0038] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0039] Additionally, throughout this specification, the terms "neural network," "neural network," and "network function" may be used interchangeably. A neural network may be comprised of a set of interconnected computational units, which may generally be referred to as "nodes." These "nodes" may also be referred to as "neurons." A neural network comprises at least two or more nodes. The nodes (or neurons) constituting the neural networks may be interconnected by one or more "links."
[0040] Artificial Intelligence (AI)
[0041] Artificial intelligence (AI) is the study of artificial intelligence or the methodologies for creating it, while machine learning (ML) defines various problems in the field of AI and studies the methodologies for solving them. Machine learning is also defined as an algorithm that improves performance on a task through consistent experience.
[0042] An artificial neural network (ANN) is a model used in machine learning. It can refer to a model with problem-solving capabilities, comprised of artificial neurons (nodes) formed by the connection of synapses. An ANN can be defined by the connection patterns between neurons in different layers, the learning process that updates model parameters, and the activation function that generates output values.
[0043] An artificial neural network may include an input layer, an output layer, and optionally one or more hidden layers. Each layer may contain one or more neurons, and the artificial neural network may include synapses connecting neurons. In an artificial neural network, each neuron can output a function value of an activation function based on input signals, weights, and biases received through the synapses.
[0044] Model parameters are parameters determined through learning, including synaptic connection weights and neuron biases. Hyperparameters are parameters that must be set before learning in machine learning algorithms, including the learning rate, number of iterations, mini-batch size, and initialization function.
[0045] The goal of artificial neural network training can be seen as determining model parameters that minimize a loss function. The loss function can be used as an indicator for determining optimal model parameters during the artificial neural network training process.
[0046] Machine learning can be classified into supervised learning, unsupervised learning, and reinforcement learning depending on the learning method.
[0047] Supervised learning refers to a method for training an artificial neural network given labels for training data. Labels can refer to the correct answer (or output) that the artificial neural network must infer when inputting training data. Unsupervised learning refers to a method for training an artificial neural network without given labels for the training data. Reinforcement learning refers to a learning method in which an agent defined within a given environment is trained to select actions or action sequences that maximize cumulative rewards in each state.
[0048] Among artificial neural networks, machine learning implemented with a deep neural network (DNN) containing multiple hidden layers is sometimes called deep learning, and deep learning is a subset of machine learning. Hereinafter, "machine learning" is used to encompass deep learning.
[0049] Robot
[0050] A robot can be defined as a machine that automatically performs or operates a given task based on its own capabilities. Specifically, a robot capable of perceiving its environment, making independent judgments, and performing actions can be called an intelligent robot.
[0051] Robots can be classified into industrial, medical, household, and military types depending on their purpose or field of use.
[0052] Robots are equipped with actuators or motors that enable them to perform various physical actions, such as moving robot joints. Furthermore, mobile robots include wheels, brakes, propellers, and other actuators within their actuators, enabling them to move on the ground or fly in the air.
[0053] Self-Driving
[0054] Autonomous driving refers to the technology of driving on its own, and an autonomous vehicle refers to a vehicle that drives without user intervention or with minimal user intervention.
[0055] For example, autonomous driving can include technologies that maintain the driving lane, technologies that automatically adjust speed such as adaptive cruise control, technologies that automatically drive along a set route, and technologies that automatically set a route and drive when a destination is set.
[0056] Vehicles include vehicles equipped only with internal combustion engines, hybrid vehicles equipped with both internal combustion engines and electric motors, and electric vehicles equipped only with electric motors, and may include not only automobiles but also trains, motorcycles, etc.
[0057] At this time, autonomous vehicles can be viewed as robots with autonomous driving functions.
[0058] Extended Reality (XR)
[0059] Extended reality is a general term for virtual reality (VR), augmented reality (AR), and mixed reality (MR). VR technology presents real-world objects and backgrounds as CG images only, AR technology presents virtual CG images over images of real objects, and MR technology is a computer graphics technology that blends and combines virtual objects with the real world.
[0060] MR technology is similar to AR in that it presents both real and virtual objects simultaneously. However, while AR uses virtual objects to complement real objects, MR uses virtual and real objects on an equal footing.
[0061] XR technology can be applied to HMD (Head-Mount Display), HUD (Head-Up Display), mobile phones, tablet PCs, laptops, desktops, TVs, digital signage, etc., and devices to which XR technology is applied can be called XR devices.
[0062] Figure 1 illustrates an AI device (100) according to one embodiment of the present disclosure.
[0063] The AI device (100) can be implemented as a fixed device or a movable device, such as a TV, a projector, a mobile phone, a smart phone, a desktop computer, a laptop, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation device, a tablet PC, a wearable device, a set-top box (STB), a DMB receiver, a radio, a washing machine, a refrigerator, a desktop computer, digital signage, a robot, a vehicle, etc.
[0064] Referring to FIG. 1, the AI device (100) may include a communication unit (110), an input unit (120), a learning processor (130), a sensing unit (140), an output unit (150), a memory (170), and a processor (180).
[0065] The communication unit (110) can transmit and receive data with external devices such as other AI devices (100a to 100e) or AI servers (200) using wired or wireless communication technology. For example, the communication unit (110) can transmit and receive sensor information, user input, learning models, control signals, etc. with external devices.
[0066] At this time, the communication technologies used by the communication unit (110) include GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth (Bluetooth), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), etc.
[0067] The input unit (120) can obtain various types of data.
[0068] At this time, the input unit (120) may include a camera for inputting a video signal, a microphone for receiving an audio signal, a user input unit for receiving information from a user, etc. Here, the camera or microphone may be treated as a sensor, and the signal obtained from the camera or microphone may be referred to as sensing data or sensor information.
[0069] The input unit (120) can obtain input data to be used when obtaining output using learning data and learning models for model learning. The input unit (120) can also obtain unprocessed input data, in which case the processor (180) or learning processor (130) can extract input features as preprocessing for the input data.
[0070] The learning processor (130) can train a model composed of an artificial neural network using learning data. Here, the trained artificial neural network may be referred to as a learning model. The learning model can be used to infer result values for new input data other than the learning data, and the inferred values can be used as a basis for making decisions regarding certain actions.
[0071] At this time, the running processor (130) can perform AI processing together with the running processor (240) of the AI server (200) of FIG. 2.
[0072] At this time, the running processor (130) may include a memory integrated or implemented in the AI device (100). Alternatively, the running processor (130) may be implemented using a memory (170), an external memory directly coupled to the AI device (100), or a memory maintained in an external device.
[0073] The sensing unit (140) can obtain at least one of internal information of the AI device (100), information about the surrounding environment of the AI device (100), and user information using various sensors.
[0074] At this time, the sensors included in the sensing unit (140) include a proximity sensor, a light sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a lidar, a radar, etc.
[0075] The output unit (150) can generate output related to vision, hearing, or touch.
[0076] At this time, the output unit (150) may include a display unit that outputs visual information, a speaker that outputs auditory information, a haptic module that outputs tactile information, etc.
[0077] The memory (170) can store data that supports various functions of the AI device (100). For example, the memory (170) can store input data, learning data, learning models, learning history, etc. obtained from the input unit (120).
[0078] The processor (180) may determine at least one executable operation of the AI device (100) based on information determined or generated using a data analysis algorithm or a machine learning algorithm. Then, the processor (180) may control components of the AI device (100) to perform the determined operation.
[0079] To this end, the processor (180) may request, search, receive, or utilize data from the running processor (130) or memory (170), and control components of the AI device (100) to execute at least one of the executable operations, a predicted operation, or an operation determined to be desirable.
[0080] At this time, if connection of an external device is required to perform a determined operation, the processor (180) can generate a control signal for controlling the external device and transmit the generated control signal to the external device.
[0081] The processor (180) can obtain intent information for user input and determine the user's requirement based on the obtained intent information.
[0082] At this time, the processor (180) can obtain intent information corresponding to the user input by using at least one of an STT (Speech To Text) engine for converting voice input into a string or a natural language processing (NLP) engine for obtaining intent information of natural language.
[0083] At this time, at least one of the STT engine or the NLP engine may be configured with an artificial neural network, at least in part, trained according to a machine learning algorithm. Furthermore, at least one of the STT engine or the NLP engine may be trained by the learning processor (130), the learning processor (240) of the AI server (200), or through distributed processing thereof.
[0084] The processor (180) can collect history information including the operation details of the AI device (100) or the user's feedback on the operation, and store the information in the memory (170) or the learning processor (130), or transmit the information to an external device such as an AI server (200). The collected history information can be used to update the learning model.
[0085] The processor (180) can control at least some of the components of the AI device (100) to drive an application program stored in the memory (170). Furthermore, the processor (180) can operate two or more of the components included in the AI device (100) in combination to drive the application program.
[0086] FIG. 2 illustrates an AI server (200) according to one embodiment of the present disclosure.
[0087] Referring to FIG. 2, the AI server (200) may refer to a device that trains an artificial neural network using a machine learning algorithm or utilizes a trained artificial neural network. Here, the AI server (200) may be composed of multiple servers to perform distributed processing, and may be defined as a 5G network. In this case, the AI server (200) may be included as part of the AI device (100) and may perform at least a portion of the AI processing.
[0088] The AI server (200) may include a communication unit (210), memory (230), a learning processor (240), and a processor (260).
[0089] The communication unit (210) can transmit and receive data with an external device such as an AI device (100).
[0090] The memory (230) may include a model storage unit (231). The model storage unit (231) may store a model (or artificial neural network, 231a) that is being learned or has been learned through the learning processor (240).
[0091] The learning processor (240) can train an artificial neural network (231a) using learning data. The learning model can be used while mounted on the AI server (200) of the artificial neural network, or can be mounted on an external device such as an AI device (100).
[0092] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (230).
[0093] The processor (260) can infer a result value for new input data using a learning model and generate a response or control command based on the inferred result value.
[0094] Figure 3 shows an AI system (1) according to one embodiment of the present invention.
[0095] Referring to FIG. 3, an AI system (1) is connected to a cloud network (10) by at least one of an AI server (200), a robot (100a), an autonomous vehicle (100b), an XR device (100c), a smartphone (100d), or an appliance (100e). Here, a robot (100a), an autonomous vehicle (100b), an XR device (100c), a smartphone (100d), or an appliance (100e) to which AI technology is applied may be referred to as an AI device (100a to 100e).
[0096] A cloud network (10) may refer to a network that constitutes part of a cloud computing infrastructure or exists within a cloud computing infrastructure. Here, the cloud network (10) may be configured using a 3G network, a 4G or LTE (Long Term Evolution) network, a 5G network, etc.
[0097] That is, each device (100a to 100e, 200) constituting the AI system (1) can be connected to each other via a cloud network (10). In particular, each device (100a to 100e, 200) can communicate with each other via a base station, but can also communicate with each other directly without going through a base station.
[0098] The AI server (200) may include a server that performs AI processing and a server that performs operations on big data.
[0099] The AI server (200) is connected to at least one of the AI devices constituting the AI system (1), such as a robot (100a), an autonomous vehicle (100b), an XR device (100c), a smartphone (100d), or a home appliance (100e), through a cloud network (10), and can assist at least part of the AI processing of the connected AI devices (100a to 100e).
[0100] At this time, the AI server (200) can train an artificial neural network according to a machine learning algorithm on behalf of the AI devices (100a to 100e), and can directly store the learning model or transmit it to the AI devices (100a to 100e).
[0101] At this time, the AI server (200) can receive input data from the AI devices (100a to 100e), infer a result value for the received input data using a learning model, and generate a response or control command based on the inferred result value and transmit it to the AI devices (100a to 100e).
[0102] Alternatively, the AI device (100a to 100e) may infer a result value for input data using a direct learning model and generate a response or control command based on the inferred result value.
[0103] Below, various embodiments of AI devices (100a to 100e) to which the above-described technology is applied are described. Here, the AI devices (100a to 100e) illustrated in FIG. 3 can be viewed as specific embodiments of the AI device (100) illustrated in FIG. 1.
[0104] <AI+로봇>
[0105] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, etc. by applying AI technology.
[0106] The robot (100a) may include a robot control module for controlling movement, and the robot control module may mean a software module or a chip that implements the same in hardware.
[0107] The robot (100a) can obtain status information of the robot (100a), detect (recognize) the surrounding environment and objects, generate map data, determine a movement path and driving plan, determine a response to user interaction, or determine an action using sensor information obtained from various types of sensors.
[0108] Here, the robot (100a) can use sensor information acquired from at least one sensor among lidar, radar, and camera to determine a movement path and driving plan.
[0109] The robot (100a) can perform the above-described operations using a learning model comprised of at least one artificial neural network. For example, the robot (100a) can recognize its surroundings and objects using the learning model, and determine operations using the recognized surrounding environment information or object information. Here, the learning model may be learned directly by the robot (100a) or by an external device such as an AI server (200).
[0110] At this time, the robot (100a) may perform an action by generating a result using a direct learning model, but may also perform an action by transmitting sensor information to an external device such as an AI server (200) and receiving the result generated accordingly.
[0111] The robot (100a) can determine a movement path and a driving plan using at least one of map data, object information detected from sensor information, or object information acquired from an external device, and control a driving unit to drive the robot (100a) according to the determined movement path and driving plan.
[0112] Map data may include object identification information for various objects positioned in the space where the robot (100a) moves. For example, map data may include object identification information for fixed objects such as walls and doors, as well as movable objects such as flower pots and desks. Furthermore, object identification information may include name, type, distance, location, etc.
[0113] Additionally, the robot (100a) can perform actions or drive by controlling the driving unit based on the user's control / interaction. At this time, the robot (100a) can acquire intention information regarding the interaction based on the user's actions or voice utterances, and determine a response based on the acquired intention information to perform the action.
[0114] <AI+자율주행>
[0115] An autonomous vehicle (100b) can be implemented as a mobile robot, vehicle, or unmanned aerial vehicle by applying AI technology.
[0116] The autonomous vehicle (100b) may include an autonomous driving control module for controlling autonomous driving functions. The autonomous driving control module may refer to a software module or a chip implementing the same as hardware. The autonomous driving control module may be included internally as a component of the autonomous vehicle (100b), but may also be configured as separate hardware and connected to the exterior of the autonomous vehicle (100b).
[0117] An autonomous vehicle (100b) can obtain status information of the autonomous vehicle (100b), detect (recognize) the surrounding environment and objects, generate map data, determine a movement path and driving plan, or determine an action using sensor information obtained from various types of sensors.
[0118] Here, the autonomous vehicle (100b) can use sensor information acquired from at least one sensor among lidar, radar, and camera, similar to the robot (100a), to determine a movement path and driving plan.
[0119] In particular, the autonomous vehicle (100b) can recognize the environment or objects in an area where the field of view is obstructed or an area beyond a certain distance by receiving sensor information from external devices, or can receive information recognized directly from external devices.
[0120] The autonomous vehicle (100b) can perform the above-described operations using a learning model comprised of at least one artificial neural network. For example, the autonomous vehicle (100b) can recognize its surroundings and objects using the learning model, and determine a driving route using the recognized surrounding environment information or object information. Here, the learning model may be learned directly by the autonomous vehicle (100b) or by an external device such as an AI server (200).
[0121] At this time, the autonomous vehicle (100b) may perform an action by generating a result using a direct learning model, but may also perform an action by transmitting sensor information to an external device such as an AI server (200) and receiving the result generated accordingly.
[0122] An autonomous vehicle (100b) can determine a movement path and a driving plan using at least one of map data, object information detected from sensor information, or object information acquired from an external device, and control a driving unit to drive the autonomous vehicle (100b) according to the determined movement path and driving plan.
[0123] Map data may include object identification information for various objects located in the space (e.g., a road) where the autonomous vehicle (100b) travels. For example, map data may include object identification information for fixed objects such as streetlights, rocks, and buildings, as well as movable objects such as vehicles and pedestrians. Furthermore, object identification information may include name, type, distance, location, and the like.
[0124] Additionally, the autonomous vehicle (100b) can perform actions or drive by controlling the driving unit based on the user's control / interaction. At this time, the autonomous vehicle (100b) can acquire intention information regarding the interaction based on the user's actions or voice utterances, and determine a response based on the acquired intention information to perform the action.
[0125] <AI+XR>
[0126] The XR device (100c) can be implemented as an HMD (Head-Mount Display), a HUD (Head-Up Display) installed in a vehicle, a television, a mobile phone, a smart phone, a computer, a wearable device, a home appliance, digital signage, a vehicle, a fixed robot, or a mobile robot by applying AI technology.
[0127] The XR device (100c) can obtain information about surrounding space or real objects by analyzing 3D point cloud data or image data acquired through various sensors or from an external device to generate location data and attribute data for 3D points, and can render and output an XR object to be output. For example, the XR device (100c) can output an XR object including additional information about a recognized object in correspondence with the recognized object.
[0128] The XR device (100c) can perform the above-described operations using a learning model composed of at least one artificial neural network. For example, the XR device (100c) can recognize a real-world object from 3D point cloud data or image data using the learning model, and provide information corresponding to the recognized real-world object. Here, the learning model may be learned directly in the XR device (100c) or learned from an external device such as an AI server (200).
[0129] At this time, the XR device (100c) may perform an operation by generating a result using a direct learning model, but may also perform an operation by transmitting sensor information to an external device such as an AI server (200) and receiving the result generated accordingly.
[0130] <AI+로봇+자율주행>
[0131] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, etc. by applying AI technology and autonomous driving technology.
[0132] A robot (100a) to which AI technology and autonomous driving technology are applied may refer to a robot itself with autonomous driving function, or a robot (100a) that interacts with an autonomous vehicle (100b).
[0133] A robot (100a) with an autonomous driving function can be a general term for devices that move on their own along a given path without user control or move by determining the path on their own.
[0134] A robot (100a) with autonomous driving capabilities and an autonomous vehicle (100b) may use a common sensing method to determine one or more of a movement path or a driving plan. For example, a robot (100a) with autonomous driving capabilities and an autonomous vehicle (100b) may use information sensed via lidar, radar, and cameras to determine one or more of a movement path or a driving plan.
[0135] A robot (100a) interacting with an autonomous vehicle (100b) may exist separately from the autonomous vehicle (100b), and may be linked to autonomous driving functions within the autonomous vehicle (100b) or perform actions linked to a user riding in the autonomous vehicle (100b).
[0136] At this time, the robot (100a) interacting with the autonomous vehicle (100b) can control or assist the autonomous driving function of the autonomous vehicle (100b) by acquiring sensor information on behalf of the autonomous vehicle (100b) and providing it to the autonomous vehicle (100b), or by acquiring sensor information and generating surrounding environment information or object information and providing it to the autonomous vehicle (100b).
[0137] Alternatively, a robot (100a) interacting with an autonomous vehicle (100b) may monitor a user riding in the autonomous vehicle (100b) or control functions of the autonomous vehicle (100b) through interaction with the user. For example, if the robot (100a) determines that the driver is drowsy, it may activate the autonomous driving function of the autonomous vehicle (100b) or assist in controlling the driving unit of the autonomous vehicle (100b). Here, the functions of the autonomous vehicle (100b) controlled by the robot (100a) may include not only the autonomous driving function, but also functions provided by a navigation system or audio system installed inside the autonomous vehicle (100b).
[0138] Alternatively, a robot (100a) interacting with an autonomous vehicle (100b) may provide information to the autonomous vehicle (100b) or assist functions from outside the autonomous vehicle (100b). For example, the robot (100a) may provide traffic information, including signal information, to the autonomous vehicle (100b), such as a smart traffic light, or may interact with the autonomous vehicle (100b) to automatically connect an electric charger to a charging port, such as an automatic electric charger for an electric vehicle.
[0139] <AI+로봇+XR>
[0140] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, drone, etc. by applying AI technology and XR technology.
[0141] A robot (100a) to which XR technology is applied may refer to a robot that is the subject of control / interaction within an XR image. In this case, the robot (100a) is distinct from the XR device (100c) and can be linked with each other.
[0142] When a robot (100a) that is the target of control / interaction within an XR image obtains sensor information from sensors including a camera, the robot (100a) or the XR device (100c) can generate an XR image based on the sensor information, and the XR device (100c) can output the generated XR image. In addition, the robot (100a) can operate based on a control signal input through the XR device (100c) or a user's interaction.
[0143] For example, a user can check an XR image corresponding to the viewpoint of a remotely connected robot (100a) through an external device such as an XR device (100c), and through interaction, adjust the autonomous driving path of the robot (100a), control the operation or driving, or check information on surrounding objects.
[0144] <AI+자율주행+XR>
[0145] Autonomous vehicles (100b) can be implemented as mobile robots, vehicles, unmanned aerial vehicles, etc. by applying AI technology and XR technology.
[0146] An autonomous vehicle (100b) to which XR technology is applied may refer to an autonomous vehicle equipped with a means for providing XR images, an autonomous vehicle that is the subject of control / interaction within an XR image, etc. In particular, an autonomous vehicle (100b) that is the subject of control / interaction within an XR image is distinct from an XR device (100c) and can be linked with each other.
[0147] An autonomous vehicle (100b) equipped with a means for providing XR images can acquire sensor information from sensors including cameras and output XR images generated based on the acquired sensor information. For example, the autonomous vehicle (100b) can be equipped with a HUD to output XR images, thereby providing passengers with XR objects corresponding to real objects or objects on the screen.
[0148] At this time, when the XR object is output to the HUD, at least a part of the XR object may be output so as to overlap with an actual object toward which the passenger's gaze is directed. On the other hand, when the XR object is output to a display provided inside the autonomous vehicle (100b), at least a part of the XR object may be output so as to overlap with an object on the screen. For example, the autonomous vehicle (100b) may output XR objects corresponding to objects such as a road, another vehicle, a traffic light, a traffic sign, a two-wheeled vehicle, a pedestrian, a building, etc.
[0149] When an autonomous vehicle (100b) that is the target of control / interaction within an XR image obtains sensor information from sensors including a camera, the autonomous vehicle (100b) or the XR device (100c) can generate an XR image based on the sensor information, and the XR device (100c) can output the generated XR image. In addition, the autonomous vehicle (100b) can operate based on a control signal input through an external device such as the XR device (100c) or a user's interaction.
[0150] FIG. 4 is a diagram for explaining the operation of an artificial intelligence device according to one embodiment of the present disclosure.
[0151] As illustrated in FIG. 4, the artificial intelligence device (100) of the present disclosure may include a memory (170) that stores a character-specific learning dataset, and a processor (180) that generates a personalized 3D agency using the character-specific learning dataset.
[0152] Here, the processor (180) preprocesses a two-dimensional image (20) including a character (22) to extract a learning data set including motion information corresponding to the appearance of the character, inputs the learning data set into an agency generation model to learn the motion characteristics of the character, and when receiving user setting information including a personalized character from a user client, inputs the user setting information into the pre-learned agency generation model to generate a three-dimensional agency (40) expressing motion characteristics corresponding to the appearance of the personalized character.
[0153] When receiving user setting information, the processor (180) requests setting information for agency creation from the user client when an agency creation request is received from the user client, and when receiving setting information for agency creation from the user client, it can check whether basic image data corresponding to a personalized character is included in the setting information.
[0154] Here, the processor (180) can request the user client for appearance information of a personalized character for agency creation if basic image data is not included.
[0155] For example, the appearance information of a personalized character may include the personal appearance information of a specific celebrity, mixed appearance information of multiple celebrities, modified appearance information of a specific celebrity, the user's personal appearance information, appearance information of a basic character, appearance information of a new character, etc.
[0156] In addition, when receiving setup information for agency creation, the processor (180) can check whether at least one of the user client's system information and character creation mode information is included in the setup information in addition to the basic image data.
[0157] Here, the processor (180) can generate a 3D agency that expresses motion characteristics corresponding to the appearance of a personalized character based on basic image data included in the setting information, and determine a transmission method of the 3D agency generated based on at least one of system information and character creation mode information included in the setting information.
[0158] At this time, the processor (180) can transmit 3D agency image data and motion data to the user client in streaming mode if at least one of the system information and character creation mode information of the user client is not included.
[0159] In some cases, the processor (180) may determine either a first transmission mode that transmits only 3D agency image data or a second transmission mode that transmits only 3D agency motion data based on system information when system information of the user client is included.
[0160] For example, the processor (180) may determine, based on the system information included in the system information, a first transmission mode for transmitting only 3D agency image data when the specifications of the user client are less than a set value, and a second transmission mode for transmitting only 3D agency motion data when the specifications of the user client are greater than or equal to the set value.
[0161] Here, the user client can generate a 3D agency that expresses motion characteristics corresponding to the motion data based on the motion data by receiving only the 3D agency motion data.
[0162] In another case, the processor (180) may determine either a first transmission mode that transmits only 3D agency image data or a second transmission mode that transmits only 3D agency motion data based on the character generation mode information when the character generation mode information of the user client is included.
[0163] Here, the processor (180) may determine, based on the character creation mode information of the user client, that the character creation mode is a large mode that displays characters on the entire display screen, and transmits only 3D agency image data in a first transmission mode, and, based on the character creation mode information, determine, based on the character creation mode information, that the character creation mode is a small mode that displays characters on a part of the display screen, and transmits only 3D agency motion data in a second transmission mode.
[0164] In addition, the processor (180) determines the first transmission mode for transmitting only 3D agency image data when the character creation mode is a large mode for displaying characters on the entire display screen, checks the specifications of the user client, and if the specifications of the user client are equal to or greater than the set value, changes the determined first transmission mode to a second transmission mode for transmitting only 3D agency motion data.
[0165] At this time, the user client can generate a 3D agency that expresses motion characteristics corresponding to the motion data based on the motion data by receiving only the 3D agency motion data.
[0166] Next, the processor (180) obtains user setting information including a personalized character from a user client when generating a 3D agency, obtains text information for generating motion data of the personalized character, inputs the obtained user setting information and text information into the pre-learned agency generation model to generate motion data and motion data-based image data corresponding to the appearance of the personalized character, and transmits only the motion data-based image data to the user client corresponding to the user setting information, or transmits only the motion data to the user client, thereby generating a 3D agency.
[0167] Here, the processor (180) can determine either a first transmission mode that transmits only 3D agency image data or a second transmission mode that transmits only 3D agency motion data based on the system information when the user client's system information is included in the user setting information.
[0168] For example, the processor (180) may analyze system information and determine a first transmission mode that transmits only 3D agency image data if the specifications of the user client are less than a set value, and may determine a second transmission mode that transmits only 3D agency motion data if the specifications of the user client are greater than or equal to the set value.
[0169] At this time, the user client can generate a 3D agency that expresses motion characteristics corresponding to the motion data based on the motion data by receiving only the 3D agency motion data.
[0170] In some cases, the processor (180) may determine either a first transmission mode that transmits only 3D agency image data or a second transmission mode that transmits only 3D agency motion data based on the character generation mode information when the character generation mode information of the user client is included in the user setting information.
[0171] Here, the processor (180) may analyze the character creation mode information and determine, if the character creation mode is a large mode that displays characters on the entire display screen, to be a first transmission mode that transmits only 3D agency image data, and if the character creation mode is a small mode that displays characters on a portion of the display screen, to be a second transmission mode that transmits only 3D agency motion data.
[0172] In addition, the processor (180) determines the first transmission mode for transmitting only 3D agency image data when the character creation mode is a large mode for displaying characters on the entire display screen, checks the specifications of the user client, and if the specifications of the user client are equal to or greater than the set value, changes the determined first transmission mode to a second transmission mode for transmitting only 3D agency motion data.
[0173] At this time, the user client can generate a 3D agency that expresses motion characteristics corresponding to the motion data based on the motion data by receiving only the 3D agency motion data.
[0174] Next, when generating a 3D agency, the processor (180) analyzes the appearance of the personalized character and, if it is predicted that a 3D agency with a high similarity to a specific celebrity will be generated, collects content information related to the specific celebrity, and recommends content related to the specific celebrity based on the collected content information.
[0175] For example, when the appearance of a personalized character is selected, the processor (180) can check whether the personalized character is a celebrity, analyze the appearance of the personalized character if the personalized character is not a celebrity, select a specific celebrity with the highest similarity to the appearance of the personalized character, collect content information related to the selected specific celebrity, and expose the content information related to the specific celebrity collected based on the user's content usage pattern to content frequently used by the user.
[0176] Here, when confirming whether a personalized character is a celebrity, the processor (180) can confirm whether character information is included in the setting information for agency creation received from the user client, and recognize the personalized character as a celebrity based on an identifier corresponding to a specific celebrity included in the character information.
[0177] At this time, if the character information does not include an identifier corresponding to a specific celebrity, the processor (180) can analyze the appearance of the personalized character and select a specific celebrity with the highest similarity to the appearance of the personalized character.
[0178] That is, if the character information does not include an identifier corresponding to a specific celebrity, the processor (180) can vectorize the appearance of the personalized character using a feature extraction algorithm, and select a specific celebrity most similar to the appearance of the vectorized personalized character from a data pool that collects appearance image data of multiple celebrities, including singers, actors, entertainers, etc.
[0179] And, when collecting content information related to a specific celebrity, the processor (180) can search and collect content information including hairstyles, makeup, clothes, songs, movies, etc. related to the specific celebrity.
[0180] Additionally, the processor (180) can extract the user's content usage pattern by searching the user's content usage history information when collecting content information related to a specific celebrity.
[0181] Here, the processor (180) can expose content information related to a specific celebrity collected based on the user's content usage pattern to content frequently used by the user.
[0182] In addition, the processor (180) can check the user's interest in specific celebrity-related content information through content in which specific celebrity-related content information is exposed, and if the user's interest in specific celebrity-related content information is above a preset value, the specific celebrity-related content information can be registered in content frequently used by the user.
[0183] In some cases, the processor (180) may exclude content information related to a specific celebrity from content frequently used by the user if the user interest in the content information related to a specific celebrity is below a preset value.
[0184] Meanwhile, when a processor (180) obtains a two-dimensional image (20), if a user input for selecting a personalized character (22) is received, the processor (180) obtains basic information of the subject corresponding to the selected personalized character (22), and based on the basic information of the subject, obtains a two-dimensional image (20) including the selected personalized character (22) from at least one of an internal server and an external server.
[0185] Here, the two-dimensional image (20) obtained from the internal server and the external server may be a video including a target image corresponding to a personalized character (22) and his / her voice, but this is only one example and is not limited thereto.
[0186] As another example, when acquiring a two-dimensional image (20), the processor (180) controls the camera unit to capture a subject corresponding to the selected personalized character when a user input for selecting a personalized character (22) is received, and can acquire a two-dimensional image (20) including the personalized character (22) captured from the camera unit.
[0187] Here, the two-dimensional image (20) captured from the camera unit may be a video including an image of a subject corresponding to a personalized character (22) and his / her voice, but this is only one example and is not limited thereto.
[0188] At this time, the camera unit may include an RGB camera that captures an RGB image of the subject corresponding to the selected personalized character, and a depth camera that acquires a 3D point cloud of the subject corresponding to the selected personalized character.
[0189] Additionally, the camera unit may further include a microphone for acquiring audio data of the subject corresponding to the selected personalized character.
[0190] Next, the processor (180), when preprocessing a two-dimensional image (20), checks whether a pre-selected personalized character (22) is included in the acquired two-dimensional image (20), and if the two-dimensional image (20) includes a personalized character (22), assigns an identification code to the personalized character (22) included in the two-dimensional image (20), extracts a learning dataset including voice information, facial information, and motion information corresponding to the personalized character (22) to which the identification code is assigned, and stores the learning dataset in the memory (170) for each identification code of the personalized character (22).
[0191] Here, when extracting a learning dataset including facial information, the processor (180) extracts a facial landmark from the face of a personalized character (22) included in a two-dimensional image (20), and inputs the two-dimensional image (20) and the facial landmark of the personalized character (22) into a first neural network model that has been pre-learned to extract facial information including facial movement data of the personalized character (22).
[0192] As an example, the first neural network model may include a 3D Morphable Model (3DMM) algorithm, but this is only an example and is not limited thereto.
[0193] In addition, when extracting a learning dataset including motion information, the processor (180) extracts three-dimensional keypoints (3D keypoints) for joint positions from the body of a personalized character (22) included in a two-dimensional image (20), and inputs the three-dimensional keypoints of the two-dimensional image (20) and the personalized character (22) into a pre-learned second neural network model to extract motion information including three-dimensional rotation parameters of joints corresponding to the motion of the personalized character (22).
[0194] As an example, the second neural network model may include a 3D Rotation Model algorithm, but this is only an example and is not limited thereto.
[0195] In addition, when the three-dimensional rotation parameters of the joints are extracted, the processor (180) can analyze the degree of movement of each joint based on the three-dimensional rotation parameters and extract a motion control reference value for each body part of the personalized character (22) for each frame of the two-dimensional image (20).
[0196] For example, the motion control reference values may include motion control reference values for the hand position and its movement speed, the head position and its movement speed, the foot position and its movement speed, the neck position and its movement speed, the arm position and its movement speed, the leg position and its movement speed, and the waist position and its movement speed among the body parts of the personalized character (22), but this is only one example and is not limited thereto.
[0197] Here, the motion control reference value can have different values for each personalized character.
[0198] For example, the motion control threshold may vary depending on the physical condition of the personalized character.
[0199] In addition, when extracting a learning data set including voice information, the processor (180) can extract audio data of a personalized character (22) included in a two-dimensional image (20), input the audio data into a pre-trained third neural network model, and extract voice information including a sentence corresponding to the voice of the personalized character (22), timing information of words within the sentence, and voice features of the personalized character (22).
[0200] For example, the third neural network model may include a STT (Speech To Text) model, a Forced Alignment Model, and a voice model, but this is only an example and is not limited thereto.
[0201] Here, the processor (180) can extract sentences corresponding to the voice of the personalized character (22) by converting audio data into text when extracting voice information, extract timing information of words in the sentences extracted from the voice of the personalized character (22), and extract voice features of the personalized character (22) from audio data.
[0202] In addition, the processor (180) can extract a learning dataset including special motion feature information from a personalized character (22) included in a two-dimensional image (20), and store the learning dataset including the special motion feature information in the memory (170) according to the identification code of the personalized character (22).
[0203] Here, the processor (180), when extracting a learning dataset including special movement feature information, selects a specific section from a two-dimensional image (20) including a personalized character (22), vectorizes facial information and movement information extracted from the personalized character (22) of the two-dimensional image (20) corresponding to the selected specific section, analyzes the distribution of the vectorized data to determine whether there is data whose occurrence frequency is higher than a reference frequency, and if there is data whose occurrence frequency is higher than the reference frequency, determines the corresponding data as a special movement feature, and extracts special movement feature information based on the facial information and movement information of the specific section corresponding to the data determined to be the special movement feature.
[0204] For example, when selecting a specific section, the processor (180) may use a keyframe extraction algorithm to select a section in which the personalized character (22) performs a special motion among the entire section of the two-dimensional image (20) including the personalized character (22), as the specific section.
[0205] Here, the processor (180) can recognize at least one of the unique facial expression, unique gesture, and unique body movement shape of the personalized character (22) as a special movement of the personalized character (22).
[0206] In some cases, when a specific section is selected from a two-dimensional image (20), the processor (180) extracts a face landmark from the face of a personalized character (22) included in the two-dimensional image (20) corresponding to the selected specific section, inputs the two-dimensional image (20) corresponding to the specific section and the face landmark of the personalized character (22) into a first neural network model that has been pre-learned to extract facial information including facial movement data of the personalized character (22), extracts three-dimensional keypoints for joint positions from the body of the personalized character (22) included in the two-dimensional image (20) corresponding to the selected specific section, and inputs the two-dimensional image (20) corresponding to the specific section and the three-dimensional keypoints of the personalized character (22) into a second neural network model that has been pre-learned to extract motion information including three-dimensional rotation parameters of joints corresponding to the motion of the personalized character (22).
[0207] For example, the first neural network model may include a 3D Morphable Model (3DMM) algorithm, and the second neural network model may include a 3D Rotation Model algorithm, but this is only an example and is not limited thereto.
[0208] In another case, when a specific section is selected from a two-dimensional image (20), the processor (180) checks whether facial information and motion information extracted from a personalized character (22) of the two-dimensional image (20) corresponding to the selected specific section exist in the memory (170), and if facial information and motion information of the personalized character (22) corresponding to the specific section exist, the processor can vectorize the facial information and motion information of the personalized character (22) corresponding to the specific section.
[0209] Next, as an example, when learning the motion characteristics of a personalized character (22), the processor (180) inputs a dataset including an identification code, voice information, facial information, and motion information corresponding to the personalized character (22) into an agency generation model, thereby learning the facial movements of the personalized character (22) based on the identification code, voice information, and facial information, and learning the three-dimensional rotation parameters of joints corresponding to the personalized character (22) based on the identification code, voice information, and motion information.
[0210] Here, the processor (180) can learn the facial movements of the personalized character (22) by using an identification code and a feature vector to identify the personalized character (22), and learn the facial movements of the personalized character (22) to be synchronized with the voice timing of the personalized character (22) based on the voice information including the voice features of the personalized character (22), the timing information of words within the sentence, and the facial information including the facial movement data of the personalized character (22), corresponding to the voice of the personalized character (22).
[0211] And, when learning the three-dimensional rotation parameters of the joints corresponding to the personalized character (22), the processor (180) can identify the personalized character (22) using an identification code and a feature vector, and learn the three-dimensional rotation parameters of the joints corresponding to the personalized character (22) so as to be synchronized with the voice timing of the personalized character (22) based on the voice information including the sentence corresponding to the voice of the personalized character (22), the timing information of the words in the sentence, the voice features of the personalized character (22), the three-dimensional rotation parameters of the joints corresponding to the movement of the personalized character (22), and the motion information including the motion control reference values for each body part of the personalized character (22).
[0212] Here, the agency generation model may include a face motion model that learns facial movements of a personalized character based on an identification code, voice information, and face information, and a body motion model that learns three-dimensional rotation parameters of joints corresponding to the personalized character based on an identification code, voice information, and motion information.
[0213] As another embodiment, when learning the motion characteristics of a personalized character (22), the processor (180) inputs a dataset including an identification code, voice information, face information, motion information, and text information corresponding to the personalized character (22) into an agency generation model, so as to learn the voice of the personalized character (22) based on the identification code and text information, learn the facial movement of the personalized character (22) based on the identification code, voice information, and face information, and learn the three-dimensional rotation parameters of the joints corresponding to the personalized character (22) based on the identification code, voice information, and motion information.
[0214] Here, the processor (180) can learn the voice of the personalized character (22) by using an identification code and a feature vector to identify the personalized character (22) and convert text information into audio data.
[0215] And, when learning the facial movement of the personalized character (22), the processor (180) can identify the personalized character (22) using an identification code and a feature vector, and learn the facial movement of the personalized character (22) to be synchronized with the voice timing of the personalized character (22) based on the voice information including the voice features of the personalized character (22), the timing information of words in the sentence, and the facial information including the facial movement data of the personalized character (22) corresponding to the voice of the personalized character (22).
[0216] Next, when learning the three-dimensional rotation parameters of the joints corresponding to the personalized character (22), the processor (180) can identify the personalized character (22) using an identification code and a feature vector, and learn the three-dimensional rotation parameters of the joints corresponding to the personalized character (22) so as to be synchronized with the voice timing of the personalized character (22) based on voice information including sentences corresponding to the voice of the personalized character (22), timing information of words within the sentences, voice features of the personalized character (22), and motion information including three-dimensional rotation parameters of the joints corresponding to the motion of the personalized character (22) and motion control reference values for each body part of the personalized character (22).
[0217] Here, the agency generation model may include a Text To Speech (TTS) model that learns the voice of a personalized character based on an identification code and text information, a face motion model that learns the facial movements of a personalized character based on an identification code, voice information, and face information, and a body motion model that learns three-dimensional rotation parameters of joints corresponding to a personalized character based on an identification code, voice information, and motion information.
[0218] As another embodiment, when learning the motion characteristics of a personalized character (22), the processor (180) inputs a dataset including an identification code, voice information, face information, motion information, text information, and special motion feature information corresponding to the personalized character (22) into an agency generation model, so as to learn the voice of the personalized character (22) based on the identification code and text information, learn the facial motion of the personalized character (22) based on the identification code, voice information, and face information, learn the three-dimensional rotation parameters of the joints corresponding to the personalized character (22) based on the identification code, voice information, and motion information, and learn the unique specific motion of the personalized character (22) based on the identification code and special motion feature information.
[0219] Here, the processor (180) can learn the voice of the personalized character (22) by using an identification code and a feature vector to identify the personalized character and converting text information into audio data.
[0220] And, when learning the facial movement of the personalized character (22), the processor (180) can identify the personalized character (22) using an identification code and a feature vector, and learn the facial movement of the personalized character (22) to be synchronized with the voice timing of the personalized character (22) based on the voice information including the voice features of the personalized character (22), the timing information of words in the sentence, and the facial information including the facial movement data of the personalized character (22) corresponding to the voice of the personalized character (22).
[0221] Next, when learning the three-dimensional rotation parameters of the joints corresponding to the personalized character (22), the processor (180) can identify the personalized character (22) using an identification code and a feature vector, and learn the three-dimensional rotation parameters of the joints corresponding to the personalized character (22) so as to be synchronized with the voice timing of the personalized character (22) based on voice information including sentences corresponding to the voice of the personalized character (22), timing information of words within the sentences, voice features of the personalized character (22), and motion information including three-dimensional rotation parameters of the joints corresponding to the motion of the personalized character (22) and motion control reference values for each body part of the personalized character (22).
[0222] In addition, when learning a unique specific movement of a personalized character (22), the processor (180) can identify the personalized character (22) using an identification code and a feature vector, and learn the unique specific movement of the personalized character (22) by adjusting the weights to lower the weight for a specific movement with a negative meaning and to raise the weight for a specific movement with a positive meaning based on the special movement feature information of the personalized character (22).
[0223] Here, the agency generation model may include a Text To Speech (TTS) model that learns the voice of the personalized character based on an identification code and text information, a face motion model that learns the facial movements of the personalized character based on the identification code, voice information, and face information, a body motion model that learns three-dimensional rotation parameters of joints corresponding to the personalized character based on the identification code, voice information, and motion information, and a special motion feature model that learns unique specific movements of the personalized character based on the identification code and special motion feature information.
[0224] Next, as an example, when generating a 3D agency (40), the processor (180) checks whether the user input data includes a feature vector, audio-related basic data, motion control level, face mesh, and body length of a personalized character when user input data for generating a 3D agency is received, and inputs the user input data including the feature vector, audio-related basic data, motion control level, face mesh, and body length into a pre-learned agency generation model to output voice, facial movement, and body movement of the personalized character, and generates a 3D agency (40) that expresses the motion characteristics of the personalized character and user-applied general voice characteristics based on the voice, facial movement, and body movement.
[0225] Here, the basic audio-related data may include a general voice that the user wants to apply to the 3D agency, a sentence corresponding to the general voice, and timing information required for synchronizing the general voice and the sentence.
[0226] Additionally, if the motion control level is not included in the user input data, the processor (180) can input the motion control reference values for each body part of the personalized character pre-stored in the memory (170) into the agency generation model.
[0227] Additionally, the processor (180) can generate a 3D agency (40) that expresses the general voice characteristics that the user wants to apply instead of the voice characteristics of the personalized character, if the general voice that the user wants to apply to the 3D agency is included in the audio-related basic data.
[0228] In addition, the processor (180) may, when receiving a user input requesting a personalized character list before receiving user input data for creating a 3D agency, provide a personalized character list table including a representative name and an identification code thereof corresponding to a personalized character pre-stored in a memory, and, when receiving a user input for selecting a predetermined personalized character from the personalized character list table, obtain an identification code of the selected personalized character, and recognize the obtained identification code as user input data for creating a 3D agency.
[0229] Here, the personalized character list table may further include representative images corresponding to personalized characters pre-stored in the memory (170).
[0230] In another embodiment, when generating a 3D agency (40), the processor (180) may, when receiving user input data for generating a 3D agency, check whether the user input data includes a feature vector, text, motion control level, face mesh, and body length of a personalized character, input the user input data including the feature vector, text, motion control level, face mesh, and body length into a pre-learned agency generation model to output voice, facial movement, and body movement of the personalized character, and generate a 3D agency (40) that expresses the motion characteristics and voice characteristics of the personalized character based on the voice, facial movement, and body movement.
[0231] Here, if the motion control level is not included in the user input data, the processor (180) can input the motion control reference values for each body part of the personalized character pre-stored in the memory (170) into the agency generation model.
[0232] Additionally, the processor (180) can generate a three-dimensional agency (40) that expresses the content of the text as the voice characteristics of a personalized character when the user input data includes text.
[0233] In addition, the processor (180) may, when receiving a user input requesting a personalized character list before receiving user input data for creating a 3D agency, provide a personalized character list table including a representative name and its identification code corresponding to a personalized character pre-stored in the memory (170), and, when receiving a user input for selecting a predetermined personalized character from the personalized character list table, obtain an identification code of the selected personalized character, and recognize the obtained identification code as user input data for creating a 3D agency.
[0234] Here, the personalized character list table may further include representative images corresponding to personalized characters pre-stored in the memory (170).
[0235] In this way, the present disclosure generates a 3D agency that expresses motion characteristics corresponding to the appearance of the personalized character by inputting the user setting information including the personalized character into the pre-learned agency generation model when receiving the user setting information from the user client, thereby generating a personalized 3D agency having the appearance and motion characteristics desired by the user through learning of a 2D image.
[0236] In addition, the present disclosure can provide fun and interest to customers by extracting unique and special movement characteristic information for each personalized character and learning the unique specific movements of the personalized character, thereby generating a three-dimensional agency that expresses the unique specific movements of the personalized character.
[0237] In addition, the present disclosure can provide various services to customers by generating a three-dimensional agency that expresses not only the unique voice of a personalized character but also the user's applied voice by learning the actual voice of a personalized character or the voice that the user wishes to apply.
[0238] FIG. 5 and FIG. 6 are diagrams for explaining a two-dimensional image acquisition process including a personalized character of an artificial intelligence device according to one embodiment of the present disclosure.
[0239] The present disclosure can obtain a two-dimensional image (20) including a personalized character (22) from a first user client (32) and a second user client (34), and generate a three-dimensional agency expressing the motion characteristics of the personalized character (22) through learning of the two-dimensional image (20).
[0240] As illustrated in FIG. 5, when the artificial intelligence device (100) of the present disclosure acquires a two-dimensional image (20), when a user input for selecting a personalized character (22) is received, the device acquires basic information of a subject corresponding to the selected personalized character (22), and acquires a two-dimensional image (20) including the selected personalized character (22) from at least one of an internal server and an external server (30) based on the basic information of the subject.
[0241] Here, the personalized character may be a new character created by the user through the user client, an improved character with the body length and some appearance adjusted from the basic character of a specific celebrity, or a mixed character that combines characters from multiple celebrities.
[0242] Additionally, personalized characters can include a variety of people, such as regular celebrities, famous influencers, specially created virtual characters, and general audiences.
[0243] In addition, a two-dimensional image (20) including a personalized character (22) can be obtained from various data images distributed by online media and public institutions.
[0244] For example, a two-dimensional image (20) obtained from an internal server and an external server (30) may be a video including a target image corresponding to a personalized character (22) and his / her voice, but this is only an example and is not limited thereto.
[0245] As illustrated in FIG. 6, the artificial intelligence device (100) of the present disclosure, which may be a server or a user client, may control the camera unit (191) to capture a subject (10) corresponding to the selected personalized character when a user input for selecting a personalized character (22) is received when acquiring a two-dimensional image (20), and may acquire a two-dimensional image (20) including the personalized character (22) captured from the camera unit (191).
[0246] Here, the two-dimensional image (20) captured from the camera unit (191) may be a video including an image of a subject corresponding to a personalized character (22) and his / her voice, but this is only one example and is not limited thereto.
[0247] At this time, the camera unit (191) may include an RGB camera (191a) that captures an RGB image of a subject corresponding to a selected personalized character, and a depth camera (191b) that acquires a 3D point cloud of a subject corresponding to a selected personalized character.
[0248] Additionally, the artificial intelligence device (100) may further include a microphone (192) that acquires audio data of a subject corresponding to a selected personalized character.
[0249] For example, the target (10) corresponding to the personalized character may include various people such as general celebrities, famous influencers, specially created virtual characters, general targets, etc.
[0250] FIG. 7 is a diagram for explaining a processor of an artificial intelligence device according to one embodiment of the present disclosure.
[0251] As illustrated in FIG. 7, the processor of the present disclosure may include a preprocessing unit (510) for preprocessing a two-dimensional image, a learning dataset generation unit (520) for generating a learning dataset, and an agency generation model (530) for generating a three-dimensional agency.
[0252] Here, the preprocessing unit (510) may include an identification code assignment unit (512) that assigns an identification code to a personalized character included in a two-dimensional image, a facial information extraction unit (514) that extracts facial information corresponding to the personalized character to which the identification code has been assigned, a motion information extraction unit (516) that extracts motion information corresponding to the personalized character to which the identification code has been assigned, and a voice information extraction unit (518) that extracts voice information corresponding to the personalized character to which the identification code has been assigned.
[0253] The preprocessing unit (510) can check whether a pre-selected personalized character is included in the 2D image, and if the 2D image includes a personalized character, assign an identification code to the personalized character included in the 2D image, and extract voice information, facial information, and motion information corresponding to the personalized character to which the identification code is assigned.
[0254] Next, the learning dataset generation unit (520) can generate a learning dataset including voice information, facial information, and motion information corresponding to a personalized character.
[0255] Here, the learning dataset generation unit (520) can store the learning dataset in memory according to the identification code of the personalized character.
[0256] Next, when a dataset including an identification code, voice information, facial information, and motion information corresponding to a personalized character is input, the agency generation model (530) learns facial movements of the personalized character based on the identification code, voice information, and facial information, and learns 3D rotation parameters of joints corresponding to the personalized character based on the identification code, voice information, and motion information.
[0257] In some cases, when a dataset including an identification code, voice information, face information, motion information, and text information corresponding to a personalized character is input, the agency generation model (530) can learn the voice of the personalized character based on the identification code and text information, learn the facial movement of the personalized character based on the identification code, voice information, and face information, and learn the three-dimensional rotation parameters of the joints corresponding to the personalized character based on the identification code, voice information, and motion information.
[0258] In another case, when a dataset including an identification code, voice information, face information, motion information, text information, and special motion feature information corresponding to a personalized character is input, the agency generation model (530) learns the voice of the personalized character based on the identification code and text information, learns the facial motion of the personalized character based on the identification code, voice information, and face information, learns the three-dimensional rotation parameters of the joints corresponding to the personalized character based on the identification code, voice information, and motion information, and learns the unique specific motion of the personalized character based on the identification code and special motion feature information.
[0259] Next, the pre-learned agency generation model (530) outputs the voice, facial movements, and body movements of a personalized character when user input data including an identification code, audio-related basic data, and motion control levels are input, and can generate a 3D agency (40) that expresses the motion characteristics of the personalized character and user-applied general voice characteristics based on the voice, facial movements, and body movements.
[0260] Here, the basic audio-related data may include a general voice that the user wants to apply to the 3D agency, a sentence corresponding to the general voice, and timing information required for synchronizing the general voice and the sentence.
[0261] The pre-learned agency generation model (530) can input the body part-specific motion control reference values of a personalized character pre-stored in memory into the agency generation model when the motion control level is not included in the user input data.
[0262] That is, the pre-learned agency generation model (530) can generate a 3D agency that expresses the general voice characteristics that the user wants to apply to the 3D agency instead of the voice characteristics of the personalized character, if the general voice that the user wants to apply to the 3D agency is included in the audio-related basic data.
[0263] In some cases, the pre-trained agency generation model (530) can output voice, facial movements and body movements of a personalized character when user input data including an identification code, text and motion control level is input, and can generate a 3D agency that expresses motion characteristics and voice characteristics of the personalized character based on the voice, facial movements and body movements.
[0264] Here, the pre-learned agency generation model (530) can input the body part-specific motion control reference values of the personalized character pre-stored in memory into the agency generation model when the motion control level is not included in the user input data.
[0265] That is, the pre-learned agency generation model (530) can generate a three-dimensional agency that expresses the content of the text as the voice characteristics of a personalized character when text is included in user input data.
[0266] FIG. 8 is a diagram for explaining a facial information extraction process of an artificial intelligence device according to an embodiment of the present disclosure.
[0267] As illustrated in FIG. 8, the present disclosure extracts a face landmark from the face of a personalized character included in a two-dimensional image, and inputs the two-dimensional image (610) and the face landmark (520) of the personalized character into a first neural network model (630) that has been pre-learned, thereby extracting face information (640) including facial movement data of the personalized character.
[0268] As an example, the first neural network model (630) may include a 3DMM (3D Morphable Model) algorithm, but this is only an example and is not limited thereto.
[0269] Here, the first neural network model (630) defines facial movements such as eye blinking, eyebrow raising, mouth opening, etc., and adjusts the value of the degree of movement for each defined part to a range of 0 to 1.0, etc., so that the face can be expressed in 3D.
[0270] In this way, the present disclosure obtains a 2D image and a 3D point cloud, extracts a facial landmark from a face in the 2D image using an algorithm that finds landmarks, synchronizes the 2D image and 3D mesh data to extract a 3D landmark from the 3D mesh, and sets only a filtering portion of the 3D point cloud so that an incorrect portion of the 3D landmark can be corrected through correction logic.
[0271] In addition, the present disclosure can perform a fitting process of extracting the most approximate value by comparing a 3D mesh generated by modifying a face combination model (3DMM) value used in a graphic engine, a 3D landmark value, and an actual 3D point cloud and 3D landmark, and after learning is performed, facial information can be extracted using a face combination model (3DMM) through input of a 2D image and a 2D landmark.
[0272] Here, the facial combination model (3DMM) expresses the movement of the face with values ranging from 0 to 1.0. If the value for the eyelids is set to 0.0, the eyes are expressed as closed, and if the value for the eyelids is set to 1.0, the eyes are expressed as wide open.
[0273] FIG. 9 is a diagram for explaining a process of extracting operation information of an artificial intelligence device according to an embodiment of the present disclosure.
[0274] As illustrated in FIG. 9, the present disclosure extracts three-dimensional keypoints (3D keypoints) for joint positions from the body of a personalized character included in a two-dimensional image, and inputs the two-dimensional image (710) and the three-dimensional keypoints (720) of the personalized character into a pre-learned second neural network model (730) to extract motion information (740) including three-dimensional rotation parameters of joints corresponding to the motion of the personalized character.
[0275] As an example, the second neural network model (730) may include a 3D Rotation Model algorithm, but this is only an example and is not limited thereto.
[0276] In addition, the present disclosure can extract a motion control reference value for each body part of a personalized character for each frame of a two-dimensional image by analyzing the degree of movement of each joint based on the three-dimensional rotation parameters when the three-dimensional rotation parameters of the joints are extracted.
[0277] For example, the motion control reference values may include motion control reference values for the hand position and its movement speed, the head position and its movement speed, the foot position and its movement speed, the neck position and its movement speed, the arm position and its movement speed, the leg position and its movement speed, and the waist position and its movement speed among the body parts of the personalized character, but this is only an example and is not limited thereto.
[0278] Here, the motion control reference value can have different values for each personalized character.
[0279] For example, the motion control threshold may vary depending on the physical condition of the personalized character.
[0280] In this way, the present disclosure extracts a three-dimensional rotation parameter of a human joint, synchronizes a two-dimensional image and the three-dimensional rotation parameter of the joint through a predetermined motion, extracts a key point using an algorithm for finding a key point in a two-dimensional image, and maps the three-dimensional rotation parameter of the corresponding part and the joint.
[0281] In addition, the present disclosure can learn to extract 3D rotation parameters of joints corresponding to each 3D key point through input of a 2D image and a 3D key point.
[0282] Here, a keypoint is a location with a specific meaning in a two-dimensional image, such as the center point of a joint.
[0283] Next, the 3D rotation parameters can be parameters that express the rotation of the joint, such as rotation matrices, Euler angle X / Y / Z values, etc.
[0284] For example, in a three-dimensional space, rotation can be defined by three Euler values, and in addition to Euler angles, it can be defined by various three-dimensional rotation expression methods such as rotation matrices and six-dimensional rotation parameters.
[0285] FIG. 10 is a diagram for explaining a voice information extraction process of an artificial intelligence device according to an embodiment of the present disclosure.
[0286] As illustrated in FIG. 10, the present disclosure extracts audio data (810) of a personalized character included in a two-dimensional image, inputs the audio data into a pre-trained third neural network model (820), and extracts voice information (830) including sentences corresponding to the voice of the personalized character, timing information of words within the sentences, and voice features of the personalized character.
[0287] For example, the third neural network model (830) may include a STT (Speech To Text) model (822), a Forced Alignment Model (824), and a voice model (826), but this is only an example and is not limited thereto.
[0288] Here, the present disclosure can extract sentences corresponding to the voice of a personalized character by converting audio data (810) into text through an STT model (822), extract timing information of words in sentences extracted from the voice of a personalized character through a post-alignment model (824), and extract voice features of a personalized character from audio data through a voice model (826).
[0289] FIG. 11 is a diagram illustrating a processor of an artificial intelligence device according to another embodiment of the present disclosure.
[0290] As illustrated in FIG. 11, the processor of the present disclosure may include a preprocessing unit (510) for preprocessing a two-dimensional image, a learning dataset generation unit (520) for generating a learning dataset, and an agency generation model (530) for generating a three-dimensional agency.
[0291] Here, the preprocessing unit (510) may include an identification code assignment unit (512) that assigns an identification code to a personalized character included in a two-dimensional image, a face information extraction unit (514) that extracts facial information corresponding to the personalized character to which the identification code has been assigned, a motion information extraction unit (516) that extracts motion information corresponding to the personalized character to which the identification code has been assigned, a voice information extraction unit (518) that extracts voice information corresponding to the personalized character to which the identification code has been assigned, and a special motion feature information extraction unit (519) that extracts special motion feature information from the personalized character included in the two-dimensional image.
[0292] The preprocessing unit (510) can check whether a pre-selected personalized character is included in the 2D image, and if the 2D image includes a personalized character, assign an identification code to the personalized character included in the 2D image, and extract voice information, facial information, motion information, and special movement feature information corresponding to the personalized character to which the identification code is assigned.
[0293] Next, the learning dataset generation unit (520) can generate a learning dataset including voice information, facial information, motion information, and special motion feature information corresponding to a personalized character.
[0294] Here, the learning dataset generation unit (520) can store the learning dataset in memory according to the identification code of the personalized character.
[0295] Next, when a dataset including an identification code, voice information, face information, motion information, text information, and special motion feature information corresponding to a personalized character is input, the agency generation model (530) learns the voice of the personalized character based on the identification code and text information, learns the facial movement of the personalized character based on the identification code, voice information, and face information, learns the three-dimensional rotation parameters of the joints corresponding to the personalized character based on the identification code, voice information, and motion information, and learns the unique specific motion of the personalized character based on the identification code and special motion feature information.
[0296] Next, the pre-learned agency generation model (530) outputs voice, facial movements, and body movements of a personalized character when user input data including an identification code, text, and motion control level are input, and can generate a three-dimensional agency that expresses motion characteristics and voice characteristics of the personalized character based on the voice, facial movements, and body movements.
[0297] Here, the pre-learned agency generation model (530) can input the body part-specific motion control reference values of the personalized character pre-stored in memory into the agency generation model when the motion control level is not included in the user input data.
[0298] That is, the pre-learned agency generation model (530) can generate a three-dimensional agency that expresses the content of the text as the voice characteristics of a personalized character when text is included in user input data.
[0299] FIG. 12 and FIG. 13 are diagrams for explaining a special movement feature information extraction process of an artificial intelligence device according to an embodiment of the present disclosure.
[0300] As illustrated in FIGS. 12 and 13, the present disclosure extracts a learning dataset including special motion feature information from a personalized character included in a two-dimensional image, and stores the learning dataset including the special motion feature information in memory according to the identification code of the personalized character.
[0301] The present disclosure can select a specific section from a two-dimensional image including a personalized character (910).
[0302] For example, the present disclosure can select a portion of the entire section of a two-dimensional image including a personalized character, in which the personalized character performs a special motion, as the specific section by using a keyframe extraction algorithm when selecting a specific section.
[0303] Here, the present disclosure can recognize at least one of a unique facial expression, a unique gesture, and a unique body movement shape of the personalized character as a special movement of the personalized character.
[0304] Next, the present disclosure can vectorize facial information and motion information extracted from a personalized character of a two-dimensional image corresponding to a selected specific section (920).
[0305] Next, the present disclosure analyzes the distribution of vectorized data to determine whether there is data whose occurrence frequency is greater than or equal to a reference frequency, and if there is data whose occurrence frequency is greater than or equal to the reference frequency, the corresponding data can be determined as a special motion feature (930).
[0306] As shown in FIG. 13, as an example, in the case of a first personalized character A corresponding to a dataset having an identification code of 001, if the frequency of occurrence of a hand gesture of raising a thumb is greater than or equal to a reference frequency, the hand gesture of raising a thumb can be determined as a special movement feature of the first personalized character.
[0307] As another example, the present disclosure can determine that a fist-clenching hand motion is a special movement feature of a second personalized character B corresponding to a dataset having an identification code of 002, if the frequency of occurrence of the fist-clenching hand motion is greater than a reference frequency.
[0308] Next, the present disclosure can extract special motion feature information based on facial information and motion information of a specific section corresponding to data determined to be a special motion feature (940).
[0309] In some cases, the present disclosure extracts a face landmark from the face of a personalized character included in a two-dimensional image corresponding to the selected specific section when a specific section is selected from a two-dimensional image, inputs the two-dimensional image corresponding to the specific section and the face landmark of the personalized character into a first neural network model that has been pre-learned to extract facial information including facial movement data of the personalized character, and extracts three-dimensional keypoints for joint positions from the body of the personalized character included in the two-dimensional image corresponding to the selected specific section, and inputs the two-dimensional image corresponding to the specific section and the 3D keypoint of the personalized character into a second neural network model that has been pre-learned to extract motion information including a three-dimensional rotation parameter of a joint corresponding to the motion of the personalized character.
[0310] In another case, the present disclosure can determine whether facial information and motion information extracted from a personalized character of a two-dimensional image corresponding to the selected specific section exist in a memory when a specific section is selected from a two-dimensional image, and if facial information and motion information of the personalized character corresponding to the specific section exist, the facial information and motion information of the personalized character corresponding to the specific section can be vectorized.
[0311] FIG. 14 is a diagram for explaining an agency generation model of an artificial intelligence device according to one embodiment of the present disclosure.
[0312] As illustrated in FIG. 14, the present disclosure inputs a dataset including an identification code, voice information, facial information, and motion information corresponding to a personalized character into an agency generation model (530) to learn facial movements of a personalized character and 3D rotation parameters of joints corresponding to the personalized character.
[0313] Here, the agency generation model (530) may include a face motion model (532) that learns facial movements of a personalized character based on an identification code, voice information, and face information, and a body motion model (534) that learns three-dimensional rotation parameters of joints corresponding to a personalized character based on an identification code, voice information, and motion information.
[0314] The facial movement model (532) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn facial movements of the personalized character to be synchronized with the voice timing of the personalized character based on voice information including sentences corresponding to the voice of the personalized character, timing information of words in the sentences, voice features of the personalized character, and facial information including facial movement data of the personalized character.
[0315] In addition, the body movement model (534) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn a 3D rotation parameter of a joint corresponding to a personalized character so as to be synchronized with the voice timing of the personalized character based on voice information including a sentence corresponding to the voice of the personalized character, timing information of words within the sentence, voice features of the personalized character, and motion information including a 3D rotation parameter of a joint corresponding to the motion of the personalized character and a motion control reference value for each body part of the personalized character.
[0316] FIG. 15 is a diagram for explaining an agency generation model of an artificial intelligence device according to another embodiment of the present disclosure.
[0317] As illustrated in FIG. 15, the present disclosure inputs a dataset including an identification code, voice information, face information, motion information, and text information corresponding to a personalized character into an agency generation model (530), thereby learning the voice of the personalized character based on the identification code and text information, learning the facial movement of the personalized character based on the identification code, voice information, and face information, and learning the three-dimensional rotation parameters of the joints corresponding to the personalized character based on the identification code, voice information, and motion information.
[0318] Here, the agency generation model (530) may include a Text To Speech (TTS) model (536) that learns the voice of a personalized character based on an identification code and text information, a face motion model (532) that learns the facial movement of a personalized character based on an identification code, voice information, and face information, and a body motion model (534) that learns the three-dimensional rotation parameters of joints corresponding to the personalized character based on the identification code, voice information, and motion information.
[0319] The TTS model (536) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn the voice of the personalized character by converting text information into audio data.
[0320] In addition, the facial movement model (532) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn facial movements of the personalized character to be synchronized with the voice timing of the personalized character based on voice information including a sentence corresponding to the voice of the personalized character, timing information of words in the sentence, voice features of the personalized character, and facial information including facial movement data of the personalized character.
[0321] In addition, the body movement model (534) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn a 3D rotation parameter of a joint corresponding to a personalized character so as to be synchronized with the voice timing of the personalized character based on voice information including a sentence corresponding to the voice of the personalized character, timing information of words within the sentence, voice features of the personalized character, and motion information including a 3D rotation parameter of a joint corresponding to the motion of the personalized character and a motion control reference value for each body part of the personalized character.
[0322] FIG. 16 is a diagram for explaining an agency generation model of an artificial intelligence device according to another embodiment of the present disclosure.
[0323] As illustrated in FIG. 16, the present disclosure inputs a dataset including an identification code, voice information, facial information, motion information, text information, and special motion feature information corresponding to a personalized character into an agency generation model (530), thereby learning the voice of the personalized character based on the identification code and text information, learning the facial motion of the personalized character based on the identification code, voice information, and facial information, learning the three-dimensional rotation parameters of the joints corresponding to the personalized character based on the identification code, voice information, and motion information, and learning the unique specific motion of the personalized character based on the identification code and special motion feature information.
[0324] Here, the agency generation model (530) may include a Text To Speech (TTS) model (536) that learns the voice of the personalized character based on an identification code and text information, a face motion model (532) that learns the facial movement of the personalized character based on an identification code, voice information, and face information, a body motion model (534) that learns the three-dimensional rotation parameters of joints corresponding to the personalized character based on the identification code, voice information, and motion information, and a special motion feature model (538) that learns the unique specific motion of the personalized character based on the identification code and special motion feature information.
[0325] The TTS model (536) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn the voice of the personalized character by converting text information into audio data.
[0326] In addition, the facial movement model (532) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn facial movements of the personalized character to be synchronized with the voice timing of the personalized character based on voice information including a sentence corresponding to the voice of the personalized character, timing information of words in the sentence, voice features of the personalized character, and facial information including facial movement data of the personalized character.
[0327] In addition, the body movement model (534) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn a 3D rotation parameter of a joint corresponding to a personalized character so as to be synchronized with the voice timing of the personalized character based on voice information including a sentence corresponding to the voice of the personalized character, timing information of words within the sentence, voice features of the personalized character, and motion information including a 3D rotation parameter of a joint corresponding to the motion of the personalized character and a motion control reference value for each body part of the personalized character.
[0328] In addition, the special movement feature model (538) of the agency generation model (530) can identify a personalized character using an identification code and a feature vector, and learn unique specific movements of the personalized character by adjusting weights to lower the weights for specific movements with negative meanings and to raise the weights for specific movements with positive meanings based on the special movement feature information of the personalized character.
[0329] FIG. 17 is a diagram for explaining the learning process of an agency generation model of an artificial intelligence device according to one embodiment of the present disclosure.
[0330] As illustrated in FIG. 17, the present disclosure can extract a feature vector for a personalized character (22) included in a two-dimensional image (20) through a feature vector extraction unit (512) when a two-dimensional image (20) including a personalized character (22) is input.
[0331] In addition, the present disclosure can extract a learning dataset including voice information, facial information, and motion information corresponding to a personalized character (22), and store the learning dataset in memory according to the identification code of the personalized character (22).
[0332] Here, the present disclosure extracts a face landmark from the face of a personalized character (22) included in a two-dimensional image (20) through a face information extraction unit (514), and inputs the two-dimensional image (20) and the face landmark of the personalized character (22) into a first neural network model (630) that has been pre-learned, thereby extracting face information including face movement data of the personalized character (22).
[0333] As an example, the first neural network model (630) may include a 3DMM (3D Morphable Model) algorithm, but this is only an example and is not limited thereto.
[0334] In addition, the present disclosure extracts 3D keypoints for joint positions from the body of a personalized character (22) included in a 2D image (20) through a motion information extraction unit (516), and inputs the 3D keypoints of the 2D image (20) and the personalized character (22) into a pre-learned second neural network model (730) to extract motion information including 3D rotation parameters of joints corresponding to the motion of the personalized character (22).
[0335] As an example, the second neural network model (730) may include a 3D Rotation Model algorithm, but this is only an example and is not limited thereto.
[0336] In addition, the present disclosure can analyze the degree of movement of each joint based on a three-dimensional rotation parameter through a motion control reference value setting unit (732) to set a motion control reference value for each body part of a personalized character (22) for each frame of a two-dimensional image (20).
[0337] For example, the motion control reference values may include motion control reference values for the hand position and its movement speed, the head position and its movement speed, the foot position and its movement speed, the neck position and its movement speed, the arm position and its movement speed, the leg position and its movement speed, and the waist position and its movement speed among the body parts of the personalized character (22), but this is only one example and is not limited thereto.
[0338] In addition, the present disclosure can set the appearance and body length of a personalized character (22) through a face mesh and body length setting unit (734).
[0339] In addition, the present disclosure extracts audio data of a personalized character (22) included in a two-dimensional image (20) through a voice information extraction unit (518), inputs the audio data into a pre-learned third neural network model (820), and extracts voice information including a sentence corresponding to the voice of the personalized character (22), timing information of words within the sentence, and voice characteristics of the personalized character (22).
[0340] For example, the third neural network model can extract sentences corresponding to the voice of a personalized character (22) by converting audio data into text through a STT (Speech To Text) model (822), extract timing information of words in sentences extracted from the voice of a personalized character (22) through a forced alignment model, and extract voice features of a personalized character (22) from audio data through a voice model (826).
[0341] Next, the present disclosure inputs a dataset including a feature vector, voice information, facial information, and motion information corresponding to a personalized character (22) into an agency generation model (530), thereby learning facial movements of the personalized character (22) based on the feature vector, voice information, and facial information, and learning three-dimensional rotation parameters of joints corresponding to the personalized character (22) based on the feature vector, voice information, and motion information.
[0342] Here, the face motion model (532) of the agency generation model (530) can learn the facial movement of the personalized character (22) to be synchronized with the voice timing of the personalized character (22) based on voice information including a sentence corresponding to the voice of the personalized character (22), timing information of words within the sentence, a feature vector and voice features of the personalized character (22), and face information including facial movement data of the personalized character (22).
[0343] In addition, the body motion model (534) of the agency generation model (530) can learn the 3D rotation parameters of the joints corresponding to the personalized character (22) so as to be synchronized with the voice timing of the personalized character (22) based on the voice information including the sentence corresponding to the voice of the personalized character (22), the timing information of the words in the sentence, the voice characteristics of the personalized character (22), the 3D rotation parameters of the joints corresponding to the motion of the personalized character (22), and the motion information including the motion control reference value and the body length for each body part of the personalized character (22).
[0344] FIG. 18 and FIG. 19 are drawings for explaining a three-dimensional agency generation process of an artificial intelligence device according to one embodiment of the present disclosure.
[0345] As illustrated in FIGS. 18 and 19, the present disclosure can verify whether, when user input data for creating a 3D agency is received, the user input data includes a feature vector (62) of a personalized character, audio-related basic data (64), and motion control level (65) (S10).
[0346] Here, the audio-related basic data (64) may include a general voice (64a) that the user wants to apply to the 3D agency, a sentence corresponding to the general voice (64b), and timing information (64c) required for synchronization of the general voice and the sentence.
[0347] In addition, the present disclosure can input the body part-specific motion control reference values of a personalized character stored in memory into the agency generation model (530) when the motion control level (65) is not included in the user input data.
[0348] Next, the present disclosure can input user input data including a feature vector (62), audio-related basic data (64), and motion control level (65) into a pre-trained agency generation model (530) (S12).
[0349] Next, the present disclosure can output voice, facial movements, and body movements of a personalized character through a pre-learned agency generation model (530) (S14).
[0350] For example, the present disclosure can output facial movement parameters and three-dimensional rotation parameters of joints corresponding to audio data of a personalized character (22) through a dataset of a personalized character corresponding to a celebrity identification code 00001.
[0351] In addition, the present disclosure can generate a three-dimensional agency that expresses the motion characteristics of a personalized character and user-applied general voice characteristics based on voice, facial movement, and body movement through a graphic engine render (S16).
[0352] Here, the present disclosure can generate a 3D agency that expresses the general voice characteristics that the user wants to apply to the 3D agency instead of the voice characteristics of the personalized character, if the general voice that the user wants to apply to the 3D agency is included in the audio-related basic data.
[0353] For example, if the present disclosure creates a 3D agency for a personalized character that frequently performs a thumbs-up gesture with an active voice, the 3D agency can provide responses to user inquiries with an active voice and unique motion characteristics such as a thumbs-up gesture.
[0354] Here, the voice of the 3D agency can have a user-specified voice, but can also have a personalized character voice intonation through learning.
[0355] FIG. 20 and FIG. 21 are drawings for explaining a three-dimensional agency generation process of an artificial intelligence device according to another embodiment of the present disclosure.
[0356] As illustrated in FIGS. 20 and 21, the present disclosure can verify whether, when user input data for creating a 3D agency is received, the user input data includes a feature vector (62), text (66), and motion control level (65) of a personalized character (S20).
[0357] Here, in the present disclosure, if the motion control level (65) is not included in the user input data, the motion control reference values for each body part of the personalized character pre-stored in the memory can be input into the agency generation model (530).
[0358] Next, the present disclosure can input user input data including a feature vector (62), text (66), and motion control level (65) into a pre-trained agency generation model (530) (S22).
[0359] Next, the present disclosure can output voice, facial movements, and body movements of a personalized character through a pre-learned agency generation model (530) (S24).
[0360] For example, the present disclosure can output facial movement parameters and three-dimensional rotation parameters of joints corresponding to audio data of a personalized character (22) through a dataset of a personalized character corresponding to a celebrity identification code 00001.
[0361] In addition, the present disclosure can generate a three-dimensional agency that expresses the motion characteristics and voice characteristics of a personalized character based on voice, facial movement, and body movement through a graphic engine render (S26).
[0362] Here, the present disclosure can generate a three-dimensional agency that expresses the content of the text as the voice characteristics of a personalized character when the user input data includes text.
[0363] For example, if the present disclosure creates a 3D agency for a personalized character that frequently performs a thumbs-up gesture with an active voice, the 3D agency can provide responses to user inquiries with an active voice and unique motion characteristics such as a thumbs-up gesture.
[0364] Here, the voice of the 3D agency can have the same voice, intonation and mood as the personalized character through learning.
[0365] FIG. 22 is a diagram for explaining the overall operation flow of an artificial intelligence device according to one embodiment of the present disclosure.
[0366] As illustrated in FIG. 22, the present disclosure can create a personalized 3D agency through a network connection between an artificial intelligence device (100) of a server and a first user client (32).
[0367] First, the first user client (32) can check whether the appearance of the personalized character desired by the user is set when the agency creation app is executed (S110).
[0368] Here, the first user client (32) can provide a character appearance setting window and activate a camera, etc., if the appearance of a personalized character is not set.
[0369] Next, the first user client (32) can receive character appearance information through the character appearance setting window (S120).
[0370] For example, the character's appearance information may include personal appearance information of a specific celebrity, mixed appearance information of multiple celebrities, modified appearance information of a specific celebrity, personal appearance information of the user, appearance information of a basic character, appearance information of a new character, etc.
[0371] Next, when the character's appearance setting is completed, the first user client (32) can transmit an agency creation request including the character's appearance information, which has been set, to the server (S130).
[0372] Next, the first user client (32) can check the system specifications and system status to collect system information and transmit the collected system information to the server (S140).
[0373] And, the first user client (32) can decide whether to perform random streaming based on the collected system information (S150).
[0374] Next, the first user client (32) can drive web streaming if it decides to use render streaming (S210), and can drive a 3D agency set as a 3D graphics engine if it does not decide to use render streaming (S160).
[0375] Here, the first user client (32) can drive web streaming by deciding on render streaming if the system's specifications are below the set value, and can drive a 3D agency set as a 3D graphics engine if the system's specifications are above the set value.
[0376] Next, the first user client (32) can drive web streaming and transmit character creation mode information to the server (S210).
[0377] Here, the first user client (32) can transmit character creation mode information to the server, including a large mode that displays characters on the entire display screen or a small mode that displays characters on a portion of the display screen.
[0378] Next, the server's artificial intelligence device (100) can set the appearance of a personalized character or a personal target persona based on the character's appearance information received from the first user client (32) (S180).
[0379] In addition, the artificial intelligence device (100) of the server can set a method of transmitting operation data of a personalized agency based on system information received from the first user client (32) (S190).
[0380] In addition, the artificial intelligence device (100) of the server can set a method of transmitting motion data of a personalized agency based on character creation mode information received from the first user client (32) (S200).
[0381] In addition, the artificial intelligence device (100) of the server can obtain text information for generating motion data of a personalized character (S210) and collect data required for generating motion data (S220).
[0382] Next, the artificial intelligence device (100) of the server can input user setting information and text information including appearance information, system information, and character creation mode information of a personalized character into the pre-learned agency creation model to generate motion data corresponding to the appearance of a personalized character and image data based on the motion data (S230).
[0383] Next, the server's artificial intelligence device (100) can determine whether to perform render streaming based on system information and character creation mode information (S240).
[0384] Here, the artificial intelligence device (100) of the server can determine either a first transmission mode that transmits only 3D agency image data or a second transmission mode that transmits only 3D agency motion data based on the system information of the first user client (32).
[0385] For example, the artificial intelligence device (100) of the server may analyze system information and determine a first transmission mode that transmits only 3D agency image data if the specifications of the first user client (32) are less than a set value, and may determine a second transmission mode that transmits only 3D agency motion data if the specifications of the first user client (32) are greater than or equal to the set value.
[0386] At this time, the first user client (32) can generate a 3D agency that expresses motion characteristics corresponding to the motion data based on the motion data by receiving only the 3D agency motion data.
[0387] In some cases, the artificial intelligence device (100) of the server may determine either a first transmission mode that transmits only 3D agency image data or a second transmission mode that transmits only 3D agency motion data based on the character creation mode information of the first user client (32).
[0388] Here, the artificial intelligence device (100) of the server may analyze the character creation mode information and determine, if the character creation mode is a large mode that displays the character on the entire display screen, to be a first transmission mode that transmits only 3D agency image data, and if the character creation mode is a small mode that displays the character on a part of the display screen, to be a second transmission mode that transmits only 3D agency motion data.
[0389] In addition, the artificial intelligence device (100) of the server, when the character creation mode is a large mode that displays the character on the entire display screen, determines only the 3D agency image data as the first transmission mode, checks the specifications of the first user client (32), and if the specifications of the first user client (32) are equal to or greater than the set value, changes the determined first transmission mode to a second transmission mode that transmits only the 3D agency motion data.
[0390] At this time, the first user client (32) can generate a 3D agency that expresses motion characteristics corresponding to the motion data based on the motion data by receiving only the 3D agency motion data.
[0391] Next, the first user client (32) can output the streamed 3D agency if the server decides to render streaming (S260), and can output the 3D agency as a 3D graphic engine animation if the server does not decide to render streaming (S250).
[0392] In addition, the artificial intelligence device (100) of the server analyzes the appearance of the personalized character and, when it is predicted that a 3D agency with a high similarity to a specific celebrity will be created, collects content information related to the specific celebrity, recommends content related to the specific celebrity based on the collected content information, and provides the recommended content to the first user client (32).
[0393] FIG. 23 is a diagram for explaining an operation flow for providing recommended content related to a specific celebrity by an artificial intelligence device according to an embodiment of the present disclosure, and FIGS. 24 and 25 are diagrams showing the provision of recommended content related to a specific celebrity by an artificial intelligence device according to an embodiment of the present disclosure.
[0394] As illustrated in FIGS. 23 to 25, the present disclosure can determine whether the personalized character is a celebrity (S320) when the appearance of the personalized character is selected (S310).
[0395] Here, the present disclosure can confirm whether character information is included in the setting information for agency creation received from the user client, and recognize a personalized character as a celebrity based on an identifier corresponding to a specific celebrity included in the character information.
[0396] Next, the present disclosure analyzes the appearance of the personalized character if the personalized character is not a celebrity (S330) and selects a specific celebrity with the highest similarity to the appearance of the personalized character (S340).
[0397] Here, the present disclosure can analyze the appearance of a personalized character and select a specific celebrity with the highest similarity to the appearance of the personalized character if the character information does not include an identifier corresponding to a specific celebrity.
[0398] That is, the present disclosure can vectorize the appearance of a personalized character using a feature extraction algorithm when the character information does not include an identifier corresponding to a specific celebrity, and select a specific celebrity that is most similar to the appearance of the vectorized personalized character from a data pool that collects appearance image data of multiple celebrities, including singers, actors, entertainers, etc.
[0399] Next, the present disclosure can collect content information related to a selected specific celebrity (S350).
[0400] Here, the present disclosure can search and collect content information including hairstyles, makeup, clothing, songs, movies, etc. related to a specific celebrity.
[0401] In addition, the present disclosure can expose content information related to specific celebrities collected based on the user's content usage pattern to content frequently used by the user (S360).
[0402] That is, the present disclosure can extract the user's content usage pattern by searching the user's content usage history information by collecting content information related to a specific celebrity.
[0403] Here, the present disclosure can expose content information related to specific celebrities collected based on the user's content usage pattern to content frequently used by the user.
[0404] Next, the present disclosure can check whether the user interest in content information related to a specific celebrity is greater than a preset value through content in which content information related to a specific celebrity is exposed (S370).
[0405] Next, the present disclosure can register content information related to a specific celebrity in content frequently used by the user if the user's interest in the content information related to a specific celebrity is greater than a preset value (S380).
[0406] In addition, the present disclosure can exclude content information related to a specific celebrity from content frequently used by the user if the user's interest in the content information related to a specific celebrity is below a preset value (S390).
[0407] As illustrated in FIG. 24, when the personalized character is a specific celebrity (510), the present disclosure can retrieve information about clothes (520) and shoes (530) worn by the specific celebrity (510).
[0408] Next, as illustrated in FIG. 25, content information (540) related to clothes (520) worn by a specific celebrity (510) and content information (540) related to shoes (530) worn by a specific celebrity (510) can be provided.
[0409] In this way, the present disclosure generates a 3D agency that expresses motion characteristics corresponding to the appearance of the personalized character by inputting the user setting information including the personalized character into the pre-learned agency generation model when receiving the user setting information from the user client, thereby generating a personalized 3D agency having the appearance and motion characteristics desired by the user through learning of a 2D image.
[0410] In addition, the present disclosure can provide fun and interest to customers by extracting unique and special movement characteristic information for each personalized character and learning the unique specific movements of the personalized character, thereby generating a three-dimensional agency that expresses the unique specific movements of the personalized character.
[0411] In addition, the present disclosure can provide various services to customers by generating a three-dimensional agency that expresses not only the unique voice of a personalized character but also the user's applied voice by learning the actual voice of a personalized character or the voice that the user wishes to apply.
[0412] The above-described present disclosure can be implemented as computer-readable code on a program-recorded medium. The computer-readable medium includes all types of recording devices that store data that can be read by a computer system. Examples of computer-readable media include hard disk drives (HDDs), solid-state disk drives (SSDs), silicon disk drives (SDDs), read-only memory (ROM), random access memory (RAM), CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, the computer may include a processor (180) of an artificial intelligence device.
[0413] According to the artificial intelligence device according to the present disclosure, when receiving user setting information including a personalized character from a user client, the user setting information is input into the pre-learned agency generation model to generate a 3D agency that expresses motion characteristics corresponding to the appearance of the personalized character, thereby generating a personalized 3D agency having the appearance and motion characteristics desired by the user through learning of a 2D image, and thus has significant industrial applicability.
Claims
1. Memory for storing learning datasets for each character; and, A processor for generating a personalized 3D agency using the character-specific learning dataset is included. The above processor, An artificial intelligence device characterized in that a two-dimensional image including the character is preprocessed to extract a learning dataset including motion information corresponding to the appearance of the character, the learning dataset is input into an agency generation model to learn the motion characteristics of the character, and when receiving user setting information including a personalized character from a user client, the user setting information is input into the pre-trained agency generation model to generate a three-dimensional agency expressing motion characteristics corresponding to the appearance of the personalized character.
2. In paragraph 1, The above processor, An artificial intelligence device characterized in that when receiving the user setting information, if an agency creation request is received from the user client, the device requests setting information for creating the agency to the user client, and when receiving the setting information for creating the agency from the user client, the device checks whether basic image data corresponding to a personalized character is included in the setting information.
3. In paragraph 2, The above processor, An artificial intelligence device characterized in that, if the above basic image data is not included, it requests the user client for appearance information of a personalized character for creating the agency.
4. In paragraph 2, The above processor, An artificial intelligence device characterized in that, when receiving the setting information for creating the above agency, it is confirmed that at least one of the user client's system information and character creation mode information is included in the setting information in addition to the basic image data.
5. In paragraph 4, The above processor, Create a 3D agency that expresses motion characteristics corresponding to the appearance of a personalized character based on the basic image data included in the above setting information, An artificial intelligence device characterized in that it determines a transmission method of a 3D agency generated based on at least one of system information and character creation mode information included in the above setting information.
6. In paragraph 1, The above processor, An artificial intelligence device characterized in that, when generating the three-dimensional agency, user setting information including a personalized character is obtained from the user client, text information for generating motion data of the personalized character is obtained, the obtained user setting information and text information are input into the pre-learned agency generation model to generate motion data corresponding to the appearance of the personalized character and image data based on the motion data, and only the motion data-based image data is transmitted to the user client corresponding to the user setting information, or only the motion data is transmitted to the user client to generate the three-dimensional agency.
7. In paragraph 6, The above processor, An artificial intelligence device characterized in that, when the user client's system information is included in the user setting information, it determines either a first transmission mode that transmits only 3D agency image data or a second transmission mode that transmits only 3D agency motion data based on the system information.
8. In paragraph 7, The above processor, By analyzing the above system information, if the specifications of the user client are less than the set value, the first transmission mode is determined to transmit only the 3D agency image data, An artificial intelligence device characterized in that it determines to use a second transmission mode that transmits only the 3D agency operation data when the specifications of the user client are greater than or equal to the set value.
9. In paragraph 8, The above user client, An artificial intelligence device characterized in that, when receiving only the above-mentioned 3D agency motion data, it generates a 3D agency that expresses motion characteristics corresponding to the motion data based on the above-mentioned motion data.
10. In paragraph 6, The above processor, An artificial intelligence device characterized in that, when the character creation mode information of the user client is included in the user setting information, one of a first transmission mode that transmits only the 3D agency image data or a second transmission mode that transmits only the 3D agency motion data is determined based on the character creation mode information.
11. In paragraph 10, The above processor, By analyzing the above character creation mode information, if the character creation mode is a large mode that displays the character on the entire display screen, the first transmission mode is determined to transmit only the 3D agency image data, An artificial intelligence device characterized in that, if the above character creation mode is a small mode that displays a character on a part of the display screen, it is determined as a second transmission mode that transmits only the 3D agency motion data.
12. In paragraph 11, The above processor, An artificial intelligence device characterized in that, when the character creation mode is a large mode that displays the character on the entire display screen, the first transmission mode is determined to transmit only the 3D agency image data, the specifications of the user client are checked, and if the specifications of the user client are equal to or greater than a set value, the determined first transmission mode is changed to a second transmission mode to transmit only the 3D agency motion data.
13. In paragraph 11, The above user client, An artificial intelligence device characterized in that, when receiving only the above-mentioned 3D agency motion data, it generates a 3D agency that expresses motion characteristics corresponding to the motion data based on the above-mentioned motion data.
14. In paragraph 1, The above processor, An artificial intelligence device characterized in that when generating the above 3D agency, the appearance of the personalized character is analyzed and, if the generation of a 3D agency having a high similarity to a specific celebrity is predicted, content information related to the specific celebrity is collected, and content related to the specific celebrity is recommended based on the collected content information.
15. In a method for creating a personalized agency of an artificial intelligence device, A step of obtaining a two-dimensional image including a character; A step of preprocessing a two-dimensional image including the above character; A step of extracting a learning dataset containing motion information corresponding to the appearance of the character; A step of inputting the above learning dataset into an agency generation model to learn the motion characteristics of the character; A step of receiving user-defined information including a personalized character; and A method for creating a personalized agency, characterized by comprising a step of inputting the user setting information into the pre-learned agency creation model to create a three-dimensional agency that expresses motion characteristics corresponding to the appearance of the personalized character.
Citation Information
Patent Citations
Universal multimedia access system
KR1020060070923A
C Typed-noodle and Manufacture Method Thereof
KR1020200024202A
Particulate matter treatment device and particulate matter treatment method using the same
KR1020240073426A
Electrolyte for aluminum air battery, and aluminum air battery comprising the same
KR102854230B1
KR20200087338A