Artificial intelligence-based avatar service system and method

The AI-based avatar service system efficiently generates customized avatar videos by editing 2D images and synthesizing voice and motion, addressing the limitations of traditional methods with cost-effective and flexible production of personalized avatars.

WO2025225791A1PCT designated stage Publication Date: 2025-10-30DEEPBRAIN AI INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/012304
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2024-08-20
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing methods for creating personalized avatars lack efficiency and flexibility, particularly in generating customized avatar videos that meet user-specific conditions regarding image, voice, dialogue, and movement, often requiring high production costs and complex 3D rendering processes.

Method used

An artificial intelligence-based avatar service system and method that utilizes AI models to determine, process, and generate customized avatar videos by editing 2D images, synthesizing voice and motion, and performing post-processing to create a natural and immersive experience, allowing users to input their images, styles, or random generation, and control dialogue and motion.

Benefits of technology

Enables the rapid and cost-effective production of customized avatar videos that match user preferences, utilizing 2D images and short voice samples, with natural voice synthesis and motion correction, resulting in an immersive and high-quality video experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024012304_30102025_PF_FP_ABST
    Figure KR2024012304_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an artificial intelligence-based avatar service system and method. An avatar service system according to one embodiment of the present invention comprises: a video determination unit for determining a first avatar image according to a preset criterion and receiving dialogue corresponding to the first avatar image; a video processing unit for generating a second avatar image by matching voice and motion to the first avatar image, synthesizing voice corresponding to the dialogue with the first avatar image, and generating motion for video on the basis of the avatar image, the motion, and the voice; and a video generation unit for generating an avatar video by performing post-processing on the second avatar image, wherein motion for each body part of the second avatar image is blended, and the blended second avatar image is supplemented according to a preset quality criterion.
Need to check novelty before this filing date? Find Prior Art

Description

Artificial intelligence-based avatar service system and method

[0001] The disclosed embodiments relate to an artificial intelligence-based avatar service system and method.

[0002] Business cards allow users to introduce their companies or affiliations for business purposes or as a means of self-promotion. Advances in technology have led users to experiment with various methods of self-introduction, moving beyond the traditional paper-based business cards consisting of text-based personal information and simple 2D images.

[0003] Additionally, users may wish to create videos for themselves or other users, even if they are not for business purposes or self-promotion purposes.

[0004] The disclosed embodiments aim to provide an artificial intelligence-based avatar service system and method that generates and provides a customized avatar video that meets the user's conditions using artificial intelligence technology.

[0005] An avatar service system according to one embodiment includes: a video determination unit that determines a first avatar image according to preset criteria and receives dialogue corresponding to the first avatar image; a video processing unit that matches a voice and a motion to the first avatar image to generate a second avatar image, synthesizes a voice corresponding to the dialogue to the first avatar image, and generates a motion for a video based on the avatar image, the motion, and the voice to generate the second avatar image; and a video generation unit that performs post-processing on the second avatar image to generate an avatar video, blending motions of each body part of the second avatar image and performing supplementary processing on the blended second avatar image according to preset quality criteria.

[0006] The above video determination unit may determine an uploaded user image as the first avatar image, or, if text including a desired style is input, may generate an image of a style corresponding to the desired style according to preset image generation criteria and determine the image as the first avatar image, or, if a random number is input, may determine an image randomly generated using a first artificial intelligence model pre-trained based on the random number as the first avatar image.

[0007] The above video decision unit can edit the first avatar image according to editing items including appearance, clothing, hairstyle, age, and background using a pre-learned second artificial intelligence model.

[0008] The above video determination unit can determine lines to be applied to the first avatar image when receiving an image for motion, dialogue, and voice, or receiving a first user voice including dialogue by a user, or receiving a desired dialogue text, or receiving a desired dialogue text and a second user voice.

[0009] The video processing unit, when receiving the desired dialogue text, generates a voice matching the first avatar image using a pre-trained third artificial intelligence model and synthesizes the voice to the first avatar image according to the desired dialogue text, or synthesizes a voice selected from among a plurality of preset voices to the first avatar image according to the desired dialogue text, and when receiving the desired dialogue text and the second user voice, performs learning based on the second user voice and synthesizes the second user voice to the first avatar image according to the desired dialogue text using a fourth artificial intelligence model.

[0010] The above video processing unit, when receiving the motion image reflecting motion, dialogue, and voice, synthesizes the motion and voice into the first avatar image based on the motion, dialogue, and voice of the motion image, and when receiving a first user voice including dialogue from the user, synthesizes lip motion into the first avatar image according to the first user voice, or synthesizes body motion according to the dialogue.

[0011] The above video generation unit can perform supplementary processing, including sharpness, resolution, and color conversion, of the second avatar image according to quality standards using a pre-learned fifth artificial intelligence model.

[0012] The above video processing unit can use a pre-learned sixth artificial intelligence model to identify a motion vector including gender, gaze direction, and age based on the face of the first avatar image, and can correct a motion vector of an image for motion synthesized to the first avatar image based on the identified motion vector of the first avatar image.

[0013] According to another embodiment, an avatar service method is provided, which is performed by an avatar service system, comprising: a step in which the avatar service system determines a first avatar image according to preset criteria, and receives dialogue corresponding to the first avatar image; a step in which a voice and a motion are matched to the first avatar image to generate a second avatar image, wherein a voice corresponding to the dialogue is synthesized into the first avatar image, and a motion for a video is generated based on the avatar image, the motion, and the voice, to generate the second avatar image; and a step in which post-processing is performed on the second avatar image to generate an avatar video, wherein the motion of each body part of the second avatar image is blended, and the blended second avatar image is supplementally processed according to preset quality criteria.

[0014] The above avatar service method may determine, when determining the first avatar image, an uploaded user image as the first avatar image, or, if text including a desired style is input, an image of a style corresponding to the desired style is generated according to preset image generation criteria and determined as the first avatar image, or, if a random number is input, an image randomly generated using a first artificial intelligence model pre-learned based on the random number may be determined as the first avatar image.

[0015] In addition, a computer-readable recording medium recording a computer program for executing a method for implementing the disclosed embodiment may be further provided.

[0016] According to the disclosed embodiments, a customized avatar video that meets the conditions of a user's desired image, voice, dialogue, movement, and style can be generated and provided using artificial intelligence technology.

[0017] In addition, according to the disclosed embodiments, since the motion is implemented using the 2D image itself rather than the 3D avatar rendering work, the production cost is relatively low and the production can be done quickly.

[0018] Additionally, according to the disclosed embodiments, an avatar video can be generated by learning a user's own voice through relatively short voice samples.

[0019] Additionally, according to the disclosed embodiments, a natural voice that matches the avatar image can be synthesized into the avatar through facial appearance recognition even without a second user voice.

[0020] Additionally, according to the disclosed embodiments, an all-in-one avatar video can be created that allows a user to read a desired script or control the angle and gestures of the head using just one 2D image.

[0021] Figure 1 is a block diagram illustrating an avatar service system according to one embodiment.

[0022] Figures 2a and 2b are block diagrams showing the avatar service system of Figure 1 in more detail.

[0023] Figures 3 to 7 are exemplary diagrams for explaining an avatar service method according to one embodiment.

[0024] Figure 8 is a flowchart for explaining an avatar service method according to one embodiment.

[0025] FIG. 9 is a block diagram illustrating a computing environment including a computing device according to one embodiment.

[0026] Hereinafter, specific embodiments of the present invention will be described with reference to the drawings. The following detailed description is provided to facilitate a comprehensive understanding of the methods, devices, and / or systems described herein. However, these are merely examples and the present invention is not limited thereto.

[0027] In describing embodiments of the present invention, if a detailed description of a known technology related to the present invention is judged to unnecessarily obscure the gist of the present invention, the detailed description will be omitted. In addition, the terms described below are terms defined in consideration of their functions in the present invention, and this may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification. The terminology used in the detailed description is only for the purpose of describing embodiments of the present invention and should not be limited in any way. Unless clearly used otherwise, the singular form includes the plural form. In this description, expressions such as "comprises" or "having" are intended to indicate certain features, numbers, steps, operations, elements, parts or combinations thereof, and should not be construed to exclude the presence or possibility of one or more other features, numbers, steps, operations, elements, parts or combinations thereof other than those described.

[0028] FIG. 1 is a block diagram illustrating an avatar service system according to one embodiment, and FIGS. 2a and 2b are block diagrams illustrating the avatar service system of FIG. 1 in more detail. In this case, FIGS. 2a and 2b are sections of FIG. 2 divided for convenience of explanation, and FIG. 2b may be a drawing illustrating a portion of FIG. 2a in detail.

[0029] Hereinafter, an avatar service method according to one embodiment will be described with reference to FIGS. 3 to 7, which are exemplary diagrams for explaining the method.

[0030] Referring to FIG. 1, the avatar service system (100) includes a video determination unit (110), a video processing unit (130), and a video generation unit (150). The components illustrated in FIG. 1 are not essential for implementing the avatar service system (100) according to the present disclosure, and thus the avatar service system (100) described in this specification may have more or fewer components than the components listed above.

[0031] Referring to FIGS. 2A and 2B, the video determination unit (110) described above can be divided into an image generation unit, an image editing unit, an image preprocessing unit, and a motion and dialogue input unit, the video processing unit (130) can be divided into a voice generation unit and a motion generation unit, and the video generation unit (150) can be divided into a video postprocessing unit, an image quality enhancement unit, and an image generation unit, but is not limited thereto, and can be divided into other categories or integrated. Hereinafter, for convenience of explanation, the video determination unit (110), the video processing unit (130), and the video generation unit (150) will be described by naming them.

[0032] The components illustrated in FIG. 1 may be communicatively connected to one another via a communications network (not shown). In some embodiments, the communications network may include the Internet, one or more local area networks, a wide area network, a cellular network, a mobile network, other types of networks, or a combination of these networks.

[0033] Referring to FIGS. 1 and 3, the video determination unit (110) can determine a first avatar image according to preset criteria and receive dialogue corresponding to the first avatar image. Referring to step 3 of FIG. 3, the video determination unit (110) can receive dialogue (script) such as "Hello. Candidate 1 ~ Thank you." input by the user. To this end, the video determination unit (110) can provide an item for inputting dialogue through an application screen running on a web page or a user terminal (not shown).

[0034] The first avatar image may refer to an avatar image to be applied to an avatar video. In this case, the first avatar image, as well as the avatars described below, may be actual images of the user, such as photographs, or images created based on the actual user's appearance or fiction, such as characters. However, the present invention is not limited thereto and may be implemented as any form of image capable of representing the user. For example, the first avatar image may be determined to be a race different from the user's actual appearance, based on the user's selection or preset criteria.

[0035] For example, referring to FIGS. 2A, 2B, 4A, and 4B, the video determination unit (110) may determine an uploaded user image as the first avatar image. Referring to step 2 of FIG. 3, the video determination unit (110) may determine an uploaded user image, including a user's self-camera image, an image downloaded from the Internet, etc., as the first avatar image so that it can be applied as is to the avatar video. At this time, the number of uploaded user images may be one, but is not limited thereto. In addition, the user image may be a 2D image. The user image may mean an image uploaded by the user or another user.

[0036] When a user uploads an image, the video decision unit (110) may provide a recommended image guide on a web page or application screen, as shown in FIGS. 4a and 4b, to facilitate the user's upload. At this time, FIGS. 4a and 4b are diagrams that are separated from FIG. 4 for convenience of explanation, and FIG. 4b may be a drawing for explaining a portion of FIG. 4a in detail.

[0037] As another example, when text including a desired style is input, the video determination unit (110) may generate an image of a style corresponding to the desired style according to preset image generation criteria and determine it as the first avatar image. The preset image generation criteria may be criteria for extracting and applying an image corresponding to a specific text from data stored by pre-matching text and its corresponding image.

[0038] For example, if the video decision unit (110) receives text including a desired style such as <short-haired male with blue hair>, it can extract a matching image from pre-stored body region images to generate a first avatar image.

[0039] Additionally, the video decision unit (110) can use a pre-learned artificial intelligence model to input text containing the desired style and output an image corresponding to the desired style.

[0040] As another example, when a random number is input, the video determination unit (110) may determine an image randomly generated using a first artificial intelligence model that has been pre-trained based on the random number as the first avatar image. Specifically, the video determination unit (110) may prepare a first artificial intelligence model that is a generative AI model that has been pre-trained with a plurality of human appearance data. Since the first artificial intelligence model has been trained with specific data sets of various races, including Asians, Caucasians, and Blacks, when a specific race is selected by the user, a random image that matches the race can be generated and output. The generative AI model has a structure called Generative Adversarial Networks (GAN), and can operate in a manner that when a random number is initially input, a high-resolution human image is generated while passing through layers of the artificial intelligence model. Accordingly, when the "Generate" button is selected by the user, the video determination unit (110) may generate a random random number using the first artificial intelligence model, and generate and output an image having a human appearance while passing the random number through layers.

[0041] The video decision unit (110) can edit the first avatar image according to editing items including appearance, clothing, hairstyle, age, and background using a pre-trained second artificial intelligence model (② of FIG. 2a). At this time, the video decision unit (110) can select a part of the first avatar image that needs editing and allow the user to change it by entering a description such as a recommended style or a text description if there is a style desired by the user. To this end, the video decision unit (110) can display items that can process editing item selection, recommended style selection, and desired style text input, etc., through a web page or application screen that provides the avatar service.

[0042] Specifically, the video determination unit (110) can generate a first avatar image that corresponds to (matches) the text containing the input desired style. For example, if an old man is input, the video determination unit (110) can generate a first avatar image that reflects relatively many wrinkles and gray hair.

[0043] Afterwards, the video decision unit (110) can change the hair color in the first avatar image at the user's request, but can also recommend editing guidelines (e.g., suggestions for changing white hair to black hair, suggestions for improving wrinkles on the skin) so that the previously input desired style text can be improved externally (① in FIG. 2a).

[0044] At this time, the video determination unit (110) uses an artificial intelligence model, and when generating the first avatar image, text for the style desired by the user can be embedded in numerical form and input for a positive prompt. The video determination unit (110) can apply attenuation to the embedded numerical value and change the sign to have a negative prompt tendency, thereby generating an image of an intermediate form.

[0045] The video decision unit (110) can identify similar words and antonyms to the input text through an artificial intelligence model (e.g., an LLM (large language model) model) and apply these as positive prompts to recommend the creation of a first avatar image of a variety of or completely different styles.

[0046] The above-described technique can be applied when a user inputs text or an image containing a desired style. For example, when an image is uploaded, the video determination unit (110) can segment the hair-related portion of the uploaded user image and provide cases in which the segmented portion is synthesized with various color prompts as editing guidelines so that the user can confirm and edit the segmented portion. For example, when the video determination unit (110) receives an image of black hair, it can provide cases in which hair colors other than black, such as blue and yellow, are synthesized as editing guidelines.

[0047] The video decision unit (110) can determine the dialogue to be applied to the first avatar based on the received information described below, but is not limited thereto.

[0048] For example, the video determination unit (110) may receive an image for motion that reflects motion, dialogue, and voice. At this time, the image for motion is an image for providing dialogue, voice, and motion to be synthesized onto the first avatar image, and may be a video that the user directly produces and records desired dialogue, voice, and motion. When the image for motion is uploaded, the video determination unit (110) may extract and transmit the motion, dialogue, and voice applied to the image for motion so that they can be synthesized onto the first avatar image.

[0049] As another example, the video determination unit (110) may receive a first user voice containing dialogue from the user. In this case, the first user voice may refer to a voice recorded by reflecting the dialogue to be output by the user through the avatar in the avatar video. In other words, the first user voice refers to a state in which the dialogue is contained. The video determination unit (110) may control the lips of the avatar to move according to the first user voice and dialogue through the video processing unit (130) (Fig. 2b (a)).

[0050] As another example, the video determination unit (110) may receive a desired dialogue text. The video determination unit (110) may control the video processing unit (130) to generate a voice corresponding to the desired dialogue text by an artificial intelligence model, or to apply a pre-stored default voice. In addition, the video determination unit (110) may control the avatar's lips to move according to the desired dialogue text by the video processing unit (130) ((b) of FIG. 2b).

[0051] As another example, the video determination unit (110) may receive a desired dialogue text and a second user voice. For example, referring to step 1 of FIG. 3, the video determination unit (110) may receive a second user voice recorded for approximately 10 seconds. The second user voice is intended to recognize the voice to be applied to the avatar video, and may be used to simply recognize the user's voice. In other words, the second user voice does not include the dialogue to be output through the avatar video.

[0052] Specifically, the video decision unit (110) can recognize and synthesize the second user's voice (user voice) using a zero-shot TTS model, but the applied technology is not limited thereto. The zero-shot TTS model can be trained to generate multiple user's voices as one embedding each, and to reproduce the corresponding utterance using the generated embedding and text. Accordingly, the video decision unit (110) can perform text-to-speech (TTS) that mimics the user's voice by converting any user's voice input using the zero-shot TTS model into an embedding that captures the characteristics of the voice well.

[0053] When the first and second users upload their voices, the video decision unit (110) can provide a voice recording guide on a web page or application screen, as shown in FIGS. 5a and 5b, to facilitate the user's recording process. In this case, FIGS. 5a and 5b are separate parts of FIG. 5 for convenience of explanation.

[0054] As another example, the video decision unit (110) can recommend lines based on the appearance and costumes of the first avatar image.

[0055] Specifically, the video decision unit (110) can recommend frequently used lines by considering the user's occupation identified based on the first avatar image. At this time, the video decision unit (110) can perform area identification for the face area, area identification for clothing parts, and area identification for the background excluding objects using an artificial intelligence model with a segmentation function.

[0056] In addition, the video determination unit (110) can identify gender and age through the face region, and extract keywords for the clothing part. For example, if the video determination unit (110) determines from the first avatar image that a young man wearing a swimsuit is at the beach, it can provide a recommended line such as "The weather is nice today, so it's good for swimming." In other words, once the video determination unit (110) determines the region of interest (ROI), it can identify the key text prompt from the first avatar image based on this. The video determination unit (110) can apply the key text prompts to an artificial intelligence model such as LLM or ChatGPT to recommend customized lines (natural lines) corresponding to the first avatar image.

[0057] The video decision unit (110) can perform a preprocessing procedure to process the area in which the motion is to be generated for the first avatar image into a form that is easy to edit using an artificial intelligence model before matching the voice and motion to the first avatar image (③ of FIG. 2a).

[0058] Specifically, the video decision unit (110) uses a region detection technology to recognize pre-designated regions of the body region, background region, and face region (e.g., head, torso, arms, legs, background, etc.) in the first avatar image, and can reconfigure the data into a data behavior suitable for the avatar video using functions such as crop, resize, and realign.

[0059] The video decision unit (110) can change the attributes of a specific appearance when editing the first avatar image. To do this, since recognition of the area in the first avatar image is required, the video decision unit (110) can use an artificial intelligence model with a segmentation function (② of FIG. 2a).

[0060] The video decision unit (110) may, prior to editing the first avatar image, preliminarily perform a segmentation operation on the area of ​​the face or body where editing will be performed the most, and may calculate bounding boxes in advance to determine the composition in which the image should be cropped during preprocessing of the first avatar image (③ of FIG. 2a) so that a quick process can be performed.

[0061] The video processing unit (130) can generate a second avatar image by matching voice and motion to the first avatar image, synthesize a voice corresponding to the dialogue into the first avatar image, and generate a motion for the video based on the avatar image, motion, and voice to generate the second avatar image.

[0062] The above second avatar image may refer to an image in which voice and motion are synthesized with the first avatar image. In this case, the second avatar image may additionally synthesize various effects that can be applied to images within the avatar video in addition to voice and motion.

[0063] For example, when the video processing unit (130) receives a desired dialogue text through the video determination unit (110), it may generate a voice matching the first avatar image using a pre-learned third artificial intelligence model and synthesize the voice into the first avatar image according to the desired dialogue text, or synthesize a voice selected from among a plurality of preset voices into the first avatar image according to the desired dialogue text.

[0064] As another example, referring to FIGS. 5a and 5b, when the video processing unit (130) receives the desired dialogue text and the second user voice through the video determination unit (110), the second user voice can be synthesized into the first avatar image according to the desired dialogue text using the fourth artificial intelligence model based on the second user voice.

[0065] Specifically, the video processing unit (130) can recognize the second user's voice and use the fourth artificial intelligence model to cause the first avatar image to speak in the second user's voice according to the desired dialogue text. Thereafter, the video processing unit (130) can process the avatar's lips to move according to the desired dialogue text ((c) of FIG. 2b).

[0066] At this time, the video processing unit (130) can perform processing such as sentence concatenation, space addition, and resampling before synthesizing the generated user voice to the first avatar image.

[0067] As another example, when the video processing unit (130) receives an image for motion that reflects motion, dialogue, and voice, it can synthesize motion and voice into the first avatar image based on the motion, dialogue, and voice of the image for motion. That is, the video processing unit (130) can implement the motion and voice of the first avatar image according to the speech and movement of the video (motion image) that reflects motion, dialogue, and voice. To this end, the first avatar image can be divided into each region of the image according to the motion generation region in advance, so that individual condition changes can be performed for each region.

[0068] As another example, when the video processing unit (130) receives a first user voice including dialogue from a user, it can synthesize lip movements into the first avatar image according to the first user voice, or synthesize body movements according to the dialogue.

[0069] Meanwhile, people may have behavioral patterns, such as performing specific actions when making specific utterances. The present embodiment can apply unique patterns of humans to create avatar videos that reflect natural movements. To this end, the video processing unit (130) can calculate cosine similarity for vectors in the embedding space when embedding dialogue. If the calculation result falls within a specific threshold, the video processing unit (130) determines that the personality is similar, and can match specific behavioral patterns to dialogue with similar personalities.

[0070] When generating a movement, the video processing unit (130) can reflect speaking habits by giving a pattern to the expression (e.g., blinking, etc.) of the speaker's face when a specific sentence or word is repeatedly applied in the delivered dialogue.

[0071] Specifically, when the video processing unit (130) recognizes a word related to an expression in the transmitted dialogue, it can reflect an expression matched thereto. For example, when a dialogue (e.g., a prompt) such as "The weather is nice today, so my mood is good" is input, the video processing unit (130) can generate a smiling face reflecting happiness pre-matched to the dialogue as an expression on the speaker's face. That is, a dialogue such as "The weather is nice today, so my mood is good" may be pre-matched and stored with a smiling expression reflecting happiness, or a dialogue such as "It's nice" or "My mood is good" may be pre-matched and stored with a smiling expression reflecting happiness, but is not limited thereto.

[0072] In addition, if the video processing unit (130) recognizes a word (including a sentence) related to a head pose in the transmitted dialogue, it can reflect the head pose matching it. For example, if a dialogue showing an attitude of agreement, such as "Yes, that's right," is input, the video processing unit (130) can reflect a head pose of nodding that has been pre-matched to the dialogue as a gesture on the speaker's face.

[0073] When the video processing unit (130) randomly generates a voice that matches the image, rather than using the face of the first avatar image generated by the video determination unit (110), it may utilize an artificial intelligence model of the user's actual face (e.g., a model having gender estimation, age estimation, and headpose estimation functions).

[0074] Specifically, once the gender is determined, the video processing unit (130) can embed it and process it into a style vector. These style vectors can be input into a TTS artificial intelligence model to generate a voice of the corresponding gender. In other words, the video processing unit (130) can extract gender information from the external information of the user's face and provide a new voice (voice) that corresponds to (matches) the face.

[0075] Once the age is determined, the video processing unit (130) can embed it and process it into a style vector. These style vectors can be input into a TTS artificial intelligence model to generate a voice with a corresponding age (numerical age or degree of aging). At this time, the video processing unit (130) can transform the generated voice into a more mature and mature voice, or generate the voice of a minor whose voice has reached puberty.

[0076] In addition, the video processing unit (130) can adjust the volume of the voice according to the pitch information of the head pose. For example, the video processing unit (130) can adjust the volume of the voice to be larger than the current volume by considering the principle that sound is transmitted well when a person looks straight ahead, and can adjust the volume to implement a three-dimensional voice transmission effect by considering the principle that a relatively small voice is transmitted when the pitch angle is low because the throat is pressed.

[0077] The video processing unit (130) can generate an identification embedding vector by applying a face recognition model when recognizing a user's face. These can be used as part of the style vectors input to the TTS model to create a special voice according to the facial appearance.

[0078] The video processing unit (130) can use the pre-learned sixth artificial intelligence model to identify motion vectors, including gender, gaze direction, and age, based on the face of the first avatar image, and can correct the motion vector of the motion image synthesized with the first avatar image based on the identified motion vector of the first avatar image. The motion vector of the first avatar image or the motion image may refer to a vector applied to implement a motion applied to the avatar.

[0079] At this time, the first avatar image may be an actual image of the user, such as a photograph, or may include an image created based on the actual user's appearance or a fictional character. That is, the video processing unit (130) corrects the motion vector applied according to the motion vector of the first avatar image created based on the user's appearance or the image the user wishes to set as an avatar so that they match.

[0080] For example, the first avatar image generated by the video determination unit (110) may be a female, and the motion subject of the captured motion image may be a male. In this case, if the large and aggressive movements of the male are directly applied to the first avatar image of the female, the resulting dissonance between the appearance and the motion may hinder the immersion of the users. The video processing unit (150) of the present embodiment may perform a correction operation on the motion vector for the avatar applied to the avatar video.

[0081] For example, the video processing unit (150) can compensate for the size of a motion using gender information obtained from the face of the first avatar image. For example, if the gender of the first avatar image is male, the video processing unit (150) can compensate for the motion to be larger than the current size, and if the gender of the first avatar image is female, the motion to be smaller.

[0082] As another example, the video processing unit (150) can estimate emotions from the face of the first avatar image. Emotion estimation AI models can typically derive seven different emotions. For example, the video processing unit (150) can reflect a joyful emotion with relatively fast and large motions to create a lively effect. For negative emotions such as depression, the video processing unit (150) can reflect a relatively restrained and small motion.

[0083] As another example, the video processing unit (150) can estimate the body pose through a keypoint detector. For example, if a first avatar image is generated with one hand raised, the video processing unit (150) can guarantee robustness in the overall motion and natural movement by excluding the arm portion from the motion reflection area and applying motion to the remaining skeleton because the motion of the motion image is in a state of raising both hands, which results in a large inconsistency.

[0084] As another example, the video processing unit (150) can correct the relative scale of the key point's motion when the facial size of the first avatar image and the facial size of the motion image do not match, thereby synthesizing the motion onto the first avatar image. Using this principle, the video processing unit (150) can correct the relative motion scale synthesized onto the first avatar image according to emotion, gender, etc.

[0085] The video generation unit (150) performs post-processing on the second avatar image to generate an avatar video, and can perform blending processing on the motion of each body part of the second avatar image and supplementary processing on the blended second avatar image according to preset quality standards.

[0086] The avatar videos disclosed in these embodiments can be implemented as self-introduction videos including business cards, various greeting videos including New Year's greeting videos, digital business cards for promotional purposes, presentation materials using avatars, promotional videos, and videos for memorial photos. The examples of implementation are not limited to these. In other words, they can be applied to any field where avatar videos can be utilized.

[0087] Specifically, the video generation unit (150) can perform blending processing so that the motions of the head and each body part (motion generation area) generated by the video processing unit (130) can be naturally synthesized into the second avatar image. The video generation unit (150) enables the motions to be ergonomically positioned in the corresponding parts to which they are matched, and performs post-processing so that the motions of the avatar in the synthesized resultant avatar video appear non-discontinuous and visually natural.

[0088] The video generation unit (150) can perform supplementary processing, including sharpness, resolution, and color conversion, of the second avatar image according to quality standards using the pre-learned fifth artificial intelligence model.

[0089] Specifically, the video generation unit (150) can utilize the fifth artificial intelligence model to perform various filters and high-resolution conversions to improve the quality of the second avatar image, which has been degraded after blending processing. Through this, the video generation unit (150) can implement improvements such as increased clarity, resolution, and natural color conversion of the second avatar image.

[0090] Referring to FIG. 7, the video generation unit (150) performs procedures such as adding video effects including subtitles and banners, adding voice effects, and adding background music, and encodes the final avatar video and the voice applied thereto into a video file and exports it (e.g., download, upload to OTT or SNS, provide URL, etc.).

[0091] The video processing unit (130) generates movement within a localized area, such as the head or the arms, in the avatar image, and then performs a process of combining and synthesizing all parts to create a single video. At this time, if the head movement is significant, the entire movement may not align with the body when merged. This can cause misalignment between the neck movement generated from the face and the neck movement generated from the body movement.

[0092] The video generation unit (150) of the present embodiment can, taking into account the misalignment described above, segment the neck, which is a joint that connects the head and the body in the second avatar image, using an artificial intelligence model to determine the center of gravity of the neck. The video generation unit (150) can perform calibration so that the center of gravity of the neck determined from the face image and the center of gravity of the neck determined from the body image are aligned with each other.

[0093] When the video generation unit (150) receives audio transmitted from the video determination unit (110), it can perform STT (Speech to Text) operations in parallel while other processes are in progress, taking into account the possibility that additional options, such as subtitles, will be transmitted. This can shorten the overall processor time.

[0094] For example, the video generation unit (150) can store text for subtitles in advance and provide the pre-stored text the moment an option requesting video post-processing work is activated, so that the post-processing work can be performed quickly.

[0095] As another example, the video generation unit (150) may perform procedures such as spell checking on the dialogue transmitted from the video determination unit (110). Through this, it is expected that the reliability of users (e.g., production requesters, avatar video recipients) regarding the avatar videos provided in this embodiment can be improved.

[0096] As another example, the video generation unit (150) can recommend video effects stored in a dictionary that match the words obtained from STT.

[0097] When the video generation unit (150) performs editing processing on the first avatar image in the video determination unit (110), when the avatar desired by the user is finally completed, the video generation unit (150) can determine the atmosphere of the first avatar image and prepare for special video effects before the avatar video is completed.

[0098] Specifically, the video generation unit (150) inputs a second avatar image obtained through image editing in the video determination unit (110) into a pre-trained artificial intelligence model (e.g., a captioning model) to identify a key text prompt, and compares the prompt vector of the identified key text prompt with a music mood label prompt matched to pre-stored background music (bgm) to recommend background music having a similar cosine similarity.

[0099] The video generation unit (150) can apply the above-described method in the same or similar manner not only to background music but also to special effects that can be applied to avatar videos.

[0100] The special effects that can be applied to the above-described avatar video can be confirmed by blending the final avatar video using an alpha channel, which is a special channel that repeatedly covers or controls a specific area of ​​an image. The video generation unit (150) can prepare the above-described special effects in advance while the video processing unit (130) is performing image synthesis, so that an avatar video with special effects applied can be obtained relatively quickly. For example, in the case of an image with a dark background (using an image captioning model) or a gloomy expression of the speaker (using an emotion estimator model), the video generation unit (150) can prepare a rainy video effect in advance while the image synthesis process in the video processing unit (130) is in progress so that the above-described atmosphere can be highlighted in the avatar video, and when a second avatar image is transmitted, it can be projected onto the alpha channel to immediately recommend the pre-prepared special effects, so that the user can confirm that the second avatar image has the special effects applied. At this time, an additional description of the special effects can be labeled in advance.

[0101] The above-described avatar service may be provided, for example, in the form of a web page or through an application installed on a user terminal (not shown).

[0102] Specifically, referring to FIGS. 3 to 7, the avatar service system (100) can provide information required in relation to the overall procedure when recording a user's voice, uploading a user image, and performing various procedures for creating an avatar video (e.g., entering a title for the avatar video in FIG. 6, selecting a voice to apply, selecting an image and motion (gesture), entering a dialogue (script), and registering and specifying information on a recipient to whom the avatar video will be provided, etc.) through a web page or application screen. Referring to FIG. 7, the video generation unit (150) can store and manage information (e.g., file title, video title, dialogue (script), playback time, production date, production status, and video provision method (e.g., video download and URL copy)) for each of a plurality of avatar videos by matching them with each other. The video generation unit (150) can provide avatar video information managed in the form of a list according to a user's request and provide a selected avatar video. At this time, the video generation unit (150) can enable downloading an avatar video or copying a URL matching an avatar video according to a preset video provision method. The above video provision method is an example and can be changed according to the needs of the operator.

[0103] According to one embodiment, the video determination unit (110), the video processing unit (130), and the video generation unit (150) may be implemented using one or more physically separate devices, or may be implemented by one or more hardware processors or a combination of one or more hardware processors and software, and may not be clearly distinguished in specific operations, unlike the illustrated example.

[0104] The avatar service system (100) of the present disclosure may be configured to include one or more cores, although not illustrated, and may include a processor for data analysis and deep learning, such as a central processing unit, a general purpose graphics processing unit, and a tensor processing unit of a computing device. Such processors may be implemented within each of the video determination unit (110), the video processing unit (130), and the video generation unit (150), or may be implemented separately. The processor may read a computer program stored in a memory (not illustrated) to perform data processing for machine learning according to the present disclosure. According to the present disclosure, the processor may perform operations for learning a neural network. The processor may perform calculations for learning a neural network, such as processing input data for learning in deep learning, extracting features from the input data, calculating errors, and updating the weights of the neural network using backpropagation.

[0105] The above neural network model may be a deep neural network. In the present disclosure, neural network, network function, and neural network may be used interchangeably. A deep neural network (DNN) may refer to a neural network that includes multiple hidden layers in addition to an input layer and an output layer. Using a deep neural network, it is possible to identify latent structures of data. That is, it is possible to identify latent structures of photos, text, videos, voices, and music (e.g., what objects are in the photo, what the content and emotion of the text are, what the content and emotion of the voice are, etc.). A deep neural network may include a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a Q network, a U network, a Siamese network, etc.

[0106] Convolutional neural networks (CNNs) are a type of deep neural network that include neural networks containing convolutional layers. CNNs are a type of multilayer perceptron designed to use minimal preprocessing. CNNs can be composed of one or more convolutional layers and artificial neural network layers combined with them. CNNs can additionally utilize weight and pooling layers. This structure allows CNNs to fully utilize two-dimensional input data. CNNs can be used to recognize objects in images. CNNs can process image data by representing it as a matrix with dimensions. For example, in the case of image data encoded in RGB (red-green-blue), each of the R, G, and B colors can be represented as a two-dimensional (for example, in a two-dimensional image) matrix. That is, the color value of each pixel of the image data can be an element of a matrix, and the size of the matrix can be the same as the size of the image. Therefore, the image data can be represented as three two-dimensional matrices (a three-dimensional data array).

[0107] In a convolutional neural network, a convolutional process (input and output of a convolutional layer) can be performed by moving the convolutional filter and multiplying the matrix elements at each location of the image with the convolutional filter. The convolutional filter can be composed of an n*n matrix. The convolutional filter can generally be composed of a fixed-shape filter that is smaller than the total number of pixels in the image. That is, when an m*m image is input to a convolutional layer (for example, a convolutional layer whose convolutional filter has a size of n*n), a matrix representing n*n pixels containing each pixel of the image can be component-wise multiplied with the convolutional filter (i.e., each element of the matrix is ​​multiplied). By multiplying with the convolutional filter, a component matching the convolutional filter can be extracted from the image. For example, a 3*3 convolutional filter for extracting up and down straight line components from an image can be configured as [[0,1,0], [0,1,0], [0,1,0]]. When a 3*3 convolutional filter for extracting up and down straight line components from an image is applied to an input image, up and down straight line components matching the convolutional filter from the image can be extracted and output. A convolutional layer can apply a convolutional filter to each matrix for each channel representing an image (i.e., R, G, B colors in the case of an R, G, B coded image). A convolutional layer can extract features matching the convolutional filter from the input image by applying a convolutional filter to the input image. The filter value of the convolutional filter (i.e., the value of each element of the matrix) can be updated by backpropagation during the learning process of a convolutional neural network.

[0108] A subsampling layer can be connected to the output of a convolutional layer to simplify the output of the convolutional layer and reduce memory usage and computational amount. For example, when the output of the convolutional layer is input to a pooling layer having a 2*2 max pooling filter, the image can be compressed by outputting the maximum value included in each patch for each 2*2 patch from each pixel of the image. The above-described pooling may be a method of outputting the minimum value in a patch or the average value of a patch, and any pooling method may be included in the present disclosure.

[0109] A convolutional neural network may include one or more convolutional layers and subsampling layers. A convolutional neural network can extract features from an image by repeatedly performing convolutional and subsampling processes (e.g., the aforementioned max pooling). Through repeated convolutional and subsampling processes, the neural network can extract global features of the image.

[0110] The output of a convolutional layer or a subsampling layer can be input to a fully connected layer. A fully connected layer is a layer in which all neurons in one layer are connected to all neurons in the neighboring layer. A fully connected layer can refer to a structure in a neural network in which all nodes in each layer are connected to all nodes in other layers.

[0111] At least one of the CPU, GPGPU, and TPU of the processor can process network function learning. For example, the CPU and GPGPU can jointly process network function learning and data classification using the network function. Furthermore, in one embodiment of the present disclosure, processors of multiple computing devices can be used together to process network function learning and data classification using the network function. Furthermore, a computer program executed on a computing device according to one embodiment of the present disclosure may be a CPU, GPGPU, or TPU executable program.

[0112] FIG. 8 is a flowchart illustrating an avatar service method according to one embodiment. The method illustrated in FIG. 8 may be performed, for example, by the aforementioned avatar service system (100). While the illustrated flowchart divides the method into multiple steps and describes them, at least some of the steps may be performed in a different order, combined with other steps and performed together, omitted, divided into substeps and performed, or one or more steps not illustrated may be added and performed.

[0113] The avatar service system (100) disclosed in FIG. 8 can perform the same role as the avatar service system (100) disclosed in FIGS. 1 to 7 described above, and for the convenience of explanation, overlapping disclosures will be omitted.

[0114] At step 1100, the video determination unit (110) of the avatar service system (100) can determine a first avatar image according to preset criteria and receive dialogue corresponding to the first avatar image.

[0115] When determining the first avatar image, the video determination unit (110) may determine an uploaded user image as the first avatar image, or, if text including a desired style is input, may determine an image of a style corresponding to the desired style according to preset image generation criteria as the first avatar image, or, if a random number is input, may determine an image randomly generated using a first artificial intelligence model pre-learned based on the random number as the first avatar image.

[0116] At step 1200, the video processing unit (130) can generate a second avatar image by matching voice and motion to the first avatar image, synthesize voice corresponding to the dialogue into the first avatar image, and generate motion for the video based on the avatar image, motion, and voice to generate the second avatar image.

[0117] At step 1300, the video generation unit (150) performs post-processing on the second avatar image to generate an avatar video, and may blend the motions of each body part of the second avatar image and supplement the blended second avatar image according to preset quality standards.

[0118] FIG. 9 is a block diagram illustrating a computing environment including a computing device according to one embodiment. In the illustrated embodiment, each component may have different functions and capabilities other than those described below, and may include additional components other than those described below.

[0119] The illustrated computing environment (10) includes a computing device (12). The computing device (12) may be one or more components included in an avatar service system (100) according to one embodiment.

[0120] A computing device (12) includes at least one processor (14), a computer-readable storage medium (16), and a communication bus (18). The processor (14) may cause the computing device (12) to operate according to the exemplary embodiments mentioned above. For example, the processor (14) may execute one or more programs stored in the computer-readable storage medium (16). The one or more programs may include one or more computer-executable instructions, which, when executed by the processor (14), may be configured to cause the computing device (12) to perform operations according to the exemplary embodiments.

[0121] A computer-readable storage medium (16) is configured to store computer-executable instructions or program code, program data, and / or other suitable forms of information. A program (20) stored in the computer-readable storage medium (16) includes a set of instructions executable by the processor (14). In one embodiment, the computer-readable storage medium (16) may be a memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, any other form of storage medium that can be accessed by the computing device (12) and store desired information, or a suitable combination thereof.

[0122] A communication bus (18) interconnects various other components of the computing device (12), including the processor (14) and computer-readable storage media (16).

[0123] The computing device (12) may also include one or more input / output interfaces (22) that provide interfaces for one or more input / output devices (24) and one or more network communication interfaces (26). The input / output interfaces (22) and the network communication interfaces (26) are connected to the communication bus (18). The input / output devices (24) may be connected to other components of the computing device (12) via the input / output interfaces (22). Exemplary input / output devices (24) may include input devices such as pointing devices (such as a mouse or a trackpad), a keyboard, a touch input device (such as a touchpad or a touchscreen), a voice or sound input device, various types of sensor devices and / or photographing devices, and / or output devices such as display devices, printers, speakers and / or network cards. The exemplary input / output devices (24) may be included within the computing device (12) as a component constituting the computing device (12), or may be connected to the computing device (12) as a separate device distinct from the computing device (12).

[0124] Embodiments implemented in the above-described avatar service system (100) can apply various editing methods to the entire body area, not just a specific area such as the face, when generating an avatar video. These embodiments can generate an avatar video by separating various parts within an image, creating movements for each part, and then merging them back together. These embodiments can improve the completeness of the result by applying specialized editing and correction techniques for each area within the image (each body area, background area, etc.), and can make the avatar within the newly generated avatar video appear visually natural.

[0125] The disclosed embodiments may be implemented in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, and when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0126] While representative embodiments of the present invention have been described in detail above, those skilled in the art will appreciate that various modifications to the above-described embodiments are possible without departing from the scope of the present invention. Therefore, the scope of the present invention should not be limited to the described embodiments, but should be determined not only by the claims set forth below but also by equivalents thereof.

Claims

1. A video determination unit that determines a first avatar image according to preset criteria and receives dialogue corresponding to the first avatar image; A video processing unit that generates a second avatar image by matching voice and motion to the first avatar image, synthesizes a voice corresponding to the dialogue into the first avatar image, and generates a video motion based on the avatar image, motion, and voice to generate the second avatar image; and An avatar service system comprising a video generation unit that performs post-processing on the second avatar image to generate an avatar video, blends motions of each body part of the second avatar image, and supplements the blended second avatar image according to preset quality standards.

2. In claim 1, The above video decision unit, Determine the uploaded user image as the first avatar image, or When text including a desired style is input, an image of a style corresponding to the desired style is generated based on preset image generation criteria and determined as the first avatar image, or An avatar service system that determines an image randomly generated using a first artificial intelligence model that has been pre-learned based on the random number as the first avatar image when a random number is input.

3. In claim 1, The above video decision unit, An avatar service system that edits the first avatar image according to editing items including appearance, clothing, hairstyle, age, and background using a pre-trained second artificial intelligence model.

4. In claim 1, The above video decision unit, An avatar service system that determines lines to be applied to the first avatar image when receiving an image for action that reflects movements, lines, and voices, or receiving a first user voice including lines by a user, or receiving a desired line text, or receiving a desired line text and a second user voice.

5. In claim 4, The above video processing unit, When the desired dialogue text is received, a voice matching the first avatar image is generated using a pre-trained third artificial intelligence model and the voice is synthesized to the first avatar image according to the desired dialogue text, or a voice selected from among a plurality of preset voices is synthesized to the first avatar image according to the desired dialogue text. An avatar service system that, when receiving the desired dialogue text and the second user voice, performs learning based on the second user voice and synthesizes the second user voice into the first avatar image according to the desired dialogue text using a fourth artificial intelligence model.

6. In claim 5, The above video processing unit, When receiving the motion image reflecting the motion, dialogue and voice, the motion and voice are synthesized into the first avatar image based on the motion, dialogue and voice of the motion image, An avatar service system that, when receiving a first user voice including dialogue by the user, synthesizes lip movements into the first avatar image according to the first user voice, or synthesizes body movements according to the dialogue.

7. In claim 6, The above video generation unit, An avatar service system that performs supplementary processing, including sharpness, resolution, and color conversion, of the second avatar image according to quality standards using a pre-learned fifth artificial intelligence model.

8. In claim 5, The above video processing unit, An avatar service system that uses a pre-learned sixth artificial intelligence model to identify a motion vector including gender, gaze direction, and age based on the face of the first avatar image, and corrects the motion vector of an image for motion synthesized on the first avatar image based on the identified motion vector of the first avatar image.

9. In a method performed by the avatar service system, A step in which the above avatar service system determines a first avatar image according to preset criteria and receives dialogue corresponding to the first avatar image; A step of generating a second avatar image by matching a voice and motion to the first avatar image, synthesizing a voice corresponding to the dialogue into the first avatar image, and generating a motion for a video based on the avatar image, motion, and voice to generate the second avatar image; and An avatar service method comprising the step of performing post-processing on the second avatar image to generate an avatar video, blending motions of each body part of the second avatar image, and performing supplementary processing on the blended second avatar image according to preset quality standards.

10. In claim 9, When determining the first avatar image above, Determine the uploaded user image as the first avatar image, or When text including a desired style is input, an image of a style corresponding to the desired style is generated based on preset image generation criteria and determined as the first avatar image, or An avatar service method, wherein when a random number is input, an image randomly generated using a first artificial intelligence model pre-trained based on the random number is determined as the first avatar image.

Citation Information

Patent Citations

  • Method and apparatus for providing moving picture using 3D user avatar

    KR101306221B1

  • Method and syste for creating three dimension contents

    KR1020180097914A

  • Pack housing of extrusion type and manufacturing method for the same

    KR1020250141972A

  • Method and apparatus for generating story book which provides sticker reflecting user's face to character

    KR102318111B1

  • Electronic device and method for generating user avatar-based emoji sticker

    US20230087879A1