Method For Generating a 3D character's talking face image using generative artificial intelligence

The method addresses unnatural speech movements in virtual characters by using a generative AI model to match feature points and adjust face change values, resulting in a natural speech video suitable for diverse applications.

KR102997236B1Active Publication Date: 2026-07-29MILLENNIAL WORKS
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
MILLENNIAL WORKS
Filing Date
2024-07-22
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing technologies for generating speech videos of virtual characters often result in unnatural mouth and head movements due to mismatched feature points between the user's face and the character's face, and low accuracy in feature detection on character faces.

Method used

A method using a generative artificial intelligence model to determine identical feature points between a real user's face and a transformed character face, adjusting face change feature values to maintain accuracy, and mapping user voice to corresponding facial expressions and head movements to generate a natural speech face video of a 3D character.

Benefits of technology

The method produces a character video that naturally mimics a user's speech, applicable to various fields such as games, the metaverse, social media, chat, live commerce, movies, and real-time animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112024079037062-PAT00006_ABST
    Figure 112024079037062-PAT00006_ABST
Patent Text Reader

Abstract

The present invention is characterized by a method for generating a speech face image of a 3D character using a generative artificial intelligence model, wherein the generative artificial intelligence model determines a face change feature value in which the feature points of a real user face and a character face transformed from the real user face are recognized identically; generates a 2D character face image from a single user face image based on the determined face change feature value; estimates a 3D face structure from the 2D character face image; and maps a user voice to determine a facial expression and head movement corresponding to the user voice, and generates a speech face image of a 3D character. According to the present invention as described above, by producing a character video that naturally mimics a user's speech, it has the effect of being applicable to various fields such as games, metaverse, SNS, chat, live commerce, movies, advertisements, and real-time animation.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a method for generating a speech face video of a character using generative artificial intelligence, and more specifically, to a technology for generating a character video that moves and speaks vividly by synthesizing a user's photo image and voice. Background Technology

[0002] Face animation of virtual characters plays a significant role in computer games, reception desks, chat rooms, movies, advertisements, and real-time animations. Creating realistic face animations is a difficult task that requires a great deal of time and effort from skilled animators. Additionally, there is a growing demand for services that provide voice-synchronized lip-sync animations using humanoid characters in conversational systems.

[0003] Accordingly, many technologies have been proposed for speech videos of avatars or characters that generate a 3D user avatar using a 2D user image and transmit voice information in a three-dimensional manner by changing the mouth shape of the face character to match the user's voice information.

[0004] For example, Korean Registered Patent No. 1306221 discloses a technology that generates a 3D user avatar reflecting a user's face and generates a dynamic avatar with moving expressions and movements by reflecting the user's facial expressions captured by a camera onto the user avatar. This technology generates a 3D user avatar using a 2D user image, recognizes the user's facial expressions from a user image input in real time and presents movements corresponding to the recognized expressions, generates a dynamic avatar by applying movements selected by the user and the recognized expressions to the 3D user avatar, and generates a user animation by compositing the user's dynamic avatar with a pre-stored background video.

[0005] In addition, Korean Published Patent No. 2023-0151162 discloses a lip-sync avatar face generation device comprising: a collection unit that collects video information from one or more video platform servers; an extraction unit that performs voice extraction and face image extraction for each speaker from the video information and provides time synchronization information; a matching unit that, based on the results of voice extraction and face image extraction for each speaker and the time synchronization information, constructs voice-text matching information in which voice and text are matched using a TTS learning model and an STF learning model of a model processing unit, and constructs face-voice matching information in which the speaker's face and voice are matched; a model processing unit that respectively constructs a TTS learning model for an avatar speech service based on artificial intelligence neural network learning based on voice-text matching information and an STF learning model for an avatar speech service based on artificial intelligence neural network learning based on face-voice matching information; a service providing unit that constructs avatar speech service information using the TTS learning model and the STF learning model and provides it to a user terminal; and a voice emotion analysis-based face image constructing unit that analyzes emotions within the voice to generate an avatar speech video in which facial expressions change and outputs it to a user terminal.

[0006] Korean Registered Patent No. 1558202 discloses an animation generation device using an avatar, comprising: a model generation unit that generates a parametric model capable of deforming facial expressions using a plurality of feature points extracted from a two-dimensional face image and searches a database for facial expression parameters corresponding to a feature word extracted from text data and the plurality of feature points; an animation generation unit that generates a vector representing the positional change of a plurality of feature points when expressing text data as speech using the phonetic characteristics of a word included in the text data and generates an animation reflecting the change in facial expression while expressing text data as speech by applying the facial expression parameters and vector to the parametric model; and an output unit that outputs the animation using an avatar corresponding to the two-dimensional face image in a virtual space.

[0007] However, existing prior art all utilizes techniques that extract feature points from the actual user's face and reflect them on the avatar (character)'s face. Consequently, the feature points extracted from the actual user's face generally do not match those of the face converted into a character, often resulting in unnatural speech images of the character, such as mouth movements and head movements during utterance.

[0008] Furthermore, since existing face feature detection technologies are models trained on images of real people's faces, there is a problem in that the accuracy of feature detection is very low in images of characters' faces. Prior art literature

[0009] 1. Korean Registered Patent No. 1306221 (Device and method for producing video using a 3D user avatar) 2. Korean Published Patent No. 2023-0151162 (Device and method for generating a lip-sync avatar face based on voice emotion analysis) 3. Korean Registered Patent No. 1558202 (Device and method for generating animation using an avatar) The problem to be solved

[0010] The present invention has been devised to solve the aforementioned problems, and the objective of the present invention is to provide a natural character speech video by maintaining the feature points of the actual user's face through the selection of optimal facial change feature values. means of solving the problem

[0011] According to one aspect of the present invention for achieving the above-mentioned purpose, a method for generating a speech face image of a 3D character using a generative artificial intelligence model is provided, comprising: a step of determining a face change feature value in which a generative artificial intelligence model recognizes the feature points of a real user face and a character face transformed from the real user face as identical; a step of generating a 2D character face image from a single user face image based on the determined face change feature value; a step of estimating a 3D face structure from the 2D character face image; and a step of determining a facial expression and head movement corresponding to the user voice by mapping the user voice and generating a speech face image of a 3D character.

[0012] Here, the step of determining a face change feature value in which a generative artificial intelligence model recognizes the feature points of a real user face and a character face transformed from the real user face as identical may include changing the face change feature value until the generative artificial intelligence model recognizes the face position in a 2D character face image, and the face change feature value may include any one of a CFG scale value, a noise removal strength value, or a combination of these values.

[0013] The step of mapping the user voice to determine facial expressions and head movements corresponding to the user voice and generating a speech face video of a 3D character may include the step of generating facial expression coefficients including mouth shape through an audio encoder and a mapping process, the step of generating head motion coefficients during speech through a variational autoencoder, and the step of generating video frames from the facial expression coefficients and head motion coefficients generated through 3D keypoint extraction and warping field generation. Effects of the invention

[0014] According to the present invention, by producing a character video that naturally mimics a user's speech, it has the effect of being applicable to various fields such as games, the metaverse, social media, chat, live commerce, movies, advertisements, and real-time animation. Brief explanation of the drawing

[0015] Figure 1 is a configuration diagram of a service system for generating speech face images of a 3D character using generative artificial intelligence according to the present invention. Figure 2 is a block diagram illustrating the internal configuration of a photo booth. FIG. 3 is a figure illustrating the process of performing a method for generating a speech face image of a 3D character using generative artificial intelligence according to the present invention. Figure 4 is an example of a 3D character's speech face using generative artificial intelligence. Specific details for implementing the invention

[0016] The embodiments described in the present invention and the configurations illustrated in the drawings are merely preferred embodiments of the present invention and do not represent all of the technical concept of the present invention; therefore, the scope of the rights of the present invention should not be interpreted as being limited by the embodiments and drawings described in the text. That is, since the embodiments are subject to various modifications and may take various forms, the scope of the rights of the present invention should be understood to include equivalents capable of realizing the technical concept. Furthermore, the objectives or effects presented in the present invention do not imply that a specific embodiment must include all of them or only such effects; therefore, the scope of the rights of the present invention should not be understood as being limited by them.

[0017] Unless otherwise defined, all terms used herein have the same meaning as generally understood by those skilled in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having meanings consistent with the context of the relevant technology and should not be interpreted as having an ideal or overly formal meaning not explicitly defined in this invention.

[0018] Hereinafter, a preferred embodiment of the present invention will be described in detail with reference to the attached drawings.

[0019] FIG. 1 is a configuration diagram of a service system for generating speech face images of a 3D character using generative artificial intelligence according to the present invention, and FIG. 2 is a block diagram illustrating the internal configuration of a photo booth.

[0020] Referring to FIG. 1, the speech face image generation service system of a 3D character using generative artificial intelligence according to the present invention comprises a photo booth (10), a user terminal (20), a service server (30), and a generative artificial intelligence server (40) connected to each other through a network.

[0021] In addition, such a network may be configured to include, for example, a plurality of access networks (not shown) and a core network (not shown), and to include an external network, for example, an internet network (not shown). Here, the access network (not shown) is an access network that performs wired and wireless communication with a service provider server (100), a user terminal (200), and an artificial neural network model server (300), and may be implemented, for example, by a plurality of base stations such as a BS (Base Station), a BTS (Base Transceiver Station), a NodeB, an eNodeB, etc., and a base station controller such as a BSC (Base Station Controller) or an RNC (Radio Network Controller). Furthermore, as described above, the digital signal processing unit and the wireless signal processing unit that were integrally implemented in the base station may be separated into a digital unit (hereinafter referred to as DU) and a radio unit (hereinafter referred to as RU), respectively, and a plurality of RUs (not shown) may be installed in a plurality of areas, and the plurality of RUs (not shown) may be connected to a centralized DU (not shown) to form the network.

[0022] The photo booth (10) is a device that performs the function of taking a user's face and providing a voiced face image of a generated 3D character, and can be provided in the form of a stand-type device in places where many people come and go, such as department stores and shopping malls.

[0023] As shown in FIG. 2, the photo booth (10) may be configured to include, in detail, a shooting unit (11), a display unit (12), a communication unit (13), a printer unit (14), a payment unit (15), and a control unit (16).

[0024] The camera unit (11) may be a standard monocular camera as a device for photographing the user's face.

[0025] A normal touchscreen can be used to perform the function of displaying a video of a user's face, a 3D character voice image corresponding to the user's face, and various menu buttons for user operation.

[0026] The communication unit (13) is a device for transmitting a video of a user's face to a service server (30), and various wired and wireless communication modules may be used.

[0027] The printer unit (14) may be an image printer for outputting a three-dimensional character speech image corresponding to a user's face in the form of a color photograph, or a 3D printer device for outputting a three-dimensional character face in the form of a figure.

[0028] The payment unit (15) is a device for paying costs when a user wants to output their 3D character speech video in the form of a photo or 3D output, and a terminal device capable of ordinary card or NFC payment may be used.

[0029] The control unit (16) is a processor for controlling the overall operation of the photo booth (10) and can perform functions such as outputting a three-dimensional character speech image corresponding to the user's face or providing an image produced by the user terminal (20).

[0030] The user terminal (20) is a terminal possessed by the user and refers to the terminal of the user who intends to use the 3D character speech video service.

[0031] The user terminal (20) can be implemented as a computer capable of connecting to a remote server or terminal via a network. Here, the computer may include, for example, a laptop, desktop, or laptop equipped with a navigation system or a web browser. At this time, the user terminal (20) can be implemented as a terminal capable of connecting to a remote server or terminal via a network. The user terminal (20) may include all kinds of handheld-based wireless communication devices, such as navigation, PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), Wibro (Wireless Broadband Internet) terminal, smartphone, smartpad, tablet PC, etc.

[0032] Although the above example describes a case where the user takes a picture of their face through the photo booth (10), the user may also receive a 3D character speech video through the service server (30) by directly connecting to the service server (30) through the user terminal (20) without using the photo booth (10).

[0033] The service server (30) is a computer provided by a company that provides a 3D character speech video generation service and can provide a 3D character speech video generation service either alone or in conjunction with a generative artificial intelligence server (40). The detailed technical configuration for generating a 3D character speech video will be explained in detail in FIG. 3.

[0034] FIG. 3 illustrates the configuration of a 3D character speech video generation device. FIG. 3 illustrates the components from the perspective of the 3D character speech video generation process without distinguishing between the functions of the service server (30) and the generative artificial intelligence server (40). The 3D character speech video generation process is explained as follows through FIG. 3.

[0035] When the user characterization unit (100) receives a user photo file, it performs the function of extracting feature points from the user's face and converting them into a character.

[0036] The face position recognition unit (200) recognizes the face position in the image converted into a character, and when the face part is recognized by extracting feature points from the character face image, the cropping unit (300) cuts out the character face part from the image so that the face part appears large in the entire image.

[0037] Depending on the degree of characterization in the user characterization unit (100), that is, the difference between the actual face and the character face, there may be cases where the face position recognition unit (200) fails to recognize the character face. Since conventional face position recognition technology is trained to recognize faces by extracting feature points based on actual human faces, if the converted character face differs significantly from the actual face, the face region may not be recognized.

[0038] If the face position recognition unit (200) fails to recognize the character face, it is important to adjust the degree of characterization by changing the face change feature value in the user characterization unit to maintain a level where feature points can be extracted from the character face while sufficiently taking on the shape of a character.

[0039] In other words, a face change feature value is determined in which the generative AI model recognizes the feature points of the actual user's face and the character face transformed from the actual user's face as identical.

[0040] Here, the face change feature value may include any one of the CFG scale value, the noise removal strength value, or a combination of these values.

[0041] The Class-Free Guidance Scale (CFG) is a parameter that determines how much importance the model places on the input text (prompt) during the generation process. It is used to control how closely the generated image matches the given text description, and the following Equation 1 represents the degree to which the text prompt adjusts the diffusion process.

[0042]

[0043] εθ(xt, t, c) is a noise prediction value with text conditions, εθ(xt, c) is a noise prediction value without text conditions, and the importance of text conditions is enhanced by adjusting the CFG scale s.

[0044] Therefore, by adjusting the CFG scale to control the reflection of text conditions, the degree of characterization of the user's face can be controlled.

[0045] The noise removal strength value can be expressed as shown in Equation 2 below, and the degree of characterization of the user's face can also be controlled by adjusting the noise removal strength.

[0046]

[0047] Here, αt is the denoising strength at time step t, and

[0048] α_min, α_max, and α_min are the minimum and maximum denoising intensity values, respectively, T is the total number of time steps, and p is a control parameter that usually has a value greater than or equal to 1.

[0049] Therefore, the optimal face change feature value is determined by appropriately changing the CFG scale and noise removal strength values ​​so that the generative AI model can recognize the feature points of the actual user face and the character face transformed from the actual user face identically.

[0050] The monocular 3D face reconstruction unit (400) estimates a 3D face structure from a single 2D character face image, and a learning framework that predicts parameters of a 3D model based on a large dataset may be used.

[0051] The audio encoder (500) converts the user's voice audio signal into a digital signal, and the mapping unit (600) maps the user's audio voice to a corresponding lip shape image. The generative artificial intelligence model has learned lip shape images corresponding to the audio voice. For example, if the audio voice is "Hello," a lip shape image corresponding to "Hello" is learned, so a lip shape image matching "Hello" is taken and combined with the audio signal. If there is no learned word, a lip shape image with the highest similarity to the pronunciation of the word is extracted and combined with the audio signal.

[0052] The variational autoencoder (700) is a part that generates motion coefficients of the head during speech, and the three-dimensional face shape S in the 3D Morphable Model (3DMM) can be separated as shown in Equation 3 below.

[0053]

[0054] Here, S is the average shape of the 3D face and U id Wow U exp is a normal orthogonal basis for the identification and representation of the LSFM morphable model. Motion parameters are modeled as {β, r, t}, and after learning the head pose ρ=[r, t] and representation coefficient β individually from the guide audio, the motion coefficients are used to implicitly modulate the face rendering for final video synthesis.

[0055] VAE-based models can be designed to learn realistic and stylized head movements from actual speaking videos.

[0056] Finally, the speech video generation unit (800) generates video frames from facial expression coefficients and head motion coefficients generated through 3D keypoint extraction and warping field generation. Moving images can be smoothly connected through 3D keypoint extraction.

[0057] Figure 4 illustrates an example of a 3D character's speaking face using generative artificial intelligence, where (a) shows a character image created from a video of a user's face and (b) shows a 3D speaking face image created after cropping the face portion of the character.

[0058] Although the present invention has been described in relation to the preferred embodiments mentioned above, various modifications and variations are possible without departing from the essence and scope of the invention. Accordingly, the appended claims will include such modifications and variations that fall within the essence of the invention.

Claims

Claim 1 A method for generating a speech face image of a 3D character using a generative artificial intelligence model, wherein the generative artificial intelligence model determines a face change feature value in which the feature points of an actual user face and a character face transformed from the actual user face are recognized identically; a step of generating a 2D character face image from a single user face image based on the determined face change feature value; a step of estimating a 3D face structure from the 2D character face image; and a step of mapping a user voice to determine a facial expression and head movement corresponding to the user voice and generating a speech face image of a 3D character, wherein the face change feature value includes any one of a CFG scale value, a noise removal intensity value, or a combination thereof. Claim 2 A method for generating a speech face image of a 3D character, wherein, in claim 1, the step of determining a face change feature value in which a generative artificial intelligence model recognizes the feature points of an actual user face and a character face transformed from the actual user face as identical is characterized by changing the face change feature value until the generative artificial intelligence model recognizes the face position in a 2D character face image. Claim 3 delete Claim 4 A method for generating a speech face image of a 3D character according to claim 1, wherein the step of mapping the user voice to determine facial expressions and head movements corresponding to the user voice and generating a speech face image of a 3D character comprises: a step of generating facial expression coefficients including mouth shape through an audio encoder and a mapping process; a step of generating head motion coefficients during speech through a variational autoencoder; and a step of generating video frames from the facial expression coefficients and head motion coefficients generated through 3D keypoint extraction and warping field generation.