Apparatus and method for generating speech video

The device and method address the challenge of synchronizing non-speech motion with speech audio by encoding a person's image and recognizing emotions to generate natural speech videos with synchronized lip movements and facial expressions.

WO2026116633A1PCT designated stage Publication Date: 2026-06-04DEEPBRAIN AI INC

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DEEPBRAIN AI INC
Filing Date
2025-06-05
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

The challenge in generating speech videos is the weak correlation between spoken speech and non-speech motion, making it difficult for AI to accurately infer and synchronize facial expressions and head movements with the audio.

Method used

A device and method using an encoder to encode a person's image into an identity latent vector, a speech emotion recognition model to recognize emotions, a motion latent vector generation model to generate motion latent vectors based on speech audio and emotions, and a decoder to produce a speech video with synchronized lip movements and natural non-speech motions.

Benefits of technology

Generates speech videos where the person speaks naturally with synchronized lip movements and appropriate facial expressions, enhancing the naturalness of the speech video generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025007763_04062026_PF_FP_ABST
    Figure KR2025007763_04062026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an apparatus and method for providing a speech video. The apparatus for generating a speech video according to an embodiment comprises: an encoder for encoding a person image to generate an identity latent vector; a voice emotion recognition model for recognizing emotion from speech audio; a motion latent vector generation model which is a probabilistic generative model for generating a motion latent vector on the basis of the speech audio and the emotion; and a decoder for generating a speech video of speech spoken by the corresponding person on the basis of the identity latent vector and the motion latent vector.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for generating speech video

[0001] It is related to AI-based speech video generation technology.

[0002] With the advancement of artificial intelligence technology, the technology to generate a speech video in which the subject of a given portrait speaks naturally in sync with the audio, using just a single portrait photo and spoken audio, is also making a significant leap forward.

[0003] Key factors determining the naturalness of speech video generation include the generation of lip movements synchronized with the speech audio and the generation of non-speech motion (facial expressions, head movements, etc.) that matches the speech audio. While the former (synchronized lip movements) has recently made significant advancements in terms of accuracy and variety, the latter (generation of non-speech motion that matches the speech audio) remains a challenging problem because the inherently weak correlation between spoken speech and non-speech motion makes it difficult for AI to directly infer this relationship.

[0004] The purpose is to provide a device and method for generating speech videos based on artificial intelligence.

[0005] A speech video generation device according to one aspect includes: an encoder that encodes a person image to generate an identity latent vector; a speech emotion recognition model that recognizes emotions from speech audio; a motion latent vector generation model which is a probabilistic generation model that generates a motion latent vector based on the speech audio and the emotions; and a decoder that generates a speech video in which the person speaks based on the identity latent vector and the motion latent vector.

[0006] The motion latent vector generation model may include: a score function estimation model that estimates a score function used to generate the motion latent vector from a noise latent vector based on the speech audio and the emotion; a scale setting unit that sets a scale of the speech audio and a scale of the emotion, which indicate the degree to which the motion latent vector is considered when generating the motion latent vector, to the estimated score function; and a motion latent vector generation unit that generates a motion latent vector corresponding to the speech audio and the emotion based on the score function with the set scale.

[0007] The scale of the above-mentioned speech audio and the scale of the above-mentioned emotion can be changed according to user input.

[0008] The score function with the above scale set can be expressed by the following mathematical formula.

[0009] [Mathematical Formula]

[0010]

[0011] (s * is a scaled score function, t is a time point t∈[0, 1], k(t) is a noise-added motion latent vector at time point t, and A s is speech audio, and E s is an emotion, and θ * is the optimized parameter of the score function estimation model, and is the scale of the spoken audio, and is the scale of emotion, and p is the probability density function)

[0012] The above motion potential vector generation unit can generate the motion potential vector using the following mathematical formula.

[0013] [Mathematical Formula]

[0014]

[0015] (f is the drift coefficient describing fixed motion, and g is the diffusion coefficient determining the degree of Brownian motion, is Brownian motion following the counter-time direction)

[0016] The above encoder is an encoder of a motion autoencoder, and the above decoder may be a decoder of the above motion autoencoder.

[0017] The motion autoencoder above can be pre-trained by randomly extracting two frames from a training speech video and minimizing the error loss between the frame generated using the first frame among the two frames as input and the second frame among the two frame images.

[0018] The above-described speech video generating device further includes a combining unit that generates a combined potential vector by combining the identity potential vector and the motion potential vector, and the decoder can generate the speech video by decoding the combined potential vector.

[0019] The motion potential vector may include information about a mouth shape synchronized with the speech audio and information about non-speech motions including motions according to the emotion.

[0020] A method for generating a speech video performed by a computing device according to a different aspect comprises: a step of encoding a person image to generate an identity latent vector; a step of recognizing an emotion from speech audio; a step of generating a motion latent vector based on the speech audio and the emotion using a motion latent vector generation model, which is a probabilistic generation model; and a step of generating a speech video in which the person speaks based on the identity latent vector and the motion latent vector.

[0021] According to an exemplary embodiment, by using a motion latent vector generation model, which is a score-based probabilistic generation model, to generate motion latent vectors corresponding to speech audio and emotions recognized from the speech audio and using them for speech video generation, a speech video can be generated in which the person speaks the content of the speech audio with natural motion.

[0022] According to an exemplary embodiment, by adjusting the scale of speech audio and emotion when generating motion latent vectors, speech videos containing various emotions can be generated.

[0023] FIG. 1 is a drawing illustrating a speech video generating device according to an exemplary embodiment.

[0024] FIG. 2 is a diagram illustrating a motion latent vector generation model according to an exemplary embodiment.

[0025] FIG. 3 is a diagram illustrating a learning method of a motion autoencoder according to an exemplary embodiment.

[0026] FIG. 4 is a flowchart illustrating a method for generating a speech video according to an exemplary embodiment.

[0027] FIG. 5 is a block diagram illustrating a computing environment including a computing device according to an exemplary embodiment.

[0028] Hereinafter, an embodiment of the present disclosure will be described in detail with reference to the attached drawings. It should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description will be omitted.

[0029] The terms described below are defined in consideration of their functions within this disclosure, and these may vary depending on the intent or practices of the user or operator. Therefore, their definitions should be based on the content throughout this disclosure.

[0030] Terms such as "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. Terms are used solely for the purpose of distinguishing one component from another. A singular expression includes a plural expression unless the context clearly indicates otherwise, and terms such as "include" or "have" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0031] Furthermore, the classification of components in this disclosure is merely based on the primary function performed by each component. That is, two or more components may be combined into a single component, or a single component may be divided into two or more components based on more subdivided functions. Additionally, each component may perform some or all of the functions performed by other components in addition to its primary function, and some of the primary functions performed by each component may be exclusively performed by other components. Each component may be implemented in hardware or software, or as a combination of hardware and software.

[0032] FIG. 1 is a drawing illustrating a speech video generating device according to an exemplary embodiment, and FIG. 2 is a drawing illustrating a motion latent vector generating model according to an exemplary embodiment.

[0033] A speech video generation device (100) according to an exemplary embodiment may be an audio-driven speech video generation device based on a probabilistic generative model. The speech video generation device (100) may convert a person image of a specific person into an identity latent vector through a pre-trained motion autoencoder based on an artificial neural network (which may be referred to as a motion latent vector autoencoder or a motion latent autoencoder). The speech video generation device (100) may recognize an emotion inherent in a given speech audio (or speech voice) and probabilistically generate motion latent vectors containing information about mouth shapes synchronized with the speech audio and information about non-speech motions including motions according to the emotion inherent in the speech audio, using a motion latent vector generation model which is a probabilistic generative model. The speech video generation device (100) may generate a speech video by converting the generated motion latent vectors and identity latent vectors into frame images through a motion autoencoder. Here, the speech video can be referred to as a talking portrait video.

[0034] Referring to FIG. 1, a speech video generating device (100) according to an exemplary embodiment may include an encoder (110), a speech emotion recognition model (120), a motion latent vector generation model (130), a combination unit (140), and a decoder (150). Here, the encoder (110) and the decoder (150) may constitute a motion autoencoder. For example, the encoder (110) may be an encoder that constitutes a motion autoencoder, and the decoder (150) may be a decoder that constitutes a motion autoencoder.

[0035] According to one embodiment, the motion autoencoder, that is, the encoder (110) and the decoder (150), can be pre-trained. The training method of the encoder (110) and the decoder (150) will be described later with reference to FIG. 3.

[0036] The encoder (110) is a person image (or portrait image) (I s Encode the person image (I s The latent vector corresponding to ) (w s ) can be generated. According to one embodiment, a latent vector (w s ) is the identity latent vector (w i ) and motion latent vector(w m It may include ). Here, the identity latent vector (w i ) includes features related to the identity of a specified person, and the motion latent vector (w m ) may include features related to the motion of a specific person. The encoder (110) encodes a person image of a specific person and an identity latent vector (w) corresponding to the person image. i ) and motion latent vector(w m Can generate ).

[0037] Motion potential vector (w) generated from the encoder (110) m Identity latent vector (w excluding ) i ) can be used when creating a video of the person's speech.

[0038] The voice emotion recognition model (120) is a speech audio (or speech voice) (A s From ) emotion(E s It can recognize ). For example, the speech emotion recognition model (120) may be an artificial intelligence model pre-trained to recognize emotions inherent in the speech audio from the speech audio. For example, the speech emotion recognition model (120) may extract probability distribution values ​​of emotions such as anger, disgust, fear, happiness, neutrality, sadness, and joy from the speech audio.

[0039] The motion latent vector generation model (130) is speech audio (A s ) and emotion(E s Based on ), motion latent vector(w g_m) can be generated probabilistically. The motion latent vector (w generated at this time) g_m ) is speech audio (A s Information about the mouth shape synchronized with ), and emotion (E s Information regarding non-speaking motion including motion according to ) may be included.

[0040] According to an exemplary embodiment, the motion potential vector generation model (130) may be a probabilistic generation model, such as a score-based probabilistic generation model.

[0041] Referring to FIG. 2, the motion latent vector generation model (130) may include a score function estimation model (210), a scale setting unit (220), and a motion latent vector generation unit (230).

[0042] The score function estimation model (210) is speech audio (A s ) and speech audio (A s Emotions perceived from ) (E s The score function can be estimated based on ). In this case, the score function is derived from the noisy latent vector to the motion latent vector (w g_m It is used to generate a noise latent vector, and the noise latent vector can be extracted from an arbitrary distribution, for example, a multidimensional standard normal distribution. To this end, the score function estimation model (220) can be trained to estimate a score function that can generate a motion latent vector similar to the learning motion latent vector by using the learning speech audio, learning emotion data, and learning motion latent vector, by progressively removing noise from the noise latent vector. At this time, the learning motion latent vector can be derived from a frame (or frame image) of the learning speech video corresponding to the learning speech audio. In addition, the learning emotion data is data related to the emotion inherent in the learning speech audio, and may be emotion data recognized from the learning speech audio through a pre-trained speech emotion recognition model.

[0043]

[0044] Training of the score function estimation model

[0045] A score function estimation model (210) can be trained by gradually adding noise to a learning motion potential vector to convert it into a noise potential vector, and then gradually removing noise from the noise potential vector based on learning speech audio and learning emotion data to estimate a score function that can generate a motion potential vector similar to the learning motion potential vector.

[0046] For example, according to one embodiment, the score function estimated in the score function estimation model (210) during the learning process can be expressed by Equation 1.

[0047]

[0048] Here, t is a time point t∈[0, 1], x(t) is a noise-added learning motion latent vector at time point t created by adding noise to a learning motion latent vector (x(0)), y1 is a learning speech audio, y2 is a learning emotion data, θ is a learning parameter, and p may be a probability density function.

[0049] The score function estimation model (210) can be trained to estimate a score function based on training speech audio, training emotion data, and training motion latent vectors, and then to minimize the error loss between the estimated score function and the target score function. For example, the optimized parameters of the score function estimation model can be expressed by Equation 2.

[0050]

[0051] Here, θ * is the optimized parameter of the score function estimation model, and x(0) is the motion latent vector for training, and can be a target score function. argmin can represent a function that finds θ that minimizes the error loss between the estimated score function and the target score function.

[0052]

[0053] Application of the trained score function estimation model

[0054] The score function estimation model (210) is speech audio (A s ) and speech audio (A s Emotions perceived from ) (E s A score function can be estimated based on ). A score function estimation model (210) can estimate the speech audio (A s ) and speech audio (A s Emotions perceived from ) (E s The score function estimated based on ) can be expressed by mathematical equation 3.

[0055]

[0056] Here, k(t) is the noise-added motion potential vector at time t, and k(T) can be the noise potential vector.

[0057] The scale setting unit (220) can set the scale (or weight) of speech audio and the scale (or weight) of emotion data, which represent the degree to which they are considered when generating motion potential vectors, in the estimated score function. The score function with the scale (or weight) of speech audio and the scale (or weight) of emotion data set can be expressed by Equation 4.

[0058]

[0059] Here, and may be the scale (or weight) of speech audio and the scale (or weight) of emotion data, respectively, representing the degree to which they are considered when generating motion latent vectors. and It can be changed in various ways depending on user input, and the user and By adjusting it, the degree of reflection of speech audio and emotion data when generating motion latent vectors can be controlled.

[0060] The motion potential vector generation unit (230) is speech audio (A s ) scale (or weight) and sentiment data (E s Based on a score function with a set scale (or weight) of ), speech audio (A s ) and sentiment data(E s Motion latent vector (w) corresponding to ) g_m Can generate ).

[0061] According to one embodiment, the motion potential vector generating unit (230) uses Equation 5 to generate a motion potential vector (w g_m Can generate ).

[0062]

[0063] Here, f is the drift coefficient describing fixed motion, and g is the diffusion coefficient determining the degree of Brownian motion, and It can be Brownian motion following the reverse time direction.

[0064] For example, the motion potential vector generation unit (230) solves Equation 5 to produce speech audio (A s ) and sentiment data(E s Motion latent vector (w) corresponding to ) g_m k(0) can be generated as ). The motion potential vector generation unit (230) generates speech audio (A s Multiple motion potential vectors (w) corresponding to the length of ) g_m Can generate ).

[0065] The combination unit (140) is an identity latent vector (w) generated by the encoder (110). i ) and the motion potential vector (w) generated in the motion potential vector generation unit (230). g_mCombining ) to create a combination latent vector (w g_s ) can be generated. For example, the combination part (140) can generate an identity latent vector (w i ) and each motion potential vector(w g_m By combining ) speech audio (A s Multiple combination latent vectors (w) corresponding to the length of ) g_s Can generate ).

[0066] The decoder (150) is a combination potential vector (w g_s Decoding ) to produce a speech video (V) in which a specific person speaks g ) can be generated. For example, the decoder (150) can generate each combination potential vector (w g_s Decoding ) to obtain the utterance image (I) corresponding to each combinational latent vector g Generate ) and each generated utterance image (I g Connect the utterance video (V g Can generate ).

[0067] According to an exemplary embodiment, the speech video generating device (100) can generate an identity latent vector containing information about the identity of a person from a person image, and generate a motion latent vector corresponding to speech audio and emotion data from speech audio and emotion data recognized from speech audio using a score-based motion latent vector generating model. At this time, the user can adjust the scale of the speech audio and emotion data, and accordingly, when generating the motion latent vector, the speech video generating device (100) can generate the motion latent vector by considering the speech audio and emotion data according to each scale set by the user. The speech video generating device (100) can generate a speech video in which the person speaks the content of the speech audio in natural motion using the identity latent vector generated from the person image and the motion latent vector generated from the speech audio and emotion data using a score-based motion latent vector generating model.

[0068] FIG. 3 is a diagram illustrating a learning method of a motion autoencoder according to an exemplary embodiment.

[0069] As described above, the motion autoencoder may include an encoder (110) and a decoder (150). The motion autoencoder may convert an image into a latent vector corresponding to the image. As described above, the latent vector may be expressed as the sum of an identity vector and a motion vector.

[0070] The reason a motion autoencoder can encode in the form of a latent vector, which is the sum of an identity vector and a motion vector, may be due to the learning method of the motion autoencoder. According to one embodiment, the motion autoencoder can be trained as follows.

[0071] First, two frames (videos) are randomly extracted from a training speech video of a specific person. For example, two frames can be randomly extracted from a training speech video of a specific person as a source frame and a driving frame.

[0072] Next, one of the two extracted frames (x s A frame (x) generated by inputting )(hereinafter, source frame) into a motion autoencoder g )(hereinafter, generated frame) and the other of the two frames (x d A motion autoencoder is trained so that the error loss of )(hereinafter, driving frame) is minimized.

[0073] For example, the encoder (110) receives an input source frame (x s Encoding the source frame (x s Generates a latent vector (w) of ), and the decoder (150) generates a source frame (x s Decoding the latent vector (w) of ) to generate the frame (x g ) can be generated. The encoder (110) and decoder (150) generate a frame (x g ) and driving frame(x d It can be trained so that the error loss of ) is minimized. This can be expressed mathematically as Equation 6.

[0074]

[0075] Here may be an error loss function. According to one embodiment, the error loss function may include an L1 loss function, a perceptual loss function, etc., but is not limited thereto.

[0076] According to one embodiment, a source frame (x s ) and driving frame(x dSince ) is extracted from the same speech video and used for training, the motion autoencoder can be trained to include all motion (e.g., mouth shape, facial expression, head movement, etc.) in latent vectors to generate a driving frame from the source frame.

[0077] FIG. 4 is a flowchart illustrating a method for generating a speech video according to an exemplary embodiment. The speech video generation method of FIG. 4 can be performed by the speech video generation device (100) of FIG. 1. Although the method is described in the illustrated flowchart as being divided into a plurality of steps, at least some of the steps may be performed in a different order, combined with other steps and performed together, omitted, divided into detailed steps, or performed with one or more steps not illustrated added.

[0078] Referring to FIG. 4, in step 410, the speech video generating device can generate an identity latent vector corresponding to the person image by encoding a person image of a predetermined person. Here, the identity latent vector may include features related to the identity of the predetermined person. For example, the speech video generating device can generate an identity latent vector corresponding to the person image by using an encoder of a pre-trained motion autoencoder.

[0079] In step 420, the speech video generating device can recognize emotions inherent in the speech audio from the speech audio. For example, the speech video generating device can use a pre-trained artificial intelligence model to extract probability distribution values ​​of emotions such as anger, disgust, fear, happiness, neutrality, sadness, and joy from the speech audio.

[0080] In step 430, the speech video generating device can probabilistically generate motion latent vectors based on speech audio and emotions recognized from the speech audio. The generated motion latent vectors may include information about mouth shapes synchronized with the speech audio and information about non-speech motions including motions according to emotions.

[0081] According to an exemplary embodiment, a speech video generating device can generate motion latent vectors using a probabilistic generating model, such as a score-based probabilistic generating model.

[0082] For example, a speech video generation device can estimate a score function used to generate a motion latent vector from a noise latent vector based on speech audio and an emotion recognized from speech audio by using a score function estimation model, which is a score-based probabilistic generation model. Additionally, the speech video generation device can generate a motion latent vector corresponding to the speech audio and the emotion recognized from speech audio by setting a scale for each of the speech audio and emotion data in the estimated score function and solving Equation 5, which consists of the scaled score function.

[0083] In step 440, the speech video generating device can generate a combination potential vector by combining the identity potential vector generated in step 410 and the motion potential vector generated in step 430.

[0084] In step 450, the combination latent vector generated in step 440 can be decoded to generate a speech video in which a specific person speaks. For example, a speech video generating device can generate a speech video using a decoder of a pre-trained motion autoencoder.

[0085] FIG. 5 is a block diagram illustrating a computing environment including a computing device according to an exemplary embodiment. In the illustrated embodiment, each component may have different functions and capabilities in addition to those described below, and may include additional components in addition to those described below.

[0086] The illustrated computing environment (10) includes a computing device (12). The computing device (12) may be one or more components included in a speech video generating device (100) according to one embodiment.

[0087] The computing device (12) includes at least one processor (14), a computer-readable storage medium (16), and a communication bus (18). The processor (14) can cause the computing device (12) to operate according to the exemplary embodiment described above. For example, the processor (14) can execute one or more programs stored in the computer-readable storage medium (16). The one or more programs may include one or more computer-executable instructions, and the computer-executable instructions may be configured to cause the computing device (12) to perform operations according to the exemplary embodiment when executed by the processor (14).

[0088] A computer-readable storage medium (16) is configured to store computer-executable instructions or program code, program data and / or other suitable forms of information. A program (20) stored in the computer-readable storage medium (16) includes a set of instructions executable by a processor (14). In one embodiment, the computer-readable storage medium (16) may be memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, other forms of storage media that are accessed by a computing device (12) and capable of storing desired information, or a suitable combination thereof.

[0089] The communication bus (18) interconnects various other components of the computing device (12), including the processor (14) and the computer-readable storage medium (16).

[0090] The computing device (12) may also include one or more input / output interfaces (22) and one or more network communication interfaces (26) that provide interfaces for one or more input / output devices (24). The input / output interfaces (22) and network communication interfaces (26) are connected to a communication bus (18). The input / output devices (24) may be connected to other components of the computing device (12) through the input / output interfaces (22). An exemplary input / output device (24) may include an input device such as a pointing device (such as a mouse or trackpad), a keyboard, a touch input device (such as a touchpad or touchscreen), a voice or sound input device, various types of sensor devices and / or imaging devices, and / or an output device such as a display device, a printer, a speaker and / or a network card. An exemplary input / output device (24) may be included inside the computing device (12) as a component constituting the computing device (12), or it may be connected to the computing device (12) as a separate device distinct from the computing device (12).

[0091] The disclosed embodiments may be implemented in the form of a recording medium that stores instructions executable by a computer. The instructions may be stored in the form of program code and, when executed by a processor, may generate a program module to perform the operation of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0092] The present invention has been described above with reference to its preferred embodiments. Those skilled in the art will understand that the present disclosure may be implemented in modified forms without departing from the essential characteristics of the present disclosure. Accordingly, the scope of the present disclosure should not be limited to the aforementioned embodiments but should be interpreted to include various embodiments within the scope equivalent to those described in the claims.

Claims

1. An encoder that encodes a person image to generate an identity latent vector; Speech emotion recognition model that recognizes emotions from speech audio; A motion latent vector generation model, which is a probabilistic generation model that generates a motion latent vector based on the above-mentioned speech audio and the above-mentioned emotion; and A speech video generating device comprising a decoder that generates a speech video in which a person speaks based on the above-mentioned identity latent vector and the above-mentioned motion latent vector.

2. In Claim 1, The above motion latent vector generation model is, A score function estimation model that estimates a score function used to generate the motion latent vector from a noise latent vector based on the speech audio and the emotion; A scale setting unit that sets the scale of the speech audio and the scale of the emotion, which indicate the degree to which the motion potential vector is considered when generating the above motion potential vector, to the above estimated score function; and A speech video generating device comprising a motion latent vector generating unit that generates a motion latent vector corresponding to the speech audio and the emotion based on a score function with the scale set above.

3. In Claim 2, A speech video generating device in which the scale of the speech audio and the scale of the emotion are changed according to user input.

4. In Claim 2, A speech video generating device in which the score function with the above scale is expressed by the following mathematical formula. [Mathematical Formula] (s * is a scaled score function, t is a time point t∈[0, 1], k(t) is a noise-added motion latent vector at time point t, and A s is speech audio, and E s is an emotion, and θ * is the optimized parameter of the score function estimation model, and is the scale of the spoken audio, and is the scale of emotion, and p is the probability density function) 5. In Claim 4, A speech video generating device in which the motion latent vector generating unit generates the motion latent vector using the following mathematical formula. [Mathematical Formula] (f is the drift coefficient describing fixed motion, and g is the diffusion coefficient determining the degree of Brownian motion, is Brownian motion following the counter-time direction) 6. In Claim 1, A speech video generating device in which the encoder is an encoder of a motion autoencoder and the decoder is a decoder of the motion autoencoder.

7. In Claim 6, The motion autoencoder above randomly extracts two frames from a training speech video and is a speech video generation device pre-trained such that the error loss between a frame generated using a first frame among the two frames as input and a second frame among the two frame images is minimized.

8. In Claim 1, It further includes a combination unit that generates a combination potential vector by combining the above-mentioned identity potential vector and the above-mentioned motion potential vector, and A speech video generating device in which the above decoder decodes the above combination potential vector to generate the speech video.

9. In Claim 1, A speech video generating device, wherein the motion latent vector includes information about a mouth shape synchronized with the speech audio and information about a non-speech motion including motion according to the emotion.

10. A method for generating a speech video performed by a computing device, A step of encoding a person image to generate an identity latent vector; Step of recognizing emotions from spoken audio; A step of generating a motion latent vector based on the utterance audio and the emotion using a motion latent vector generation model, which is a probabilistic generation model; and A method for generating a speech video, comprising the step of generating a speech video in which a person speaks based on the above-mentioned identity latent vector and the above-mentioned motion latent vector.