Image generation system and method based on facial feature
Patent Information
- Authority / Receiving Office
- TW · TW
- Patent Type
- Applications
- Current Assignee / Owner
- CHUNGHWA TELECOM CO LTD
- Filing Date
- 2025-01-16
- Publication Date
- 2026-08-01
AI Technical Summary
Existing AI technologies for image generation lack flexibility and require specific formats for facial keypoint coordinates, limiting their applicability and efficiency in generating images based on facial features.
An image generation system and method that includes a facial keypoint driving module, a facial keypoint coordinate transformation module, and an audio-visual concatenation module, which convert facial keypoint coordinates into a format suitable for a facial image rendering model, enabling the generation of rendered facial images and final images using driving audio.
Enhances the flexibility and efficiency of image generation by allowing the system to handle various facial landmark coordinate formats and generate high-quality images using audio-visual inputs.
Smart Images

Figure TWG2TA001069514_001 
Figure TWG2TA001069514_002 
Figure TWG2TA001069514_003
Abstract
Description
Image generation system and method based on facial features This invention relates to an image generation system and method based on facial features. Currently, AI technology is widely used. How to more effectively apply AI technology to the field of image generation is a goal that those skilled in the art should strive to achieve. The image generation system based on facial features of the present invention includes a facial keypoint driving module, a facial keypoint coordinate transformation module, a facial image rendering model, and an audio-visual concatenation module. The facial keypoint driving module uses a first set of facial keypoint coordinates and audio features to obtain a second set of facial keypoint coordinates, wherein the second set of facial keypoint coordinates corresponds to a time series; the facial keypoint coordinate transformation module uses the second set of facial keypoint coordinates to obtain a third set of facial keypoint coordinates, wherein the third set of facial keypoint coordinates includes missing facial keypoint coordinates; the facial image rendering model uses a facial image and the third set of facial keypoint coordinates to obtain a rendered facial image sequence; and the audio-visual concatenation module uses the rendered facial image sequence and driving audio to obtain the generated image. The image generation method based on facial features of the present invention is suitable for a system including a facial key point driving module, a facial key point coordinate transformation module, a facial image rendering model, and an audio-visual concatenation module. The method includes the following steps: the facial key point driving module uses a first set of facial key point coordinates and audio features to obtain a second set of facial key point coordinates, wherein the second set of facial key point coordinates corresponds to a time series; the facial key point coordinate transformation module uses the second set of facial key point coordinates to obtain a third set of facial key point coordinates, wherein the third set of facial key point coordinates includes missing facial key point coordinates; the facial image rendering model uses a facial image and the third set of facial key point coordinates to obtain a rendered facial image sequence; and the audio-visual concatenation module uses the rendered facial image sequence and driving audio to obtain a generated image. 100: Image generation system based on facial features LD: Facial Key Point Driving Module C trans Facial landmark coordinate conversion module R: Face image rendering model 40: Audio / Video Serialization Module M: Facial landmark coordinate detection module E: Audio Feature Extraction Module S21, S22, S23, S24, S25, S26: Steps Figure 1 is a schematic diagram of an image generation system based on facial features according to an embodiment of the present invention. Figure 2 is an example of the operation of the system shown in Figure 1. Figure 1 is a schematic diagram of an image generation system 100 based on facial features according to an embodiment of the present invention. The system 100 may include a facial landmark driving module (LD) and a facial landmark coordinate transformation module. C trans The system 100 includes a face image rendering model R and an audio-visual concatenation module 40. In other embodiments, the system 100 may further include a face key point coordinate detection module M. In other embodiments, the system 100 may further include an audio feature extraction module E. In one embodiment, a face key point driving module LD and a face key point coordinate transformation module are included. C trans The facial image rendering model R, the audio-visual concatenation module 40, the facial key point coordinate detection module M, and the audio feature extraction module E can be software and / or firmware code executed by a processor. In one embodiment, the system 100 can be used to generate a virtual anchor; however, the invention is not limited thereto. Figure 2 is an example of the operation of the system 100 shown in Figure 1. Please refer to both Figure 1 and Figure 2. In step S21, the facial landmark coordinate detection module M can use the facial image (i) to detect the first set of facial landmark coordinates (i) L M = M( i)). The facial landmark coordinate detection module M is, for example, MediaPipe, but the present invention is not limited thereto. Specifically, the first set of facial landmark coordinates... ,in n M The number of facial landmark coordinates detected for MediaPipe. In step S22, the audio feature extraction module E can extract audio features using the driving audio (a). In one embodiment, the audio feature extraction module E can be composed of a two-dimensional convolutional neural network. Specifically, the audio feature extraction module E can first convert the driving audio (a) to a Mel spectrum to obtain a. mel Then, the audio feature extraction module E can analyze a. mel Extracting embedded audio features, i.e., embedded audio feature a embedding = E(a mel ). In step S23, the face key point driving module LD can use the first face key point coordinate set and audio features to obtain the second face key point coordinate set (i.e., the second face key point coordinate set). L serial = LD( L M ,a mel The second set of facial key point coordinates corresponds to a time series. In one embodiment, the facial key point driving module LD can be a Transformer encoder architecture. Specifically, the second set of facial key point coordinates... L serial ={ L gen,1 , L gen,2 ,... L gen,t } ,t=1,2,..., T, where T represents the number of frames generated. Specifically, the set of second facial landmark coordinates. L serial The lip landmarks within the face can reflect the mouth shape corresponding to the speaking sound. It should be noted that the facial landmark-driven module (LD) can be trained using a dataset of speaking faces. The loss function used during training can be expressed as Equation 1 below. in, For the first The facial landmark coordinates generated in t-frames L M,t Then it is the first Ground-truth coordinates of facial landmarks generated in t-frames. In step S24, the facial landmark coordinate conversion module C trans The coordinates of the third facial landmark can be obtained using the coordinates of the second facial landmark. L all = C trans ( L serial The third set of facial landmark coordinates includes missing facial landmark coordinates. In one embodiment, a facial landmark coordinate conversion module... C trans It can be composed of convolutional neural networks and fully connected neural networks. Specifically, it includes a facial landmark coordinate transformation module. C trans The set of coordinates of key points of the second face can be calculated. L serial Other facial landmark coordinates. In particular, the third set of facial landmark coordinates. L all This may include the number of facial landmark coordinates required for the facial image rendering model R. In other words, after performing step S24, the third set of facial landmark coordinates... L all This will conform to the input format of the face image rendering model R. It should be noted that, due to the face key point coordinate transformation module of this invention... C trans The second set of facial landmark coordinates can be used L serial Converted to a third set of facial landmark coordinates (conforming to the input format of the face image rendering model R). L all This invention does not require a specific set of second facial landmark coordinates. L serial The present invention uses a specific second facial landmark coordinate format and a corresponding facial image rendering model R. In other words, the present invention can improve the versatility between the facial landmark driving module LD and the facial image rendering model R. Furthermore, it should be noted that, in order to enable the facial landmark coordinate conversion module C trans Capable of generating missing facial landmark coordinates, facial landmark coordinate conversion module C trans It can be trained using facial datasets to meet various facial landmark coordinate calculation requirements. The facial landmark coordinate transformation module will be explained below. C trans The training method. First, system 100 can process facial landmark datasets. Specifically, since the number and location of facial landmarks marked in each facial dataset may differ, system 100 can first train different facial landmark prediction neural network models M_i (i=1~n) for each facial landmark dataset D_i (i=1~n). Next, system 100 can aggregate the facial images from each facial landmark dataset D_i to obtain a set of facial images D_all. Then, system 100 can extract each image from the set of facial images D_all and predict facial landmarks using each facial landmark prediction neural network model M_i. At this point, system 100 can obtain the landmarks predicted by different facial landmark prediction neural network models M_i for a given face. Then, system 100 can overlay these landmarks onto a single facial image to obtain the landmarks predicted by different facial landmark models for that face image. Next, the system 100 can repeat the above process until all face images are labeled with the key points predicted by different facial key point prediction neural network models M_i, in order to obtain the set of facial key points D_all_key. After processing the facial landmark dataset, the system can train a facial landmark coordinate transformation module. C trans The loss function used during training can be expressed as Equation 2 below. in, p i Facial landmark coordinate conversion module C trans Predicted facial landmark coordinates These are the actual coordinates of facial landmarks. N is the number of facial landmark coordinates. In detail, system 100 can extract a face image from the face image set D_all and predict facial landmarks using a facial landmark prediction neural network model M_i. Then, system 100 can input these facial landmarks into a facial landmark coordinate transformation module. C trans With facial landmark coordinate conversion module C trans Predict the remaining facial landmarks p i The prediction result can be obtained from the answers in the facial landmark set D_all_key. as well as Loss 2. Update the facial landmark coordinate conversion module. C trans The model weights. Then, system 100 can extract another facial landmark prediction neural network model from the facial landmark prediction neural network model M_i and repeat the above steps until the facial landmark coordinate transformation module... C trans The goal is to generate complete facial landmarks when different missing facial landmarks are received. Please refer to Figure 2. In step S25, the face image rendering model R can use the face image (i) and the third set of face key point coordinates to obtain the rendered face image sequence (i.e., the rendered face image sequence). I render = R( I,L all )={ i render,1 , i render , 2,... i render,T In one embodiment, the face image rendering model R may consist of convolutional neural networks and fully connected neural networks. In other words, the face image rendering model R can determine the coordinates of a third set of facial key points based on the features of the face image. L all Rendering is performed to generate a sequence of rendered face images. I render . In step S26, the audio-visual concatenation module 40 can utilize the rendered face image sequence ( I render (a) and driving audio (a) to obtain the generated image. In other words, the audio-visual concatenation module 40 can generate a sequence of rendered face images. I render The signals are connected in series and the drive audio (a) is closed to obtain the final generated image. The present invention also provides an image generation method based on facial features, wherein the method can be implemented by the system 100 shown in FIG1. The method includes the following steps: (a) The facial landmark driving module uses a first set of facial landmark coordinates and audio features to obtain a second set of facial landmark coordinates, wherein the second set of facial landmark coordinates corresponds to a time series. (b) The facial landmark coordinate conversion module uses the second set of facial landmark coordinates to obtain the third set of facial landmark coordinates, wherein the third set of facial landmark coordinates includes the missing facial landmark coordinates. (c) The face image rendering model uses face images and a third set of face key point coordinates to obtain a sequence of rendered face images. (d) The audio-visual concatenation module uses the rendered face image sequence and driving audio to obtain the generated video. In summary, the image generation system and method based on facial features of the present invention, after obtaining a second set of facial key point coordinates, converts the second set of facial key point coordinates into a third set of facial key point coordinates (conforming to the input format of a facial image rendering model) using a facial key point coordinate conversion module. Then, a rendered facial image sequence can be obtained, and the generated image can be obtained using the rendered facial image sequence and driving audio. In particular, the present invention does not require the use of a corresponding facial image rendering model based on a specific format of the second facial key point coordinates (in a specific second set of facial key point coordinates). Therefore, the flexibility of image generation can be improved. S21, S22, S23, S24, S25, S26: Steps
Claims
1. An image generation system based on facial features, comprising: Facial landmark driver module; Facial landmark coordinate conversion module; Human face image rendering model; The system includes an audio-visual concatenation module, wherein the facial key point driving module uses a first set of facial key point coordinates and audio features to obtain a second set of facial key point coordinates, wherein the second set of facial key point coordinates corresponds to a time series; the facial key point coordinate conversion module uses the second set of facial key point coordinates to obtain a third set of facial key point coordinates, wherein the third set of facial key point coordinates includes missing facial key point coordinates; the facial image rendering model uses a facial image and the third set of facial key point coordinates to obtain a rendered facial image sequence; and the audio-visual concatenation module uses the rendered facial image sequence and driving audio to obtain a generated image. The system as described in claim 1 further includes a facial landmark coordinate detection module, wherein the facial landmark coordinate detection module uses the facial image to detect the first set of facial landmark coordinates. The system as described in claim 1 further includes an audio feature extraction module, wherein the audio feature extraction module uses the driving audio to extract the audio features. The system as described in Request 1, wherein the facial keypoint driving module is a Transformer encoder architecture. The system as described in claim 1, wherein the facial landmark coordinate transformation module is composed of a convolutional neural network and a fully connected layer neural network. The system as described in claim 1, wherein the face image rendering model is composed of a convolutional neural network and a fully connected neural network. A method for image generation based on facial features is suitable for a system including a facial keypoint driving module, a facial keypoint coordinate transformation module, a facial image rendering model, and an audio-visual concatenation module. The method includes the following steps: the facial keypoint driving module uses a first set of facial keypoint coordinates and audio features to obtain a second set of facial keypoint coordinates, wherein the second set of facial keypoint coordinates corresponds to a time series; the facial keypoint coordinate transformation module uses the second set of facial keypoint coordinates to obtain a third set of facial keypoint coordinates, wherein the third set of facial keypoint coordinates includes missing facial keypoint coordinates; the facial image rendering model uses a facial image and the third set of facial keypoint coordinates to obtain a rendered facial image sequence; and the audio-visual concatenation module uses the rendered facial image sequence and driving audio to obtain a generated image.