3D digital human facial expression synthesis method and visualization system based on audio driving
By constructing joint loss function and multimodal feature optimization, using the KAN model and graph convolutional neural network encoder, the problems of incoherence and innocence of 3D digital human facial expressions are solved, and a more vivid expression generation is achieved.
Patent Information
- Application Number
- CN202510267116.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art generates 3D digital facial expressions inconsistent and unrealistic movements, making it difficult to accurately restore the expressions of real characters.
By constructing a joint loss function that integrates audio features, 3D facial expression parameters and lip features, the KAN model and graph convolutional neural network encoder are used to realize the coordinated optimization and model training of multimodal features to generate accurate 3D digital human facial expression movements.
It improves the consistency of the sound and lips and the restoration of expressions, making the generated 3D digital facial expressions more vivid and close to the expressions of real characters.
Smart Images

Figure CN120355822A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly to a method for synthesizing 3D digital human facial expressions driven by audio, a visualization system, a device, a terminal device, and a computer-readable storage medium. Background Art
[0002] The present invention relates to the technical field of deep learning, and particularly to a method for synthesizing 3D digital human facial expressions driven by audio, a visualization system, a device, a terminal device, and a computer-readable storage medium. Summary of the Invention
[0003] In order to solve the problems of incoherence and unreality in generating 3D digital human facial expression actions in the prior art, the present invention provides a method for synthesizing 3D digital human facial expressions driven by audio, a visualization system, a device, a terminal device, and a computer-readable storage medium, which can make the generated 3D digital human facial expression actions more vivid and closer to real human expressions.
[0004] The first object of the present invention is to provide a method for synthesizing 3D digital human facial expressions driven by audio.
[0005] The second object of the present invention is to provide a visualization system for 3D digital human facial expressions driven by audio.
[0006] The third object of the present invention is to provide a device for synthesizing 3D digital human facial expressions driven by audio.
[0007] The fourth object of the present invention is to provide a terminal device.
[0008] The fifth object of the present invention is to provide a computer-readable storage medium.
[0009] The first object of the present invention can be achieved by adopting the following technical solutions: A method for synthesizing 3D digital human facial expressions driven by audio, the method comprising: Obtaining a data set; the data set includes English sentences recorded by multiple volunteers, and each English sentence corresponds to an original audio; Inputting the original audio in the data set into an audio encoder in the model to extract audio features; Inputting the audio features into a decoder based on KAN in the model to generate 3D digital human facial expression actions; Inputting the 3D digital human facial expression actions into a graph encoder in the model to extract lip features; By constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features, realizing the collaborative optimization of multi-modal features and model training; Input the original audio to be tested into the trained model, and successively pass through the audio encoder and the KAN decoder to generate corresponding 3D digital human facial expression actions.
[0010] Furthermore, each English sentence also corresponds to the real text and the 3D vertex coordinate sequence y; By constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features, the collaborative optimization of multi-modal features and model training are realized, including: According to the audio encoder and lip features, use the BPR loss to construct a cross-modal consistency loss function; In order to distinguish the importance differences of each region in the face animation generation task, construct a sub-region loss function according to the 3D vertex coordinate sequence y and the 3D digital human facial expression actions; To solve the problem of jitter output frames when only using the reconstruction loss, according to the 3D vertex coordinate sequence y at time t t and the 3D vertex coordinate sequence y at time t-1 t-1 and the difference between the 3D digital human facial expression actions at time t and the 3D digital human facial expression actions at time t-1 construct a speed loss function; To ensure that the generated lip movements are perfectly synchronized with the original audio, construct a text consistency loss function according to the real text and the predicted text; among them, the predicted text is obtained by inputting the lip shape corresponding to the lip features into the text decoder; Use the cross-modal consistency loss function, the sub-region loss function, the speed loss function, and the text consistency loss function and their corresponding weights to construct a total loss function; among them, the weights are obtained by parameter tuning using the chaos theory and the particle swarm optimization algorithm; Use the total loss function to train the model to optimize the parameters in the model and obtain the trained model.
[0011] Furthermore, the cross-modal consistency loss function, the sub-region loss function, the speed loss function, and the text consistency loss function and their corresponding weights to construct a total loss function are collectively referred to as four loss functions; The weights corresponding to the four loss functions are obtained through the following process: Use the chaos mapping to initialize the initial positions and velocities of multiple particles; and establish a search space, including the weight preset values of the four loss functions; design a fitting function according to the four loss functions and their corresponding weight preset values; Update the best position of a single particle and the global best position according to the fitness; If the velocity and position of the particles remain unchanged after multiple iterations, the model converges; when the model converges, the weights corresponding to the global best position are used as the optimization weights for the four loss functions.
[0012] Further, the text decoder is the wav2vec2-large-xlsr-53-english model with open-source parameters.
[0013] Further, according to the coordinates of each vertex in the patch structure of the 3D digital human facial expression action, the coordinate sequence of the lip vertices is extracted; the coordinates of all lip vertices are denoted as the feature representation F (0) ; The input of the 3D digital human facial expression action into the graph encoder of the model to extract lip features includes:
[0014] In the formula, F (l+1) is the feature representation output by the (l + 1)-th layer of the graph convolutional neural network, which fuses the structural and temporal information of the expression of the 3D model; F (l) is the feature representation output by the l-th layer of the graph convolutional neural network after F (0) ; W (l) is the weight of the l-th layer, is the adjacency matrix of the lip vertices in the patch structure of the 3D digital human facial expression action, and σ() is the activation function; The graph encoder has K layers of graph convolutional neural networks, and the feature representation output by the K-th layer is used as the lip feature; where K is a positive integer greater than or equal to 2.
[0015] Further, each volunteer has a corresponding three-dimensional facial template; The input of the audio feature into the KAN-based decoder in the model to generate the 3D digital human facial expression action includes: Input the audio feature into the KAN-based decoder, and output the offset sequence B of the three-dimensional coordinates of each vertex in the 3D model T ; Add the offset sequence B T to the three-dimensional facial template of the corresponding volunteer to generate the 3D digital human facial expression action.
[0016] Further, the audio encoder is the wav2vec model with frozen parameters.
[0017] The second object of the present invention can be achieved by adopting the following technical solutions: A 3D digital human facial expression visualization system based on audio driving, the system includes a front end and a back end; the front end includes a visualization display and rendering module, an audio upload module for interacting with users, and a rendering and debugging module; among them, the display and rendering module is used to render and display the generated 3D digital human facial expression actions; the audio upload module is used to provide an interface for users to upload the original audio file to drive the 3D model into 3D digital human facial expression actions; the rendering and debugging module is used to provide users with setting functions; the back end includes a deep learning algorithm module; The front and back ends use WebSocket and Request response connections: users input the original audio data through the audio upload module of the front end, and the system sends the audio to the back end in the form of a network request. The back end calls the deep learning module to execute the 3D digital human facial expression synthesis method described in any one of claims 1 to 7 to infer the 3D animation, and then responds and transmits it back to the display and rendering module for display; among them, the framework used by the front end is React and ThreeJS, and the framework used by the back end is FastAPI.
[0018] The third object of the present invention can be achieved by adopting the following technical solutions: A 3D digital human facial expression synthesis device, the device includes: A data acquisition unit for acquiring a data set; the data set includes English sentences recorded by multiple volunteers, and each English sentence corresponds to the original audio; A first extraction unit for inputting the original audio in the data set into the audio encoder in the model to extract audio features; A generation unit for inputting the audio features into the KAN-based decoder in the model to generate 3D digital human facial expression actions; A second extraction unit for inputting the 3D digital human facial expression actions into the graph encoder in the model to extract lip features; A model training unit for realizing the collaborative optimization and model training of multi-modal features by constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features; A 3D animation generation unit for inputting the original audio to be tested into the trained model, and successively passing through the audio encoder and the KAN decoder to generate corresponding 3D digital human facial expression actions.
[0019] The fourth object of the present invention can be achieved by adopting the following technical solutions: A terminal device includes a processor and a memory for storing the executable program of the processor. When the processor executes the program stored in the memory, it realizes the above-mentioned 3D digital human facial expression synthesis method based on audio driving.
[0020] The fifth object of the present invention can be achieved by adopting the following technical solutions: A computer-readable storage medium stores a program, which when executed by a processor, implements the above-mentioned 3D digital human facial expression synthesis method based on audio driving.
[0021] The present invention has the following beneficial effects compared with the prior art: The 3D digital human facial expression synthesis method, visualization system, device, terminal device and computer-readable storage medium based on audio driving provided by the present invention obtain a data set; the data set includes English sentences recorded by multiple volunteers, and each English sentence corresponds to an original audio; input the original audio in the data set into the audio encoder in the model to extract audio features; input the audio features into the KAN-based decoder in the model to generate 3D digital human facial expression actions; input the 3D digital human facial expression actions into the graph encoder in the model to extract lip features; by constructing a joint loss function that fuses audio features, 3D facial expression parameters and lip features, realize the collaborative optimization and model training of multi-modal features; input the original audio to be tested into the trained model, and successively pass through the audio encoder and the KAN decoder to generate corresponding 3D digital human facial expression actions. The present invention uses the KAN model encoder in the model to generate accurate three-dimensional facial expression actions; uses the encoder based on the graph convolutional neural network (GCN) in the model to obtain the relationship between vertices and their adjacent vertices to update the feature representation; further constrains the gap between the generated animation and the real animation by using cross-modal loss calculation of lip features and audio features, and realizes the collaborative optimization and model training of multi-modal features to narrow the gap between the generated animation and the real animation by constructing a joint loss function that fuses audio features, 3D facial expression parameters and lip features, thereby improving the audio-lip consistency and the reduction degree of expressions, making the generated 3D digital human facial expression actions more vivid and closer to real human expressions. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0023] Figure 1 It is a flowchart of the 3D digital human facial expression synthesis method based on audio driving according to Embodiment 1 of the present invention; Figure 2 It is a schematic diagram of the 3D digital human facial expression synthesis method based on audio driving according to Embodiment 1 of the present invention; Figure 3 Schematic diagram of the lip feature extraction process in Embodiment 1 of the present invention; Figure 4 Structural diagram of the 3D digital human facial expression visualization system based on audio driving in Embodiment 1 of the present invention; Figure 5 Structural schematic diagram of the 3D digital human facial expression action synthesis device based on audio driving in Embodiment 1 of the present invention; Figure 6 Structural block diagram of the terminal device in Embodiment 3 of the present invention. Detailed implementation manners
[0024] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. It should be understood that the specific embodiments described are only used to explain the present application and are not used to limit the present application.
[0025] Embodiment 1: As Figure 1 、 2 shown, this embodiment provides a 3D digital human facial expression synthesis method based on audio driving, including the following steps: S101. Obtain a data set and divide the data set into a training set and a test set.
[0026] Obtain the VOCASET data set from the Internet. The VOCASET data set includes English sentences recorded by 12 volunteers. Each sentence includes the original audio (in WAV format) and the corresponding real text S and the three-dimensional vertex coordinate sequence y. Each volunteer has their own three-dimensional facial template, that is, the modeling representation when not speaking. The three-dimensional modeling representation is unified as the modeling representation of FLAME, that is, the number of all vertices and the relative connection method remain unchanged, and each vertex has a region division such as the face, lips, etc., and the vertex coordinate sequence of a certain region part can be extracted separately.
[0027] In the division of the data set, all the sentences of 8 volunteers are used as the training set, all the sentences of 2 volunteers are used as the validation set in the supervised training stage, and all the sentences of the remaining 2 volunteers are used as the test set.
[0028] S102. Input the original audio in the training set into the audio encoder in the model to extract audio features.
[0029] Input the original audio in the training set into the audio encoder to output the audio feature E1.
[0030] Specifically, the audio encoder is a wav2vec model with frozen parameters. This model needs to import pre-trained parameters from the website and set the model parameters to be frozen.
[0031] S103. Input the audio feature into the KAN decoder in the model to generate 3D digital human facial expression actions.
[0032] Input the audio feature output in step S102 into the decoder based on KAN (Kolmogorov - Arnold Networks) to output the offset sequence B of the three-dimensional coordinates of each vertex in the 3D model. T , and then add the offset sequence B T to the three-dimensional facial template of the corresponding volunteer role to obtain the 3D vertex coordinate sequence generated by the model , that is, the 3D digital human facial expression actions generated by the model. In subsequent steps, supervised training will be carried out to improve the accuracy and restoration degree of the sequence generated by the model.
[0033] In this embodiment, by using the KAN model encoder, accurate three-dimensional facial expression actions are generated according to the original audio, that is, the output is the three-dimensional vertex coordinate sequence of the entire face.
[0034] S104. Input the 3D digital human facial expression actions into the graph encoder in the model to extract lip features.
[0035] Obtain lip features for multi-modal alignment supervision.
[0036] Denote the coordinates E of each vertex v in the patch structure of the 3D digital human facial expression actions v as the feature representation F v .
[0037] According to the three-dimensional coordinates of each vertex, extract the coordinate sequence of the lip vertices from the 3D digital human facial expression actions, and denote the coordinates of the lip vertex v as the feature representation , and denote the coordinates of all lip vertices as the feature representation .
[0038] Input the feature representation of the lip vertices into the encoder composed of a multi-layer graph convolutional neural network to output the lip feature E2:
[0039] In the formula, F (l+1) is the feature representation output by the (l + 1)-th layer of the graph convolutional neural network, which fuses the structural and temporal information of the expression of the 3D model; F (l) is F(0) The feature representation output by the l-th layer of the graph convolutional neural network; W (l) is the weight of the l-th layer, is the adjacency matrix of the lip vertices in the patch structure of the 3D digital human expression actions, and σ() is the activation function.
[0040] The encoder based on the graph convolutional neural network (GCN) has a total of K layers of graph convolutional neural networks, and the feature representation F output by the K-th layer (K) is used as the lip feature E2.
[0041] In the 3D digital human head model, the feature transfer between the vertices and their adjacent vertices in the FLAME topology that makes up the 3D human head model is related to their relative positions, and the connection relationship between individual vertices is non-Euclidean. Existing other methods often ignore the relationship between vertices and their adjacent vertices and only use a multi-layer perceptron to update the feature representation according to the coordinate adjacency relationship of the vertices. Refer to Figure 3 , consider the first orange vertex as the active vertex, surrounded by blue inactive vertices. The active vertex will promote the feature change of the surrounding inactive vertices and transform them into active vertices respectively.
[0042] In this embodiment, an encoder based on a graph convolutional neural network is used to extract the features of the lips in the generated animation of the model. Subsequently, cross-modal loss calculation is performed using the lip features and audio features to further constrain the gap between the generated animation and the real animation.
[0043] S105. By constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features, collaborative optimization of multi-modal features and model training are achieved.
[0044] Use a module similar to the BPR loss to calculate the cross-modal consistency loss of each vertex. It makes the distance calculation between the audio features of the speech and the 3D lip region spatial information generated by the encoder based on the graph neural network in S103 closer, as shown in the following formula:
[0045] where E2 represents the lip features output by the encoder based on the graph neural network in step S103; E1 represents the audio features output by the pre-trained wav2vec in step S101, ε is a small value, V represents the total number of vertices in the FLAME structure, and ⊙ is the Hadamard product.
[0046] To distinguish the importance differences of each region in the face animation generation task, a sub-region loss is set to distinguish the key points of each region:
[0047] Wherein, T represents the total length of the sequence; V represents the total number of vertices; k represents each different region in the three-dimensional human head topology, such as the face, lips, etc.; W k is the weight of each region respectively, used to distinguish the different degrees of emphasis of each region; represents the real three-dimensional coordinate of the v-th vertex of the real three-dimensional vertex coordinate sequence y at the t-th moment. Similarly, represents the predicted three-dimensional coordinate of the v-th vertex of the three-dimensional vertex coordinate sequence generated by the model at the t-th moment.
[0048] To solve the problem of jittery output frames when only using the reconstruction loss, a velocity loss is introduced to encourage smooth and natural lip movement over time. The velocity loss used can optimize the coherence between generated animation frames, resulting in the animation not appearing torn or overly variable. The velocity loss can be expressed as:
[0049] To ensure that the generated lip movements are perfectly synchronized with the original audio, this embodiment uses a text consistency loss calculation method. The audio features are input into an open-source and pre-trained lip-reading decoder (wav2vec2-large-xlsr-53-english), and the subtle differences between the output data and the extracted lip features are compared.
[0050] Specifically, the transformers library of huggingface is used to read the wav2vec2-large-xlsr-53-english model with open-source parameters on the huggingface website, and its parameters are frozen. The lip shape π corresponding to the lip features is input into the wav2vec2-large-xlsr-53-english model to obtain the predicted text , and the text consistency loss is calculated through the following text consistency loss formula. By incorporating the text consistency loss into the training process, the deviation between the predicted character sequence and the actual text can be accurately measured, thereby optimizing the performance of the KAN decoder. The text consistency loss calculation formula is as follows:
[0051] Wherein, is the mapping function from π to the text S of the corresponding sentence in the dataset (provided by the wav2vec2-large-xlsr-53-english model), is the predicted text recognized by the pre-trained lip-reading decoder model according to π, T is the total length of the sequence, e 1:T is the probability parameter which is a constant, p() is the conditional probability function; L ctcIt is a text consistency loss function and can be used to optimize the lip shape of a 3D digital human.
[0052] For the above four loss functions, dynamic weights are used to fuse them. Inspired by the chaotic particle swarm optimization algorithm, a Lip-PSO optimization module is added to the algorithm part. It combines chaotic theory with traditional PSO, overcoming the disadvantages that particles are prone to falling into local optimal solutions and aggregating. Specifically, first, the initial positions and velocities of multiple particles are initialized using chaotic mapping, and a search space S is established, which includes the preset values W1, W2, W3, and W4 of the weights of the four loss functions. Then, the fitness of each iteration is calculated, and the fitness function is designed as:
[0053] In the formula, ϵ is a small constant to ensure that f opt >0; L i is the above four loss functions L ctc , L vel , L reg and L bpr , and W i are the corresponding variable weights.
[0054] Update the best position Pbest() of a single particle and the global best position Gbest() according to the fitness. All particles search for the optimal solution in the search space according to the corrected velocity and position, that is:
[0055] In the formula, v j (t + 1) is the velocity of the j-th particle at time t + 1; λ is the inertia weight adjusted by chaotic perturbation and is a constant; c1 is the self-learning factor, which determines the influence of the best position of the particle on the current particle velocity; c2 is the social learning factor, which determines the influence of the global best position on the particle's current velocity; r1 and r2 are random numbers within the interval; Pbest j (t) and Gbest j (t) represent the best position of the j-th particle and the global best position at time t respectively, and x j (t + 1) is the position of the j-th particle at time t + 1.
[0056] The first equation is the update formula of the velocity from time t to time t + 1, and the second equation is the update formula of the position from time t to time t + 1. Among them, [t, t + 1] is a continuous time period. After multiple iterations and stabilization, that is, when v j (t) and x j (t) do not change, it means that the model converges, and the respective weights w i corresponding to the global best position Gbest can be obtained as the optimized weights.
[0057] Final total loss function , where Loss is the combined loss function.
[0058] In each of the previous steps, there is its own loss function. In order to obtain the final loss function in this model, this embodiment does not directly use fixed weights to add these loss functions, but uses Lip-PSO to optimize the weight parameters, and designs several loss functions to regularize each emphasized generation aspect to improve the generation performance of the model.
[0059] In the training phase, the training data and validation data in the dataset are used for division, and supervised training is carried out following the general method of deep learning model training. Each supervised training only selects one sentence from the training set, that is, the batch_size is 1. Each round of training traverses all the sentences in the entire training set, and at most 100 rounds of training are performed. The loss is calculated using the predicted output of the model and the real sequence according to the total loss function, and then deep learning gradient descent is performed to update each parameter of the model. At the same time, in order to ensure the effect of the training phase, a validation set is used and the cross-validation principle is followed for validation during the training phase. After multiple rounds of training, when the model converges or 100 training rounds are completed, the training is stopped and the model parameters are saved to obtain the trained model weights.
[0060] S106. Input the original audio in the test set into the trained model, and successively pass through the audio encoder and the KAN decoder to generate the corresponding 3D digital human facial expression actions.
[0061] Randomly select an audio file of a sentence, the real three-dimensional vertex coordinate sequence, and the template of the volunteer corresponding to this sentence from the training set. Input the audio into the audio encoder in the trained model to obtain audio features, then input the audio features into the KAN decoder to obtain the three-dimensional vertex coordinate offset, and finally add this offset to the template to obtain the predicted three-dimensional vertex coordinate sequence of the model, that is, the corresponding 3D digital human facial expression actions.
[0062] In order to verify the effect of the model constructed in this embodiment (referred to as this model), the following comparative experiments were carried out between this model and the comparative model: Extract a sentence from the test set to obtain the original audio, three-dimensional vertex coordinate sequence, and the three-dimensional facial template of the volunteer of this sentence. Input the original audio sentence into this model and other comparative models FaceFormer, CodeTalker, and SelfTalk, and output the three-dimensional vertex sequence ; Calculate The "average lip vertex coordinate error" between the generated sequence and the real sequence y (the coordinate distance between the vertexes of the lip area of the 3D model in the generated sequence and the real sequence) (unit: 10 -5 mm), and the "average facial vertex coordinate error" (the coordinate distance between the vertexes of the facial area of the 3D model in the generated sequence and the real sequence) (unit: 10 -7 mm). The lower these two calculation results are, the higher the expression restoration degree of the model is, and the more similar the generated action expression is to the real action expression. In this embodiment, all sentences in the test set are selected, and the average result is calculated and filled into Table 1. It can be seen from Table 1 that the "average lip vertex coordinate error" and the "average facial vertex coordinate error" of this model are the lowest, so the expression restoration degree is the highest. Therefore, this model has a positive optimization effect on action expression synthesis.
[0063] Table 1 Average errors calculated by the model provided in this embodiment and other comparison models
[0064] It can be seen from Table 1 that the average lip vertex coordinate error of this model has dropped to 2.8271, and the average facial vertex coordinate error of this model has dropped to 6.3263. According to the data of the comparison models and the data of this model, the optimization effect of this model is illustrated. Compared with the multi-layer perceptron, KAN provides stronger interpretability because each weight is replaced by a non-linear kernel function. Due to its architecture, KAN can achieve higher accuracy compared with the multi-layer perceptron. This shows that this model has more excellent 3D virtual digital human generation ability.
[0065] This embodiment also provides an audio-driven 3D digital human facial expression visualization system, including a backend and a frontend. The frontend includes a visual frontend display and rendering module, an audio upload module for interacting with users, and a rendering and debugging module. These three modules independently undertake corresponding functions. The frontend display and rendering module is used to render and display the generated 3D animation; the audio upload module is used to provide an interface for users to upload audio files, and this audio file is used to drive the 3D model into a 3D animation; the rendering and debugging module is used to provide setting functions for users, such as the coordinates of the light and whether to display the coordinate axes, etc. The backend responsible for inference includes a connection component that maintains WebSocket and Request responses, a deep learning algorithm module, and a data persistence module. The frontend and backend are connected using WebSocket and Request responses. The running process is as follows: The user inputs audio data at the frontend, and the system sends the audio to the backend in the form of a network request. The backend calls the deep learning module to infer the data representation of the 3D animation, and then responds and transmits it back to the frontend for visualization. At the same time, the data is persisted. The data persistence includes saving the audio input by the user and the meta-information for generating the visualization, such as time, etc.
[0066] The framework used in the front-end is React + ThreeJS. React is a JavaScript library for building user interfaces, which can efficiently manage component states and update the interface to ensure the responsiveness and maintainability of the system. The declarative programming model of React can build complex user interfaces more clearly and interact seamlessly with the backend through HTTP requests and WebSocket to achieve real-time data transmission and display. ThreeJS is a 3D graphics library based on WebGL, which can create and render highly complex 3D graphics in the browser. By integrating ThreeJS, the front-end interface can process and display 3D facial expression animation data generated by the backend.
[0067] In this embodiment, ThreeJS is mainly used to render and display the facial motion data obtained through deep learning inference to provide users with a real-time interactive experience. In addition, the component structure provided by React makes the development and maintenance of front-end code more modular and efficient.
[0068] The framework used in the backend is FastAPI. FastAPI is a modern web framework based on the Python programming language, designed specifically for building efficient and scalable API services. FastAPI has many features such as high-performance asynchronous processing capabilities, automatic API documentation generation, type checking, and data validation mechanisms, and can build applications with high concurrency and low latency. FastAPI is fully compatible with Python type hints. In backend development, the advantages of FastAPI are particularly prominent. In application scenarios that require handling high-concurrency requests, it can handle a large number of I / O-intensive tasks such as database queries, file uploads, and external API calls.
[0069] In this embodiment, the high performance of FastAPI is used to build the backend system, which is responsible for receiving and processing the audio data uploaded by users. Specifically, FastAPI is used to implement the connection components for WebSocket and HTTP requests. The audio files uploaded by the front-end are transmitted to the backend through HTTP requests, and FastAPI immediately processes them and calls the deep learning model for inference. The 3D digital human facial expression animation data generated after inference is then sent back to the front-end through WebSocket for visual display. Based on the FastAPI framework, the system can achieve efficient interaction and processing of front-end and backend data while ensuring high concurrency and low latency, so as to improve the user experience and satisfaction. The architecture of the system can be referred to Figure 4 。
[0070] The system development environment is NodeJs18; the backend environment is Python version: 3.8.20; the operating system version is Windows11 22631.4602, and the FastAPI version is 0.108.0.
[0071] Before performing the visual display, it is necessary to start the visualization system first. The steps to start are as follows: (1) Place the pth file saved by the trained deep learning model in the corresponding position of the backend, and start the backend program using Docker. At the same time, it is necessary to note that port 5000 is allowed through the system firewall. At this time, Request listening and WebSocket listening will be enabled simultaneously to listen for network requests from the front end.
[0072] (2) Package the front-end code using npm, place the packaged product in the relevant position of the docker container, and then use the Docker command to run the container to start the front end, so that the front end can enable Request listening.
[0073] (3) Open the website using the browser of a personal computer or mobile phone to access the visualization system, and then perform tasks such as audio file upload and display.
[0074] Compared with the existing visualization systems that require running code to perform visualization, the visualization system provided in this embodiment can directly open the visualization page using a browser due to different implementation methods, which is more convenient to operate; the ThreeJS framework is used in the system. Due to the optimization of the ThreeJS framework itself and its use of more advanced WebGL, the software and hardware requirements are reduced. This is mainly reflected in that other similar methods must use the Linux system and compile and run the psbody-mesh library, resulting in difficult installation and high CPU occupancy during operation; while the visualization system provided by the present invention can be installed without using the above library and has low CPU occupancy during use. Compared with the existing visualization systems, the visualization system provided in this embodiment is more flexible on the client side. Any device equipped with a browser can run the visualization part of this system. This is mainly reflected in that existing similar systems need to use a Linux server for visual rendering, while the system provided in the present invention can implement a B-S architecture, where the inference model runs on the server and the front-end rendering program can run on any device equipped with a browser, such as edge devices like personal laptops and mobile phones.
[0075] Those skilled in the art can understand that all or part of the steps in the method for implementing the above embodiments can be completed by instructing relevant hardware through a program, and the corresponding program can be stored in a computer-readable storage medium.
[0076] It should be noted that although the method operations of the above embodiments are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the depicted steps can be changed in the order of execution. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0077] Embodiment 2: As Figure 5 shown, this embodiment provides a 3D digital human facial expression synthesis device based on audio drive. The device includes a data acquisition unit 501, a first extraction unit 502, a generation unit 503, a second extraction unit 504, a model training unit 505, and a 3D animation generation unit 506, where: The data acquisition unit 501 is used to acquire a data set; the data set includes English sentences recorded by multiple volunteers, and each English sentence corresponds to an original audio. The first extraction unit 502 is used to input the original audio in the data set into the audio encoder in the model to extract audio features. The generation unit 503 is used to input the audio features into the KAN-based decoder in the model to generate 3D digital human facial expression actions. The second extraction unit 504 is used to input the 3D digital human facial expression actions into the graph encoder in the model to extract lip features. The model training unit 505 is used to achieve the collaborative optimization and model training of multi-modal features by constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features. The 3D animation generation unit 506 is used to input the original audio to be tested into the trained model, and successively pass through the audio encoder and the KAN decoder to generate corresponding 3D digital human facial expression actions.
[0078] For the specific implementation of each module in this embodiment, reference can be made to Embodiment 1 above, and details will not be repeated here; it should be noted that the system provided in this embodiment is only illustrated by the above division of each functional module. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above.
[0079] Embodiment 3: This embodiment provides a terminal device, which can be a computer, such as Figure 6As shown in the figure, it includes a processor 602, a memory, an input device 603, a display 604, and a network interface 605 connected through a system bus 601. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 606 and an internal memory 607. The non-volatile storage medium 606 stores an operating system, a computer program, and a database. The internal memory 607 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 602 executes the computer program stored in the memory, the method for synthesizing 3D digital human facial expressions based on audio driving in the above-mentioned Embodiment 1 is implemented as follows: Obtain a data set; the data set includes English sentences recorded by multiple volunteers, and each English sentence corresponds to an original audio; Input the original audio in the data set into the audio encoder in the model to extract audio features; Input the audio features into the KAN-based decoder in the model to generate 3D digital human facial expression actions; Input the 3D digital human facial expression actions into the graph encoder in the model to extract lip features; By constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features, realize the collaborative optimization of multi-modal features and model training; Input the original audio to be tested into the trained model, and successively pass through the audio encoder and the KAN decoder to generate corresponding 3D digital human facial expression actions.
[0080] Embodiment 4: This embodiment provides a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, the method for synthesizing 3D digital human facial expressions based on audio driving in the above-mentioned Embodiment 1 is implemented as follows: Obtain a data set; the data set includes English sentences recorded by multiple volunteers, and each English sentence corresponds to an original audio; Input the original audio in the data set into the audio encoder in the model to extract audio features; Input the audio features into the KAN-based decoder in the model to generate 3D digital human facial expression actions; Input the 3D digital human facial expression actions into the graph encoder in the model to extract lip features; By constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features, realize the collaborative optimization of multi-modal features and model training; Input the original audio to be tested into the trained model, and successively pass through the audio encoder and the KAN decoder to generate corresponding 3D digital human facial expression actions.
[0081] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0082] As mentioned above, only the preferred embodiments of this invention patent are described, but the protection scope of this invention patent is not limited thereto. Any person skilled in the art within the scope disclosed by this invention patent, according to the technical solution and inventive concept of this invention patent, makes equivalent substitutions or changes, and all belong to the protection scope of this invention patent.
Claims
1. A 3D digital human facial expression synthesis method based on audio driving, characterized in that, The method includes: Obtaining a dataset; the dataset includes English sentences recorded by multiple volunteers, and each English sentence corresponds to an original audio; Inputting the original audio in the dataset into the audio encoder in the model to extract audio features; Inputting the audio features into the KAN-based decoder in the model to generate 3D digital human facial expression actions; Inputting the 3D digital human facial expression actions into the graph encoder in the model to extract lip features; By constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features, collaborative optimization of multi-modal features and model training are achieved; Inputting the original audio to be tested into the trained model, and successively passing through the audio encoder and the KAN-based decoder to generate corresponding 3D digital human facial expression actions.
2. The 3D digital human facial expression synthesis method according to claim 1, wherein Each English sentence also corresponds to a real text and a three-dimensional vertex coordinate sequence y; The achieving of collaborative optimization of multi-modal features and model training by constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features includes: According to the audio encoder and lip features, constructing a cross-modal consistency loss function using the BPR loss; In order to distinguish the importance differences of each region in the face animation generation task, constructing a sub-region loss function according to the three-dimensional vertex coordinate sequence y and the 3D digital human facial expression actions; To solve the problem of jittery output frames when only using the reconstruction loss, based on the sequence of 3D vertex coordinates y at time t t and the sequence of 3D vertex coordinates y at time t-1 t-1 the difference between them, and the 3D digital human facial expression actions at time t and the 3D digital human facial expression actions at time t-1 the difference between them, construct a velocity loss function; In order to ensure that the generated lip movements are perfectly synchronized with the original audio, constructing a text consistency loss function according to the real text and the predicted text; wherein, the predicted text is obtained by inputting the lip shape corresponding to the lip features into the text decoder; Using the cross-modal consistency loss function, the sub-region loss function, the speed loss function, and the text consistency loss function and their corresponding weights to construct a total loss function; wherein, the weights are obtained by parameter tuning using the chaos theory and the particle swarm optimization algorithm; Using the total loss function to train the model to optimize the parameters in the model and obtain a trained model.
3. The 3D digital human facial expression synthesis method according to claim 2, wherein The cross-modal consistency loss function, the sub-region loss function, the speed loss function, and the text consistency loss function and their corresponding weights to construct a total loss function are collectively referred to as four loss functions; The weights corresponding to the four loss functions are obtained through the following process: Using the chaos mapping to initialize the initial positions and velocities of multiple particles; and establishing a search space, including the weight preset values of the four loss functions; designing a fitting function according to the four loss functions and their corresponding weight preset values; Updating the best position of a single particle and the global best position according to the fitness; After multiple iterations, if the velocities and positions of the particles remain unchanged, the model converges; when the model converges, the respective weights corresponding to the global best position are used as the optimized weights of the four loss functions.
4. The 3D digital human facial expression synthesis method according to claim 2, wherein The text decoder is the wav2vec2-large-xlsr-53-english model with open-source parameters.
5. The 3D digital human facial expression synthesis method according to any one of claims 1 to 4, characterized in that, Extract the coordinate sequence of the lip vertices according to the coordinates of each vertex in the patch structure of the 3D digital human facial expression movements; Denote the coordinates of all lip vertices as the feature representation F (0) ; The inputting of the 3D digital human facial expression actions into the graph encoder in the model to extract lip features includes: where F (l+1) is the feature representation output by the l +1-th layer of the graph convolutional neural network, which fuses the structural and temporal information of the expressions of the 3D model; F (l) is the feature representation output by (0) F l after passing through the (l) -th layer of the graph convolutional neural network; W l is the weight of the σ -th layer and is the adjacency matrix of the lip vertices, σ (·) is the activation function; Among them, the graph encoder has a total of K graph convolutional neural networks, and the feature representation output by the Kth layer is used as the lip features; where K is a positive integer greater than or equal to 2.
6. The 3D digital human facial expression synthesis method according to any one of claims 1 to 4, characterized in that, Each volunteer has a corresponding three-dimensional facial template; Inputting the audio features into the KAN-based decoder in the model to generate 3D digital human facial expression actions, including: Input the audio features into the KAN-based decoder to output a sequence B of the offsets of the three-dimensional coordinates of each vertex in the 3D model T ; Add the offset sequence B T to the 3D facial template of the corresponding volunteer to generate 3D digital human facial expression actions.
7. The 3D digital human facial expression synthesis method according to any one of claims 1 to 4, characterized in that, The audio encoder is a wav2vec model with frozen parameters.
8. A 3D digital human facial expression visualization system based on audio drive, characterized in that, The system includes a front end and a back end; the front end includes a visual display and rendering module, an audio upload module for interacting with the user, and a rendering and debugging module; among them, the display and rendering module is used to render and display the generated 3D digital human facial expression actions; the audio upload module is used to provide an interface for the user to upload the original audio file to drive the 3D model into 3D digital human facial expression actions; the rendering and debugging module is used to provide setting functions for the user; the back end includes a deep learning algorithm module; The front and back ends use WebSocket and Request responses to connect: the user inputs the original audio data through the audio upload module of the front end, the system sends the audio to the back end in the form of a network request, the back end calls the deep learning module to execute the 3D digital human facial expression synthesis method described in any one of claims 1 to 7 to infer the 3D animation, and then responds and transmits it back to the display and rendering module for display; among them, the framework used by the front end is React and ThreeJS, and the framework used by the back end is FastAPI.
9. A 3D digital human facial expression synthesis device based on audio drive, characterized in that, The device includes: A data acquisition unit for acquiring a data set; the data set includes English sentences recorded by multiple volunteers, and each English sentence corresponds to original audio; A first extraction unit for inputting the original audio in the data set into the audio encoder in the model to extract audio features; A generation unit for inputting the audio features into the KAN-based decoder in the model to generate 3D digital human facial expression actions; A second extraction unit for inputting the 3D digital human facial expression actions into the graph encoder in the model to extract lip features; A model training unit for realizing the collaborative optimization and model training of multi-modal features by constructing a joint loss function that fuses audio features, 3D facial expression parameters, and lip features; A 3D animation generation unit for inputting the original audio to be tested into the trained model, passing through the audio encoder and the KAN-based decoder in sequence, and generating corresponding 3D digital human facial expression actions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it realizes the 3D digital human facial expression synthesis method described in any one of claims 1 to 7.
Citation Information
Cited By
Multi-modal face restoration and expression recognition system and method based on semantic guidance of facial action unit
CN121998875A
A multi-modal face restoration and expression recognition system and method based on facial action unit semantic guidance
CN121998875B
Deep pseudo voice active defense method and system based on perception characteristic and disturbance decoupling
CN122266390A