A virtual anchor generation method and system based on neural radiance field and hidden attributes
By using neural radiation fields and latent property technology to generate virtual anchors, the problem of insufficient facial expression vividness in virtual anchors has been solved, achieving high-quality and low-cost virtual anchor generation, which is suitable for applications in multiple fields.
Patent Information
- Application Number
- CN202311094348.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-08-28
AI Technical Summary
Current virtual anchor technology lacks vividness and naturalness in facial expressions and implicit attribute information, resulting in virtual anchors that lack realism and fail to resonate with the audience's emotions.
A virtual anchor generation method based on neural radiation fields and latent attributes is adopted. Through facial feature extraction and construction, speech synthesis, explicit and implicit attribute feature extraction and an improved NeRF network module, combined with a background replacement algorithm, high-quality virtual anchor videos are generated.
It enhances the realism and entertainment value of virtual anchors, reduces production time and costs, is applicable to multiple fields, meets the needs of different scenarios and users, and provides a brand-new viewing experience.
Smart Images

Figure CN117171392B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence, and relates to text, voice and video analysis and synthesis, in particular to a virtual anchor generation method and system based on neural radiance field and hidden attributes. BACKGROUND
[0002] With the continuous development of Internet technology, network live broadcast is increasingly widely used in social, entertainment and education fields. The anchor is the core figure of live broadcast and the main attraction of live broadcast. The emergence of virtual anchor technology enables the generation of virtual anchor voice, video and action through artificial intelligence and machine learning technology, realizing a real live broadcast experience. In the field of radio live broadcast, this technology can generate virtual broadcasters; in the field of television live broadcast, it can generate virtual hosts; in the field of entertainment live broadcast, it can generate virtual network anchors; and in the field of education live broadcast, it can generate virtual teachers to provide online teaching for students.
[0003] Virtual anchor generation technology is a new type of digital content generation technology. Based on artificial intelligence, this technology uses machine learning and deep learning algorithms to input a large amount of audio, video and text data into a model to generate content similar to the original data. Its field involves artificial intelligence, machine learning, speech recognition, image processing and natural language processing.
[0004] Virtual anchors use artificial intelligence natural language processing technology to analyze input text data and capture relevant contextual information to generate voices with the same tone, intonation, speech rate and language rhythm.
[0005] However, the virtual anchor's character image is currently a difficult problem to overcome. The current virtual anchor generation, the lively, natural and harmonious facial expressions of the characters have always been missing in existing virtual anchors. Because human facial expressions are extremely rich, the character image generated using existing intelligent algorithms often has quite rigid facial expressions, lacks agility, and makes people feel a sense of estrangement. It is difficult for the audience to produce heartfelt emotional resonance when facing such a lifeless doll face. Some seemingly insignificant details, referred to as hidden attributes, affect these factors.
[0006] The hidden attribute refers to the hidden attribute of the character feature, which plays an important role in generating a virtual anchor. The apparent attributes of the character mainly include: age, gender, posture angle, expression, glasses wearing and race, etc. The hidden attributes of the character are some unobvious features mainly including: emotion, eye movement, blink frequency, body and gesture, micro-expression, etc. These hidden attributes may be guided by other deeper dimensional information, providing more rich information for character synthesis, so that the generated virtual anchor is more real and natural, and has important value for the application of the generated virtual anchor.
[0007] NeRF (Neural Radiance Fields) is a neural radiance field, which is a technology for three-dimensional reconstruction using a deep neural network model. It can model the facial expressions, body movements and other features of the anchor, thereby generating a lively virtual anchor model. The NeRF technology also has strong scalability, and can improve the generation effect of the model through continuous training and improvement of the model. Therefore, the hidden features of the character are used as input, and the NeRF technology is used for virtual anchor rendering, and the body posture, expression and the like of the virtual anchor can be finely adjusted, thereby improving the effect of the generated virtual anchor. However, the original NeRF needs to learn a large amount of data and a long time to generate a three-dimensional image. SUMMARY
[0008] In order to improve the realism of the virtual anchor and expand the application range and application scenarios of the virtual anchor, a virtual anchor generation method and system based on neural radiation field and hidden attribute are provided, which learns and trains the apparent and hidden attributes of the character based on the NeRF neural radiation field technology, to realize a more lively and realistic virtual anchor.
[0009] The virtual anchor generation system based on neural radiation field and hidden attribute includes a face feature extraction and construction module, a speech synthesis module, a speech feature extraction module, an apparent and hidden attribute feature extraction module, an improved NeRF network module and a background replacement module. The background replacement module includes a background segmentation module and an image harmonization module.
[0010] The virtual anchor generation method based on neural radiation field and hidden attribute specifically includes the following steps:
[0011] Step 1: Construct the character image of the virtual anchor according to the actual needs, including character video data, speech and text data and background data, as the input of the virtual anchor generation system.
[0012] Determine the language used by the virtual anchor, and obtain the speech and text data set of the corresponding language according to different languages, so as to serve as the model training data for semantic analysis and speech synthesis of the virtual anchor.
[0013] Step two, the face feature extraction and construction module extracts and constructs the face features of the character video data to generate a three-dimensional face of the virtual anchor;
[0014] The face feature extraction and construction process of the virtual anchor includes face analysis, 3DMM face feature extraction, and face reconstruction, specifically:
[0015] Step 201, face analysis is to decompose the character video data into face components through deep learning technology, and obtain the corresponding facial features;
[0016] The face components include skin, hair, eyes, eyebrows, nose, and mouth, etc.
[0017] Step 202, through 3DMM face feature extraction, the facial features of different parts of the face are three-dimensionally coded, and the combination of the reference shape and texture map that best represents the required face is selected to reconstruct the face;
[0018] Step 203, the reference shape and texture map of the face are combined by weighting in the database to generate a reconstructed three-dimensional face.
[0019] Step three, the speech synthesis network is trained through speech data, and at the same time, the text data is input into the text transcription module for text front-end processing, and then the processed text is input into the trained speech synthesis network model, after speech synthesis, the synthesized speech of the virtual anchor is obtained.
[0020] Text front-end processing refers to removing punctuation, numbers, spaces, etc. in the input text, dividing the text into individual words according to semantic understanding, and converting each character into a phoneme, and labeling the speech tags such as tone and sound, etc.
[0021] Step four, the speech feature extraction module extracts the features of the synthesized speech of the virtual anchor, and the explicit and implicit attribute feature extraction module extracts the explicit and implicit attribute feature information in the video data in combination with the three-dimensional face, and outputs each feature information extracted to the improved NeRF network module.
[0022] Speech feature extraction: extracting the features of the synthesized speech of the virtual anchor, extracting the spectral, tone, and sound features, etc. and mapping them to the corresponding discrete values.
[0023] Explicit attribute feature extraction: using 3DMM to extract the lip movement, facial movement, and expression feature data that have strong correlation with speech data from the video data, as explicit attributes output to the improved NeRF network module;
[0024] Hidden attribute feature extraction: attributes with weak correlation with speech data, i.e. other attributes related to speech context or personalized conversation style, including head movement and blinking, etc. Through the constructed three-dimensional face model, the motion of the relevant parts in the video data is extracted as hidden attribute output to the improved NeRF network module.
[0025] Step five, the improved NeRF network module models the static scene, dynamic head and dynamic torso of the virtual anchor according to the speech feature information and the hidden and explicit attribute feature information, and obtains the synthesized video of the virtual anchor;
[0026] Specifically:
[0027] (1) When the improved NeRF network is used for static scene modeling, the MLP multi-layer perceptron is reduced, and the MLP uses linear interpolation instead, so as to keep the reconstructed static information at each static 3D position, and thus the features of the 3D scene are stored in the static scene trainable grid structure.
[0028] (2) When the improved NeRF network is used for dynamic head modeling, the high-dimensional audio-video processing network of the person is decomposed into three low-dimensional trainable feature grids, i.e. lip shape motion model, person head motion model and eye blinking model; in order to realize the synchronization of audio and each motion model, the audio-spatial coding module is decomposed into 3D space grid and 2D audio grid, and the audio and spatial representation are decomposed into two grids. When each motion model keeps the static spatial coordinates in 3D, the audio dynamics is coded as low-dimensional "coordinates".
[0029] When constructing the relationship between the explicit attribute lip shape motion and the audio, the mouth motion of the auditory speech is directly synchronized and embedded. Specifically, the CNN audio encoder E a extracts phoneme features f a from the input audio, and the expression is as follows:
[0030] f a =E a (a)
[0031] where a represents the input audio data.
[0032] A contrast learning strategy is used to align the audio features and the features of the mouth. The timely aligned audio and mouth features (f a , f m ) are regarded as positive pairs, while the unaligned pairs are regarded as negative pairs. Binary cross-entropy loss is used for contrast learning, in which the distance between the timely aligned positive pairs is closer than the unaligned negative pairs.
[0033]
[0034] τcon The binary cross-entropy loss, d(f), represents the loss between lip shape and speech. m ,f a () represents the cosine distance directly opposite. It represents the cosine distance of the negative pair.
[0035] The synchronization process between audio and latent attribute blink frequency and head posture is as follows:
[0036] A controlled probabilistic model is used for blinking and head posture movements, with a sequence h of facial attributes (head posture or blinking) of length T. 1:T and a conditioning audio sequence a of length T′ 1:T′ It is necessary to embed the predicted facial attribute sequence h into the image from T to T′. T+1:T′ Facial attribute sequence h T+1:T′ It consists of two parts: (1) latent attribute space construction. Transformer-VAE is trained on a large dataset using Gaussian Process (GP) to establish a mapping between input facial attribute sequences and a latent attribute space Z. (2) Pose and blink space construction, fine-tuning the cross-modal encoder to embed two head BOPs and blink frequency audio embeddings into the latent attribute space Z on selected characters.
[0037] In obtaining the generated head pose and blinking features f e and synchronized audio features f a Next, neural radiation fields are used to generate the final image with these conditions. First, the synchronized audio features f... a and blinking features f e Connect into a new feature f c Then, using this new feature as input, a conditional radiation field is proposed. After transforming the head pose from camera space to canonical space, the head pose is directly used to replace the view direction d of the conditional radiation field. Finally, the feature f in canonical space, the view direction d, and the 3D position x constitute the implicit function F. θ The input. Actually, F θ It is implemented using a multilayer perceptron. For all input vectors, the implicit function F... θ It can estimate the color value c associated with density σ and the distributed light.
[0038] Implicit function F θ Expressed as a formula:
[0039] F θ :(f,d,x)→(c,σ)
[0040] (3) Improved NeRF network for dynamic torso modeling, another 2D mesh is used to simulate the dynamic characteristics of the torso in a lightweight pseudo-3D deformable module, and a natural torso image matching the head is synthesized.
[0041] The torso deformation is conditioned on the head pose p, so that the torso motion is synchronized with the head motion.
[0042] An MLP is used to predict the torso deformation:
[0043] Δx = MLP(x t ,p)
[0044] x t is the pixel coordinate sampled from the image space. Δx is the pixel coordinate after the torso deformation.
[0045] The coordinates of the deformed torso are fed into a two-dimensional feature mesh encoder to obtain the torso features f t :
[0046]
[0047] Another MLP is used to generate the torso RGB color and alpha value:
[0048] c t ,α t = MLP(f t ,i t )
[0049] where i t is the embedded hidden feature learned by the model, c t is the torso RGB color, and α t is the alpha value.
[0050] (4) Synthesize the above separately rendered head and torso model with the static model to obtain the synthesized video of the virtual anchor.
[0051] Step six, the background replacement module replaces the background of the virtual anchor synthesis video according to the background data, and fuses the virtual anchor character image, background and audio to synthesize the final virtual anchor.
[0052] Step 601, input the virtual anchor synthesis video into the background segmentation module, extract the alpha channel of the foreground object in the image through the Background-Matting background segmentation model, so that the image of the virtual anchor in the synthesis video is completely separated from the background.
[0053] Background-Matting Background segmentation model includes a base network and a refinement network, wherein the base network predicts an alpha matte and a foreground layer at a low resolution and outputs an error prediction tile indicating the area that needs high-resolution refinement.The refinement network uses the low-resolution result and the original image as input and generates a high-resolution output only in the indicated area to segment out the person picture in the video.
[0054] The background segmentation problem is modeled by representing each pixel of the image as a combination of foreground and background:
[0055] C=F*alpha+B*(1-alpha)
[0056] C is a given image, F is the foreground calculated for each pixel, B is the background calculated for each pixel, and alpha is the transparency of each pixel.
[0057] Step 602, synthesize the segmented virtual anchor image into another background image to obtain a synthesis image, and perform harmonization processing on the synthesis image through an image harmonization module to complete background replacement.
[0058] Harmonization processing includes sequentially performing color and light adjustment, brightness linear transformation, gray histogram equalization, color correction and local contrast enhancement on the foreground picture.
[0059] The image after background replacement is:
[0060]
[0061] Wherein the background image is I b , the foreground image is I f , the foreground image mask is M, and the combined image is I c . It is a Hadamard product.
[0062] Step 603, input the virtual anchor video after background replacement and the virtual anchor synthesis voice into the FFmpeg tool for audio and video combination to synthesize the final virtual anchor.
[0063] The advantages and positive effects of the present application are:
[0064] (1) The present application generates a virtual anchor through an automatic and intelligent method, which is a low-cost and efficient virtual anchor generation method, can greatly reduce the time and cost of virtual anchor production, can be applied to virtual anchor production in different fields, and enriches the selection of audiences; and the virtual anchor generated by the present application has high efficiency and stability, and can produce positive effects in large-scale applications.
[0065] (2) The improved NeRF network generates a virtual anchor with high-quality speech lip synchronization, making the virtual anchor more realistic and interesting, and better conveying information.
[0066] (3) The present application can adjust and optimize the image of the generated virtual anchor through the learning of hidden attributes, so that the audience has a strong perception.
[0067] (4) The present application realizes real-time replacement of video background through the combination of background replacement algorithm and image harmonization algorithm, better meets the needs of different use scenarios and different users, can bring a new viewing experience to the audience, and has a wide application in more fields. BRIEF DESCRIPTION OF DRAWINGS
[0068] Figure 1 a schematic diagram of the virtual anchor generation system of the present application;
[0069] Figure 2 a schematic diagram of the face feature extraction and construction process of the virtual anchor generation system of the present application;
[0070] Figure 3 a schematic diagram of the virtual anchor speech generation process of the present application;
[0071] Figure 4 a schematic diagram of the virtual anchor video generation process of the present application;
[0072] Figure 5 a schematic diagram of the virtual anchor background replacement process of the present application. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme of the present application will be further described in detail below in combination with the drawings and examples.
[0074] The virtual anchor generation system based on neural radiation field and hidden attributes, as shown in Figure 1 The face feature extraction and construction module extracts features from the face and models it, the speech synthesis module synthesizes speech, the speech feature extraction module extracts speech features, the explicit and implicit attribute extraction module extracts explicit and implicit features such as lip shape, blink frequency and head posture, the improved NeRF network module synthesizes virtual anchor video, and the background replacement module replaces the application scenario, thereby realizing a more realistic and lively virtual anchor.
[0075] The face feature extraction and construction module refers to a method of restoring a three-dimensional shape of a two-dimensional face image through a 3DMM model. The 3DMM here is a relatively basic three-dimensional face statistical model, which can load a BFM database, effectively expanding the applicable scenarios of the 3DMM. The BFM can fit any three-dimensional face and save 3DMM parameters.
[0076] The speech synthesis module refers to taking the input text of the virtual anchor as the input of speech synthesis, using the text front-end processing of the TTS model to transcribe the text into phoneme information, and synthesizing the transcribed phonemes into the speech of the virtual anchor through the Fastspeech acoustic synthesis model.
[0077] The speech feature extraction module and the explicit and implicit attribute extraction module refer to extracting the speech, lip shape and other features of the virtual anchor. First, the frequency spectrum features of the speech are extracted through DeepSpeech, i.e., the speech signal is decomposed into a series of short-time Fourier transform features, and then these features are input into the CNN to extract the frequency spectrum features. Then, the head posture and expression parameters are obtained using 3DMM, and then the eye blinking and mouth movement are decoupled using the face attribute decoupling module to obtain the explicit and implicit features of the virtual anchor.
[0078] The improved NeRF network module refers to using an improved NeRF network to adapt to dynamic audio-driven person modeling through efficient static scene modeling capability. The improved NeRF network here is to decompose high-dimensional audio into three low-dimensional trainable feature grids for portrait modeling. For dynamic head modeling, a decomposed audio-spatial coding module is proposed, which mainly decomposes and synthesizes eye blinking, head posture and lip movement. This decomposed audio-spatial coding realizes the NeRF network of speech-driven dynamic head modeling.
[0079] The background replacement module includes a background segmentation module and an image harmonization module, and the implementation of virtual anchor background replacement refers to segmenting the virtual anchor and the background through Background-Matting to generate a cut-out virtual anchor foreground object, and taking the background to be replaced as the background of the cut-out character, thereby realizing the background replacement of the virtual anchor.
[0080] The method for generating a virtual anchor using a neural radiation field and implicit attribute-based virtual anchor generation system includes the following steps:
[0081] Step one, according to the actual needs, construct the virtual anchor's character image, including character video data, speech and text data and background data, as the input of the virtual anchor generation system.
[0082] The virtual anchor's character video and virtual anchor background picture are determined as inputs for generating the virtual anchor action video.
[0083] The virtual anchor's language is determined, and the corresponding language speech and text data sets are obtained as training data for the virtual anchor's semantic analysis and speech synthesis model.
[0084] Step two, the face feature extraction and construction module extracts and constructs the face features of the character video data to generate a three-dimensional face of the virtual anchor.
[0085] As shown in Figure 2 The face feature extraction and construction process of the virtual anchor includes face analysis, 3DMM face feature extraction, and face reconstruction, specifically:
[0086] Step 201, face analysis is to decompose the character video data into face components through deep learning technology, and obtain the corresponding facial features.
[0087] Face analysis is a special case of semantic image segmentation. Face analysis is a technology in the field of computer vision, aiming to decompose face images or videos into their components, such as skin, hair, eyes, eyebrows, nose, mouth, etc. Face analysis algorithms usually use deep learning techniques, such as convolutional neural networks (CNN), to analyze images or videos and identify different facial features. Face analysis algorithms are trained on large annotated face image datasets, allowing them to learn the relationships between different facial features and distinguish them based on visual cues.
[0088] Step 202, encode the facial features of different parts of the face through the 3DMM network, and select the combination of the reference shape and texture map that best represents the desired face to synthesize the face.
[0089] Here, 3DMM (3D Morphable Model) is a statistical model used to represent the shape and texture variations of human faces and heads, which can be used to synthesize new faces or analyze existing ones. It is built based on a large dataset of real human faces from 3D scans or photographs. 3DMM usually consists of a set of reference shapes, each representing a different aspect of the face shape, such as the overall shape of the head, the nose, the mouth, etc. These reference shapes are combined with a set of texture maps representing variations in skin color, texture, and other details. Using 3DMM, a face can be synthesized by selecting the combination of reference shapes and texture maps that best represents the desired face.
[0090] Step 203, the reference shape and texture map of the face are combined by the BFM database to generate a reconstructed three-dimensional face.
[0091] Here, the BFM-2017 database (Blendshape Face Model 2017) is used, which is an improved version based on the Basel Face Model (BFM) publicly released in 2017 and has more shape and texture variables. BFM-2017 contains a set of shape bases representing different shape features of a human face, and corresponding texture maps representing detailed information of a human face. By weightedly combining these shape bases and texture maps, different human face shapes and textures can be synthesized, so as to generate a reconstructed three-dimensional face, effectively simulate the real shape of a human face, and effectively improve the accuracy of face reconstruction.
[0092] The facial expression S can be expressed as a combination of a 3DMM expression and a geometric parameter, as follows:
[0093]
[0094] is the mean face mesh, B id refers to the PCA (primary component analysis) geometric basis, B exp refers to the PCA expression, F id refers to the coefficient geometric basis, F exp refers to the coefficient expression.
[0095] Step three, the speech synthesis network is trained by voice data, and the text data is input into the text transcription module for text front-end processing, and then the processed text is input into the trained speech synthesis network model, after speech synthesis, the synthesized speech of the virtual anchor is obtained, so that it has real speech performance.
[0096] As Figure 3 shown, the speech synthesis network is a FastSpeech model based on Transformer architecture, which is a deep learning model based on attention mechanism. The Transformer architecture can better process speech features and generate high-quality speech. At the same time, the FastSpeech model uses a bidirectional prediction and alignment loss function method to simultaneously predict the length and pronunciation of the speech, effectively improving the generation efficiency and quality of the model.
[0097] Through the TTS (Text-to-Speech) model, the text can be effectively converted into a form that the FastSpeech model can handle, so that the FastSpeech model can better generate speech and better handle special characters and words, thereby generating more accurate and smooth speech.
[0098] Text front-end processing refers to removing punctuation marks, numbers, spaces and the like in text data, dividing the text into individual words according to semantic understanding, and converting each character into a phoneme and labeling it with a phonetic tag such as tone and pitch.
[0099] Step four, the speech feature extraction module extracts features from the synthesized speech of the virtual anchor, and the explicit and implicit attribute feature extraction module extracts explicit and implicit attribute feature information from the video data in combination with the three-dimensional face, and outputs the extracted feature information to the improved NeRF network module.
[0100] Speech feature extraction: Speech feature extraction network is a speech recognition system using deep learning technology to extract speech feature information in speech. The speech feature extraction model DeepSpeech extracts features such as frequency spectrum, tone and pitch from the synthesized speech data of the virtual anchor, and maps them to the corresponding discrete values.
[0101] Explicit and implicit attribute feature extraction network is an attribute feature extraction network based on deep learning technology.
[0102] Explicit attribute feature extraction: 3DMM is used to extract lip movement, facial movement and expression that are strongly related to speech data from video data as explicit attributes and output to the improved NeRF network module;
[0103] Implicit attribute feature extraction: Implicit attributes are attributes that have weak correlation with speech data, i.e. other attributes related to speech context or personalized conversation style, including head movement and blinking, etc. The motion of the relevant parts in the video data is extracted through the constructed three-dimensional face model as implicit attributes and output to the improved NeRF network module.
[0104] The speech feature extraction network and the implicit attribute feature extraction network can be used in combination to generate a realistic synthesis effect according to the feature information of the speech text and the implicit attribute information.
[0105] Step five, the improved NeRF network module models the static scene, dynamic head and dynamic torso of the virtual anchor according to the speech feature information and explicit and implicit attribute feature information to obtain the synthesized video of the virtual anchor.
[0106] As shown in Figure 4 The present application uses an improved NeRF network to model the static scene, the dynamic head and the dynamic torso, as follows:
[0107] NeRF neural radiance field has achieved success in high-fidelity 3D modeling, but the NeRF network is slow in training and inference, which seriously affects its efficiency. In the present invention, the NeRF network is optimized, which is an improvement of the neural radiance field, effectively realizing real-time synthesis of sound-driven tasks and faster convergence of training, and can generate high-quality virtual anchor images of lip-synchronized speech information by using less data and learning different speaker's lip shape, blinking frequency and head posture features.
[0108] The improved NeRF network here refers to decomposing the high-dimensional audio-video processing network into three low-dimensional feature grids, specifically, the decomposed audio-space encoding module is used to model the dynamic head with a 3D space grid and a 2D audio grid, and the torso is processed with another 2D grid in a lightweight pseudo 3D deformable module.
[0109] Among them, for static scene modeling, the improved NeRF is used for new view synthesis to achieve an unprecedented realistic effect. Because the synthesis efficiency of the NeRF network is relatively low, in order to improve the model efficiency of the NeRF, the present invention reduces the cost of MLP multi-layer perceptron, and uses linear interpolation instead of MLP to maintain the reconstructed static information at each static 3D position, so as to store the features of the 3D scene in the static scene trainable grid structure, realizing low-cost static scene reconstruction.
[0110] Among them, for dynamic head modeling, the present invention proposes an improved NeRF model for real-time audio spatial decomposition, which can effectively train and synthesize the audio-driven human head in real time. The present invention uses inherent high-dimensional audio to synthesize human head information, which is explicitly decomposed into three low-dimensional trainable feature grids, namely lip motion model, human head motion model and eye blinking model. This model combining the explicit features and implicit features of the virtual anchor can more effectively synthesize natural and realistic virtual anchors. For dynamic head modeling, in order to realize the synchronization of audio and lip shape, expression and emotion, the present invention proposes a decomposed audio-space encoding module, which decomposes audio and spatial representation into two grids. When the static spatial coordinates are maintained in 3D, the audio dynamics are encoded as low-dimensional "coordinates". The advantage of this is that the synthesis of video is not querying the audio and spatial coordinates in a high-dimensional feature grid, but is divided into two independent low-dimensional feature grids, thereby reducing the cost of interpolation.
[0111] The relationship between lip motion and input audio is constructed. Here, the mouth motion of the embedded auditory speech is directly synchronized instead of using visual cues. Specifically, a CNN audio encoder E a The phoneme feature f is extracted from the input audio a , the expression is as follows:
[0112] f a = E a (a)
[0113] where a denotes the input audio data.
[0114] A contrastive learning strategy is adopted to align audio features and mouth features to seek their synchronization. Specifically, timely aligned audio and mouth features (f a , f m ) are considered as positive pairs, while misaligned pairs are considered as negative pairs. Binary cross-entropy loss is used for contrastive learning, where the distance between timely aligned positive pairs should be closer than misaligned negative pairs.
[0115]
[0116] τ con denotes the binary cross-entropy loss of lip shape and speech, d(f m , f a ) denotes the cosine distance of positive pairs, denotes the cosine distance of negative pairs.
[0117] A controllable probabilistic model is used for eye blinking and head pose motion, a facial attribute (head pose or eye blink) sequence h 1:T of length T and a longer conditioning audio sequence a 1:T′ of length T' are used, where T' > T; since T+1:T' of audio exists, it is necessary to generate a facial attribute sequence h T+1:T′ by embedding prediction from frame T' to T. The prediction of facial attribute sequence h T+1:T′ includes: (1) latent attribute space construction. A Transformer-VAE is trained on a large dataset using Gaussian Process (GP) to construct the mapping between input facial attribute sequence and a latent attribute space Z. (2) head pose and eye blink space construction, where a cross-modal encoder is fine-tuned to embed two BOP (Beginning of Pose) and eye blink frequency audio embeddings into the latent attribute space Z on selected characters.
[0118] After obtaining the generated head pose, eye blink feature f e and synchronized audio feature f a , a neural radiance field is used to generate the final image with these conditions. First, the synchronized audio embedding f a and eye blink embedding f e are concatenated into a new embedding f cThen, a conditional radiance field is proposed with this new embedding as input. After converting the head pose from camera space to canonical space, the head pose is directly used to replace the viewing direction d of the conditional radiance field. Finally, the embedding f in canonical space, the viewing direction d and the 3D position x form the input of the implicit function F θ . In fact, F θ is implemented by a multi-layer perceptron. For all connected input vectors, the implicit function F θ will estimate the color value c and the assigned radiance σ of the accompanying density. The whole implicit function can be formulated as:
[0119] F θ :(f,d,x)→(c,σ)
[0120] In this way, the synchronization process of audio and implicit attributes of blink frequency and head pose is realized.
[0121] Where the modeling of the dynamic torso is concerned, the present application proposes a model for the torso part motion, which pursues lower computational cost in a lightweight manner. Because the torso motion of the virtual anchor is less, the present application proposes a lightweight pseudo 3D deformable module to model the torso with a 2D feature grid. The module based on dynamic torso modeling can successfully simulate the dynamic characteristics of the torso and synthesize natural torso images that match the head. More importantly, the pseudo 3D representation of the 2D feature grid is very lightweight and efficient. The separately rendered head and torso images can be harmoniously synthesized with any provided background image to obtain the final output virtual anchor video.
[0122] Because the torso is almost static, only containing slight motion without topological changes, the method of the present application can be regarded as a two-dimensional version of NeRF based on deformation. The torso deformation is conditioned on the head pose p, so that the torso motion is synchronized with the head motion.
[0123] An MLP (Multi-Layer Perceptron) is used to predict the torso deformation:
[0124] Δx=MLP(x t ,p)
[0125] x t is a pixel coordinate sampled from the image space. Δx is the pixel coordinate after the torso deformation.
[0126] The pixel coordinates of the torso deformation are fed into a two-dimensional feature grid encoder to obtain the torso feature f t :
[0127]
[0128] Another MLP is used to generate the torso's RGB color and alpha value:
[0129] c t ,α t =MLP(f t i t )
[0130] Where i t It incorporates embedded latent features learned by the model, c t It is the RGB color of the torso, α t It is the alpha value.
[0131] The separately rendered head and torso models are combined with the static model to obtain the composite video of the virtual anchor.
[0132] Step six: The background replacement module replaces the background of the virtual anchor's synthesized video based on the background data.
[0133] like Figure 5 As shown, the background replacement module uses the background segmentation module to separate the foreground and background of the synthesized video, and the image harmonization module to fuse the foreground and the replacement background. Specifically:
[0134] (1) Foreground and background separation of the synthesized video utilizes the Background-Matting background segmentation model. This human segmentation model is first trained on two self-made large databases with significant diversity in human poses to learn robust prior knowledge. Then, it is trained on publicly available datasets that are manually managed to learn fine texture details.
[0135] Background-Matting is an image processing technique that replaces background objects by identifying the boundaries between them. First, the algorithm extracts the alpha channel of the foreground object, completely separating its transparency from the background. This alpha channel is used to identify the boundaries between foreground and background objects. Then, the separated foreground object, representing the virtual anchor, is synthesized into the background image using an image harmonization algorithm to complete background replacement. Combining voice and video creates a vivid virtual anchor. The model consists of two networks: a base network predicts an alpha mask and foreground layer at a lower resolution and outputs an incorrectly predicted patch indicating areas that may require high-resolution refinement; a refinement network takes the low-resolution result and the original image as input and generates high-resolution output only in selected areas, thus segmenting the person image from the video.
[0136] Background-Matting Background segmentation model is a real-time, high-resolution background replacement technology, which can achieve real-time segmentation on 4K (3840x2160) resolution at 30fps and HD (1920x1080) resolution at 60fps.
[0137] The background segmentation problem is modeled, and each pixel of the image is represented as a combination of foreground and background:
[0138] C = F * alpha + B * (1 - alpha)
[0139] C is the given image, F is the foreground of each pixel, B is the background of each pixel, and alpha is the transparency of each pixel.
[0140] (2) The fusion of foreground and replacement background refers to pasting the foreground of one picture onto another background picture through image harmonization technology to obtain a composite image.
[0141] However, the composite image obtained by simple splicing may have many problems, such as unreasonable size and position of the foreground, unreasonable perspective angle of the foreground, unnatural connection between the foreground and the background, color and lighting inharmony between the foreground and the background, etc. These factors will cause the quality of the composite image to decline and look unrealistic.
[0142] Therefore, image harmonization can effectively solve the problem of color and lighting inharmony between the foreground and the background, including sequentially adjusting the color and lighting of the foreground image, linearly transforming the brightness, histogram equalization, color correction, and local contrast enhancement. Among them, by adjusting the color and lighting of the foreground, it is more suitable for the background; by linearly transforming the brightness of the foreground image, the gray value distribution range of the image is more extensive; by histogram equalization, the gray histogram of the foreground image is equalized, the gray distribution of the image is more uniform, and the contrast and details of the image are enhanced; then through color correction, the color parameters (such as hue, saturation, color temperature, etc.) of the foreground image are adjusted to improve the color balance and coordination of the image; finally, through local contrast enhancement, the local area of the foreground image is enhanced, and the details of the image are more abundant; ultimately, the foreground and background replacement of the composite video is realized.
[0143] The image after background replacement is:
[0144]
[0145] Where the background image is I b , the foreground image is I f , the foreground image mask is M, and the combined image is I c , is the Hadamard product.
[0146] Step seven, audio and video fusion of the virtual anchor synthesized voice and the virtual anchor video after background replacement to synthesize the final virtual anchor.
[0147] The virtual anchor synthesized voice and the virtual anchor video after background replacement are taken as inputs and transmitted into the FFmpeg tool to combine audio and video and synthesize the final virtual anchor. The FFmpeg contains very advanced audio / video codec library and is an open source tool that can be used to record, convert digital audio, video and convert them into streams.
[0148] It can be known from the above method that the customer only needs to take the video containing the character, the voice and the background of the virtual anchor application scene as inputs, and the whole application can automatically generate a lively and natural virtual anchor, greatly reducing the time and cost of virtual anchor production, and being applicable to virtual anchor production in different fields.
[0149] It should be noted and understood that various modifications and improvements can be made to the application described in detail above without departing from the spirit and scope of the application required by the appended claims. Therefore, the scope of the claimed technical solution is not limited by any specific exemplary teaching given.
Claims
1. A virtual anchor generation method based on neural radiance fields and latent attributes, characterized in that, Specifically comprising the following steps: Step one, according to the actual needs to build the virtual anchor person image, including video data, voice and text data and background data, as the input of the virtual anchor generation system; Step two, the face feature extraction and construction module extracts and constructs the face features of the video data to generate a three-dimensional face of the virtual anchor; Step three, the voice data is used to train the voice synthesis network, and the text data is input into the text conversion module for text front-end processing, and then the processed text is input into the trained voice synthesis network model, and after voice synthesis, the synthesized voice of the virtual anchor is obtained; Step four, the voice feature extraction module extracts the features of the synthesized voice of the virtual anchor, and the explicit and implicit attribute feature extraction module extracts the explicit and implicit attribute feature information of the video data, and outputs the extracted feature information to the improved NeRF network module; Voice feature extraction: extracting the frequency spectrum, tone and pitch feature information of the synthesized voice of the virtual anchor, and mapping them to the corresponding discrete values; Explicit attribute feature extraction: using 3DMM to extract the lip movement, facial movement and expression feature data that are strongly related to the voice data from the video data, and outputting them to the improved NeRF network module as explicit attributes; Implicit attribute feature extraction: attributes that are weakly related to voice data, i.e. other attributes related to voice context or personalized conversation style, including head movement and blinking, which are extracted from the relevant parts of the video data through the constructed three-dimensional face model, and output to the improved NeRF network module as implicit attributes; Step five, the improved NeRF network module models the static scene, dynamic head and dynamic torso of the virtual anchor according to the voice feature information and explicit and implicit attribute feature information to obtain the synthesized video of the virtual anchor; Specifically: (1) When the improved NeRF network is used for static scene modeling, the MLP multi-layer perceptron is reduced, and the linear interpolation is used instead of the MLP, so as to keep the reconstructed static information at each static 3D position, and store the features of the 3D scene in the static scene trainable grid structure; (2) When the improved NeRF network is used for dynamic head modeling, the high-dimensional audio and video processing network of the virtual anchor is decomposed into three low-dimensional trainable feature grids, i.e. lip movement model, head movement model and eye blinking model; in order to realize the synchronization of audio and each movement model, the audio-spatial coding module is decomposed into 3D space grid and 2D audio grid, and the audio and spatial representation are decomposed into two grids; when each movement model keeps the static spatial coordinates in 3D, the audio dynamics is coded as low-dimensional "coordinates"; In constructing the relationship between the overt lip shape motion and the audio, the mouth motion of the auditory speech is directly synchronized and embedded; specifically, a CNN audio encoder is used extracting phoneme features from the input audio , which is expressed as follows: Where a represents the input audio data; A contrastive learning strategy is used to align audio features with mouth features, and the aligned audio and mouth features are then displayed in real time. , () is considered as a straight pair, not an aligned pair. , ) are considered negative pairs; contrastive learning is performed using binary cross-entropy loss, where timely aligned positive pairs are closer together than unaligned negative pairs; binary cross-entropy loss representing lip shape and speech, cosine distance representing positive pairs, cosine distance representing negative pairs; The synchronization process of audio with implicit attribute blinking frequency and head posture is as follows: A controllable probabilistic model is used for eye blink and head pose motion, a facial attribute sequence of length T and a conditioning audio sequence of length The facial attributes include head pose or eye blink; the need is to embed the predicted generated facial attribute sequence into a frame The prediction of the facial attribute sequence includes: (1) latent attribute space construction, training a Transformer-VAE on a large dataset to build a mapping between input facial attribute sequences and a latent attribute space Z using a Gaussian Process; (2) head pose and eye blink space construction, fine-tuning a cross-modal encoder to embed both head BOP and eye blink frequency audio embeddings into the latent attribute space Z on selected characters; After obtaining the generated head pose, blinking feature and synchronized audio feature , a neural radiance field is used to generate the final image with these conditions; first, the synchronized audio feature and the blinking feature are connected into a new feature ; then, with this new feature as input, a conditional radiance field is proposed; after converting the head pose from camera space to canonical space, the head pose is directly used to replace the observation direction d of the conditional radiance field; finally, the features f in the canonical space, the observation direction d and the 3D position x constitute the input of the implicit function ; for all input vectors, the implicit function can estimate the color value c accompanied by the density and the assigned light rays; Implicit function In formula: (3) When the improved NeRF network is used for dynamic torso modeling, another 2D grid is used to simulate the dynamic characteristics of the torso in a lightweight pseudo 3D deformable module, and natural torso images matching the head are synthesized; (4) Synthesize the separately rendered head and torso model with the static model to obtain a synthesized video of the virtual anchor; Step six, the background replacement module replaces the background of the virtual anchor synthesis video according to the background data, and fuses the virtual anchor character image, background and audio to synthesize the final virtual anchor; Step 601, input the virtual anchor synthesis video into the background segmentation module, and extract the Alpha channel of the foreground object in the image through the Background-Matting background segmentation model, so that the image of the virtual anchor in the synthesis video is completely separated from the background; Step 602, synthesize the segmented virtual anchor image into another background image to obtain a synthesis image, and perform harmonization processing on the synthesis image through the image harmonization module to complete the background replacement; The image after background replacement is: where the background image is , the foreground image is , the foreground image mask is M, and the combined image is , is a Hadamard product. Step 603, input the virtual anchor video after background replacement and the virtual anchor synthesis audio into the FFmpeg tool for audio and video combination to synthesize the final virtual anchor.
2. The method of claim 1, wherein, The face feature extraction and construction process in step two includes face analysis, 3DMM face feature extraction and face reconstruction, specifically: Step 201, face analysis is to decompose the video data of the person into face components through deep learning technology, and obtain the corresponding facial features; Step 202, 3DMM face feature extraction encodes the facial features of different parts of the face in three dimensions, and selects the combination of the reference shape and texture map that best represents the required face to synthesize the face; Step 203, the reference shape and texture map of the face are combined by weighting in the database to generate a reconstructed three-dimensional face.
3. The method of claim 2, wherein, The face components include skin, hair, eyes, eyebrows, nose and mouth.
4. The method of claim 1, wherein, The text front-end processing in step three refers to removing punctuation, numbers and spaces in the input text, then dividing the text into individual words according to semantic understanding, and converting each character into a phoneme and labeling a speech tag.
5. The method of claim 1, wherein, The torso modeling in step five is specifically: Torso morphs to head pose Conditional on the torso motion being synchronized with the head motion; An MLP is used to predict the deformation of the torso: refers to sampling pixel coordinates from the image space, refers to pixel coordinates after the torso deformation; The coordinates of the torso deformation are fed to a two-dimensional feature grid encoder to obtain torso features : Another MLP is used to generate the RGB color and alpha value of the torso: wherein is an embedding hidden feature added to the model learning, is a torso RGB color, is an alpha value.
6. The method of claim 1, wherein, The Background-Matting background segmentation model in step six includes a base network and a refinement network, where the base network predicts the Alpha mask and the foreground layer at low resolution, and outputs an error prediction block indicating the area that needs high-resolution refinement; The refinement network takes the low-resolution result and the original image as input, and only generates a high-resolution output in the indicated area to segment the video character image; The background segmentation problem is modeled by representing each pixel of the image as a combination of foreground and background: C is the given image, F is the foreground computed for each pixel, B is the background computed for each pixel, is the transparency of each pixel.
7. The method of claim 1, wherein, The harmonization processing in step six includes color and light adjustment, brightness linear transformation, gray histogram equalization, color correction and local contrast enhancement on the foreground image in turn.
8. A virtual anchor generation system based on neural radiance field and implicit attribute, which refers to the virtual anchor generation method of claim 1, characterized in that, The face feature extraction and construction module, the speech synthesis module, the speech feature extraction module, the explicit and implicit attribute extraction module, the improved NeRF network module and the background replacement module are included; The face feature extraction and construction module restores the three-dimensional shape of the two-dimensional face image through the 3DMM model; The speech synthesis module takes the input text of the virtual host as the input of speech synthesis, converts the text into phoneme information, and synthesizes the converted phonemes into the speech of the virtual host through an acoustic synthesis model; The speech feature extraction module and the explicit and implicit attribute extraction module extract the speech, blinking, and lip shape features of the virtual host; The improved NeRF network module decomposes the high-dimensional audio into three low-dimensional trainable feature grids using an improved NeRF network, and synthesizes eye blinking, head posture, and lip movement; The background replacement module includes a background segmentation module and an image harmonization module, which realizes the background replacement of the virtual host.
Citation Information
Patent Citations
High-quality face voice driving method based on neural radiation field
CN112887698A
Construction method and device of deformable neural radiation field network
CN115909015A