Face video communication method and device, electronic equipment and storage medium
By classifying facial video frame sequences into key reference frames and inter-frames, and mapping the inter-frames to a three-dimensional parameter space to filter out semantic parameters for encoding, the need for low bit rate transmission and high-quality reconstruction in video communication is solved, realizing efficient encoding of facial dynamics and real-time interactive control at the decoding end.
Patent Information
- Application Number
- CN202610607341.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies are insufficient to meet the demands of communication scenarios such as video conferencing and virtual live streaming for ultra-low bit rate transmission, semantic-level interactive control, and high-quality reconstruction.
The face video frame sequence is classified into key reference frames and inter-frames. The key reference frames are compressed and encoded. The inter-frames are mapped from two-dimensional image frames to three-dimensional parameter space to filter out three-dimensional facial semantic representations such as mouth movement, eye blinking, head rotation, and head position, and then encoded to generate a transmission bitstream.
It achieves a balance between low bit rate transmission and high-quality reconstruction, supports real-time interaction of facial expressions and postures directly controlled by the decoding end, reduces the amount of transmitted data and improves reconstruction quality.
Smart Images

Figure CN122640554A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of communication technology, specifically to the field of semantic communication technology, and in particular to face video communication methods, devices, electronic devices, and storage media. Background Technology
[0002] In the field of semantic communication and face video coding technology, existing technologies mainly cover three types of solutions: general video coding, model coding for speaking faces, and 3D face modeling based on deep learning. However, all of them have significant technical limitations and cannot meet the needs of communication scenarios such as video conferencing and virtual live streaming for ultra-low bit rate transmission, semantic-level interactive control, and high-quality reconstruction. Summary of the Invention
[0003] This disclosure provides a face video communication method, apparatus, electronic device, and storage medium.
[0004] According to one aspect of this disclosure, a face video communication method is provided, applied at an encoder end, the method comprising: The input face video frame sequence is classified to obtain key reference frames and inter-frames, and the key reference frames are compressed and encoded to obtain key reference frame encoded data. For the inter-frame, the inter-frame is mapped from a two-dimensional image frame to a three-dimensional parameter space, and a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters and head position parameters is obtained by filtering. The correlation between different parameters satisfies the preset independence condition. The three-dimensional facial semantic representation is encoded to obtain inter-frame encoded data; The key reference frame encoded data and the inter-frame encoded data are integrated into a transmission stream, and the transmission stream is sent to the decoder. The decoder decodes the received transmission stream to obtain the 3D facial semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. Based on the 3D facial semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, an inter-frame reconstructed facial image frame is generated. In response to an interaction command, the parameters in the 3D facial semantic representation of the reconstructed inter-frame are modified to generate an inter-frame reconstructed facial image frame corresponding to the interaction command.
[0005] According to another aspect of this disclosure, a face video communication method is provided, applied at a decoder end, the method comprising: The system receives a transmission stream from an encoder, which classifies the input face video frame sequence to obtain key reference frames and inter-frames. The key reference frames are then compressed and encoded to obtain key reference frame encoded data. For the inter-frames, they are mapped from two-dimensional image frames to a three-dimensional parameter space, and a three-dimensional face semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters is obtained. The correlation between different parameters satisfies a preset independence condition. The three-dimensional face semantic representation is encoded to obtain inter-frame encoded data. The key reference frame encoded data and the inter-frame encoded data are integrated into a transmission stream, which is then sent to the decoder. The received transmission stream is decoded to obtain the three-dimensional face semantic representation of the reconstructed key reference frame and the inter-reconstruction frame; Based on the reconstructed key reference frame and the reconstructed inter-frame 3D face semantic representation, an inter-frame reconstructed face image frame is generated; In response to the interaction command, the parameters in the three-dimensional face semantic representation of the reconstructed inter-frame are modified to generate an inter-frame reconstructed face image frame corresponding to the interaction command.
[0006] According to another aspect of this disclosure, a face video communication method is provided, comprising: The encoder classifies the input facial video frame sequence to obtain key reference frames and inter-frames, and compresses and encodes the key reference frames to obtain key reference frame encoded data. For the inter-frames, the inter-frames are mapped from two-dimensional image frames to a three-dimensional parameter space, and a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters is obtained. The correlation between different parameters satisfies a preset independence condition. The three-dimensional facial semantic representation is encoded to obtain inter-frame encoded data. The key reference frame encoded data and the inter-frame encoded data are integrated into a transmission stream, and the transmission stream is sent to the decoder. The decoder decodes the received transmission stream to obtain the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame; based on the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, it generates an inter-frame reconstructed face image frame; in response to an interaction command, it modifies the parameters in the 3D face semantic representation of the reconstructed inter-frame to generate an inter-frame reconstructed face image frame corresponding to the interaction command.
[0007] According to a fourth aspect of this disclosure, a face video communication device is provided, applied at an encoder end, the device comprising: The classification module is used to classify the input face video frame sequence to obtain key reference frames and inter-frames, and to compress and encode the key reference frames to obtain key reference frame encoded data. The mapping module is used to map the inter-frame from a two-dimensional image frame to a three-dimensional parameter space for the inter-frame, and to filter to obtain a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters and head position parameters, wherein the correlation between different parameters satisfies a preset independence condition. The encoding module is used to encode the three-dimensional facial semantic representation to obtain inter-frame encoded data; The transmission module is used to integrate the key reference frame encoded data and the inter-frame encoded data into a transmission bitstream, and send the transmission bitstream to the decoder. The decoder is used to decode the received transmission bitstream to obtain the 3D facial semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. Based on the 3D facial semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, an inter-frame reconstructed facial image frame is generated. In response to an interaction command, the parameters in the 3D facial semantic representation of the reconstructed inter-frame are modified to generate an inter-frame reconstructed facial image frame corresponding to the interaction command.
[0008] According to a fifth aspect of this disclosure, a face video communication device is provided, applied at a decoder end, the device comprising: A receiving module is used to receive the transmission bitstream from the encoder. The encoder classifies the input face video frame sequence to obtain key reference frames and inter-frames, and compresses and encodes the key reference frames to obtain key reference frame encoded data. For the inter-frames, the inter-frames are mapped from two-dimensional image frames to a three-dimensional parameter space, and a three-dimensional face semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters is obtained. The correlation between different parameters satisfies a preset independence condition. The three-dimensional face semantic representation is encoded to obtain inter-frame encoded data. The key reference frame encoded data and the inter-frame encoded data are integrated into a transmission bitstream, and the transmission bitstream is sent to the decoder. The decoding module is used to decode the received transmission stream to obtain the three-dimensional face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. The reconstruction module is used to generate inter-frame reconstructed face image frames based on the three-dimensional face semantic representation of the reconstruction key reference frame and the reconstruction inter-frame; An interaction module is used to modify the parameters in the three-dimensional face semantic representation of the reconstructed inter-frame in response to an interaction command, and generate an inter-frame reconstructed face image frame corresponding to the interaction command.
[0009] According to a sixth aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in any of the above technical solutions.
[0010] According to a seventh aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any one of the methods described above.
[0011] According to the eighth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in any one of the above technical solutions.
[0012] This disclosure provides a face video communication method, apparatus, device, and storage medium. By classifying the input face video frame sequence into key reference frames and inter-frames and encoding them separately, a significant reduction in transmission bitrate and a balance between transmission quality and reconstruction quality are achieved. Specifically, the key reference frames are compressed and encoded to retain complete texture information, providing a high-quality identity and texture benchmark for the entire video. The inter-frames are mapped from two-dimensional image frames to a three-dimensional parameter space, filtering to obtain a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters. The correlation between different parameters meets a preset independence condition, allowing the inter-frames to transmit only low-dimensional semantic parameters instead of complete image data, significantly compressing the amount of transmitted data. Simultaneously, because the parameters meet the preset independence condition, each semantic parameter can independently represent the dynamic changes of different facial regions, avoiding redundant transmission caused by parameter coupling. Next, the decoder decodes the received transmission stream to obtain the 3D facial semantic representation of the reconstructed key reference frame and the inter-frame reconstruction. Based on the texture reference of the reconstructed key reference frame and the dynamic driving of the inter-frame reconstruction semantic representation, it generates inter-frame reconstructed facial image frames, ensuring that the reconstructed frames retain the high-fidelity texture of the key reference frame and accurately reproduce the facial dynamics of the inter-frame reconstruction. Furthermore, since the transmitted parameters are decoupled and independent, the decoder can directly respond to interactive commands to modify the parameters in the 3D facial semantic representation of the inter-frame reconstruction and regenerate the inter-frame reconstructed facial image frames without modifying the encoder or adding additional bitstreams. This achieves real-time interactive capability of directly controlling facial expressions and poses at the decoder, achieving a unified technical effect of low bitrate transmission, high-quality reconstruction, and semantic-level interactive control.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic diagram of the steps of a face video communication method according to an embodiment of the present disclosure; Figure 2 This is a schematic diagram of the encoder end process of a face video communication method according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the steps of a face video communication method in another embodiment of this disclosure; Figure 4 This is a schematic diagram of the decoder-side process of a face video communication method according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the overall process of a face video communication method according to an embodiment of this disclosure; Figure 6 This is a schematic block diagram of the face video communication device in the embodiments of this disclosure; Figure 7 This is a schematic block diagram of the face video communication device in the embodiments of this disclosure; Figure 8 This is a block diagram of an electronic device used to implement the face video communication method of the embodiments of this disclosure. Detailed Implementation
[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0016] This disclosure provides a face video communication method, which is applied to the encoder end. See [link to relevant documentation]. Figure 1 As shown, it includes: Step S101: Classify the input face video frame sequence to obtain key reference frames and inter-frames, and compress and encode the key reference frames to obtain key reference frame encoded data.
[0017] Specifically, key reference frames refer to specific frames in a video sequence that serve as the foundation for texture and identity anchoring. Their selection typically involves calculating the differences in 3D facial semantic representation between adjacent frames. The proportion of key reference frames in the entire video is usually kept low (e.g., no more than 5%) to ensure transmission efficiency. Interframes refer to the remaining video frames between adjacent key reference frames, whose facial dynamic changes are relatively subtle and do not require complete image data transmission. For the key reference frames obtained from classification, intra-frame compression coding standards (such as VVC intra-frame coding standard) are used to compress and encode them. Spatial redundancy is eliminated through transform coding, data precision is reduced through quantization, and the bitstream is further compressed through entropy coding, ultimately yielding the encoded key reference frame data. This data retains high-quality texture information, identity coefficients, reflectivity coefficients, and illumination coefficients of the key reference frames, providing a stable texture template and identity benchmark for the reconstruction of subsequent interframes. It also allows the decoder to extract fixed parameters from the reconstructed key reference frames and store them in a buffer for reuse during 3D face mesh reconstruction.
[0018] Step S102: For the inter-frame, the inter-frame is mapped from the two-dimensional image frame to the three-dimensional parameter space, and a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters and head position parameters is obtained. The correlation between different parameters satisfies the preset independence condition.
[0019] Specifically, inter-frames refer to video frames in a face video sequence located between adjacent key reference frames, primarily carrying the dynamic changes of the face. They exist in the form of a two-dimensional pixel matrix, containing rich texture and geometric information but with high redundancy. The three-dimensional parameter space refers to a low-dimensional parameterized representation space constructed based on a three-dimensional deformation model (such as 3DMM). This space decouples the geometric shape, texture attributes, and dynamic expressions of the face into a set of independently manipulated parameters. The specific implementation process of this scheme includes: firstly, using a pre-trained three-dimensional face reconstruction model (such as the WM3DR model) to perform parameterized prediction on the two-dimensional image frames of the inter-frames, regressing to obtain the shape parameters, texture parameters, and expression parameters of the three-dimensional face; simultaneously, a facial behavior analysis model (such as OpenFace) is introduced to specifically predict the blinking parameters of the eyes to compensate for the shortcomings of the three-dimensional face reconstruction model in representing subtle eye movements. Furthermore, core parameter dimensions closely related to facial dynamic semantics were selected from the above prediction results to form a three-dimensional facial semantic representation, which specifically includes: mouth movement parameters (6-dimensional, taken from the first 6 dimensions of the lip movement components corresponding to up-down, left-right, and forward-backward movements in the expression parameters, accurately representing the changes in mouth shape during speech), eye blinking parameters (1-dimensional, representing the continuous state of the eyes from fully open to fully closed), head rotation parameters (3-dimensional, corresponding to the rotation arcs in the X, Y, and Z axes in three-dimensional space, representing the head turning, nodding, and tilting postures), head translation parameters (3-dimensional, corresponding to the pixel translation amount in the X, Y, and Z axes in three-dimensional space, representing the position movement of the head in the image), and head position parameters (1-dimensional, representing the relative height position of the head in the entire image, assisting in the accurate positioning of the head area), totaling 14 dimensions. To ensure that the semantic parameters do not interfere with each other and can be edited independently when representing facial dynamics, a preset independence condition constraint is applied to the correlation between different parameters. This ensures the high degree of independence of different semantic dimensions such as mouth movement, eye blinking, and head posture, laying the foundation for low redundancy in subsequent encoding and transmission, as well as independent modification and interactive control of semantic parameters at the decoder end.
[0020] Step S103: Encode the three-dimensional face semantic representation to obtain inter-frame encoded data.
[0021] Specifically, 3D facial semantic representation refers to a low-dimensional, compact set of parameters extracted by mapping inter-frames from 2D image frames to a 3D parameter space. This set includes mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters. It exists as a floating-point vector, exhibiting high data precision, and its dimensions are decoupled through pre-defined independence conditions. Inter-frame encoded data refers to the binary bitstream generated after compression encoding, which can be directly integrated into the transmission stream for network transmission.
[0022] Step S104: Integrate the key reference frame encoded data and inter-frame encoded data into a transmission bitstream, and send the transmission bitstream to the decoder. The decoder decodes the received transmission bitstream to obtain the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. Based on the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, generate an inter-frame reconstructed face image frame. In response to the interaction command, modify the parameters in the 3D face semantic representation of the reconstructed inter-frame to generate the inter-frame reconstructed face image frame corresponding to the interaction command.
[0023] Specifically, the transport bitstream refers to a structured binary data stream carrying compressed video data. Its structure includes a key reference frame index table (recording the position of key reference frames in the video sequence), a fixed parameter block (stores identity coefficients, reflectivity coefficients, and illumination coefficients extracted from the key reference frames; this block is transmitted only once throughout the entire video), and an inter-frame residual coding block (stores the prediction residual coding data for each inter-frame). It supports random access and semantic editing tags, facilitating rapid location and interactive processing by the decoder. Upon receiving the transport bitstream, the decoder first performs separate decoding to obtain the 3D facial semantic representation of the reconstructed key reference frames and the reconstructed inter-frames. Subsequently, based on the 3D facial semantic representation of the reconstructed key reference frames and the reconstructed inter-frames, the decoder generates inter-frame reconstructed facial image frames. Furthermore, since the transport bitstream transmits decoupled independent semantic parameters rather than complete image data, the decoder can directly respond to interactive commands to modify the parameters in the 3D facial semantic representation of the reconstructed inter-frames, enabling independent and controllable adjustments to mouth shape, eye state, or head posture.
[0024] This disclosure provides a face video communication method, apparatus, device, and storage medium. By classifying the input face video frame sequence into key reference frames and inter-frames and encoding them separately, a significant reduction in transmission bitrate and a balance between transmission quality and reconstruction quality are achieved. Specifically, the key reference frames are compressed and encoded to retain complete texture information, providing a high-quality identity and texture benchmark for the entire video. The inter-frames are mapped from two-dimensional image frames to a three-dimensional parameter space, filtering to obtain a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters. The correlation between different parameters meets a preset independence condition, allowing the inter-frames to transmit only low-dimensional semantic parameters instead of complete image data, significantly compressing the amount of transmitted data. Simultaneously, because the parameters meet the preset independence condition, each semantic parameter can independently represent the dynamic changes of different facial regions, avoiding redundant transmission caused by parameter coupling. Next, the decoder decodes the received transmission stream to obtain the 3D facial semantic representation of the reconstructed key reference frame and the inter-frame reconstruction. Based on the texture reference of the reconstructed key reference frame and the dynamic driving of the inter-frame reconstruction semantic representation, it generates inter-frame reconstructed facial image frames, ensuring that the reconstructed frames retain the high-fidelity texture of the key reference frame and accurately reproduce the facial dynamics of the inter-frame reconstruction. Furthermore, since the transmitted parameters are decoupled and independent, the decoder can directly respond to interactive commands to modify the parameters in the 3D facial semantic representation of the inter-frame reconstruction and regenerate the inter-frame reconstructed facial image frames without modifying the encoder or adding additional bitstreams. This achieves real-time interactive capability of directly controlling facial expressions and poses at the decoder, achieving a unified technical effect of low bitrate transmission, high-quality reconstruction, and semantic-level interactive control.
[0025] In some optional embodiments, the inter-frames are mapped from two-dimensional image frames to a three-dimensional parameter space, and a three-dimensional facial semantic representation including mouth motion parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters is obtained, including: A pre-trained 3D face reconstruction model is used to predict the shape parameters, texture parameters, and dynamic expression parameters of the 3D face from the 2D face images between frames, thus obtaining the initial 3D face parameters. A facial behavior analysis model is used to predict eye blinking parameters from inter-frame two-dimensional face images; The parameters representing the lip movement components are extracted from the dynamic expression parameters in the initial three-dimensional face parameters to obtain the mouth movement parameters; The head rotation parameters, head translation parameters, and head position parameters are extracted from the initial 3D face parameters. The head rotation parameters correspond to the rotation radians in each coordinate axis direction in 3D space, the head translation parameters correspond to the pixel translation amount in each coordinate axis direction in 3D space, and the head position parameters represent the relative height position of the head in the image. By integrating mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters, a three-dimensional facial semantic representation is obtained.
[0026] Specifically, this technical solution maps two-dimensional face images from inter-frame sequences to a three-dimensional parameter space and integrates them to obtain a three-dimensional face semantic representation through the collaborative work of pre-trained models. Inter-frame sequences refer to video frames in a face video sequence that lie between adjacent key reference frames and primarily carry facial dynamic changes; they exist in the form of a two-dimensional pixel matrix. The three-dimensional face semantic representation refers to a low-dimensional parameter set selected from the three-dimensional parameter space to compactly represent the dynamic semantics of the face.
[0027] The specific implementation process of this scheme includes: First, a pre-trained 3D face reconstruction model (such as the WM3DR model) is used to parametrically predict the 2D face images between frames. This model is based on an encoder-decoder architecture, extracting multi-scale features of the 2D image through a convolutional neural network, and then regressing the shape parameters, texture parameters, and dynamic expression parameters of the 3D face through a fully connected layer to obtain the initial 3D face parameters. Among them, the shape parameters control the face geometry, the texture parameters control skin color and reflectivity, and the dynamic expression parameters control facial muscle movement. Simultaneously, a facial behavior analysis model (such as the OpenFace model) is used to perform specialized eye region detection and behavior analysis on the 2D face images of the same frame. By locating key points of the eye contour and tracking the degree of eyelid opening and closing, the blinking parameters are predicted. This parameter, represented by a single-dimensional scalar, depicts the continuous state of the eyes from fully open to fully closed, compensating for the shortcomings of 3D face reconstruction models in representing subtle eye movements. Furthermore, dimensional filtering is performed on the dynamic expression parameters from the initial 3D face parameters to extract the first six dimensions representing the movement components of the lips in different directions (up, down, left, right, forward, and backward), thus obtaining the mouth movement parameters. This is used to accurately depict changes in lip shape during speech. Head rotation parameters are directly extracted from the initial 3D facial parameters. Head translation parameters and head position parameters The head rotation parameter is a 3D vector, corresponding to the rotation radians along the X, Y, and Z axes in 3D space, representing changes in head posture such as turning left and right, nodding up and down, and tilting to the side. The head translation parameter is a 3D vector, corresponding to the pixel translation amount along the X, Y, and Z axes in 3D space, representing the head's forward / backward, left / right, and up / down movement within the image. The head position parameter is a 1D scalar, representing the relative height of the head within the entire image, aiding in precise head region localization. Finally, the above mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters are integrated at the parameter level to form a total of 14-dimensional 3D facial semantic representation. This representation, in the form of a compact vector, fully covers the core dimensions of facial dynamic semantics, providing a data foundation for subsequent low-bitrate encoded transmission and semantic-level interactive control at the decoder end.
[0028] In this way, a pre-trained 3D face reconstruction model predicts the shape, texture, and dynamic expression parameters of a 3D face from inter-frame 2D face images, obtaining initial 3D face parameters. This decouples the high-dimensional image data, originally existing as a pixel matrix, into interpretable geometric, texture, and expression components, significantly compressing the data dimensionality. Simultaneously, a facial behavior analysis model is introduced to specifically predict blinking parameters, compensating for the shortcomings of general 3D face reconstruction models in representing subtle eye movements and ensuring the accuracy of eye dynamic semantics. Next, mouth motion parameters representing lip movement components are extracted from the dynamic expression parameters, along with head rotation, translation, and position parameters. This decomposes complex facial dynamic changes into independent semantic dimensions such as mouth, eyes, and head posture, each parameter corresponding to a clear physical meaning and visual representation. Finally, these parameters are integrated into a 3D face semantic representation, allowing facial dynamics to be reconstructed at the decoder only by transmitting low-dimensional semantic parameters, eliminating the need to transmit complete 2D image data between frames, thus achieving ultra-low bitrate transmission. In addition, since each semantic parameter is predicted separately from different pre-trained models and integrated based on explicit physical semantics, the coupling between parameters is significantly reduced, laying the foundation for satisfying the subsequent preset independence conditions. This also enables the decoder to directly modify individual semantic parameters independently without understanding complex image processing procedures, thus achieving precise control over mouth shape, eye state, and head posture, and achieving a technical effect that unifies low bit rate, high precision, and strong interactivity.
[0029] In some optional embodiments, mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters are integrated to obtain a three-dimensional facial semantic representation, including: The initial semantic representation is obtained by integrating the mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters. Principal component analysis spatial projection processing is performed on the initial semantic representation to make the correlation between parameters less than or equal to a preset threshold, so as to meet the preset independence condition and obtain a three-dimensional face semantic representation.
[0030] Specifically, the initial semantic representation refers to the original parameter vector formed by concatenating mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters according to their dimensions. Its dimension is the sum of the dimensions of each parameter (e.g., 14 dimensions). Although this representation covers the core dimensions of facial dynamic semantics, residual correlations may still exist between parameters during the prediction process, leading to mutual interference, redundant transmission, and difficulties in independent editing. The pre-set independence condition ensures, through mathematical constraints, that each semantic parameter is statistically uncorrelated or has extremely low correlation, allowing each parameter dimension to independently represent specific facial dynamic semantics, providing a foundation for subsequent low-redundancy coding and independent interactive editing. Principal Component Analysis (PCA) spatial projection processing refers to performing a linear transformation on the initial semantic representation using the PCA algorithm. By calculating the covariance matrix between parameters and solving for its eigenvalues and eigenvectors, the direction with the largest variance is selected as a new orthogonal basis. The original parameter vector is projected onto this orthogonal basis, ensuring that the covariance between each dimension after projection is zero or close to zero.
[0031] The specific implementation process of this scheme includes: First, mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters are concatenated in a fixed order to obtain an initial semantic representation. Then, the covariance matrix of this initial semantic representation on the training dataset is calculated, and eigenvalue decomposition is used to obtain an eigenvector matrix and an eigenvalue diagonal matrix. The eigenvectors corresponding to the k largest eigenvalues are selected to form a projection matrix. The initial semantic representation is multiplied by this projection matrix to obtain the projected semantic representation. By adjusting the number of retained principal components or setting a correlation threshold (e.g., 0.1), the correlation coefficient between any two parameter dimensions after projection is made less than or equal to the preset threshold, thus satisfying the preset independence condition. After principal component analysis spatial projection processing, the initial semantic representation, which may have had residual correlation, is transformed into a new representation with highly orthogonal dimensions. Each parameter dimension independently carries specific facial dynamic information, avoiding mutual interference between mouth movement and head posture semantics.
[0032] In this way, by integrating mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters, an initial semantic representation is obtained. This allows the multidimensional semantic information of facial dynamics to be aggregated into a unified vector form, facilitating subsequent overall processing and encoding. Principal component analysis (PCA) spatial projection processing is then performed on the initial semantic representation. After PCA spatial projection processing, the coupling between parameters is completely eliminated, and each dimension independently carries specific facial dynamic semantics. This allows redundant information to be removed during the encoding stage, further compressing the bitrate. Simultaneously, the decoder can directly modify individual semantic parameters without affecting the representation effect of other parameters, achieving the dual technical effects of low-redundancy transmission and precise independent control.
[0033] In some optional embodiments, the input sequence of face video frames is classified to obtain key reference frames and inter-frames, including: Calculate the differences in three-dimensional facial semantic representations between adjacent frames in a face video frame sequence, including differences in head rotation parameters and mouth motion parameters; When the difference in head rotation parameters exceeds the first preset threshold, or the difference in mouth movement parameters exceeds the second preset threshold, the current frame is marked as a critical reference frame. The remaining frames that were not marked as key reference frames are identified as interframes.
[0034] Specifically, this technical solution achieves adaptive classification of facial video frame sequences by calculating the differences in 3D facial semantic representations between adjacent frames and setting a dual-threshold judgment mechanism to distinguish key reference frames from inter-frame frames. The 3D facial semantic representation difference refers to the distance metric between semantic parameter vectors in the 3D parameter space of two adjacent frames, reflecting the intensity of facial dynamic changes. Specifically, the head rotation parameter difference refers to the absolute value or Euclidean distance of the difference in head rotation parameters (corresponding to the rotation radians in the X, Y, and Z axes of 3D space) between adjacent frames, used to quantify the amplitude of head posture changes. The mouth motion parameter difference refers to the absolute value or Euclidean distance of the difference in mouth motion parameters (corresponding to the 6-dimensional vectors of lip movement components in different directions such as up, down, left, right, front, and back) between adjacent frames, used to quantify the amplitude of lip shape changes during speech.
[0035] The specific implementation of this scheme includes: First, calculating the 3D facial semantic representation differences between adjacent frames in chronological order for the input facial video frame sequence. Specifically, this involves extracting the head rotation parameters of the current frame and the previous frame and calculating their differences to obtain the head rotation parameter difference; and extracting the mouth motion parameters of the current frame and the previous frame and calculating their differences to obtain the mouth motion parameter difference. Then, the head rotation parameter difference is compared with a first preset threshold, and the mouth motion parameter difference is compared with a second preset threshold. When the head rotation parameter difference exceeds the first preset threshold (e.g., significant head turning, nodding, or tilting) or the mouth motion parameter difference exceeds the second preset threshold (e.g., significant changes in lip movements), it indicates that the current frame has undergone drastic facial dynamic changes relative to the previous frame. If the semantic residual is still transmitted in inter-frame format, it will lead to the accumulation of reconstruction errors or prediction failure. Therefore, the current frame is marked as a key reference frame, and it is fully encoded using an intra-frame compression coding standard to provide a new high-quality texture benchmark. Conversely, when the difference in head rotation parameters does not exceed the first preset threshold and the difference in mouth movement parameters does not exceed the second preset threshold, it indicates that the facial dynamics in the current frame are smooth and can be efficiently encoded through semantic residual prediction; therefore, it is not marked as a key reference frame. Finally, the remaining frames in the video sequence that are not marked as key reference frames are determined as interframes. Interframes employ residual coding based on 3D facial semantic representation, using the semantic parameters of adjacent key reference frames or preceding interframes as the prediction benchmark to calculate residuals, achieving low bitrate transmission.
[0036] In this way, by calculating the differences in 3D facial semantic representations between adjacent frames, including differences in head rotation parameters and mouth motion parameters, the evaluation of inter-frame changes is directly based on low-dimensional semantic parameters rather than the original pixels, significantly reducing computational complexity. When the difference in head rotation parameters exceeds a first preset threshold or the difference in mouth motion parameters exceeds a second preset threshold, the current frame is marked as a key reference frame. This allows for the timely insertion of fully encoded key reference frames when there are significant changes in head posture or drastic changes in lip movements, providing new high-quality texture benchmarks and prediction starting points for subsequent inter-frames, effectively suppressing prediction failures and error accumulation caused by drastic facial dynamics. The remaining frames not marked as key reference frames are determined as inter-frames, allowing frames with smooth facial dynamics to use semantic residual-based predictive coding. The difference is calculated based on the semantic parameters of adjacent key reference frames or previous inter-frames, requiring only the transmission of low-dimensional residual data instead of the complete image. Since the selection of key reference frames takes into account both head rotation parameter differences and mouth motion parameter differences—the two most influential dynamic factors in visual perception—rather than using fixed intervals, the distribution of key reference frames adaptively matches the intensity of facial dynamic changes. When the dynamics are intense, the density of key reference frames is increased to ensure reconstruction accuracy, while the proportion of time frames in the dynamics are smooth is increased to reduce the bit rate. This avoids the waste of redundant key frames caused by fixed interval selection or the error spread caused by insufficient key frames, and achieves optimal allocation of bit rate resources and stable guarantee of reconstruction quality.
[0037] In some optional embodiments, the three-dimensional facial semantic representation is encoded to obtain inter-frame encoded data, including: Using the 3D face semantic representation of the key reference frame or the 3D face semantic representation of the preceding inter-frame as the prediction benchmark, the 3D face semantic representation of the current inter-frame is predicted to obtain the prediction residual. The predicted residuals are quantized and converted into binary code; Binary codes are encoded using a context-based arithmetic coding model to generate inter-frame encoded data.
[0038] Specifically, the prediction baseline refers to the prior reference data used to estimate the semantic parameters of the current inter-frame. This is specifically the 3D facial semantic representation of a key reference frame or the 3D facial semantic representation of a preceding inter-frame. This baseline carries the dynamic state of the face closest to the current frame, ensuring that the prediction residual only reflects subtle changes between frames rather than the complete semantics. The prediction residual is the difference vector between the 3D facial semantic representation of the current inter-frame and the corresponding parameters of the prediction baseline. Its numerical distribution is typically concentrated near zero, and its data redundancy is significantly lower than the original semantic representation. The context arithmetic coding model is an entropy coding model based on the Predictive Partial Matching (PPM) algorithm. By constructing a context window and adaptively adjusting the symbol probability distribution using historical data statistical patterns, it achieves compression efficiency close to the information theory limit.
[0039] The specific implementation process of this scheme includes: First, determining the prediction reference source based on the position of the current inter-frame in the video sequence: if it is the first inter-frame, the 3D face semantic representation of the key reference frame is used as the prediction reference, and the difference between the semantic representation of the current inter-frame and the semantic representation of the key reference frame is calculated to obtain the prediction residual; if it is a subsequent inter-frame, the 3D face semantic representation of the preceding inter-frame is used as the prediction reference, and the difference between the semantic representation of the current inter-frame and the semantic representation of the preceding inter-frame is calculated to obtain the prediction residual. Then, quantization processing is performed on the prediction residual, using quantization strategies such as the zero-order exponential Golomb algorithm to map the floating-point residual to discrete integer values. The quantization step size is adaptively set according to the residual distribution characteristics to ensure that the quantization error is controllable. Subsequently, the quantized integer values are converted into binary code representations. Finally, entropy encoding is performed on the binary code using a context arithmetic coding model, constructing a context window (e.g., setting the window size to 8), statistically analyzing the frequency of symbols appearing in historical encoded data, dynamically updating the probability model, allocating fewer bits to short codes with high occurrence probabilities, and more bits to long codes with low occurrence probabilities, ultimately generating inter-frame encoded data. This encoding process fully utilizes the temporal correlation of inter-frame semantic parameters and the statistical characteristics of residual distribution, enabling inter-frame encoded data to compactly store facial dynamic changes at a bit rate much lower than that of the original image encoding, while maintaining the complete reversibility that the decoding end can accurately recover the original semantic parameters by superimposing the prediction benchmark and residual.
[0040] In this way, by using the 3D facial semantic representation of the key reference frame or the 3D facial semantic representation of the preceding inter-frame as the prediction benchmark, the prediction residual is obtained by predicting the 3D facial semantic representation of the current inter-frame. This allows the current inter-frame to transmit only the difference with the benchmark of the neighboring frame, without needing to transmit the complete semantic parameters. The numerical distribution of the prediction residual is usually concentrated near zero, resulting in significantly lower data redundancy than the original semantic representation. Next, the prediction residual is quantized and converted into binary code. By mapping the floating-point residual to discrete integer values and converting them into binary representation, the data precision is further compressed and adapted to digital transmission. Finally, the binary code is encoded using a contextual arithmetic coding model to generate inter-frame coded data. The symbol probability distribution is adaptively adjusted using historical data statistical patterns, allocating fewer bits to high-frequency short codes and more bits to low-frequency long codes, approaching the information theory compression limit. Because the prediction benchmark is selected from the semantic representation of the reference frame closest to the current frame's dynamic state, the prediction residual energy is minimized, the quantization error is controllable, and contextual arithmetic coding can fully utilize the statistical characteristics of the residual distribution to achieve efficient compression. Compared to directly encoding the original two-dimensional image frame or complete semantic parameters, the bit rate of inter-frame coded data is significantly reduced. At the same time, the decoding end can accurately recover the original semantic parameters by superimposing the prediction benchmark and the decoding residual, ensuring reconstruction accuracy and achieving a balance between ultra-low bit rate transmission and high-quality reconstruction.
[0041] In some optional embodiments, the 3D face semantic representation of the current inter-frame is predicted based on the 3D face semantic representation of the key reference frame or the 3D face semantic representation of the preceding inter-frame, to obtain the prediction residual, including: When processing the first inter-frame, the three-dimensional face semantic representation of the key reference frame is obtained as the prediction benchmark. The difference between the three-dimensional face semantic representation of the first inter-frame and the three-dimensional face semantic representation of the key reference frame is calculated to obtain the first prediction residual. When processing subsequent frames between the first and second frames, the 3D face semantic representation of the previous frame is obtained as the prediction benchmark. The difference between the 3D face semantic representation of the current frame and the 3D face semantic representation of the previous frame is calculated to obtain the second prediction residual.
[0042] Specifically, the first prediction residual refers to the difference vector between the semantic parameters of the current inter-frame and the semantic parameters of the key reference frame when processing the first inter-frame in the video sequence; the second prediction residual refers to the difference vector between the semantic parameters of the current inter-frame and the semantic parameters of the previous inter-frame when processing subsequent inter-frames after the first inter-frame.
[0043] The specific implementation process of this scheme includes: When processing the first inter-frame, since this frame is the first dynamic frame after the key reference frame, and there are no encoded preceding inter-frames for reference, the 3D facial semantic representation of the key reference frame is obtained as the prediction benchmark. This benchmark carries the most complete static and dynamic facial information of the beginning segment of the video. The difference between the 3D facial semantic representation of the first inter-frame and the 3D facial semantic representation of the key reference frame is calculated to obtain the first prediction residual. This residual reflects the initial dynamic changes of the face from the key reference frame to the first inter-frame. For example, when processing the first inter-frame (denoted as...) When using the 3D face semantic representation of the key reference frame, As a prediction baseline, calculate 3D facial semantic representation and The difference is used to obtain the predicted residual. .
[0044] When processing frames following the first inter-frame, since facial dynamics between adjacent frames are usually smoother and temporally continuous, using the key reference frame as the baseline would cause the residual energy to increase over time. Therefore, the 3D facial semantic representation of the previous frame is used as the prediction baseline, and the difference between the 3D facial semantic representation of the current inter-frame and the 3D facial semantic representation of the previous frame is calculated to obtain the second prediction residual. This residual reflects the subtle dynamic increments between adjacent inter-frames. For example, when processing the second frame and subsequent inter-frames (denoted as...),... , 2) When, the 3D face semantic representation of the previous frame. As a prediction baseline, calculate 3D facial semantic representation and The difference is used to obtain the predicted residual. .
[0045] Thus, when processing the first inter-frame, the 3D facial semantic representation of the key reference frame is obtained as the prediction benchmark. The difference between the 3D facial semantic representation of the first inter-frame and the 3D facial semantic representation of the key reference frame is calculated to obtain the first prediction residual. Since the key reference frame provides a high-quality texture benchmark for the entire video and its semantic representation is complete and reliable, using it as the benchmark ensures that the prediction residual of the first dynamic frame accurately reflects the initial facial changes, avoiding prediction failure due to the lack of a benchmark. When processing subsequent inter-frames, the 3D facial semantic representation of the previous frame is obtained as the prediction benchmark. The difference between the 3D facial semantic representation of the current inter-frame and the 3D facial semantic representation of the previous frame is calculated to obtain the second prediction residual. Since the facial dynamic changes between adjacent inter-frames are temporally continuous and the change amplitude is usually more gradual, using the neighboring previous frame as the benchmark can make full use of temporal correlation, making the residual energy significantly lower than that of long-distance prediction based on the key reference frame. By selecting differentiated references, the first inter-frame uses a key reference frame as an anchor point to ensure initial prediction accuracy, and subsequent inter-frames use the neighboring previous frame as a reference to minimize residual energy frame by frame. This ensures that the prediction residuals of inter-frames at each position maintain a low amplitude distribution, which reduces the impact of quantization error on reconstruction quality and allows contextual arithmetic coding to fully utilize the statistical characteristics of residuals to achieve efficient compression. At the same time, the decoding end can recover the original semantic parameters of each inter-frame by accurately superimposing the corresponding references, avoiding the accumulation and diffusion of errors over time, and achieving a balance between prediction efficiency, compression performance and reconstruction stability.
[0046] To better understand the overall workflow of the encoder, see [link to documentation]. Figure 2 , Figure 2 This is a flowchart illustrating the encoder-side process of a face video communication method according to one embodiment of this disclosure. The flowchart demonstrates the complete workflow of the encoder-side in the Interactive Face Video Semantic Transmission (IFVC) framework proposed in this application, covering the entire link process of input processing, frame type determination, dual-track encoding, and integrated bitstream transmission. The following describes each functional module in detail: Input and Frame Type Determination: The input source is 2D face video frames, encompassing a continuous sequence of two-dimensional face images. Input frames first enter the frame type determination module. This module classifies video frames into two categories based on the difference in 3D face semantic representation between adjacent frames (e.g., whether changes in head rotation parameters or mouth movement parameters exceed a preset threshold): if the difference exceeds the threshold, it is determined to be a key reference frame; if the difference does not exceed the threshold, it is determined to be an inter-frame. Key reference frames typically account for no more than 5% and serve to provide high-quality texture templates; inter-frames account for more than 95% and carry the main dynamic semantic information.
[0047] For the key reference frame coding branch: Input frames identified as key reference frames directly enter the encoding end: the VVC intra-frame coding module. It employs the Multifunctional Video Coding (VVC) intra-frame coding standard for compression encoding. This standard achieves efficient compression while preserving high-quality texture information of the key reference frame through advanced block partitioning structure, multiple types of intra-frame prediction modes, and context-adaptive entropy coding technology. After encoding, an encoded key reference frame bitstream is generated. This bitstream contains the complete pixel data of the key reference frame and the extracted identity coefficients. Reflectivity coefficient Illumination coefficient Fixed parameter information provides a basic texture template and identity anchor for subsequent inter-frame reconstruction.
[0048] For the inter-frame coding branch: Input frames determined to be inter-frames enter the encoding end: the IDI semantic extraction and processing module. IDI (Intrinsic Dimensionality Improvement) is the core semantic compression scheme proposed in this application. This module first performs 3D face semantic extraction: based on the WM3DR / OpenFace model, the shape, texture, and dynamic semantic parameters of the 3D face are regressed from the 2D face image through the pre-trained 3D face reconstruction model WM3DR. At the same time, the facial behavior analysis model OpenFace is introduced to specifically predict the blinking parameters of the eyes, and then 14-dimensional core semantic parameters are obtained from the original 3D parameter space. Specifically, this includes 6-dimensional mouth motion parameters (corresponding to the vertical, horizontal, and forward / backward motion components of the lips), 1-dimensional eye blinking parameters (values ranging from 0 to 5), 3-dimensional head rotation parameters (corresponding to rotation radians along the X / Y / Z axes), 3-dimensional head translation parameters (corresponding to pixel translation amounts along the X / Y / Z axes), and 1-dimensional head position parameters (representing the relative height of the head). Then, semantic parameter inter-frame prediction is performed: residuals are calculated. That is, for each inter-frame, using the 3D face semantic representation of the key reference frame or the previous inter-frame as the prediction benchmark, the difference between the 14-dimensional semantic parameters of the current inter-frame and the corresponding parameters of the prediction benchmark is calculated to obtain the prediction residual. The residual represents the change in semantic parameters between adjacent frames rather than their absolute value, fully utilizing the temporal semantic continuity of face videos to achieve data compression. Finally, residual quantization + context entropy coding (PPM algorithm) is performed. The predicted residual is quantized using the zero-order exponential Golomb algorithm, converting continuous floating-point values into discrete integer values. The quantization step size is adaptively set according to the residual distribution characteristics. Then, the quantized residual is entropy-coded using a context arithmetic coding model based on the Predictive Partial Matching (PPM) algorithm. The PPM algorithm constructs a context window of size 8 and optimizes coding efficiency using historical data statistical patterns, ultimately generating the encoded inter-frame semantic bitstream.
[0049] Bitstream transmission: The encoded key reference frame bitstream and the encoded inter-frame semantic bitstream are integrated at the decoding module after transmission. They are then integrated into the final transmission bitstream according to a specific structure (including a key reference frame index table, fixed parameter blocks, and inter-frame residual coding blocks), and sent to the decoder via the communication network. The core advantage of this dual-track encoding architecture is that the key reference frame uses VVC intra-frame coding to ensure texture quality, while the inter-frames use 14-dimensional compact semantic parameters to replace the original pixel data. Compared to traditional frame-by-frame pixel coding schemes, this achieves orders-of-magnitude bitrate compression while preserving semantic-level editability, laying the data foundation for real-time interactive control at the decoder.
[0050] This disclosure provides a face video communication method, which is applied at the decoder end. See [link to relevant documentation]. Figure 3 As shown, it includes: Step S301: Receive the transmission bitstream from the encoder. The encoder classifies the input face video frame sequence to obtain key reference frames and inter-frames, and compresses and encodes the key reference frames to obtain key reference frame encoded data. For the inter-frames, map them from two-dimensional image frames to a three-dimensional parameter space, and filter to obtain a three-dimensional face semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters. The correlation between different parameters satisfies a preset independence condition. Encode the three-dimensional face semantic representation to obtain inter-frame encoded data. Integrate the key reference frame encoded data and the inter-frame encoded data into a transmission bitstream, and send the transmission bitstream to the decoder.
[0051] Specifically, this feature describes the complete generation process of the transmission stream received by the decoder. Its core lies in the encoder's use of a dual-track encoding strategy of "key reference frames + inter-frame semantic representations" to achieve efficient compression and semantically decoupled transmission of face videos. Since the encoder's operation has been detailed in the preceding embodiments, it will not be repeated here. The encoder integrates the key reference frame encoded data and inter-frame encoded data into a transmission stream according to a specific structure and sends it to the decoder. After receiving this stream, the decoder can achieve high-quality reconstruction of the face video and semantic-level interactive manipulation based on the texture template of the key reference frames and the compact semantic parameters of the inter-frames.
[0052] Step S302: Decode the received transmission stream to obtain the three-dimensional face semantic representation of the reconstruction key reference frame and the reconstruction inter-frame.
[0053] Specifically, the decoder first performs structural analysis on the transmitted bitstream, separating the key reference frame encoded data and the inter-frame encoded data. The decoder reconstructs the key reference frame (as a texture template and source of fixed parameters) and the reconstructed inter-frame 3D facial semantic representation (containing 14 independent semantic parameters: 6-dimensional mouth motion parameters, 1-dimensional eye blinking parameters, 3-dimensional head rotation parameters, 3-dimensional head translation parameters, and 1-dimensional head position parameters) through the decoding output. Together, these two constitute the complete input data for subsequent 3D face mesh reconstruction, mesh basis motion estimation, and inter-frame generation, enabling the decoder to recover semantically editable face video content under extremely low bitrate conditions.
[0054] Step S303: Based on the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, generate the inter-frame reconstructed face image frame.
[0055] Specifically, this feature describes the complete technical implementation of the decoder based on the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. It is achieved through a multi-stage cascaded process of "3D face mesh reconstruction - 2D projection - mesh basis motion estimation - neural network generation" to finally output the generated inter-frame reconstructed face image frame. Its core lies in transforming compact semantic parameters into high-fidelity 2D face images, realizing accurate mapping from abstract semantic space to pixel-level visual content.
[0056] Step S304: In response to the interaction command, modify the parameters in the 3D face semantic representation of the reconstructed inter-frame to generate the inter-frame reconstructed face image frame corresponding to the interaction command.
[0057] Specifically, this feature describes the technical implementation whereby, after receiving a user interaction command, the decoder independently edits the semantic parameters in the 3D facial semantic representation of the reconstructed inter-frame, and re-executes the complete inter-frame generation process based on the modified semantic parameters, ultimately outputting an inter-frame reconstructed facial image frame that meets the requirements of the interaction command. Its core lies in utilizing the high independence of each parameter dimension in the 3D facial semantic representation to achieve precise control and real-time response of the decoder to the dynamic semantics of the face, without the need for re-encoding or transmission of additional bitstreams.
[0058] In the embodiments of this disclosure, the encoder uses VVC intra-frame encoding to preserve high-quality texture templates for key reference frames, while mapping inter-frames from two-dimensional image frames to a 14-dimensional three-dimensional parameter space. The original pixel data is replaced by a compact representation of mouth motion parameters (6-dimensional), eye blinking parameters (1-dimensional), head rotation parameters (3-dimensional), head translation parameters (3-dimensional), and head position parameters (1-dimensional). PCA projection is used to ensure the independence of each parameter, with the absolute value of the correlation coefficient not exceeding 0.1, thus eliminating redundancy. This significantly reduces bitrate and bandwidth usage compared to traditional frame-by-frame pixel encoding schemes. Furthermore, the decoder uses a VVC decoder to reconstruct the high-quality texture of the key reference frames and extracts fixed parameters, storing them in a buffer for reuse across all inter-frames, ensuring identity consistency and texture fidelity. Regarding interactive flexibility, thanks to the high independence of each parameter dimension in the 3D face semantic representation, the decoder can directly respond to user interaction commands to independently modify specific semantic parameters without re-encoding or transmitting additional bitstreams, achieving a semantic-level interaction mode of "transmit once, edit infinitely." In summary, this solution, through a closed-loop design of semantic decoupling compression at the encoding end and semantic-level remanipulation at the decoding end, significantly reduces bandwidth consumption while ensuring reconstruction quality, and endows the decoder with unprecedented facial dynamic semantic editing capabilities, achieving the integration of efficient compression, high-quality reconstruction and real-time interaction.
[0059] In some optional embodiments, the received transmission bitstream is decoded to obtain the three-dimensional face semantic representation of the reconstructed key reference frame and the inter-reconstruction frames, including: For the encoded data of the key reference frame in the transmitted bitstream, perform key reference frame decoding and reconstruction to obtain the reconstructed key reference frame; For the inter-frame encoded data in the transmission bitstream, perform inter-frame semantic representation decoding and reconstruction to obtain the three-dimensional face semantic representation of the reconstructed inter-frame.
[0060] Specifically, for the key reference frame encoded data in the transmitted bitstream, the decoder calls the VVC (Versatile Video Coding) decoder to perform key reference frame decoding and reconstruction. For the inter-frame encoded data in the transmitted bitstream, the decoder performs inter-frame semantic representation decoding and reconstruction. Finally, the decoder outputs the reconstructed key reference frame (as a high-quality texture template and fixed parameter source) and the 3D face semantic representation of the reconstructed inter-frame (as a dynamic semantic driving source). Together, they constitute the input data for the subsequent complete process of "3D face mesh reconstruction - 2D projection - mesh basis motion estimation - neural network generation". This enables the decoder to recover semantically editable face video content under extremely low bitrate conditions and lays the data foundation for independently modifying semantic parameters in response to interactive commands.
[0061] In this way, the differentiated decoding and reconstruction strategy enables the divide-and-conquer recovery of key reference frames and inter-frame data, achieving a synergistic improvement in compression efficiency, reconstruction accuracy, and timing stability.
[0062] In some optional embodiments, for the key reference frame encoded data in the transport bitstream, key reference frame decoding and reconstruction are performed to obtain the reconstructed key reference frame, including: The key reference frame is reconstructed by decoding the encoded data of the key reference frame using a decoder. Fixed parameters are extracted from the reconstructed key reference frame and stored in the fixed parameter buffer. The fixed parameters include identity coefficient, reflectivity coefficient and illumination coefficient.
[0063] Specifically, the decoder first performs decoding operations on the key reference frame encoded data in the transmitted bitstream: since the encoder uses the VVC (Versatile Video Coding) intra-frame coding standard to compress and encode the key reference frames, the decoder calls the matching VVC decoder to sequentially perform standard decoding processes such as entropy decoding, inverse quantization, inverse transform, and intra-frame prediction compensation. Pixel-level reconstructed key reference frames are recovered from the binary bitstream. These reconstructed key reference frames retain the high-resolution texture information of the original key reference frames, including visual features such as facial skin pore details, hair texture, and lighting levels, providing a unified texture reference template for the entire video sequence. After obtaining the reconstructed key reference frames, the decoder further extracts fixed parameters from them. These fixed parameters refer to a set of parameters that characterize the inherent attributes of facial identity and do not change over time. These fixed parameters specifically include three types of coefficients: identity coefficients... This refers to the identity feature vector regressed from 2D face images using a pre-trained 3D face reconstruction model, WM3DR. This vector is linearly combined with the identity basis vectors in a parametric 3D deformation model (3DMM) to determine the basic contours, skeletal structure, and facial geometry of the face. The identity coefficients of different individuals exhibit significant differences and are the core identifiers for distinguishing different facial identities. Reflectivity coefficient This refers to the reflectance feature vector, which characterizes the texture and material properties of facial skin. This vector is weighted and superimposed with the reflectance basis vectors in 3DMM to determine surface texture details such as skin tone, pore distribution, and wrinkle direction, reflecting the appearance and material characteristics of the face. Illumination coefficient. The illumination feature vector, representing the scene's lighting conditions, is combined with the illumination basis vectors in the 3DMM to determine the facial brightness distribution, shadow position, and highlight intensity, reflecting the lighting configuration information under the current shooting environment. The decoder stores the extracted identity coefficients, reflectivity coefficients, and illumination coefficients in a fixed parameter buffer. This buffer is a persistent storage area specifically allocated by the decoder for the entire video sequence. Its core function is to achieve "one-time extraction and global reuse" of fixed parameters. Since these parameters are only related to individual identity and shooting environment, they usually remain stable during video calls or live broadcasts. Therefore, there is no need to repeatedly transmit or re-extract them for each frame. The 3D face mesh reconstruction of all subsequent frames can be directly read and called from this buffer, thereby significantly reducing the size of the transmission bitstream and bandwidth usage.
[0064] In this way, the decoder performs VVC standard decoding on the key reference frame encoded data to recover the reconstructed key reference frame that retains high-resolution texture information. This reconstructed key reference frame serves as the basic texture template for the entire video, carrying visual features such as facial skin pore details, hair texture, and lighting levels. It provides a unified texture benchmark and identity anchor for all subsequent inter-frames, ensuring the visual coherence and identity consistency of the entire video reconstruction result. Simultaneously, the decoder further extracts three types of fixed parameters from this reconstructed key reference frame: identity coefficients, reflectivity coefficients, and illumination coefficients, and stores them in a fixed parameter cache for reuse throughout the video. This scheme achieves complete decoupling of static identity attributes and dynamic semantic parameters. Specifically, since the identity coefficients, reflectivity coefficients, and illumination coefficients are only related to individual identity and shooting environment, they remain stable during video calls or live broadcasts. The caching mechanism enables "extraction once, global reuse," avoiding the repeated transmission of identity and texture information in traditional frame-by-frame encoding. This allows inter-frames to drive complete reconstruction by transmitting only the prediction residual of a 14-dimensional compact semantic representation, significantly compressing the bitstream size and bandwidth usage. Meanwhile, the fixed parameter cache, serving as a persistent storage area on the decoder side, separates the "heavy asset" texture template from the "light asset" semantic parameters at the system architecture level. This allows the inter-frame reconstruction process to directly read and call fixed parameters from the cache, eliminating the need to re-execute complex identity regression and lighting estimation operations, significantly reducing decoding computational overhead and processing latency. Furthermore, this decoupled architecture provides a natural technical foundation for advanced interactive functions such as virtual character replacement. When a user issues a virtual character replacement command, only the identity coefficients and reflectivity coefficients in the cache need to be replaced with a new virtual character image. This generates virtual character animation frames while maintaining the original dynamic semantics (mouth movements, eye blinks, head posture), without modifying the encoder or retransmitting the bitstream. This achieves a flexible balance between privacy protection and personalized content creation, effectively solving the technical problems of insufficient privacy protection mechanisms and the inability to effectively hide the user's real identity information in existing generative compression schemes.
[0065] In some optional embodiments, for the inter-frame encoded data in the transmission bitstream, inter-frame semantic representation decoding and reconstruction are performed to obtain the reconstructed three-dimensional face semantic representation of the inter-frame, including: Entropy decoding is performed on the inter-frame encoded data to obtain the quantized prediction residual; Perform an inverse quantization operation on the quantized prediction residuals to recover the prediction residuals; Obtain the 3D face semantic representation corresponding to the prediction benchmark, where the prediction benchmark is the key reference frame for reconstruction or the preceding reconstructed frame. The predicted residuals are semantically compensated with the 3D face semantic representations corresponding to the prediction baseline to calculate the 3D face semantic representations of the reconstructed frames.
[0066] Specifically, the decoder first performs entropy decoding on the inter-frame encoded data: Since the encoder uses contextual arithmetic coding based on the Prediction by Partial Matching (PPM) algorithm to compress the prediction residuals, the decoder performs a corresponding inverse entropy decoding process. The PPM algorithm optimizes coding efficiency by constructing a context window of size 8 and utilizing historical data statistical patterns. Based on this, the decoder parses the quantized prediction residuals from the binary bitstream. These prediction residuals are the difference data obtained by the encoder subtracting the original inter-frame 3D face semantic representation from the corresponding 3D face semantic representation of the prediction baseline dimension by dimension, and then quantizing using the zero-order exponential Columbus algorithm. This represents the change in semantic parameters between adjacent frames rather than their absolute values, thus significantly compressing the data volume. Subsequently, the decoder performs inverse quantization on the quantized prediction residuals: Inverse quantization uses a quantization step size matched to that of the encoder. This step size is adaptively set according to the residual distribution characteristics to ensure that the quantization error is controllable. By mapping the quantized discrete integer values back to continuous floating-point values, the prediction residuals with the same accuracy as the encoder's residual calculation are recovered. The prediction residual preserves subtle changes in the inter-frame relative to the prediction benchmark in the 14-dimensional semantic space, including mouth shape offsets for mouth movement parameters, changes in blink closure for eye blinking parameters, radian offsets for head rotation parameters, pixel displacements for head translation parameters, and height adjustments for head position parameters. Next, the decoder obtains the 3D facial semantic representation corresponding to the prediction benchmark. The selection of the prediction benchmark follows a differentiated rule: for the first inter-frame in the video sequence, the prediction benchmark is the 3D facial semantic representation extracted from the reconstructed key reference frame. This representation includes 14 initial semantic parameters (mouth movement parameters, blinking parameters, head rotation parameters, head translation parameters, and head position parameters) regressed from the key reference frame, serving as the starting anchor point for the semantic evolution of the entire video sequence. For the second frame and subsequent inter-frames, the prediction benchmark is the 3D facial semantic representation of the previously reconstructed inter-frame, i.e., the 14-dimensional semantic parameters output from the previous frame after a complete decoding and reconstruction process. This selection strategy fully leverages the temporal semantic continuity of face videos, resulting in relatively gradual semantic changes between adjacent frames. This ensures that the numerical distribution of the prediction residuals is concentrated and easy to compress. Finally, the decoder performs a semantic compensation operation: it adds the recovered prediction residuals to the corresponding 3D face semantic representations of the prediction baseline dimension by dimension, that is, it sums the corresponding elements of the 14-dimensional residual vector and the 14-dimensional baseline vector to calculate the 3D face semantic representations of the reconstructed frames.
[0067] In this way, the scheme realizes the design concept of "differential transmission, reference reuse, and controllable error": inter-frame transmission only needs to transmit lightweight prediction residuals, the prediction reference is directly obtained from the local buffer, and the semantic buffer ensures timing stability. The three work together to enable the decoder to stably recover the semantically editable 3D face semantic representation of inter-frames under extremely low bit rate conditions. This lays the data foundation for the decoder to directly respond to interactive commands and independently modify semantic parameters, effectively solving the technical problems of block artifacts and distortions in the reconstruction process and the lack of semantic-level manipulation capabilities in traditional video coding schemes.
[0068] In some optional embodiments, based on the reconstructed key reference frame and the 3D face semantic representation of the inter-frame reconstruction, inter-frame reconstructed face image frames are generated, including: Based on the reconstructed key reference frame, fixed parameters are extracted, including: identity coefficient, reflectivity coefficient, and illumination coefficient. Based on fixed parameters and the semantic representation of the 3D face between reconstructed frames, 3D face mesh reconstruction is performed to obtain the 3D face mesh; Based on the 3D face mesh, the 3D face semantic representation of the reconstructed key reference frame, and the 3D face semantic representation of the reconstructed inter-frame, the 2D face mesh projection and eye motion calibration are performed to obtain the 2D face mesh of the reconstructed key reference frame, the 2D face mesh of the inter-frame, and the eye blinking motion map. Based on the reconstructed key reference frame, the inter-frame 2D face mesh, the eye blinking motion map, and the reconstructed key reference frame, mesh-based motion estimation is performed to obtain a fine-grained dense motion field and a face attention map. Based on the fine-grained dense motion field, face attention map, and reconstruction key reference frames, inter-frame generation is performed to obtain inter-frame reconstructed face image frames.
[0069] Specifically, this feature describes the complete technical implementation of the decoder based on the reconstructed key reference frame and the reconstructed inter-frames, through a five-stage cascaded process of "fixed parameter extraction - 3D face mesh reconstruction - 2D projection and eye calibration - mesh basis motion estimation - inter-frame generation", which ultimately outputs the generated inter-frame reconstructed face image frame. Its core lies in transforming abstract semantic parameters into concrete 3D geometric structures, and then generating high-fidelity 2D face images through motion estimation and neural network generation, thus achieving accurate reconstruction from semantic space to pixel space.
[0070] Specifically, the first stage involves extracting fixed parameters: the decoder extracts identity coefficients, reflectivity coefficients, and illumination coefficients from the reconstructed key reference frames, and stores these fixed parameters in a buffer for reuse throughout the video, ensuring identity consistency and texture stability.
[0071] The second stage involves 3D face mesh reconstruction: The decoder reads the identity coefficient, reflectivity coefficient, and illumination coefficient from the fixed parameter buffer, and simultaneously reads the mouth motion parameters (6-dimensional vectors, corresponding to the vertical, horizontal, and forward / backward motion components of the lips) from the 3D face semantic representation between reconstructed frames. Zero values are added to these mouth motion parameters to expand them into complete expression coefficients to match the 3DMM expression basis dimension. Then, the parametric 3D deformation model (3DMM) template provided by the pre-trained WM3DR model decoder is used. The 3D face shape (composed of the weighted superposition of the average neutral shape, identity basis and identity coefficient, and expression basis and expression coefficient) and the 3D face texture (composed of the weighted superposition of the average neutral texture, reflectivity basis and reflectivity coefficient, and illumination basis and illumination coefficient) are calculated using a linear combination formula. These two components together constitute a 3D face mesh containing geometric vertex coordinates and appearance color information. This mesh organizes the vertex set in the form of triangular facets, accurately depicting the 3D curved surface structure of the face.
[0072] The third stage involves 2D face mesh projection and eye motion calibration. The decoder extracts head rotation parameters (3D vectors, corresponding to rotation radians in the X / Y / Z axes) and head translation parameters (3D vectors, corresponding to pixel translation in the X / Y / Z axes) from the 3D face semantic representations of the reconstructed key reference frames and inter-frames, respectively. The head rotation parameters are converted into a 3×3 head rotation matrix through Rodrigues transformation, and the head translation parameters are converted into 3D translation vectors. Combined with the preset camera intrinsic parameter matrix (focal length 256, principal point coordinates (128,128)), the 3D face mesh vertex set is mapped to a 2D plane using the perspective projection formula, resulting in the 2D face mesh of the reconstructed key reference frames (as a static pose reference, reflecting the head pose of the key reference frames) and the 2D face mesh of the inter-frames (as a dynamic pose target, reflecting the head pose and facial expression of the inter-frames). Simultaneously, based on the blinking parameters (1-dimensional scalar, value range 0~5) in the reconstructed 3D face semantic representation between frames, the eye region in the 2D face mesh between frames is located (determined by the predefined range of eye vertex indices). The original highest and lowest points of the eye region are found, and the new highest point position is calculated according to the linear interpolation formula. The vertical coordinates of the eye mesh vertices are dynamically adjusted to generate a blinking motion map, ensuring that the degree of eye closure is accurately matched with the semantic parameter values.
[0073] The fourth stage performs mesh-based motion estimation. The decoder calculates the vertex position differences (by subtracting vertex coordinates) between the 2D face mesh of the reconstructed key reference frame and the 2D face mesh of the inter-frame, obtaining a coarse-grained vertex displacement vector. This vertex difference is then interpolated into a coarse-grained mesh motion flow consistent with the video resolution using a bilinear interpolation algorithm. The coarse-grained mesh motion flow and the reconstructed key reference frame are input into a U-type encoder-decoder network (UNet). The UNet encoder extracts multi-scale spatial features from the reconstructed key reference frame, performs an optical flow-based pixel resampling feature distortion operation in conjunction with the coarse-grained mesh motion flow, and then decodes it through the UNet decoder to generate a coarse-grained deformed frame. A Spatial Adaptive Normalization (SPADE) mechanism is then introduced, stitching together the reconstructed key reference frame, coarse-grained deformed frame, coarse-grained mesh motion flow, and eye blinking motion map to form a multi-feature input. The SPADE module adaptively adjusts the batch normalization parameters according to the spatial distribution of the input features, better preserving semantic details. By using a dual-branch prediction output structure, a fine-grained dense motion field (representing the sub-pixel displacement vector of each pixel to achieve accurate transfer of facial micro-expressions) and a face attention map (a spatial attention mask with high weights for key areas such as the mouth and eyes and low weights for background areas, guiding the generator to focus on the core facial regions) are obtained respectively.
[0074] The fifth stage involves inter-frame generation: The decoder inputs the reconstructed key reference frame into the neural network encoder (UNet encoder) to extract spatial feature maps at four scales (resolutions of the original video resolution, half the original resolution, one-quarter of the original resolution, and one-eighth of the original resolution), resulting in multi-scale spatial features. Combining the fine-grained dense motion field and the face attention map, an attention-based feature distortion operation is performed on the multi-scale spatial features (using the attention map as weights and the dense motion field as displacement, achieving spatial deformation of the feature map through inverse distortion). Subsequently, affine transformation parameters (scaling and offset coefficients) are generated from the multi-scale spatial features through a 3×3 convolutional layer. Based on these parameters, affine transformation modulation is performed on the distorted spatial features to obtain transformed face features. Finally, the distorted face spatial features and the transformed face features are concatenated and input into the generator network (a deep generative network composed of multiple convolutional layers and activation functions). Through a generative adversarial mechanism, the generated inter-frame reconstructed face image frames are output. While maintaining consistency with the identity of the key reference frame for reconstruction, the output frame accurately reproduces the dynamic semantics of mouth movement, eye blinking, and head posture corresponding to the inter-frame, achieving complete reconstruction from 14-dimensional compact semantic parameters to a high-fidelity two-dimensional face image.
[0075] In this way, through a five-stage cascaded reconstruction process of "fixed parameter extraction - 3D face mesh reconstruction - 2D projection and eye calibration - mesh basis motion estimation - inter-frame generation", the accurate mapping of face video from abstract semantic parameters to high-fidelity pixel-level images is achieved, resulting in a synergistic improvement in compression efficiency, reconstruction quality and interactive flexibility.
[0076] In some optional embodiments, based on fixed parameters and the 3D face semantic representation between reconstructed frames, 3D face mesh reconstruction is performed to obtain a 3D face mesh, including: Based on the identity coefficient, reflectivity coefficient, and illumination coefficient, combined with the average neutral shape, identity basis vector, average neutral texture, reflectivity basis vector, and illumination basis vector, the three-dimensional face shape and three-dimensional face texture are calculated. Based on the mouth motion parameters in the reconstructed 3D facial semantic representation between frames, zero values are added to form expression coefficients; Based on the 3D face shape, 3D face texture, and expression coefficients, a 3D face mesh is calculated using a parametric 3D deformation model template.
[0077] Specifically, this feature describes the semantic representation of a 3D face based on fixed parameters and reconstructed frames at the decoder end. Through a three-stage process of "3D face shape and texture calculation - expression coefficient construction - 3D face mesh synthesis," a complete technical implementation of the 3D face mesh is finally obtained. Its core lies in utilizing the linear combination mechanism of the parametric 3D deformation model (3DMM) to transform compact semantic parameters into a 3D face mesh with geometric and appearance information, laying the geometric foundation for subsequent 2D projection and frame generation.
[0078] Specifically, the first stage involves calculating the 3D face shape and texture: the decoder reads the identity coefficient, reflectance coefficient, and illumination coefficient from a fixed parameter buffer. Simultaneously, the decoder calls the parametric 3D deformation model (3DMM) template provided by the pre-trained WM3DR model decoder. This template contains basic components obtained through statistical learning: average neutral shape (i.e., the average set of 3D face vertex coordinates calculated from large-scale face scan data, representing the geometric baseline of the face under neutral expression), identity basis vector (i.e., the principal components of identity changes extracted from face shape data through principal component analysis, representing the direction and magnitude of facial contour differences between different individuals), average neutral texture (i.e., the average facial color and material distribution calculated from large-scale face texture data, representing the facial appearance baseline under neutral illumination), reflectance basis vector (i.e., the principal components of texture changes extracted from face reflectance data through principal component analysis, representing the direction and magnitude of skin color and material differences between different individuals), and illumination basis vector (i.e., the principal components of illumination changes extracted from face illumination data through principal component analysis, representing the direction and magnitude of facial brightness differences under different illumination conditions). The decoder calculates the 3D face shape (i.e., the set of 3D vertex coordinates reflecting the geometric topology of a specific individual's face) by multiplying the identity coefficient with the identity basis vector and then superimposing the result onto the average neutral shape, based on the linear combination formula. Simultaneously, it calculates the 3D face texture (i.e., the texture mapping reflecting the color and material distribution of a specific individual's face under specific lighting conditions) by multiplying the reflectivity coefficient with the reflectivity basis vector and then superimposing the result onto the average neutral texture.
[0079] For example, according to the formula Calculate the 3D face shape, where S is the 3D face shape. It is an average neutral shape. For identity coefficient, For identity basis vectors, For expression basis vectors, The expression coefficients between frames are derived from the 3D facial semantic representation of the frames. Mouth motion parameters Supplement the zero-value composition. According to the formula... Calculate the 3D face texture, where T is the 3D face texture. For average neutral texture, The reflectivity coefficient, Let reflectivity be the basis vector. This is the illumination coefficient. is the illumination basis vector.
[0080] The second stage involves constructing expression coefficients: The decoder reads the mouth motion parameters (6-dimensional vectors, corresponding to the vertical, horizontal, and forward / backward motion components of the lips, accurately representing the changes in mouth shape during speech) from the 3D facial semantic representation of the reconstructed frames. Since the expression basis vectors of the parameterized 3D deformation model usually have higher dimensions (e.g., 64 dimensions), and this scheme extracts only the first 6 dimensions of mouth motion parameters for compressed transmission, the decoder fills these 6-dimensional mouth motion parameters with zero values to expand them into complete expression coefficients (i.e., in the high-dimensional expression vector, the first 6 dimensions are filled with mouth motion parameter values, and the remaining dimensions are filled with zeros to match the complete dimensional requirements of the 3DMM expression basis). These expression coefficients represent the dynamic changes of the face relative to a neutral expression in the frames, especially the degree and direction of deformation in the lip region.
[0081] The third stage involves 3D face mesh synthesis: The decoder inputs the 3D face shape (reflecting the geometric structure of an individual's identity), 3D face texture (reflecting the material properties of an individual's appearance), and expression coefficients (reflecting the deformation parameters of dynamic expressions between frames) calculated in the first stage into a parameterized 3D deformation model template. The expression coefficients are multiplied by the expression basis vectors using a linear combination formula and then superimposed onto the 3D face shape to obtain a dynamic 3D face shape reflecting a specific individual under a specific expression. This is then combined with the color and material information of the 3D face texture to finally calculate the 3D face mesh. This mesh organizes the vertex set in the form of triangular patches, with each vertex containing 3D spatial coordinates (X / Y / Z) and color / texture coordinates (U / V). It accurately depicts the geometric structure and appearance attributes of the 3D facial surface under specific identities, lighting conditions, and expressions, providing a complete 3D geometric foundation for subsequent 2D face mesh projection (mapping 3D vertices to a 2D image plane), mesh basis motion estimation (calculating the difference in vertex positions under different poses), and inter-frame generation (pixel-level image synthesis driven by geometric motion).
[0082] In this way, the decoder calculates the 3D face shape and texture reflecting the static attributes of a specific individual based on the identity coefficients and reflectivity coefficients extracted and cached from the key reference frames for reconstruction, combined with two statistical benchmarks: average neutral shape and average neutral texture. This design ensures that all frames in the entire video share the same set of identity-related geometric and appearance bases, guaranteeing the identity stability and visual coherence of the reconstruction results across frame sequences, effectively solving the identity drift problem caused by frame-by-frame independent reconstruction in traditional generative compression schemes. Simultaneously, the combination of illumination coefficients and illumination basis vectors allows the 3D face texture to adaptively reproduce the lighting conditions of the original shooting environment, avoiding facial distortion caused by lighting estimation bias and ensuring the realism of the reconstruction results. In terms of extreme bitrate compression, this three-stage process fully utilizes the statistical prior knowledge of the parametric 3D deformation model (3DMM): the average neutral shape, identity basis vector, average neutral texture, reflectivity basis vector, and illumination basis vector are all stored locally on the decoder as pre-trained model templates, without needing to be transmitted from the encoder. Between frames, only compact coefficients such as identity coefficients, reflectivity coefficients, illumination coefficients, and mouth motion parameters need to be transmitted to drive the complete 3D face mesh reconstruction. In particular, the mouth motion parameters are only 6-dimensional (corresponding to the vertical, horizontal, and forward / backward motion components of the lips), which is significantly compressed compared to the original 64-dimensional expression vector. By supplementing zero values to form complete expression coefficients, the expression parameters are extremely simplified without losing core lip shape information. This reduces the overall bitrate by an order of magnitude compared to the traditional frame-by-frame pixel encoding scheme, effectively solving the technical problems of high cost of 3D parameter representation and its unfavorability to low bitrate transmission.
[0083] In some optional embodiments, based on the 3D face mesh, the 3D face semantic representation of the reconstructed key reference frame, and the 3D face semantic representation of the reconstructed inter-frames, a 2D face mesh projection and eye motion calibration are performed to obtain the 2D face mesh of the reconstructed key reference frame, the 2D face mesh of the inter-frames, and the eye blinking motion map, including: The first head rotation parameter and the first head translation parameter are extracted from the 3D face semantic representation of the reconstructed key reference frame. The first head rotation parameter is converted into the first head rotation matrix and the first head translation parameter is converted into the first translation vector. Combined with the camera intrinsic parameter matrix, the 3D face mesh is projected onto the 2D plane based on the first head rotation matrix and the first translation vector to obtain the 2D face mesh of the reconstructed key reference frame. The second head rotation parameter and the second head translation parameter are extracted from the 3D face semantic representation of the reconstructed inter-frames. The second head rotation parameter is converted into a second head rotation matrix and the second head translation parameter is converted into a second translation vector. Combined with the camera intrinsic parameter matrix, the 3D face mesh is projected onto the 2D plane based on the second head rotation matrix and the second translation vector to obtain the 2D face mesh of the inter-frames. Based on the blinking parameters in the 3D face semantic representation of the reconstructed frames, the eye region in the 2D face mesh of the frames is located, the original highest and lowest points of the eye region are determined, the new highest point is calculated based on the blinking parameters, the position of the eye mesh vertex is adjusted, and the blinking motion map is generated.
[0084] Specifically, this feature describes the complete technical implementation of the decoder based on the same 3D face mesh, driven by differentiated head pose parameters to perform dual-track 2D projection, and combined with eye blink parameters for local vertex calibration. Its core lies in the design concept of "one 3D template, two 2D projections, and local dynamic calibration", which not only ensures geometric consistency but also accurately depicts the differences in pose between frames and changes in eye micro-expressions.
[0085] Specifically, the first stage involves reconstructing the 2D face mesh projection of the key reference frame: The decoder extracts the first head rotation parameters (3D vectors, corresponding to the rotation radians in the X / Y / Z axes, representing the head's rotational posture relative to the camera coordinate system at the time the key reference frame was captured, such as head turning, nodding, and tilting angles) and the first head translation parameters (3D vectors, corresponding to the pixel translation amounts in the X / Y / Z axes, representing the spatial position offset of the head in the image at the time the key reference frame was captured) from the 3D face semantic representation of the reconstructed key reference frame. The first head rotation parameters are then transformed into the first head rotation matrix (a 3×3 matrix used to describe rotation transformations around arbitrary axes in 3D space) through the Rodrigues transformation (a mathematical method for converting rotation vectors into rotation matrices, converting the 3D rotation axis-angle representation into a 3×3 orthogonal rotation matrix through exponential mapping). The first head translation parameter is directly converted into the first translation vector (a 3D vector used to describe translation transformation in 3D space). Combined with the preset camera intrinsic matrix (a 3×3 matrix representing the internal geometry of the camera, containing the focal length parameter (set to 256, which determines the projection scaling ratio) and the principal point coordinates (set to (128,128), which determines the projection center position), the points in the 3D camera coordinate system are mapped to the 2D image plane. The 3D face mesh is projected onto the 2D plane using the perspective projection formula (i.e., first transforming the vertices of the 3D face mesh from the model coordinate system to the camera coordinate system through the first head rotation matrix and the first translation vector, then using the camera intrinsic matrix to divide the 3D coordinates in the camera coordinate system by the depth value Z to obtain the normalized 2D coordinates, and finally mapping them to the pixel coordinate system). This results in the 2D face mesh of the reconstructed key reference frame (a set of 2D vertex coordinates reflecting the head pose at the time of the key reference frame shooting, serving as a static reference for subsequent motion estimation).
[0086] The second stage involves projecting the 2D face mesh between frames: The decoder extracts the second head rotation parameter (a 3D vector representing the head rotation change in the inter-frame relative to the key reference frame) and the second head translation parameter (a 3D vector representing the head translation change in the inter-frame relative to the key reference frame) from the reconstructed 3D face semantic representation of the inter-frame. Similarly, the second head rotation parameter is converted into a second head rotation matrix and the second head translation parameter into a second translation vector using the Rodrigues transform. Using the same camera intrinsic matrix (assuming the camera parameters remain constant during shooting), the same 3D face mesh is projected onto a 2D plane based on the second head rotation matrix and the second translation vector using the same perspective projection formula, resulting in the 2D face mesh of the inter-frame (reflecting the 2D vertex coordinate set of the head pose at the moment of inter-frame shooting, serving as the dynamic target for subsequent motion estimation). Since the two sets of 2D face meshes originate from the same 3D face mesh template, differing only in head pose parameters, the difference in vertex positions purely reflects the pixel-level displacement caused by head rotation and translation, eliminating interference from identity geometric changes and providing a clean geometric input for accurate mesh-based motion estimation.
[0087] For example, taking inter-frames as an example, from the 3D facial semantic representation of inter-frames Extracting head rotation parameters With head translation parameters ,Will The transformation is converted to a head rotation matrix R using the Rodrigues transformation. Convert to a translation vector T; combine this with the camera intrinsic parameter matrix Ψ (focal length set to 256, principal point coordinates set to (128, 128)), and use the formula... ( Projecting a 3D face mesh onto a 2D plane (using a set of 3D face vertices) yields a 2D face mesh. .
[0088] The third stage performs eye movement calibration: The decoder uses the blinking parameters (1-dimensional scalar, ranging from 0 to 5, where 0 represents fully open eyes, 5 represents fully closed eyes, and intermediate values correspond to different degrees of half-open state, predicted from 2D images by the OpenFace facial behavior analysis model, with accuracy sufficient for semantic interaction) in the reconstructed 3D facial semantic representation of the inter-frames. It locates the eye region in the 2D face mesh of the inter-frames, achieved through a predefined range of eye vertex indices (i.e., the set of vertex numbers in the 3D face mesh template already labeled with the left and right eye regions, projected onto the corresponding eye contour region in 2D). Within the located eye region, it determines the original highest point (the original highest vertex of the upper eyelid contour, reflecting the eyelid position when the eyes are fully open) and the original lowest point (the original lowest vertex of the lower eyelid contour, reflecting the lower eyelid position when the eyes are open), and calculates a new highest point based on the blinking parameters. Based on this new highest point position, the positions of the eye mesh vertices are dynamically adjusted. This involves scaling the ordinates of all vertices within the eye region according to an interpolation ratio, causing the upper eyelid contour to smoothly decrease with the blinking parameter value, generating an eye blinking motion map (representing the dynamically adjusted two-dimensional coordinate distribution of the eye region vertices, reflecting the eyelid geometry under a specific blinking degree). This eye blinking motion map, as local micro-expression information independent of head posture movement, compensates for the shortcomings of general 3D deformation models in representing fine eye movements. It provides precise geometric constraints on the eye region for subsequent mesh-based motion estimation, ensuring accurate matching between the degree of eye closure and semantic parameter values in the final inter-frame reconstructed face image frames, achieving a precise mapping from abstract blinking parameters to concrete eye geometry.
[0089] For example, based on the eye blink parameters between frames Locate the eye region in the 2D face mesh M (determined by a predefined range of eye vertex indices) and find the original highest point of the eye region. Compared with the original lowest point According to the formula Calculate the new highest point Adjust the position of the eye grid vertices to generate an eye blinking motion map to ensure that the eye blinking effect is realistic and accurate.
[0090] In this way, through a differentiated processing architecture of "dual-track projection of the same 3D template + local dynamic calibration of the eyes," this feature achieves independent decoupling and accurate reconstruction of head pose changes and micro-expressions of the eyes, resulting in both geometric consistency and enhanced realism of the eyes. Specifically, regarding geometric consistency, the decoder extracts head rotation and translation parameters from the 3D facial semantic representations of the reconstructed key reference frame and the inter-frame reconstruction, respectively. These parameters are then converted into rotation matrices using Rodrigues transformation and combined with the camera intrinsic matrix to perform perspective projection. The same 3D face mesh is projected to obtain the 2D face mesh of the reconstructed key reference frame and the 2D face mesh of the inter-frame reconstruction. Since the two sets of 2D meshes originate from the same identity coefficient-driven 3D template, differing only in head pose parameters, the interference of identity geometric changes on motion estimation is eliminated, ensuring that the difference in vertex positions between frames purely reflects head pose motion. This provides a geometrically pure motion benchmark for subsequent mesh-based motion estimation, effectively solving the technical problems of limited viewpoint freedom and incomplete expression separation in traditional 2D keypoint methods. In terms of enhancing eye realism, a local vertex calibration mechanism driven by eye blink parameters (1-dimensional scalar, 0~5) compensates for the inherent defects of general 3D deformation models in expressing fine eye movements. It locates the eye region by predefined vertex index ranges, uses the original highest and lowest points as boundaries, calculates the new highest point position based on the blink parameters using a linear interpolation formula, and dynamically adjusts the ordinates of the eye mesh vertices to generate an eye blink motion map. This motion map is independent of head posture movement, achieving precise matching between the degree of eye closure and semantic parameter values. When the parameter is 0, the eyes are fully open; when the parameter is 5, the eyes are fully closed; intermediate values correspond to a naturally transitioning half-open state. This local dynamic calibration ensures the realism and detail of the eye region in the final inter-frame reconstructed face image, avoiding visual flaws such as "dull gaze" or "unnatural blinking" caused by ignoring micro-expressions in traditional 3D face modeling schemes. It effectively solves the technical problems of insufficient generation quality and robustness, especially the tendency to distort reconstruction under complex expressions.
[0091] In some optional embodiments, based on the 2D face mesh of the reconstructed key reference frame, the 2D face mesh of the inter-frame, the eye blinking motion map, and the reconstructed key reference frame, mesh-based motion estimation is performed to obtain a fine-grained dense motion field and a face attention map, including: Calculate the vertex position difference between the 2D face mesh of the reconstructed key reference frame and the 2D face mesh of the inter-frame; The vertex position difference is interpolated into a coarse-grained grid motion flow consistent with the video resolution using a grid data interpolation function; The coarse-grained mesh motion flow and the reconstructed key reference frame are input into the U-shaped encoder-decoder network, and coarse-grained deformed frames are generated through feature warping operations. A spatial adaptive normalization mechanism is introduced to stitch together and reconstruct key reference frames, coarse-grained deformable frames, coarse-grained grid motion flow, and eye blinking motion map to form a multi-feature input. Fine-grained dense motion field and face attention map are obtained through bi-branch prediction.
[0092] Specifically, this feature describes the complete technical implementation of fine-grained dense motion field and face attention map at the decoder end, based on the vertex differences of two sets of 2D face meshes. It employs a five-stage cascaded process: "vertex difference calculation—interpolation expansion—coarse neural network deformation—multi-feature fusion—dual-branch fine prediction," ultimately achieving this. Its core lies in starting from sparse geometric vertex displacements and progressively advancing to dense pixel-level motion field and spatial attention distribution, achieving accurate transfer of facial dynamic semantics and focus on key regions. Specifically, the first stage performs vertex position difference calculation: the decoder subtracts the vertex position difference from the reconstructed 2D face mesh of the key reference frame (reflecting the 2D vertex coordinate set under the head pose of the key reference frame, serving as a static reference) and the 2D face mesh of the inter-frame (reflecting the 2D vertex coordinate set under the head pose and facial expression of the inter-frame, serving as a dynamic target) vertex-by-vertex coordinate subtraction. This difference is a sparse vector set defined only at the vertex positions of the 3D face mesh template, representing the coarse-grained geometric displacement caused by head pose changes and facial expression movements, but not yet covering all pixels at full video resolution.
[0093] The second stage involves interpolation expansion: The decoder uses a grid data interpolation function (emphasizing bilinear interpolation, which calculates displacement estimates for non-vertex positions by linearly weighting the displacement values of four adjacent vertices according to distance) to interpolate the sparse vertex position differences into a coarse-grained grid motion stream consistent with the video resolution. This coarse-grained grid motion stream is a dense two-dimensional vector field, with each pixel position corresponding to a displacement vector, representing the global pixel motion trend from the reconstructed key reference frame to the inter-frame frame. However, its resolution and accuracy are limited by the sparsity of the vertex grid, reflecting only large-scale head movements and significant facial expression changes, and not capturing the fine-grained dynamics of facial micro-expressions.
[0094] The third stage involves generating coarse-grained deformable frames: The decoder inputs the coarse-grained mesh motion flow and the reconstructed key reference frame into a U-shaped encoder-decoder network (UNet, a deep neural network with a symmetric encoder-decoder structure and skip connections. The encoder extracts multi-scale features through downsampling, the decoder restores spatial resolution through upsampling, and the skip connections directly pass the encoder features to the decoder to preserve detail information). The UNet encoder extracts multi-scale spatial features (including hierarchical visual information such as edges, textures, and semantics) from the reconstructed key reference frame, and performs feature warping operations (pixel resampling based on optical flow, i.e., spatial deformation of the feature map according to the motion flow vector, mapping the features of the reconstructed key reference frame to the corresponding positions in the inter-frame) using the coarse-grained mesh motion flow. This is then decoded by the UNet decoder to generate coarse-grained deformable frames. These deformable frames are intermediate results after global motion deformation of the reconstructed key reference frame, roughly presenting the facial layout and pose of the inter-frame, but blurring and misalignment still exist in detailed areas such as the mouth and eyes.
[0095] For the first to third stages, for example, computing and reconstructing the 2D face mesh of the key reference frame. 2D face mesh with inter-frame Vertex position difference ,in, for vertex coordinates, for The vertex coordinates. Interpolated using the grid data interpolation function. (Using bilinear interpolation) the vertex interpolation is converted into a coarse-grained mesh motion flow consistent with the video resolution. The coarse-grained mesh motion flow With the reconstruction of key reference frames Input U-type encoder-decoder network (UNet), UNet encoder extracts Multi-scale features, combined The feature warping operation (pixel resampling based on optical flow) is performed, and then decoded by the UNet decoder to generate coarse-grained deformed frames. The formula is ,in,( For the UNet encoder feature extraction process, For the UNet decoder feature decoding process, (This is a reverse twist operation).
[0096] The fourth stage performs multi-feature fusion: The decoder introduces a Spatially-Adaptive Normalization (SPADE) mechanism (a conditional normalization technique that adaptively adjusts batch normalization parameters based on the spatial distribution of the input semantic mask or feature map. By independently learning scaling and offset parameters for each spatial location, it avoids the erasure of semantic information by traditional batch normalization and better preserves the spatial structure and semantic details of the input features). Four types of features are then concatenated to form a multi-feature input: a reconstructed key reference frame (providing original texture and identity anchoring), a coarse-grained deformation frame (providing intermediate results for global deformation), a coarse-grained mesh motion flow (providing geometric constraints for global motion trends), and an eye blinking motion map (providing geometric constraints for local dynamic calibration of the eye region). The SPADE module adaptively adjusts the normalization parameters based on the spatial distribution of this multi-feature input, enabling the network to retain higher-precision feature responses in semantically critical regions such as the mouth and eyes, while moderately smoothing non-critical regions such as the background to reduce noise interference.
[0097] The fifth stage executes a dual-branch fine-grained prediction: the decoder processes multi-feature inputs through two independent prediction output branches. The first branch focuses on the prediction of fine-grained dense motion fields. This branch takes SPADE-modulated multi-features as input and outputs sub-pixel displacement vectors for each pixel through multi-layer convolution and upsampling operations. This represents the subtle pixel movements caused by facial micro-expressions (such as the upward movement of the corners of the mouth and the contraction of the nostrils, etc., sub-pixel deformations), achieving a precise transition from coarse-grained global motion to fine-grained local motion. The second branch focuses on the prediction of the face attention map. This branch also takes SPADE-modulated multi-features as input and outputs a spatial weight mask consistent with the video resolution through a spatial attention mechanism. Core facial semantic regions such as the mouth and eyes are given high weights, while non-core regions such as the background and hair are given low weights. This attention map achieves automatic localization and importance classification of key facial regions. Finally, the fine-grained dense motion field and the face attention map serve as the core driving inputs for the subsequent inter-frame generation stage. The dense motion field guides the spatial deformation of multi-scale spatial features, and the attention map regulates the degree of regional focus during the deformation process. The two work together to ensure that the final output of the reconstructed face image frames between frames accurately reproduces the dynamic semantics and facial details of the frames while maintaining identity consistency. This effectively solves the technical limitations of traditional analysis-synthesis models, which rely on manual models, have poor reconstruction quality, and are difficult to capture micro-expressions.
[0098] For the fourth and fifth stages, for example, a spatial adaptive normalization (SPADE) mechanism is introduced to stitch together and reconstruct key reference frames. coarse-grained deformable frames Coarse-grained mesh motion flow The system uses an image of blinking motion (ℇ) to form a multi-feature input; the SPADE module adaptively adjusts the normalization parameters based on the spatial distribution of the input features to better preserve semantic information; and fine-grained dense motion fields are obtained through dual prediction output branches. Face attention map (Key areas such as the mouth and eyes have high weights, while background areas have low weights), the formulas are as follows: ; ; in, , For two different prediction output branches, For feature splicing operations, This is the SPADE normalization process.
[0099] Thus, through a five-stage cascaded motion estimation process of "vertex difference—interpolation expansion—coarse neural network deformation—multi-feature fusion—dual-branch fine prediction," a precise progressive mapping of facial dynamic semantics from sparse geometric displacement to dense pixel motion is achieved, resulting in synergistic optimization of motion estimation accuracy and generation quality. Specifically, regarding motion estimation accuracy, the decoder first calculates the vertex position difference between the 2D face mesh of the reconstructed key reference frame and the 2D face mesh of the inter-frame. This difference directly quantifies the geometric displacement of the same identity template under two head poses, eliminating interference from identity changes and ensuring the purity of the motion signal. Then, the sparse vertex difference is expanded into a coarse-grained mesh motion flow consistent with the video resolution through a bilinear interpolation algorithm, achieving dense motion coverage from a finite set of vertices to the entire pixel domain, providing a global motion prior for subsequent neural network processing. Next, the coarse-grained mesh motion flow and the reconstructed key reference frame are input into UNet to perform feature warping, generating a coarse-grained deformed frame. This deformed frame, as an intermediate result, roughly presents the facial layout of the inter-frame, but details such as the mouth and eyes are still blurred. To address this, a Spatial Adaptive Normalization (SPADE) mechanism is introduced, stitching together key reference frames, coarse-grained deformable frames, coarse-grained grid motion flow, and eye blinking motion maps to form a multi-feature input. SPADE adaptively adjusts the normalization parameters based on the spatial distribution of each feature, preserving high-precision feature responses in semantically critical regions such as the mouth and eyes, while moderately smoothing and reducing noise in the background region. Finally, through bi-branch prediction, a fine-grained dense motion field (sub-pixel displacement vector for each pixel, accurately capturing micro-expressions such as upturned corners of the mouth and contracted nostrils) and a face attention map (a spatial mask with high weights for the mouth and eyes and low weights for the background) are output respectively. This progressive motion estimation strategy, from coarse to fine and from global to local, effectively solves the technical problems of insufficient granularity of motion estimation and distortion of micro-expression transfer in traditional methods, achieving sub-pixel-level accurate reproduction of facial dynamic semantics.
[0100] In some optional embodiments, based on a fine-grained dense motion field, a face attention map, and key reference frames for reconstruction, inter-frame generation is performed to obtain inter-frame reconstructed face image frames, including: The reconstructed key reference frames are input into a neural network encoder to extract multi-scale spatial features. By combining a fine-grained dense motion field with a face attention map, attention-based feature distortion operations are performed on multi-scale spatial features to obtain distorted face spatial features. Affine transformation parameters are generated from multi-scale spatial features using a neural network to modulate distorted facial spatial features, thereby obtaining transformed facial features. The generator network splices together distorted facial spatial features and transformed facial features, inputs them into a generator network, and outputs reconstructed facial image frames between frames.
[0101] Specifically, this scheme describes the complete technical implementation of the decoder based on the reconstructed key reference frame, fine-grained dense motion field, and face attention map. It employs a four-stage cascaded process of "multi-scale feature extraction—attention-based feature distortion—affine transformation modulation—feature stitching generation" to ultimately output a reconstructed face image frame. Its core lies in combining motion-guided spatial deformation with adaptive appearance modulation to achieve high-fidelity synthesis from static texture templates to dynamic face images. Specifically, the first stage performs multi-scale spatial feature extraction: the decoder inputs the reconstructed key reference frame into the neural network encoder (using the encoder part of a U-shaped encoder-decoder network, UNet). This encoder extracts spatial feature maps at four scales from the input image through multi-layer downsampling convolution operations, with resolutions of original video resolution (preserving fine texture and edge details), 1 / 2 original resolution (capturing local structure and mid-frequency information), 1 / 4 original resolution (extracting regional semantics and component relationships), and 1 / 8 original resolution (grasping global layout and pose context), thus obtaining multi-scale spatial features. The multi-scale features are organized in a pyramid shape, with the low-resolution layer containing rich semantic and contextual information, and the high-resolution layer preserving delicate texture and edge information, providing a hierarchical visual representation for subsequent multi-resolution deformation and synthesis.
[0102] The second stage performs attention-based feature distortion: The decoder combines a fine-grained dense motion field (predicted by the first branch of a two-branch system, representing the sub-pixel displacement vector of each pixel, accurately depicting spatial position changes caused by facial micro-expressions) with a face attention map (predicted by the second branch of a two-branch system, representing the spatial importance distribution of key regions such as the mouth and eyes with high weights and background regions with low weights) to perform attention-based feature distortion on multi-scale spatial features. Specifically, the face attention map is used as a weight mask, and the fine-grained dense motion field is used as a displacement field. Inverse distortion (i.e., spatial sampling of the feature map based on the inverse vector of the motion field, mapping the pixel values of the target position back to the source position) is used to achieve spatial deformation of multi-scale spatial features, resulting in distorted facial spatial features. This distorted feature retains the detailed structure after precise deformation of the motion field in high-weight regions of the attention map (such as the mouth and eyes). In low-weight regions (such as the background), moderate smoothing is applied to avoid noise amplification, achieving motion-guided semantic-aware feature deformation.
[0103] The third stage performs affine transformation modulation: The decoder generates affine transformation parameters from the multi-scale spatial features using a neural network (typically a 3×3 convolutional layer, predicting transformation parameters pixel-by-pixel from the multi-scale spatial features). These parameters include scaling factors (controlling the scaling ratio of local brightness and contrast) and offset factors (controlling the offset of local color and hue). The decoder uses these affine transformation parameters to perform modulation on the distorted facial spatial features, obtaining transformed facial features. This transformed feature compensates for the appearance distortion caused by changes in illumination and differences in material reflection during feature distortion, ensuring that the deformed features match the appearance attributes of the target frames in terms of color and brightness, achieving adaptive local hue correction.
[0104] The fourth stage involves feature stitching and output generation: The decoder stitches the distorted facial spatial features (carrying geometric structural information after motion field deformation and attention weighting) and the transformed facial features (carrying appearance attribute information after affine modulation) along the channel dimension (i.e., stacking the two feature maps in the depth direction to form a richer feature representation), and inputs it into the generator network (a deep generative network composed of multiple convolutions, activation functions, and residual connections, typically using a Generative Adversarial Network (GAN) architecture, which learns the distribution characteristics of real facial images through adversarial training). The generator network performs layer-by-layer convolution and nonlinear transformation on the stitched multi-channel features, gradually upsampling to restore them to the target resolution, and finally outputs a reconstructed inter-frame facial image frame. This output frame, while maintaining identity consistency with the key reference frame for reconstruction (sharing the same set of identity coefficients and reflectivity coefficients), accurately reproduces the dynamic semantics of mouth movements, eye blinks, and head poses corresponding to the inter-frame, achieving complete reconstruction from 14-dimensional compact semantic parameters to a high-fidelity two-dimensional facial image.
[0105] This solution could, for example, involve reconstructing the key reference frame. Inputting the UNet encoder extracts spatial feature maps at four scales (resolutions of the original video resolution, half the original resolution, quarter of the original resolution, and eighth of the original resolution), resulting in multi-scale spatial features. Combining fine-grained, dense sports fields Face attention map For multi-scale spatial features Perform the attention-based feature warping operation, the formula is: Among them, (⊙ represents the Hadamarda product, (For inverse warping operation). Then, multi-scale spatial features are extracted through a 3×3 convolutional layer. Generate affine transformation parameters (scaling factor α, offset factor β) according to the formula. Modulation distorts facial spatial features Obtain transformed facial features Finally, the spatial features of the distorted face are pieced together. With changing facial features Input generator network Through multi-layer convolution and activation functions, the inter-frame reconstructed frames are output. The formula is: .
[0106] In this way, the feature is generated through a four-stage cascaded process of "multi-scale feature extraction—attention-based feature distortion—affine transformation modulation—feature splicing generation," achieving high-fidelity synthesis from static texture templates to dynamic face images. Specifically, in terms of motion accuracy, the decoder inputs the reconstructed key reference frame into the neural network encoder to extract multi-scale spatial features. These features are organized in a pyramid-like manner with four resolution levels: the original resolution preserves fine texture, half resolution captures local structure, quarter resolution extracts regional semantics, and eighth resolution grasps the global pose, providing a hierarchical visual representation for subsequent deformation. Combining a fine-grained dense motion field with a face attention map, attention-based feature distortion is performed on the multi-scale spatial features. Using the attention map as a weight mask and the dense motion field as a displacement field, semantically perceptual feature deformation is achieved through inverse distortion. This ensures accurate motion transfer in key areas such as the mouth and eyes, while the background area is moderately smoothed to avoid noise interference. This effectively solves the technical problem of insufficient motion estimation granularity leading to micro-expression distortion in traditional methods, achieving sub-pixel-level accurate reproduction of facial dynamic semantics. In terms of appearance realism, while simple feature distortion can achieve geometric structure transfer, it is difficult to compensate for appearance distortion caused by changes in lighting and differences in material reflection. To address this, the decoder uses a neural network (3×3 convolutional layer) to generate affine transformation parameters from multi-scale spatial features (scaling coefficient controls brightness contrast, and offset coefficient controls color tone) to modulate and distort facial spatial features to obtain transformed facial features. This modulation operation performs local tone correction pixel by pixel and channel by channel, ensuring that the deformed features accurately match the appearance attributes of the target frames in terms of color and brightness. This avoids the "texture look" or "tone drift" caused by inconsistent lighting, significantly improving the realism and naturalness of the generated results. It effectively solves the technical problems of insufficient generation quality and robustness in existing 3D face modeling schemes, especially the tendency to distort under complex expressions and lighting conditions.
[0107] In some optional embodiments, in response to an interaction command, parameters in the 3D facial semantic representation of the reconstructed inter-frame are modified to generate an inter-frame reconstructed facial image frame corresponding to the interaction command, including: In response to the semantic parameter editing command, the target semantic parameters in the 3D face semantic representation of the reconstructed inter-frame are modified to obtain the modified 3D face semantic representation of the inter-frame, and the inter-frame reconstructed face image frame corresponding to the semantic parameter editing command is generated based on the modified 3D face semantic representation of the inter-frame. In response to the virtual character replacement command, the key reference frame is reconstructed using the virtual character image, and virtual character animation frames are generated.
[0108] Specifically, the scheme describes the differentiated processing flow of independent editing of semantic parameters and replacement of virtual characters after the decoder receives two different types of interactive instructions. Its core lies in using the high independence of each parameter dimension in the three-dimensional facial semantic representation and the replaceability of the fixed parameter buffer to realize semantic-level manipulation and identity-level re-creation of video content on the decoder side without re-encoding or transmitting additional bitstreams.
[0109] Specifically, the first mode is semantic parameter editing: when the decoder receives a semantic parameter editing instruction from the user (such as "increase the degree of mouth opening to match more exaggerated lip movements", "adjust the eye blinking parameter from 0.3 to 4.0 to make the eyes change from half-open to almost fully closed", "adjust the X-axis head rotation parameter from 0.2rad to -0.3rad to achieve the posture switching of the head from right to left", etc.), it first parses the instruction to determine the target semantic parameter to be modified and the target modification value. The target semantic parameters can be precisely selected from a 14-dimensional compact 3D facial semantic representation, including 6-dimensional mouth motion parameters (corresponding to the vertical, horizontal, and forward / backward motion components of the lips, precisely controlling the mouth shape during speech), 1-dimensional eye blinking parameters (values ranging from 0 to 5, where 0 represents fully open and 5 represents fully closed), 3-dimensional head rotation parameters (corresponding to the rotation arcs in the X / Y / Z axes, controlling head turning, nodding, and tilting postures), 3-dimensional head translation parameters (corresponding to the pixel translation amounts in the X / Y / Z axes, controlling the head's position movement in the image), or 1-dimensional head position parameters (representing the relative height position of the head in the entire image). The decoder replaces the original values of the corresponding parameters in the currently reconstructed 3D facial semantic representation with the target modified values, obtaining the modified 3D facial semantic representation of the inter-frame. This modification operation only involves the numerical replacement of specific semantic dimensions and does not affect the original values of other semantic parameters, fully benefiting from the independence condition that the absolute value of the Pearson correlation coefficient is no greater than 0.1 after PCA projection processing among the parameters. Subsequently, based on the modified inter-frame 3D facial semantic representation, the decoder re-executes the complete 3D facial mesh reconstruction process. Utilizing the modified mouth motion parameters (supplemented with zero-value expansion to complete expression coefficients) and the identity coefficients, reflectivity coefficients, and illumination coefficients in the fixed parameter buffer, it recalculates the 3D facial shape and texture using a parametric 3D deformation model (3DMM), obtaining a 3D facial mesh reflecting the modified expression state. Then, it re-executes 2D facial mesh projection (based on modified head rotation / translation parameters), eye motion calibration (based on modified blinking parameters), mesh basis motion estimation (calculating new vertex position differences and performing UNet feature warping and SPADE dual-branch prediction), and inter-frame generation (multi-scale feature extraction, attention-based feature warping, affine transformation modulation, and generator network output), ultimately generating an inter-frame reconstructed facial image frame corresponding to the semantic parameter editing instructions. This output frame, while maintaining the original identity features and texture quality, only presents changes conforming to the instructions in the user-specified semantic dimension, achieving real-time semantic manipulation with "one-time transmission, unlimited editing." The second mode is virtual character replacement.
[0110] When the decoder receives a virtual character replacement instruction from the user (such as "replace the current character with an anime character image" or "use a cartoon character as a virtual digital human"), the decoder obtains the user-defined virtual character image (supporting multiple resolutions from 256×256 to 1024×1024 to adapt to different display scene requirements). The decoder replaces the original reconstruction key reference frame with this virtual character image. Specifically, the virtual character image is input into the pre-trained 3D face reconstruction model WM3DR, from which virtual identity coefficients (identity feature vectors representing the geometric topology of the virtual character's face, replacing the identity coefficients in the original cache) and virtual reflectivity coefficients (reflectivity feature vectors representing the skin texture and material properties of the virtual character, replacing the reflectivity coefficients in the original cache) are extracted, while the original illumination coefficients are retained to maintain scene illumination consistency. The decoder updates the fixed parameter buffer based on the virtual identity coefficient and virtual reflectivity coefficient, while preserving the original reconstructed inter-frame 3D facial semantic representation (including dynamic semantic parameters such as mouth movement, eye blinking, and head pose). Then, using the updated fixed parameters (virtual identity coefficient, virtual reflectivity coefficient, and original illumination coefficient) and the original dynamic semantic parameters, it re-executes the complete 3D face mesh reconstruction, 2D projection, mesh basis motion estimation, and inter-frame generation process, ultimately generating a virtual character animation frame. This virtual character animation frame, while maintaining complete consistency with the original video in terms of mouth movement, eye blinking, and other dynamic semantics, replaces the facial identity with a user-specified virtual character image, achieving the dual goals of privacy protection (hiding the user's real identity information) and personalized content creation (custom virtual avatar).
[0111] In this way, the two interaction modes share the same decoupled architecture of "fixed parameters + dynamic semantic parameters". Editing semantic parameters only modifies the dynamic semantic parameters while keeping the fixed parameters unchanged. Replacing virtual characters only modifies the identity and reflectivity coefficients in the fixed parameters while keeping the dynamic semantic parameters unchanged. Neither of them requires re-requesting the encoding end or transmitting additional bitstreams. This fully demonstrates the flexibility and efficiency of semantic-level manipulation at the decoder end. It effectively solves the technical problems of traditional video encoding schemes not supporting semantic-level interaction and not being able to directly edit facial expressions or postures in the bitstream, as well as the technical problems of existing generative compression schemes having insufficient privacy protection mechanisms and not being able to effectively hide the user's real identity information.
[0112] In some optional embodiments, in response to a semantic parameter editing instruction, the target semantic parameters in the 3D face semantic representation of the reconstructed inter-frame are modified to obtain a modified 3D face semantic representation of the inter-frame, and an inter-frame reconstructed face image frame corresponding to the semantic parameter editing instruction is generated based on the modified 3D face semantic representation of the inter-frame, including: The target semantic parameter to be modified and the target modification value are determined from the semantic parameter editing instructions. The target semantic parameter is at least one of the following: mouth movement parameter, eye blinking parameter, head rotation parameter, head translation parameter, or head position parameter. The current value of the target semantic parameter in the reconstructed 3D face semantic representation of the inter-frame is replaced with the target modified value to obtain the modified 3D face semantic representation of the inter-frame. Based on the reconstructed key reference frame, the fixed parameters in the fixed parameter buffer, and the modified inter-frame 3D face semantic representation, 3D face mesh reconstruction, 2D face mesh projection, eye motion calibration, mesh basis motion estimation, and inter-frame generation are performed to generate inter-frame reconstructed face image frames with corresponding semantic parameter editing instructions.
[0113] Specifically, the scheme describes the complete technical implementation of the decoder after receiving the semantic parameter editing instruction, through a three-stage closed-loop process of "target parameter parsing - numerical replacement - full-link regeneration", and finally outputting the inter-frame reconstructed face image frame corresponding to the semantic parameter editing instruction. Its core lies in utilizing the high independence of each parameter dimension in the three-dimensional face semantic representation to achieve accurate editing and real-time visual feedback of specific semantic dimensions without re-encoding or transmitting additional bitstreams.
[0114] Specifically, the first stage involves determining the target semantic parameters and target modification values: the decoder parses the semantic parameter editing instructions issued by the user and identifies the target semantic parameters and target modification values to be modified from the instruction text or interactive interface input. The target semantic parameters can be precisely selected from at least one of the 14-dimensional compact 3D facial semantic representations, specifically including 6-dimensional mouth movement parameters (corresponding to the vertical, horizontal, and forward / backward movement components of the lips, accurately representing changes in mouth shape during speech, such as increasing the degree of mouth opening to match more exaggerated speech shapes) and 1-dimensional eye blinking parameters (values ranging from 0 to 5, where 0 represents fully open eyes, 5 represents fully closed eyes, and intermediate values correspond to different degrees of half-open state, such as adjusting from 0.3 to 4.0 to achieve a change in eye state from half-open to nearly fully closed). The system includes 3D head rotation parameters (corresponding to rotation radians in the X / Y / Z axes, representing head movements such as turning left and right, nodding up and down, and tilting to the side; for example, adjusting the X-axis rotation parameter from 0.2 rad to -0.3 rad achieves a head movement switch from right to left), 3D head translation parameters (corresponding to pixel translation amounts in the X / Y / Z axes, representing the head's forward / backward, left / right, and up / down movement within the image), or 1D head position parameters (representing the head's relative height position within the entire image, assisting in precise head region positioning). The target modification value is set directly by the user through the interactive interface. The decoder performs syntax parsing and numerical verification of the command to ensure that the target modification value is within the valid range of the corresponding parameters.
[0115] The second stage involves numerical replacement: The decoder reads the 3D facial semantic representation of the current reconstructed frame (containing 14 independent semantic parameters, each parameter satisfying the independence condition of an absolute Pearson correlation coefficient not exceeding 0.1 after PCA projection), locates the current value storage position of the target semantic parameter, and replaces this current value with the target modified value. The remaining 13 semantic parameters retain their original values, resulting in the modified 3D facial semantic representation of the frame. The independence of this replacement operation is guaranteed by PCA projection processing. The initial semantic representation (such as 68-dimensional facial behavior unit parameters) is filtered after PCA projection to obtain a 14-dimensional compact representation. The filtering criterion is that the absolute value of the Pearson correlation coefficient between each dimension after projection is not greater than 0.1, ensuring that semantic dimensions such as mouth movement, eye blinking, and head posture do not interfere with each other. This ensures that the modification of a single parameter will not cause unexpected linkage changes in other parameters, laying a statistical foundation for precise control. The third stage involves full-link regeneration. The decoder, based on reconstructed key reference frames (serving as high-quality texture templates and identity anchors), fixed parameters in the fixed parameter buffer (including identity coefficients, reflectivity coefficients, and illumination coefficients, determining the basic contours, skin texture, and lighting conditions of the face), and the modified 3D facial semantic representation of the inter-frame (containing updated target semantic parameter values and the original values of other unmodified parameters), sequentially executes the complete technical process: First, it performs 3D face mesh reconstruction, using modified mouth motion parameters (if the target semantic parameters include mouth motion parameters, then modified 6D vectors are used to supplement zero values to expand into complete expression coefficients; otherwise, the original values are used) and identity coefficients, reflectivity coefficients, and illumination coefficients in the fixed parameter buffer, to recalculate the 3D face shape and texture through a parametric 3D deformation model (3DMM) template, obtaining a 3D face mesh reflecting the modified expression state. Next, it performs 2D face mesh projection, performing Rodrigues transformation and perspective projection based on modified head rotation and head translation parameters (if modified, new values are used; otherwise, the original values are used), obtaining a 2D face mesh reflecting the modified head pose. Next, eye motion calibration is performed. Based on the modified blink parameters (if modified, the new values are used), the eye region is located and vertex positions are adjusted to generate an updated blink motion map. Then, mesh-based motion estimation is performed, calculating the vertex position difference between the 2D face mesh of the reconstructed key reference frame and the 2D face mesh of the modified inter-frame. Through interpolation, UNet feature warping, SPADE spatial adaptive normalization, and bi-branch prediction, a fine-grained dense motion field and face attention map are re-obtained. Finally, inter-frame generation is performed. The reconstructed key reference frame is input into a neural network encoder to extract multi-scale spatial features. Combined with the updated dense motion field and attention map, attention-based feature warping is performed. After affine transformation modulation, the features are concatenated and input into the generator network, outputting an inter-frame reconstructed face image frame with corresponding semantic parameter editing instructions.While maintaining the original identity features and texture quality, the output frame only presents changes that meet the instructions in the semantic dimension specified by the user, realizing an instant closed-loop response from abstract parameter editing to concrete visual feedback.
[0116] Thus, the technical advantage of this three-stage closed-loop process lies in achieving a direct mapping between "parameter-level control and pixel-level feedback": users do not need to understand the underlying neural network architecture or the principles of 3D deformation models, but only need to adjust the intuitive semantic parameters to control facial dynamics in real time. The decoder automatically completes the subsequent full-link regeneration with extremely low response latency, fully meeting the interactive needs of low-latency communication scenarios such as real-time video conferencing and virtual live streaming. It effectively solves the technical problems of traditional video encoding schemes that do not support semantic-level interaction and cannot directly edit facial expressions or postures in the bitstream.
[0117] In some optional embodiments, in response to a virtual character replacement instruction, a virtual character animation frame is generated by replacing the reconstructed key reference frame with a virtual character image, including: Get user-defined virtual character images; The virtual character image is input into a pre-trained face reconstruction model to extract the virtual identity coefficient and virtual reflectivity coefficient. The virtual fixed parameters are obtained by replacing the identity coefficient and reflectivity coefficient in the fixed parameter buffer with the virtual identity coefficient and virtual reflectivity coefficient. Based on virtual fixed parameters and the 3D face semantic representation of the reconstructed inter-frames, 3D face mesh reconstruction, 2D face mesh projection, eye movement calibration, mesh basis motion estimation and inter-frame generation are performed to generate virtual character animation frames. The mouth movement and eye blinking dynamic semantics of the virtual character animation frames are consistent with the 3D face semantic representation of the reconstructed inter-frames.
[0118] Specifically, this scheme describes the complete technical implementation of generating virtual character animation frames by the decoder after receiving a virtual character replacement instruction, through a three-stage closed-loop process of "virtual character image acquisition—fixed parameter extraction and replacement—end-to-end regeneration". Its core lies in utilizing the substitutability of identity coefficients and reflectivity coefficients in the fixed parameter buffer to achieve flexible switching of facial identities while maintaining the original dynamic semantics, thus achieving the dual goals of privacy protection and personalized content creation.
[0119] Specifically, the first stage involves acquiring the virtual character image: the decoder receives the user's virtual character replacement command and retrieves the user-defined virtual character image from local storage or a network interface. This virtual character image can be anime characters, cartoon figures, virtual digital humans, or other digital images that are not real human faces. Resolutions support multiple sizes from 256×256 to 1024×1024 to adapt to different display scenarios. Users can upload images through an interactive interface or select from a preset character library. The decoder performs format verification and size normalization on the input image to ensure consistency in subsequent model inputs.
[0120] The second stage involves fixed parameter extraction and replacement: The decoder inputs the virtual character image into a pre-trained face reconstruction model, extracting virtual identity coefficients (identity feature vectors representing the virtual character's facial geometry, linearly combined with 3DMM identity basis vectors to determine the virtual character's basic outline, skeletal structure, and facial geometry, such as large eyes, pointed chins, and other iconic anime-style geometric features) and virtual reflectivity coefficients (reflectivity feature vectors representing the virtual character's skin texture and material properties, superimposed with 3DMM reflectivity basis vectors to determine the virtual character's skin tone, texture details, and material gloss, such as solid color fills in cartoon style or glowing materials in cyberpunk style). The decoder replaces the original identity coefficients in the fixed parameter buffer with virtual identity coefficients (the original identity coefficients are extracted from real reconstruction key reference frames, representing the user's real facial geometry), and replaces the original reflectivity coefficients in the fixed parameter buffer with virtual reflectivity coefficients (the original reflectivity coefficients represent the user's real skin texture and material properties), while retaining the original illumination coefficients unchanged (to maintain the consistency of the original shooting scene's lighting conditions and ensure the natural integration of the virtual character with the background environment), thus obtaining virtual fixed parameters. This replacement operation only involves overwriting two coefficients in the buffer, without affecting the illumination coefficients in the buffer or the 3D face semantic representation (containing 14-dimensional dynamic semantic parameters) between reconstructed frames, thus achieving complete decoupling between static identity attributes and dynamic semantic parameters.
[0121] The third stage involves full-link regeneration: The decoder, based on virtual fixed parameters (virtual identity coefficients, virtual reflectivity coefficients, and original illumination coefficients) and the reconstructed 3D facial semantic representation from the inter-frames (preserving the original mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters to ensure that the dynamic semantics are completely consistent with the original video), sequentially executes the complete technical process: First, 3D facial mesh reconstruction is performed, using virtual identity coefficients and virtual reflectivity coefficients to replace the original fixed parameters. Combined with the mouth movement parameters in the reconstructed 3D facial semantic representation (supplementing zero values to expand into complete expression coefficients), the 3D face shape and texture are recalculated using a parametric 3D deformation model (3DMM) template, resulting in a 3D facial mesh that reflects the virtual character's identity features and the original dynamic expressions (e.g., the large eye contour of an anime character matching the lip movements of the original video). Next, 2D facial mesh projection is performed, using Rodrigues transformation and perspective projection based on the head rotation and translation parameters in the reconstructed 3D facial semantic representation to obtain a 2D facial mesh reflecting the original head posture of the virtual character. Next, eye motion calibration is performed. Based on the original blinking parameters, the eye region is located and the vertex positions are adjusted to generate an eye blinking motion map, ensuring that the degree of eye closure of the virtual character accurately matches the original semantic parameter values. Subsequently, mesh-based motion estimation is performed, calculating the vertex position difference between the 2D face mesh of the virtual character and the 2D face mesh of the reconstructed key reference frame. After interpolation, UNet feature warping, SPADE spatial adaptive normalization, and bi-branch prediction, a fine-grained dense motion field and a face attention map are obtained. Finally, inter-frame generation is performed. The reconstructed key reference frame (here, the original key reference frame is retained as a texture style reference, or optionally replaced by the stylized texture of the virtual character) is input into the neural network encoder to extract multi-scale spatial features. Combined with the dense motion field and attention map, attention-based feature warping is performed. After affine transformation modulation, the features are concatenated and input into the generator network to output the virtual character animation frames. The virtual character animation frames maintain strict consistency with the 3D facial semantic representation of the reconstructed frames in dynamic semantic dimensions such as mouth movements (the lip shape is completely consistent with the original video) and eye blinking (the degree of closure is precisely matched with the original semantic parameters). It only presents the virtual character image specified by the user in terms of facial identity, realizing flexible re-creation with "dynamic semantics unchanged and static identity changeable".
[0122] Thus, the technical advantage of this three-stage closed-loop process lies in achieving an organic unity between privacy protection and personalized creation: users can participate in video communication without exposing their real facial information, effectively solving the technical problems of insufficient privacy protection mechanisms and inability to effectively hide users' real identity information in existing encoding schemes; at the same time, the customization space for virtual character images is extremely large (from anime characters to brand IP images), providing rich personalized expression means for scenarios such as virtual live streaming, online social networking, and game interaction, significantly reducing the threshold and cost of digital human content creation, and effectively solving the technical limitations of traditional 3D face modeling, which lacks an end-to-end framework designed for interactive encoding and cannot balance low bitrate and controllability.
[0123] For a better understanding of the encoder's workflow, please refer to [link / reference]. Figure 4 , Figure 4 This is a flowchart illustrating the decoder-side process of a face video communication method according to one embodiment of this disclosure. The flowchart demonstrates the complete workflow of the decoder-side in the Interactive Face Video Semantic Transmission (IFVC) framework proposed in this application, covering the entire process of bitstream parsing, dual-track decoding and reconstruction, interactive control, 3D geometric reconstruction, motion estimation, and frame generation, ultimately outputting an interactive 2D face video stream. The following sections describe each module according to the data flow: Bitstream type determination and dual-track decoding: After receiving the transmitted bitstream from the encoder, the decoder first performs bitstream type determination to distinguish between two types of data: the key reference frame bitstream and the inter-frame semantic bitstream. For the key reference frame bitstream, it enters VVC decoding: the key reference frame reconstruction module calls the Multi-Function Video Coding (VVC) decoder to perform the standard intra-frame decoding process to obtain the reconstructed key reference frame. This frame retains high-quality texture information as the basic texture template for the entire video. At the same time, three types of fixed parameters are extracted in parallel from the reconstructed key reference frame: identity coefficient, reflectivity coefficient, and illumination coefficient. These three parameters are stored together in the fixed parameter buffer for reuse throughout the entire video. For the inter-frame semantic bitstream, semantic decoding is performed: inverse quantization + residual compensation, entropy decoding (the inverse process of PPM context arithmetic coding at the encoder end), inverse quantization (using a quantization step size matched with the encoder to recover the prediction residual) and semantic compensation (adding the residual to the prediction baseline) to obtain the reconstructed three-dimensional facial semantic representation of the inter-frame, which includes 14 compact semantic parameters (6-dimensional mouth motion parameters, 1-dimensional eye blinking parameters, 3-dimensional head rotation parameters, 3-dimensional head translation parameters and 1-dimensional head position parameters).
[0124] The interactive control branch allows the reconstructed 3D facial semantic representation to enter the semantic interactive editing module. This module supports user-issued semantic parameter editing commands, allowing users to modify the current values of target semantic parameters through adjustments to mouth movement, eye blinking, and head posture modules (e.g., adjusting mouth movement parameters to change lip shape, modifying eye blinking parameters to control eye opening and closing, and changing head rotation parameters to switch postures). This results in a modified 3D facial semantic representation. This module fully utilizes the independence condition that the absolute value of the correlation coefficient of each parameter after PCA projection processing is no greater than 0.1, enabling independent editing of specific dimensions without affecting other parameters. Simultaneously, a virtual character replacement module is supported: the user inputs a virtual character image as a new texture template. The pre-trained face reconstruction model extracts the virtual character's identity / reflectivity parameters, replacing the original identity and reflectivity coefficients in the fixed parameter buffer, while retaining the original illumination coefficients and dynamic semantic parameters, providing an alternative identity source for generating virtual character animation frames. The parameters after semantic editing or virtual character replacement, along with the original parameters, serve as complete input for subsequent 3D reconstruction.
[0125] After parameter aggregation, the process proceeds to 3D face mesh reconstruction: A shape S / texture T module is generated based on a 3DMM template. Using a parametric 3D deformation model (3DMM) template, identity coefficients and identity basis vectors, reflectivity coefficients and reflectivity basis vectors, and illumination coefficients and illumination basis vectors are linearly combined to calculate the 3D face shape S and 3D face texture T. Simultaneously, mouth motion parameters are supplemented with zero values to expand into complete expression coefficients, which are then superimposed onto the 3D face shape to obtain a 3D face mesh reflecting a specific identity and expression. Next, the process proceeds to 3D→2D projection: A 2D face mesh M module is generated using a rotation matrix R and translation vector t. Head rotation parameters (converted to a 3×3 rotation matrix R via Rodrigues transformation) and head translation parameters (converted to a 3D translation vector t) are extracted from the 3D face semantic representation. Combined with the camera intrinsic parameter matrix, perspective projection is performed to map the 3D face mesh vertex set onto a 2D plane, resulting in a 2D face mesh M that reflects the head pose of the key reconstruction reference frame and the frames between reconstructions.
[0126] Next, the eye motion calibration module adjusts the vertex position of the eye region based on blink parameters. It locates the eye region within the 2D face mesh (determined by a predefined range of eye vertex indices) according to the blink parameters, finds the original highest and lowest points of the eye region, calculates the new highest point position using linear interpolation, dynamically adjusts the ordinate of the eye mesh vertices, and generates an eye blink motion map to ensure accurate matching between the degree of eye closure and semantic parameters. Then, the mesh basis motion estimation module calculates the vertex position difference between the 2D face mesh of the reconstructed key reference frame and the 2D face mesh of the inter-frame. The coarse-grained motion field calculation module expands the vertex difference into a coarse-grained mesh motion flow consistent with the video resolution using bilinear interpolation. This coarse-grained motion flow and the reconstructed key reference frame are then input into UNet to generate coarse-grained deformed frames. The UNet encoder extracts multi-scale features and performs feature warping in conjunction with the motion flow, and the decoder outputs the coarse-grained deformed frames. Coarse deformation frames are then optimized into fine-grained features: Spatial Adaptive Normalization (SPADE) mechanism is introduced to stitch together and reconstruct key reference frames, coarse deformation frames, coarse motion flow and eye blinking motion map to form multi-feature input. SPADE adaptively adjusts the normalization parameters according to the spatial distribution, and outputs fine-grained dense motion field and face attention map through bi-branch prediction.
[0127] The fine-grained motion field and attention map are fed into the frame generation module, which, based on a generative adversarial network architecture, extracts style and texture priors from the reconstructed key reference frames. Then, the module inputs the reconstructed key reference frames into a neural network encoder to extract multi-scale spatial features. These features are then combined with the fine-grained dense motion field and face attention map to perform attention-based feature warping (using the attention map as weights and the motion field as displacement). A 3×3 convolutional layer is used to generate affine transformation parameters (scaling factor α and offset factor β) to modulate the warped features, resulting in transformed face features. After concatenating the warped face spatial features and the transformed face features, the generator network outputs the final reconstructed 2D inter-frame. Finally, the reconstructed key reference frames and the inter-frames are merged into the final output: an interactive 2D face video stream, forming a complete interactive face video sequence. The reconstructed key reference frames are directly output as texture anchors, while the inter-frames support dynamic reconstruction results after semantic editing or virtual character replacement. Together, they achieve semantic-level interactive face video communication with "one-time transmission, unlimited editing."
[0128] To facilitate a comprehensive understanding of the technical solution disclosed herein, this embodiment describes the communication process between the encoder and decoder. Specifically, the encoder classifies the input facial video frame sequence to obtain key reference frames and inter-frames, and compresses and encodes the key reference frames to obtain key reference frame encoded data. For the inter-frames, the inter-frames are mapped from two-dimensional image frames to a three-dimensional parameter space, and a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters is obtained. The correlation between different parameters satisfies a preset independence condition. The three-dimensional facial semantic representation is encoded to obtain inter-frame encoded data. The key reference frame encoded data and inter-frame encoded data are integrated into a transmission stream, and the transmission stream is sent to the decoder.
[0129] The decoder decodes the received transmission stream to obtain the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame; based on the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, it generates the inter-frame reconstructed face image frame; in response to the interaction command, it modifies the parameters in the 3D face semantic representation of the reconstructed inter-frame to generate the inter-frame reconstructed face image frame corresponding to the interaction command.
[0130] Correspondingly, see Figure 5 , Figure 5 This is a schematic diagram of the overall process of a face video communication method according to one embodiment of this disclosure. The flowchart illustrates the complete working principle of the interactive face video semantic transmission (IFVC) framework proposed in this application, covering three major modules: encoder, transmission channel, and decoder. Each functional module works together to realize the entire chain process of "semantic extraction - compression encoding - transmission - decoding reconstruction - interactive control - frame generation".
[0131] The first layer involves input source and encoder processing: A continuous sequence of face video frames serves as the input source, encompassing natural facial dynamics such as speaking, blinking, head rotation, and translation. The encoder first performs frame-level classification on the input frame sequence, distinguishing between key reference frames and inter-frames. Key reference frames are compressed using the VVC encoding module with the Multifunctional Video Coding intra-frame coding standard, preserving high-quality texture information and outputting key reference frame encoded data. For inter-frames, the 3D face semantic extraction module works in conjunction with the pre-trained WM3DR model and the OpenFace model to map 2D face images to a 3D parameter space, obtaining a 14-dimensional compact 3D face semantic representation. PCA projection is used to ensure the independence condition that the absolute value of the correlation coefficient between parameters is no greater than 0.1. Then, an adder integrates the inter-frame prediction residuals, which are then quantized using the zero-order exponential Golomb algorithm by the quantization module, forming the inter-frame encoded data.
[0132] The second layer is transmission channel integration: the encoder integrates the key reference frame encoded data and inter-frame encoded data into a transmission stream according to a specific structure, and sends it to the decoder through the communication network; the stream structure includes a key reference frame index table, a fixed parameter block (storing identity coefficients, reflectivity coefficients, and illumination coefficients, which are transmitted only once for the entire video) and an inter-frame residual encoded block, supporting random access and semantic editing tags.
[0133] The third layer is the decoder-side processing and output. The decoder first performs intra-frame decoding on the received key reference frame encoded data using the VVC decoding module to obtain the reconstructed key reference frame. Identity coefficients, reflectivity coefficients, and illumination coefficients are extracted from this reconstructed frame and stored in a fixed-parameter buffer for global reuse. Simultaneously, the dequantization module performs the inverse entropy decoding process on the inter-frame encoded data to recover the quantized prediction residual. In conjunction with the 3D face semantic extraction module, the prediction residual is semantically compensated with the prediction baseline (the reconstructed key reference frame or the 3D face semantic representation of the preceding reconstructed inter-frame) at the adder. This is then added dimension-by-dimensional to obtain the 3D face semantic representation of the reconstructed inter-frame and stored in the decoded face semantic buffer pool (i.e., a semantic buffer that stores the reconstructed semantic representation of the most recent preset number of frames to avoid error accumulation). The semantic representations in this cache pool support direct access and modification by the controllable semantic editing module: responding to user semantic parameter editing commands, the target semantic parameters and target modification values are determined, the current values of the corresponding parameters in the cache pool are replaced with the target modification values, and the modified 3D face semantic representation is obtained and re-stored in the cache pool, enabling real-time control of dynamic semantics such as mouth movements, eye blinks, and head posture at the decoder end. Subsequently, the 3D face mesh reconstruction module, based on the fixed parameters in the fixed parameter cache area and the reconstructed (or modified) 3D face semantic representations between frames, calculates the 3D face shape and texture using a parameterized 3D deformation model template, and generates a 3D face mesh by superimposing expression coefficients composed of supplemented zero values. At the same time, the virtual character reference module supports the input of user-defined virtual character images (anime characters, cartoon characters, etc.), extracts virtual identity coefficients and virtual reflectivity coefficients through a pre-trained face reconstruction model, replaces the original corresponding coefficients in the fixed parameter cache area, and retains the original illumination coefficients and dynamic semantic parameters unchanged, providing an alternative identity source for virtual character animation generation. Next, the mesh-based motion estimation module performs 2D face mesh projection and eye motion calibration: head rotation and translation parameters are extracted from the 3D face semantic representations of the reconstructed key reference frame and the inter-frame reconstruction, respectively, and converted into rotation matrices via Rodrigues transform. Perspective projection is then performed using the camera intrinsic parameter matrix to obtain two sets of 2D face meshes. The eye region is located based on the blink parameters, and vertex positions are adjusted to generate a blink motion map. The vertex position difference between the two sets of 2D face meshes is calculated, and through interpolation, UNet feature warping, SPADE spatial adaptive normalization, and bi-branch prediction, a fine-grained dense motion field (sub-pixel level displacement vector) and a face attention map (a high-weight spatial mask for key regions) are obtained. Finally, the frame generation module inputs the reconstructed key reference frame into the neural network encoder to extract multi-scale spatial features. It then performs attention-based feature warping by combining the fine-grained dense motion field and the face attention map. After affine transformation modulation, the resulting image is concatenated and input into the generator network, outputting two types of transmission results: inter-frame reconstructed face image frames generated based on the original identity parameters and virtual character animation frames generated based on virtual character parameters.
[0134] In summary, this IFVC framework achieves interactive face video communication with "one-time transmission and unlimited editing" by semantically decoupled compression at the encoder end, compact bitstream integration at the transmission channel, and semantic-level manipulation and high-quality reconstruction at the decoder end. It breaks through the limitations of the traditional scheme's separate architecture of "compression at the encoder end and passive reconstruction at the decoder end" and takes into account the integrated requirements of ultra-low bitrate transmission, real-time semantic interaction, high-fidelity reconstruction, and flexible privacy protection.
[0135] The following describes an embodiment of the apparatus described in this application, which can be used to execute the face video communication method described above in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the face video communication method described above in this application.
[0136] This disclosure also provides a face video communication device 600, such as... Figure 6 As shown, it includes: The classification module 601 is used to classify the input face video frame sequence to obtain key reference frames and inter-frames, and to compress and encode the key reference frames to obtain key reference frame encoded data. The mapping module 602 is used to map the inter-frame from the two-dimensional image frame to the three-dimensional parameter space for the inter-frame, and to filter to obtain a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters and head position parameters, wherein the correlation between different parameters satisfies the preset independence condition. Encoding module 603 is used to encode the semantic representation of a three-dimensional face to obtain inter-frame encoded data; The transmission module 604 is used to integrate the key reference frame encoded data and the inter-frame encoded data into a transmission bitstream, and send the transmission bitstream to the decoder. The decoder is used to decode the received transmission bitstream to obtain the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. Based on the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, an inter-frame reconstructed face image frame is generated. In response to the interaction command, the parameters in the 3D face semantic representation of the reconstructed inter-frame are modified to generate the inter-frame reconstructed face image frame corresponding to the interaction command.
[0137] In some optional embodiments, the mapping module 602 maps the inter-frame from a two-dimensional image frame to a three-dimensional parameter space, and filters to obtain a three-dimensional facial semantic representation including mouth motion parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters, including: A pre-trained 3D face reconstruction model is used to predict the shape parameters, texture parameters, and dynamic expression parameters of the 3D face from the 2D face images between frames, thus obtaining the initial 3D face parameters. A facial behavior analysis model is used to predict eye blinking parameters from inter-frame two-dimensional face images; The parameters representing the lip movement components are extracted from the dynamic expression parameters in the initial three-dimensional face parameters to obtain the mouth movement parameters; The head rotation parameters, head translation parameters, and head position parameters are extracted from the initial 3D face parameters. The head rotation parameters correspond to the rotation radians in each coordinate axis direction in 3D space, the head translation parameters correspond to the pixel translation amount in each coordinate axis direction in 3D space, and the head position parameters represent the relative height position of the head in the image. By integrating mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters, a three-dimensional facial semantic representation is obtained.
[0138] In some optional embodiments, the mapping module 602 integrates mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters to obtain a three-dimensional facial semantic representation, including: The initial semantic representation is obtained by integrating the mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters. Principal component analysis spatial projection processing is performed on the initial semantic representation to make the correlation between parameters less than or equal to a preset threshold, so as to meet the preset independence condition and obtain a three-dimensional face semantic representation.
[0139] In some optional embodiments, the classification module 601 classifies the input face video frame sequence to obtain key reference frames and inter-frames, including: Calculate the differences in three-dimensional facial semantic representations between adjacent frames in a face video frame sequence, including differences in head rotation parameters and mouth motion parameters; When the difference in head rotation parameters exceeds the first preset threshold, or the difference in mouth movement parameters exceeds the second preset threshold, the current frame is marked as a critical reference frame. The remaining frames that were not marked as key reference frames are identified as interframes.
[0140] In some optional embodiments, the encoding module 603 encodes the three-dimensional facial semantic representation to obtain inter-frame encoded data, including: Using the 3D face semantic representation of the key reference frame or the 3D face semantic representation of the preceding inter-frame as the prediction benchmark, the 3D face semantic representation of the current inter-frame is predicted to obtain the prediction residual. The predicted residuals are quantized and converted into binary code; Binary codes are encoded using a context-based arithmetic coding model to generate inter-frame encoded data.
[0141] In some optional embodiments, the encoding module 603 uses the 3D face semantic representation of the key reference frame or the 3D face semantic representation of the preceding inter-frame as a prediction benchmark to predict the 3D face semantic representation of the current inter-frame, obtaining a prediction residual, including: When processing the first inter-frame, the three-dimensional face semantic representation of the key reference frame is obtained as the prediction benchmark. The difference between the three-dimensional face semantic representation of the first inter-frame and the three-dimensional face semantic representation of the key reference frame is calculated to obtain the first prediction residual. When processing subsequent frames between the first and second frames, the 3D face semantic representation of the previous frame is obtained as the prediction benchmark. The difference between the 3D face semantic representation of the current frame and the 3D face semantic representation of the previous frame is calculated to obtain the second prediction residual.
[0142] This disclosure also provides a face video communication device 700, such as Figure 7 As shown, it includes: The receiving module 701 is used to receive the transmission bitstream from the encoder. The encoder classifies the input face video frame sequence to obtain key reference frames and inter-frames, and compresses and encodes the key reference frames to obtain key reference frame encoded data. For the inter-frames, the inter-frames are mapped from two-dimensional image frames to three-dimensional parameter space, and three-dimensional face semantic representations containing mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters are obtained. The correlation between different parameters satisfies a preset independence condition. The three-dimensional face semantic representations are encoded to obtain inter-frame encoded data. The key reference frame encoded data and inter-frame encoded data are integrated into a transmission bitstream, and the transmission bitstream is sent to the decoder. The decoding module 702 is used to decode the received transmission bitstream to obtain the three-dimensional face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. The reconstruction module 703 is used to generate inter-frame reconstructed face image frames based on the three-dimensional face semantic representation of the reconstruction key reference frame and the inter-frame reconstruction. The interaction module 704 is used to modify the parameters in the 3D face semantic representation of the reconstructed inter-frame in response to the interaction command, and generate the inter-frame reconstructed face image frame corresponding to the interaction command.
[0143] In some optional embodiments, the decoding module 702 decodes the received transmission bitstream to obtain the three-dimensional face semantic representation of the reconstructed key reference frame and the inter-reconstruction frames, including: For the encoded data of the key reference frame in the transmitted bitstream, perform key reference frame decoding and reconstruction to obtain the reconstructed key reference frame; For the inter-frame encoded data in the transmission bitstream, perform inter-frame semantic representation decoding and reconstruction to obtain the three-dimensional face semantic representation of the reconstructed inter-frame.
[0144] In some optional embodiments, the decoding module 702 performs key reference frame decoding and reconstruction on the key reference frame encoded data in the transmission bitstream to obtain the reconstructed key reference frame, including: The key reference frame is reconstructed by decoding the encoded data of the key reference frame using a decoder. Fixed parameters are extracted from the reconstructed key reference frame and stored in the fixed parameter buffer. The fixed parameters include identity coefficient, reflectivity coefficient and illumination coefficient.
[0145] In some optional embodiments, the decoding module 702 performs inter-frame semantic representation decoding and reconstruction on the inter-frame encoded data in the transmitted bitstream to obtain the reconstructed three-dimensional face semantic representation of the inter-frame, including: Entropy decoding is performed on the inter-frame encoded data to obtain the quantized prediction residual; Perform an inverse quantization operation on the quantized prediction residuals to recover the prediction residuals; Obtain the 3D face semantic representation corresponding to the prediction benchmark, where the prediction benchmark is the key reference frame for reconstruction or the preceding reconstructed frame. The predicted residuals are semantically compensated with the 3D face semantic representations corresponding to the prediction baseline to calculate the 3D face semantic representations of the reconstructed frames.
[0146] In some optional embodiments, the reconstruction module 703 generates inter-frame reconstructed face image frames based on the 3D face semantic representation of the reconstruction key reference frame and the inter-frame reconstruction, including: Based on the reconstructed key reference frame, fixed parameters are extracted, including: identity coefficient, reflectivity coefficient, and illumination coefficient. Based on fixed parameters and the semantic representation of the 3D face between reconstructed frames, 3D face mesh reconstruction is performed to obtain the 3D face mesh; Based on the 3D face mesh, the 3D face semantic representation of the reconstructed key reference frame, and the 3D face semantic representation of the reconstructed inter-frame, the 2D face mesh projection and eye motion calibration are performed to obtain the 2D face mesh of the reconstructed key reference frame, the 2D face mesh of the inter-frame, and the eye blinking motion map. Based on the reconstructed key reference frame, the inter-frame 2D face mesh, the eye blinking motion map, and the reconstructed key reference frame, mesh-based motion estimation is performed to obtain a fine-grained dense motion field and a face attention map. Based on the fine-grained dense motion field, face attention map, and reconstruction key reference frames, inter-frame generation is performed to obtain inter-frame reconstructed face image frames.
[0147] In some optional embodiments, the reconstruction module 703 performs 3D face mesh reconstruction based on fixed parameters and the 3D face semantic representation between reconstruction frames to obtain a 3D face mesh, including: Based on the identity coefficient, reflectivity coefficient, and illumination coefficient, combined with the average neutral shape, identity basis vector, average neutral texture, reflectivity basis vector, and illumination basis vector, the three-dimensional face shape and three-dimensional face texture are calculated. Based on the mouth motion parameters in the reconstructed 3D facial semantic representation between frames, zero values are added to form expression coefficients; Based on the 3D face shape, 3D face texture, and expression coefficients, a 3D face mesh is calculated using a parametric 3D deformation model template.
[0148] In some optional embodiments, the reconstruction module 703, based on the 3D face mesh, the 3D face semantic representation of the reconstructed key reference frame, and the 3D face semantic representation of the reconstructed inter-frames, performs 2D face mesh projection and eye motion calibration to obtain the 2D face mesh of the reconstructed key reference frame, the 2D face mesh of the inter-frames, and an eye blinking motion map, including: The first head rotation parameter and the first head translation parameter are extracted from the 3D face semantic representation of the reconstructed key reference frame. The first head rotation parameter is converted into the first head rotation matrix and the first head translation parameter is converted into the first translation vector. Combined with the camera intrinsic parameter matrix, the 3D face mesh is projected onto the 2D plane based on the first head rotation matrix and the first translation vector to obtain the 2D face mesh of the reconstructed key reference frame. The second head rotation parameter and the second head translation parameter are extracted from the 3D face semantic representation of the reconstructed inter-frames. The second head rotation parameter is converted into a second head rotation matrix and the second head translation parameter is converted into a second translation vector. Combined with the camera intrinsic parameter matrix, the 3D face mesh is projected onto the 2D plane based on the second head rotation matrix and the second translation vector to obtain the 2D face mesh of the inter-frames. Based on the blinking parameters in the 3D face semantic representation of the reconstructed frames, the eye region in the 2D face mesh of the frames is located, the original highest and lowest points of the eye region are determined, the new highest point is calculated based on the blinking parameters, the position of the eye mesh vertex is adjusted, and the blinking motion map is generated.
[0149] In some optional embodiments, the reconstruction module 703 performs mesh-based motion estimation based on the 2D face mesh of the reconstructed key reference frame, the 2D face mesh of the inter-frame, the eye blinking motion map, and the reconstructed key reference frame to obtain a fine-grained dense motion field and a face attention map, including: Calculate the vertex position difference between the 2D face mesh of the reconstructed key reference frame and the 2D face mesh of the inter-frame; The vertex position difference is interpolated into a coarse-grained grid motion flow consistent with the video resolution using a grid data interpolation function; The coarse-grained mesh motion flow and the reconstructed key reference frame are input into the U-shaped encoder-decoder network, and coarse-grained deformed frames are generated through feature warping operations. A spatial adaptive normalization mechanism is introduced to stitch together and reconstruct key reference frames, coarse-grained deformable frames, coarse-grained grid motion flow, and eye blinking motion map to form a multi-feature input. Fine-grained dense motion field and face attention map are obtained through bi-branch prediction.
[0150] In some optional embodiments, the reconstruction module 703 performs inter-frame generation based on a fine-grained dense motion field, a face attention map, and key reconstruction reference frames to obtain inter-frame reconstructed face image frames, including: The reconstructed key reference frames are input into a neural network encoder to extract multi-scale spatial features. By combining a fine-grained dense motion field with a face attention map, attention-based feature distortion operations are performed on multi-scale spatial features to obtain distorted face spatial features. Affine transformation parameters are generated from multi-scale spatial features using a neural network to modulate distorted facial spatial features, thereby obtaining transformed facial features. The generator network splices together distorted facial spatial features and transformed facial features, inputs them into a generator network, and outputs reconstructed facial image frames between frames.
[0151] In some optional embodiments, the interaction module 704, in response to an interaction command, modifies the parameters in the 3D facial semantic representation of the reconstructed inter-frame to generate an inter-frame reconstructed facial image frame corresponding to the interaction command, including: In response to the semantic parameter editing command, the target semantic parameters in the 3D face semantic representation of the reconstructed inter-frame are modified to obtain the modified 3D face semantic representation of the inter-frame, and the inter-frame reconstructed face image frame corresponding to the semantic parameter editing command is generated based on the modified 3D face semantic representation of the inter-frame. In response to the virtual character replacement command, the key reference frame is reconstructed using the virtual character image, and virtual character animation frames are generated.
[0152] In some optional embodiments, the interaction module 704, in response to a semantic parameter editing instruction, modifies the target semantic parameters in the 3D face semantic representation of the reconstructed inter-frame to obtain a modified 3D face semantic representation of the inter-frame, and generates an inter-frame reconstructed face image frame corresponding to the semantic parameter editing instruction based on the modified 3D face semantic representation of the inter-frame, including: The target semantic parameter to be modified and the target modification value are determined from the semantic parameter editing instructions. The target semantic parameter is at least one of the following: mouth movement parameter, eye blinking parameter, head rotation parameter, head translation parameter, or head position parameter. The current value of the target semantic parameter in the reconstructed 3D face semantic representation of the inter-frame is replaced with the target modified value to obtain the modified 3D face semantic representation of the inter-frame. Based on the reconstructed key reference frame, the fixed parameters in the fixed parameter buffer, and the modified inter-frame 3D face semantic representation, 3D face mesh reconstruction, 2D face mesh projection, eye motion calibration, mesh basis motion estimation, and inter-frame generation are performed to generate inter-frame reconstructed face image frames with corresponding semantic parameter editing instructions.
[0153] In some optional embodiments, the interaction module 704, in response to a virtual character replacement instruction, replaces the reconstructed key reference frame with a virtual character image to generate virtual character animation frames, including: Get user-defined virtual character images; The virtual character image is input into a pre-trained face reconstruction model to extract the virtual identity coefficient and virtual reflectivity coefficient. The virtual fixed parameters are obtained by replacing the identity coefficient and reflectivity coefficient in the fixed parameter buffer with the virtual identity coefficient and virtual reflectivity coefficient. Based on virtual fixed parameters and the 3D face semantic representation of the reconstructed inter-frames, 3D face mesh reconstruction, 2D face mesh projection, eye movement calibration, mesh basis motion estimation and inter-frame generation are performed to generate virtual character animation frames. The mouth movement and eye blinking dynamic semantics of the virtual character animation frames are consistent with the 3D face semantic representation of the reconstructed inter-frames.
[0154] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0155] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0156] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0157] like Figure 8As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0158] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0159] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the face video communication method. For example, in some embodiments, the face video communication method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the face video communication method by any other suitable means (e.g., by means of firmware).
[0160] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0161] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0163] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0164] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0165] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0166] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the disclosed technical solution can be achieved, and this is not limited herein.
[0167] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A face video communication method, wherein, Applied to the encoder end, the method includes: The input face video frame sequence is classified to obtain key reference frames and inter-frames, and the key reference frames are compressed and encoded to obtain key reference frame encoded data. For the inter-frame, the inter-frame is mapped from a two-dimensional image frame to a three-dimensional parameter space, and a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters and head position parameters is obtained by filtering. The correlation between different parameters satisfies the preset independence condition. The three-dimensional facial semantic representation is encoded to obtain inter-frame encoded data; The key reference frame encoded data and the inter-frame encoded data are integrated into a transmission stream, and the transmission stream is sent to the decoder. The decoder decodes the received transmission stream to obtain the 3D facial semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. Based on the 3D facial semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, an inter-frame reconstructed facial image frame is generated. In response to an interaction command, the parameters in the 3D facial semantic representation of the reconstructed inter-frame are modified to generate an inter-frame reconstructed facial image frame corresponding to the interaction command.
2. The method according to claim 1, wherein, The process of mapping the inter-frame from a two-dimensional image frame to a three-dimensional parameter space and filtering to obtain a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters includes: Using a pre-trained 3D face reconstruction model, the shape parameters, texture parameters, and dynamic expression parameters of the 3D face are predicted from the 2D face images in the inter-frames to obtain the initial 3D face parameters. A facial behavior analysis model is used to predict eye blinking parameters from the two-dimensional face images in the inter-frames; The parameters representing the lip movement components are extracted from the dynamic expression parameters in the initial three-dimensional face parameters to obtain the mouth movement parameters; The head rotation parameter, head translation parameter, and head position parameter are extracted from the initial three-dimensional face parameters. The head rotation parameter corresponds to the rotation radian in each coordinate axis direction in three-dimensional space, the head translation parameter corresponds to the pixel translation in each coordinate axis direction in three-dimensional space, and the head position parameter represents the relative height position of the head in the image. By integrating the mouth movement parameters, the eye blinking parameters, the head rotation parameters, the head translation parameters, and the head position parameters, a three-dimensional facial semantic representation is obtained.
3. The method according to claim 2, wherein, The process of integrating the mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters to obtain a three-dimensional facial semantic representation includes: The initial semantic representation is obtained by integrating the mouth movement parameters, the eye blinking parameters, the head rotation parameters, the head translation parameters, and the head position parameters. Principal component analysis spatial projection processing is performed on the initial semantic representation to make the correlation between parameters less than or equal to a preset threshold, so as to satisfy the preset independence condition and obtain the three-dimensional face semantic representation.
4. The method according to claim 1, wherein, The process of classifying the input face video frame sequence to obtain key reference frames and inter-frames includes: Calculate the differences in three-dimensional facial semantic representations between adjacent frames in the facial video frame sequence, including differences in head rotation parameters and mouth motion parameters; When the difference in the head rotation parameters exceeds a first preset threshold, or the difference in the mouth movement parameters exceeds a second preset threshold, the current frame is marked as the key reference frame. The remaining frames that were not marked as key reference frames are identified as inter-frames.
5. The method according to claim 1, wherein, The process of encoding the three-dimensional facial semantic representation to obtain inter-frame encoded data includes: Using the 3D face semantic representation of the key reference frame or the 3D face semantic representation of the preceding inter-frame as the prediction benchmark, the 3D face semantic representation of the current inter-frame is predicted to obtain the prediction residual. The predicted residuals are quantized and converted into binary code; The binary code is encoded using a context-based arithmetic coding model to generate the inter-frame encoded data.
6. The method according to claim 5, wherein, The step of predicting the 3D face semantic representation of the current inter-frame using the 3D face semantic representation of the key reference frame or the 3D face semantic representation of the preceding inter-frame as the prediction benchmark, and obtaining the prediction residual, includes: When processing the first inter-frame, the three-dimensional face semantic representation of the key reference frame is obtained as the prediction benchmark, and the difference between the three-dimensional face semantic representation of the first inter-frame and the three-dimensional face semantic representation of the key reference frame is calculated to obtain the first prediction residual. When processing subsequent frames of the first inter-frame, the three-dimensional face semantic representation of the previous frame is obtained as the prediction benchmark, and the difference between the three-dimensional face semantic representation of the current inter-frame and the three-dimensional face semantic representation of the previous frame is calculated to obtain the second prediction residual.
7. A face video communication method, wherein, Applied to the decoder, the method includes: The system receives a transmission stream from an encoder, which classifies the input face video frame sequence to obtain key reference frames and inter-frames. The key reference frames are then compressed and encoded to obtain key reference frame encoded data. For the inter-frames, they are mapped from two-dimensional image frames to a three-dimensional parameter space, and a three-dimensional face semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters is obtained. The correlation between different parameters satisfies a preset independence condition. The three-dimensional face semantic representation is encoded to obtain inter-frame encoded data. The key reference frame encoded data and the inter-frame encoded data are integrated into a transmission stream, which is then sent to the decoder. The received transmission stream is decoded to obtain the three-dimensional face semantic representation of the reconstructed key reference frame and the inter-reconstruction frame; Based on the reconstructed key reference frame and the reconstructed inter-frame 3D face semantic representation, an inter-frame reconstructed face image frame is generated; In response to the interaction command, the parameters in the three-dimensional face semantic representation of the reconstructed inter-frame are modified to generate an inter-frame reconstructed face image frame corresponding to the interaction command.
8. The method according to claim 7, wherein, Decoding the received transmission stream to obtain the 3D face semantic representation of the reconstructed key reference frame and the inter-reconstruction frame includes: For the key reference frame encoded data in the transmitted bitstream, key reference frame decoding and reconstruction are performed to obtain the reconstructed key reference frame; For the inter-frame encoded data in the transmitted bitstream, perform inter-frame semantic representation decoding and reconstruction to obtain the three-dimensional face semantic representation of the reconstructed inter-frame.
9. The method according to claim 8, wherein, The step of performing key reference frame decoding and reconstruction on the key reference frame encoded data in the transmitted bitstream to obtain the reconstructed key reference frame includes: The key reference frame encoded data is decoded by a decoder to obtain the reconstructed key reference frame; Fixed parameters are extracted from the reconstructed key reference frame and stored in a fixed parameter buffer. The fixed parameters include identity coefficient, reflectivity coefficient, and illumination coefficient.
10. The method according to claim 8, wherein, The step of performing inter-frame semantic representation decoding and reconstruction on the inter-frame encoded data in the transmitted bitstream to obtain the reconstructed three-dimensional face semantic representation of the inter-frames includes: Entropy decoding is performed on the inter-frame encoded data to obtain the quantized prediction residual; Perform an inverse quantization operation on the quantized prediction residual to recover the prediction residual; Obtain the three-dimensional face semantic representation corresponding to the prediction benchmark, wherein the prediction benchmark is the reconstruction key reference frame or the preceding reconstructed inter-frame; The predicted residuals are semantically compensated with the three-dimensional face semantic representations corresponding to the predicted baselines to calculate the three-dimensional face semantic representations of the reconstructed frames.
11. The method according to claim 7, wherein, The process of generating inter-frame reconstructed face image frames based on the reconstructed key reference frame and the reconstructed inter-frame 3D face semantic representation includes: Based on the reconstructed key reference frame, fixed parameters are extracted, including: identity coefficient, reflectivity coefficient, and illumination coefficient. Based on the fixed parameters and the three-dimensional face semantic representation of the reconstructed frames, three-dimensional face mesh reconstruction is performed to obtain a three-dimensional face mesh; Based on the three-dimensional face mesh, the three-dimensional face semantic representation of the reconstructed key reference frame, and the three-dimensional face semantic representation of the reconstructed inter-frame, two-dimensional face mesh projection and eye motion calibration are performed to obtain the two-dimensional face mesh of the reconstructed key reference frame, the two-dimensional face mesh of the inter-frame, and the eye blinking motion map. Based on the two-dimensional face mesh of the reconstructed key reference frame, the two-dimensional face mesh of the inter-frame, the eye blinking motion map, and the reconstructed key reference frame, mesh-based motion estimation is performed to obtain a fine-grained dense motion field and a face attention map. Based on the fine-grained dense motion field, the face attention map, and the reconstruction key reference frame, inter-frame generation is performed to obtain inter-frame reconstructed face image frames.
12. The method according to claim 11, wherein, The process of reconstructing a 3D face mesh based on the fixed parameters and the reconstructed inter-frame semantic representation of the face, and obtaining a 3D face mesh, includes: Based on the identity coefficient, reflectivity coefficient, and illumination coefficient, combined with the average neutral shape, identity basis vector, average neutral texture, reflectivity basis vector, and illumination basis vector, the three-dimensional face shape and three-dimensional face texture are calculated. Based on the mouth motion parameters in the reconstructed three-dimensional facial semantic representation, zero values are added to form expression coefficients; Based on the three-dimensional face shape, the three-dimensional face texture, and the expression coefficients, a three-dimensional face mesh is calculated using a parametric three-dimensional deformation model template.
13. The method according to claim 11, wherein, Based on the 3D face mesh, the 3D face semantic representation of the reconstructed key reference frame, and the 3D face semantic representation of the reconstructed inter-frames, a 2D face mesh projection and eye motion calibration are performed to obtain the 2D face mesh of the reconstructed key reference frame, the 2D face mesh of the inter-frames, and the eye blinking motion map, including: The first head rotation parameter and the first head translation parameter are extracted from the three-dimensional face semantic representation of the reconstructed key reference frame. The first head rotation parameter is converted into a first head rotation matrix and the first head translation parameter is converted into a first translation vector. Combined with the camera intrinsic parameter matrix, the three-dimensional face mesh is projected onto a two-dimensional plane based on the first head rotation matrix and the first translation vector to obtain the two-dimensional face mesh of the reconstructed key reference frame. The second head rotation parameter and the second head translation parameter are extracted from the 3D face semantic representation of the reconstructed inter-frame. The second head rotation parameter is converted into a second head rotation matrix and the second head translation parameter is converted into a second translation vector. Combined with the camera intrinsic parameter matrix, the 3D face mesh is projected onto a 2D plane based on the second head rotation matrix and the second translation vector to obtain the 2D face mesh of the inter-frame. Based on the blinking parameters in the three-dimensional face semantic representation of the reconstructed inter-frame, the eye region in the two-dimensional face mesh of the inter-frame is located, the original highest point and the original lowest point of the eye region are determined, the new highest point is calculated based on the blinking parameters, the position of the eye mesh vertex is adjusted, and an eye blinking motion map is generated.
14. The method according to claim 11, wherein, The two-dimensional face mesh based on the reconstructed key reference frame, the two-dimensional face mesh in the inter-frame, the eye blinking motion map, and the reconstructed key reference frame are used to perform mesh-based motion estimation to obtain a fine-grained dense motion field and a face attention map, including: Calculate the vertex position difference between the two-dimensional face mesh of the reconstructed key reference frame and the two-dimensional face mesh of the inter-frame; The vertex position difference is interpolated into a coarse-grained grid motion flow consistent with the video resolution using a grid data interpolation function; The coarse-grained mesh motion flow and the reconstructed key reference frame are input into a U-shaped encoder-decoder network, and coarse-grained deformed frames are generated through feature distortion operations; A spatial adaptive normalization mechanism is introduced to stitch together the reconstructed key reference frame, the coarse-grained deformed frame, the coarse-grained grid motion flow, and the eye blinking motion map to form a multi-feature input. Fine-grained dense motion field and face attention map are obtained through bi-branch prediction.
15. The method according to claim 11, wherein, Based on the fine-grained dense motion field, the face attention map, and the reconstructed key reference frame, inter-frame generation is performed to obtain inter-frame reconstructed face image frames, including: The reconstructed key reference frame is input into a neural network encoder to extract multi-scale spatial features; By combining the fine-grained dense motion field with the face attention map, an attention-based feature distortion operation is performed on the multi-scale spatial features to obtain distorted face spatial features; Affine transformation parameters are generated from the multi-scale spatial features using a neural network to modulate the distorted face spatial features, thereby obtaining transformed face features. The distorted facial spatial features and the transformed facial features are spliced together, input into the generator network, and the generated inter-frame reconstructed facial image frames are output.
16. The method according to claim 7, wherein, The step of modifying the parameters in the 3D facial semantic representation of the reconstructed inter-frame in response to the interaction command, and generating an inter-frame reconstructed facial image frame corresponding to the interaction command, includes: In response to the semantic parameter editing instruction, the target semantic parameters in the three-dimensional face semantic representation of the reconstructed inter-frame are modified to obtain the three-dimensional face semantic representation of the modified inter-frame, and an inter-frame reconstructed face image frame corresponding to the semantic parameter editing instruction is generated based on the three-dimensional face semantic representation of the modified inter-frame. In response to the virtual character replacement instruction, the reconstructed key reference frame is replaced with a virtual character image to generate a virtual character animation frame.
17. The method according to claim 16, wherein, The step of responding to a semantic parameter editing instruction by modifying the target semantic parameters in the 3D face semantic representation of the reconstructed inter-frame to obtain a modified 3D face semantic representation, and generating an inter-frame reconstructed face image frame corresponding to the semantic parameter editing instruction based on the modified 3D face semantic representation, includes: The target semantic parameter to be modified and the target modification value are determined from the semantic parameter editing instructions, wherein the target semantic parameter is at least one of mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters or head position parameters; The current value of the target semantic parameter in the 3D face semantic representation of the reconstructed inter-frame is replaced with the target modified value to obtain the 3D face semantic representation of the modified inter-frame. Based on the reconstructed key reference frame, the fixed parameters in the fixed parameter buffer, and the 3D face semantic representation of the modified inter-frame, 3D face mesh reconstruction, 2D face mesh projection, eye motion calibration, mesh basis motion estimation, and inter-frame generation are performed to generate inter-frame reconstructed face image frames corresponding to the semantic parameter editing instructions.
18. The method according to claim 16, wherein, The step of responding to a virtual character replacement instruction by replacing the reconstructed key reference frame with a virtual character image to generate a virtual character animation frame includes: Get user-defined virtual character images; The virtual character image is input into a pre-trained face reconstruction model to extract the virtual identity coefficient and virtual reflectivity coefficient. The virtual fixed parameters are obtained by replacing the identity coefficient and reflectivity coefficient in the fixed parameter cache with the virtual identity coefficient and the virtual reflectivity coefficient. Based on the virtual fixed parameters and the three-dimensional face semantic representation of the reconstructed inter-frame, three-dimensional face mesh reconstruction, two-dimensional face mesh projection, eye movement calibration, mesh basis motion estimation and inter-frame generation are performed to generate virtual character animation frames. The mouth movement and eye blinking dynamic semantics of the virtual character animation frames are consistent with the three-dimensional face semantic representation of the reconstructed inter-frame.
19. A face video communication method, wherein, The method includes: The encoder classifies the input facial video frame sequence to obtain key reference frames and inter-frames, and compresses and encodes the key reference frames to obtain key reference frame encoded data. For the inter-frames, the inter-frames are mapped from two-dimensional image frames to a three-dimensional parameter space, and a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters is obtained. The correlation between different parameters satisfies a preset independence condition. The three-dimensional facial semantic representation is encoded to obtain inter-frame encoded data. The key reference frame encoded data and the inter-frame encoded data are integrated into a transmission stream, and the transmission stream is sent to the decoder. The decoder decodes the received transmission stream to obtain the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame; based on the 3D face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, it generates an inter-frame reconstructed face image frame; in response to an interaction command, it modifies the parameters in the 3D face semantic representation of the reconstructed inter-frame to generate an inter-frame reconstructed face image frame corresponding to the interaction command.
20. A facial video communication device, wherein, Applied to the encoder end, the device includes: The classification module is used to classify the input face video frame sequence to obtain key reference frames and inter-frames, and to compress and encode the key reference frames to obtain key reference frame encoded data. The mapping module is used to map the inter-frame from a two-dimensional image frame to a three-dimensional parameter space for the inter-frame, and to filter to obtain a three-dimensional facial semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters and head position parameters, wherein the correlation between different parameters satisfies a preset independence condition. The encoding module is used to encode the three-dimensional facial semantic representation to obtain inter-frame encoded data; The transmission module is used to integrate the key reference frame encoded data and the inter-frame encoded data into a transmission bitstream, and send the transmission bitstream to the decoder. The decoder is used to decode the received transmission bitstream to obtain the 3D facial semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. Based on the 3D facial semantic representation of the reconstructed key reference frame and the reconstructed inter-frame, an inter-frame reconstructed facial image frame is generated. In response to an interaction command, the parameters in the 3D facial semantic representation of the reconstructed inter-frame are modified to generate an inter-frame reconstructed facial image frame corresponding to the interaction command.
21. A facial video communication device, wherein, The device, applied at the decoder end, includes: A receiving module is used to receive the transmission bitstream from the encoder. The encoder classifies the input face video frame sequence to obtain key reference frames and inter-frames, and compresses and encodes the key reference frames to obtain key reference frame encoded data. For the inter-frames, the inter-frames are mapped from two-dimensional image frames to a three-dimensional parameter space, and a three-dimensional face semantic representation including mouth movement parameters, eye blinking parameters, head rotation parameters, head translation parameters, and head position parameters is obtained. The correlation between different parameters satisfies a preset independence condition. The three-dimensional face semantic representation is encoded to obtain inter-frame encoded data. The key reference frame encoded data and the inter-frame encoded data are integrated into a transmission bitstream, and the transmission bitstream is sent to the decoder. The decoding module is used to decode the received transmission stream to obtain the three-dimensional face semantic representation of the reconstructed key reference frame and the reconstructed inter-frame. The reconstruction module is used to generate inter-frame reconstructed face image frames based on the three-dimensional face semantic representation of the reconstruction key reference frame and the reconstruction inter-frame; An interaction module is used to modify the parameters in the three-dimensional face semantic representation of the reconstructed inter-frame in response to an interaction command, and generate an inter-frame reconstructed face image frame corresponding to the interaction command.
22. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-19.
23. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-19.
24. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-19.