Video portrait semantic compression method based on key point features

Through the video portrait semantic compression method based on key point features, using FLAME model and three-plane feature rendering technology, the problems of low compression efficiency and insufficient reconstruction quality in the existing technology are solved, and efficient, high-definition and low-latency video transmission is achieved.

CN120264016APending Publication Date: 2025-07-04NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510431969.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When processing portrait videos, existing video compression technology ignores the unique semantic information of portraits, resulting in low compression efficiency and is difficult to meet the high-definition and low-latency video transmission requirements in the 5G era. In addition, existing methods have consistency problems and insufficient reconstruction quality during multi-view synthesis or reconstruction.

Method used

The video portrait semantic compression method based on key point features is adopted, and the portrait head FLAME model is generated through the encoding end, dynamic features are extracted, and combined with the three-plane feature rendering technology at the decoding end, high-fidelity portrait video reconstruction is realized.

Benefits of technology

It realizes efficient video compression and reconstruction, meets the high-definition and low-latency video transmission needs in the 5G era, improves the visual quality of portrait videos, and reduces computing complexity and hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264016A_ABST
    Figure CN120264016A_ABST
Patent Text Reader

Abstract

The invention provides a video portrait semantic compression method based on key point features, and relates to the technical field of video image processing. The invention provides a video portrait semantic compression method based on key point features. The video portrait semantic compression method comprises a coding end and a decoding end, the coding end obtains a portrait video to be compressed, and dynamic features of the head of the portrait are extracted; transmitting the transmission image and the dynamic features of the head of the portrait to a decoding end; and the decoding end receives the dynamic features of the transmission images and the head of the portrait, extracts the static features of each frame of transmission image, and obtains a reconstructed high-fidelity portrait video by combining the dynamic features of the head of the portrait. According to the method, the limitation of pixel correlation is broken through, and the core is focused on semantic information of a video portrait, so that the strict requirements of modern video application on high quality, low delay and high efficiency are met, particularly, the requirements of 5G era on high-definition and low-delay video transmission can be fully met, and the visual quality of a portrait video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video image processing, and particularly to a video portrait semantic compression method based on key point features. Background Art

[0002] With the rapid development of video technology, video services have become increasingly mature, and the application scenarios have become more and more extensive. The video data generated by applications such as remote conferencing, video calls, and live streaming e-commerce accounts for an increasing proportion of Internet traffic, posing a huge challenge to network bandwidth and storage space. The popularization of 5G technology has further raised users' requirements for video transmission speed and quality to a new level. Ultra-high definition and low latency have become the goals pursued by users. Portrait videos, as an important part of video data, play a dominant role in scenarios such as remote conferencing and video calls.

[0003] Currently, the widely used traditional video compression technologies mainly include: H.264, H.265, and AVS3 in China. The latest video compression standards are AV1 and H.266 / VVC. However, traditional video compression methods are mainly based on pixel-level redundancy removal and transform coding technologies. Treating portraits as ordinary images for processing ignores the unique semantic information of portraits, resulting in low compression efficiency and difficulty in meeting the high requirements for video transmission in the 5G era.

[0004] The video portrait reconstruction method based on key point features extracts and utilizes the semantic information of portraits (such as face key points, pose key points, etc.) to achieve more efficient video compression. For example, NVIDIA's work One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing in 2021 can be used for the encoding of portrait videos, reducing the bandwidth by 10 times. The 2D-based warping (such as TPS, etc.) uses the warping field estimated by sparse key points to warp the original image, and finally synthesizes a new image. The grid-based rendering (such as ROME, etc.) uses a 3D deformation model (3DMM) to model the source image. By combining three-dimensional information, the multi-view Figure 1 consistency problem can be effectively solved. The Nerf neural rendering method (such as HideNeRF, etc.) can generate three-dimensionally consistent results that include non-face information compared with 2D and grid-based methods.

[0005] Due to the lack of necessary 3D constraints, the 2D-based methods are difficult to maintain multi-view consistency when the pose changes significantly. In addition, these methods cannot effectively separate expressions and identities from the source portrait, resulting in unfaithful driving results. This may lead to inconsistent phenomena during multi-view synthesis or reconstruction, such as partial overlap or shape incoherence, etc.

[0006] Due to the limitations of the modeling and expression capabilities of 3DMM, the reconstructed head often lacks non-facial information such as hair, and the expressions are often unnatural. Although 3DMM provides strong prior knowledge for understanding the face, it only focuses on the facial region, unable to capture other detailed features such as hairstyles and accessories, and the fidelity is limited by the mesh resolution, resulting in an unnatural appearance in the reproduced images.

[0007] NeRF-based methods cannot be well generalized to new identities. Some of these methods require a large amount of portrait data for reconstruction, and some require time-consuming optimization processes during the inference process. In addition, during the inference process, some NeRF methods may need to perform time-consuming optimization processes to find the optimal light propagation path or adapt to the light propagation law of the new scene. This may lead to a slow inference speed, especially when dealing with complex scenes or high-resolution images. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a video portrait semantic compression method based on key point features, including an encoding end and a decoding end, aiming at the deficiencies of the above-mentioned existing technologies.

[0009] The encoding end obtains the portrait video to be compressed, generates a FLAME model of the portrait head based on the portrait video to be compressed, and extracts the dynamic features of the FLAME model of the portrait head; randomly selects several frame images from the portrait video to be compressed as transmission images, and transmits the transmission images and the dynamic features of the FLAME model of the portrait head to the decoding end.

[0010] The decoding end receives the transmission images and the dynamic features of the FLAME model of the portrait head, extracts the three-plane features that can be used for rendering from the transmission images; performs dynamic query and weighted fusion on the three-plane features that can be used for rendering of different transmission images to obtain the normalized three-plane features of each frame of transmission image; performs light sampling on each frame of transmission image, extracts the static features of each frame of transmission image based on the normalized three-plane features of each frame of transmission image, and combines the dynamic features of the FLAME model of the portrait head to obtain a reconstructed high-fidelity portrait video.

[0011] The encoding end includes:

[0012] Step 1: The encoding end obtains the portrait video to be compressed and generates a FLAME model of the portrait head based on the portrait video to be compressed.

[0013] Step 1.1: Obtain the portrait video to be compressed, and use multiple pre-trained facial capture models to extract the FLAME parameters of each frame of image in parallel.

[0014] Among them, the FLAME parameters of each frame of image include expression parameter ψ, head pose parameter θ, and shape parameter β;

[0015] Step 1.2: Optimize the FLAME parameters of each frame of image based on the multi-frame joint optimization method, and optimize the smoothness of the FLAME parameters of adjacent two frames of images by constructing a temporally related loss function; the temporally related loss function includes a shape consistency loss for constraining the shape parameter β and a pose continuity loss for constraining the head pose parameter θ;

[0016] Step 1.3: Use the FLAME Dense model to decode the optimized FLAME parameters, decode the optimized FLAME parameters into a number of 3D vertex coordinates, and perform spatial alignment on the number of 3D vertex coordinates using a vertex scaling factor and camera parameters to obtain a FLAME model of the human head; among them, the camera parameters include a focal length parameter and a principal point parameter;

[0017] Step 1.4: Use a differentiable mesh renderer to render the number of 3D vertex coordinates into an Alpha image with an alpha channel, and project the number of 3D vertex coordinates onto a 2D plane;

[0018] Step 2: The encoding end extracts the dynamic features of the FLAME model of the human head;

[0019] Step 2.1: Assign learnable feature weights to each 3D vertex of the FLAME model of the human head;

[0020] Step 2.2: Use the NeRF sampling strategy to sample the FLAME model of the human head to obtain a number of dynamic sampling points, and obtain the feature weights and position coordinates of the K nearest neighbor points of each dynamic sampling point; for the dynamic sampling point s of the FLAME model of the human head, retrieve the K nearest neighbor points of the dynamic sampling point s in the entire FLAME model of the human head, and obtain the feature weight f of the k-th nearest neighbor point among the K nearest neighbor points of the dynamic sampling point s k and the position coordinate p of the k-th nearest neighbor point k ;

[0021] Step 2.3: Regress the feature weights of the K nearest neighbor points of each dynamic sampling point to extract the dynamic features of each dynamic sampling point;

[0022] Perform linear regression on the feature weights of the K nearest neighbor points of the dynamic sampling point s, perform frequency position encoding on the position coordinates of the K nearest neighbor points of the dynamic sampling point s, and perform weighted fusion on the feature weights of the K nearest neighbor points obtained by linear regression of the dynamic sampling point s and the position coordinates of the K nearest neighbor points of the frequency position encoding according to the position weights of the K nearest neighbor points of the dynamic sampling point s to generate the dynamic feature f of the dynamic sampling point s exp, as shown in the following formula:

[0023]

[0024] where Lp(·) is a linear regression function, Fpos(·) is a frequency position encoding function, and w k is the position weight of the k-th nearest neighbor point of the dynamic sampling point s;

[0025] Step 3: Randomly select several frame images from the portrait video to be compressed as transmission images, and transmit the transmission images and the dynamic features of the portrait head FLAME model to the decoding end;

[0026] The decoding end includes:

[0027] S1: The decoding end receives the transmission images and the dynamic features of the portrait head FLAME model, and extracts the three-plane features of the transmission images that can be used for rendering;

[0028] S1.1: Use the ViT network to perform global coordinate transformation and feature extraction on each frame of the transmission image to obtain the global geometric feature map of each frame of the transmission image;

[0029] For any frame of the transmission image F k , use the ViT network to divide the transmission image F k into non-overlapping blocks NOP of a fixed size, use linear projection to convert each non-overlapping block NOP of the fixed size into an embedded vector sequence PE; perform global coordinate transformation on the embedded vector sequence PE, add position encoding to each non-overlapping block NOP of the fixed size in the embedded vector sequence PE, and retain the spatial position information of each non-overlapping block NOP of the fixed size to obtain the embedded vector sequence PE_SEQ with position encoding; use multiple SegFormer networks to perform feature extraction on the embedded vector sequence PE_SEQ with position encoding to obtain the feature sequence SFB_FEAT of the transmission image F k , and perform multi-scale feature extraction on the feature sequence SFB_FEAT of the transmission image F k to obtain the multi-resolution feature map MR_FEATS of the transmission image F k ; perform feature recombination on the multi-resolution feature map MR_FEATS of the transmission image F k and rearrange it into a spatial feature map as the global geometric feature map GG_FEAT of the transmission image F k ;

[0030] S1.2: Use the VGG network to perform layer-by-layer feature extraction on each frame of the transmission image to obtain the high-dimensional feature map of each frame of the transmission image;

[0031] For any frame of the transmission image F k, use the VGG network to extract features layer by layer from the transmitted image F k to obtain the feature map CLF after multiple layers of convolution; capture the high-frequency texture information of the feature map CLF through the local receptive field to obtain the feature map LRFF after being processed by the local receptive field; then adjust the resolution of the feature map LRFF to obtain the transmitted image F k with the same spatial resolution as the global geometric feature map GG_FEAT of the transmitted image F k of the high-dimensional feature map MRF; each convolutional layer in the VGG network will increase the degree of abstraction of the features while gradually reducing the spatial size of the feature map. Each convolutional layer captures the high-frequency texture information of the transmitted image F k and retains the fine-grained details of the transmitted image F k ;

[0032] S1.3: Concatenate the global geometric feature map and the high-dimensional feature map of each frame of the transmitted image to extract the three-plane features of each frame of the transmitted image;

[0033] For any frame of the transmitted image F k , concatenate the global geometric feature map GG_FEAT and the high-dimensional feature map MRF of the transmitted image F k along the channel dimension to obtain the fused feature FF; use 3 lightweight 3*3 convolutional layers to compress the number of channels of the fused feature FF to obtain the intermediate feature map IF, and then use a 3*3 convolutional layer to map the intermediate feature map IF to a three-plane representation on three orthogonal planes XY, XZ, and YZ to obtain the three-plane feature TPF of the transmitted image F k ;

[0034] S1.4: Normalize the three-plane features of each frame of the transmitted image to obtain the three-plane features of each frame of the transmitted image that can be used for rendering;

[0035] For the three-plane feature TPF of any frame of the transmitted image F k , associate the three orthogonal planes XY, XZ, and YZ with the predefined world coordinate system for spatial alignment to obtain the spatially aligned three-plane feature ATPF; perform normalization processing on the spatially aligned three-plane feature ATPF to obtain the normalized three-plane feature NTPF; based on the compact 3D representation of the normalized three-plane feature NTPF, use the rendering model to obtain the three-plane feature RTPF of the transmitted image F k that can be used for rendering;

[0036] S2: The decoding end performs dynamic query and weighted fusion on the three-plane features of different images that can be used for rendering to obtain the normalized three-plane features of each frame of the image;

[0037] S2.1: Obtain the shape parameters of the three - plane features available for rendering in each frame of the transmitted image, perform shape transformation on the three - plane features available for rendering in each frame of the transmitted image, and obtain the new three - plane features of each frame of the transmitted image;

[0038] Obtaining the shape parameters of the three - plane features available for rendering in each frame of the transmitted image includes: the batch size B, the number of feature groups N, the three - plane index T, the number of channels C, the feature map height H, and the feature map width W of the three - plane features RTPF available for rendering in each frame of the transmitted image;

[0039] Use the rearrange function to convert the shape of the three - plane features RTPF available for rendering in each frame of the transmitted image from B * N * T * C * H * W to (B * T * H * W) * N * C, and obtain the new three - plane features XTPF of each frame of the transmitted image;

[0040] S2.2: Introduce a set of learnable query planes, and through multi - layer non - linear transformation, map the new three - plane features XTPF of each frame of the transmitted image to the query space and the key space to obtain the query matrix and the key matrix of the three - plane features of different images;

[0041] S2.3: Calculate the correlation of the new three - plane features of different transmitted images, and generate the attention weights of each frame of the transmitted image;

[0042] For any frame of the transmitted image F k and the transmitted image F j , calculate the correlation dots k between the query matrix of the new three - plane features of the transmitted image F j and the key matrix of the new three - plane features of the transmitted image F k,j , as shown in the following formula:

[0043]

[0044] where Q k is the query matrix of the k - th frame of the transmitted image F k , K j is the key matrix of the j - th frame of the transmitted image F j , and dim is the dimension of the new three - plane features of the transmitted image;

[0045] Use the Softmax function to calculate the attention weight attn k of the transmitted image F k , as shown in the following formula:

[0046] attn k = softmax(dots k,j )

[0047] S2.4: Dynamically query and weighted fuse the three-plane features from different images to obtain the normalized three-plane features of each frame of the transmitted image;

[0048] For the transmitted image F k and the transmitted image F j perform weighted fusion on the attention weights to obtain the fused three-plane feature P of the transmitted image F k as shown in the following formula:

[0049]

[0050] where N is the number of input images, attn i is the weight of the k-th frame image F k and attn j is the weight of the j-th frame image F j and E(·) is the transmitted image encoder;

[0051] Convert the shape of the fused three-plane feature P of the transmitted image F k from (B*T*H*W)*N*C to B*N*T*C*H*W as the normalized three-plane feature GTPF of the image F k ;

[0052] S3: The decoder performs ray sampling on each frame of the transmitted image, extracts the static features of each frame of the transmitted image based on the normalized three-plane features of each frame of the transmitted image, and combines the dynamic features of the FLAME model of the human head to obtain a reconstructed high-fidelity human portrait video;

[0053] S3.1: Calculate the camera origin and ray direction according to the camera parameters of the portrait video to be compressed, align the rays to the canonical space through a transformation matrix, use the axis-aligned bounding box of the FLAME model of the human head to filter out the valid ray segments from the aligned rays, and uniformly sample a number of three-plane ray sampling points on the valid ray segments;

[0054] S3.2: Project each three-plane ray sampling point onto the normalized three-plane feature GTPF of each frame of the transmitted image, and extract the static features of each frame of the transmitted image through bilinear interpolation;

[0055] For any three-plane ray sampling point projected onto any frame of the transmitted image, project the three-plane ray sampling point onto the normalized three-plane GTPF of the transmitted image, and extract the features of the three-plane ray sampling point on each plane through bilinear interpolation. Among them, the projection point coordinates of the three-plane ray sampling point on the XY plane are (x, y), and the sampling feature of the three-plane ray sampling point is f xy ; the projection point coordinates of the three-plane ray sampling point on the XZ plane are (x, z), and the sampling feature of the three-plane ray sampling point is f xz, The coordinates of the projection point of the three-plane light sampling point on the YZ plane are (y, z), and the sampling feature of the three-plane light sampling point is f yz ; Concatenate the sampling features of the three-plane light sampling points on the three planes to obtain the static feature f of the transmitted image canonical ;

[0056] S3.3: Concatenate the dynamic feature of the FLAME model of the human head with the static feature of the transmitted image to obtain a fused feature;

[0057] The dynamic feature f of the FLAME model of the human head exp and the static feature f of the transmitted image canonical are concatenated to obtain the fused feature f fused , as shown in the following formula:

[0058] f fused = Concat(f canonical , f exp )

[0059] where Concat(·) is the concatenation operation;

[0060] S3.4: Based on the fused feature f fused and the light direction, use a multi-layer perceptron MLP to extract the density σ and color RGB of each three-plane light sampling point, and use the ray_marching function to accumulate the extraction results to obtain a reconstructed high-fidelity human portrait video

[0061] The beneficial effects of adopting the above technical solutions are as follows: A video portrait semantic compression method based on key point features provided by the present invention aims to overcome the key problems faced by traditional video compression and portrait reconstruction technologies. By breaking through the limitation of pixel correlation, it focuses on the semantic information of video portraits, so as to meet the strict requirements of modern video applications for high quality, low latency and high efficiency. In particular, it can fully meet the needs of 5G era for high-definition and low-latency video transmission, and improve the visual quality of portrait videos. A video portrait semantic compression method based on key point features achieves a more efficient compression and reconstruction process with the help of portrait semantic information. On the one hand, it strongly supports emerging application scenarios; on the other hand, it greatly reduces the computational complexity and hardware cost, injects innovative vitality into video compression technology, creates a higher quality, more efficient and more flexible video portrait processing solution, and presents a better visual experience for users Description of the Drawings

[0062] Figure 1 is a framework diagram of a video portrait semantic compression method based on key point features provided by an embodiment of the present invention;

[0063] Figure 2Schematic diagram of the method for the decoding end to extract the three - plane features of the transmitted image provided by the embodiment of the present invention;

[0064] Figure 3 Comparison chart of evaluation indexes between the video portrait semantic compression method based on key - point features provided by the embodiment of the present invention and HEVC at the same bitrate, where (a) is the PSNR comparison chart and (b) is the SSIM comparison chart. Detailed implementation manners

[0065] The following combines the drawings and embodiments to further describe in detail the detailed implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0066] Existing video portrait semantic compression models can simulate common expressions to a certain extent, but there are shortcomings in accurately capturing and restoring micro - expressions. For example, for subtle emotional changes such as a slight upward curl of the corners of the mouth or an imperceptible contraction of the corners of the eyes, deep - learning models are difficult to accurately reproduce, resulting in the lack of realism and fineness in the expressions of the reconstructed portraits. Therefore, a dynamic point - based expression field is designed, and through the dynamic binding and weighted fusion mechanism of FLAME point - cloud vertex features, high - precision expression control is achieved.

[0067] When there are occlusions (such as eyes closed, hair covering the face) or extreme poses (such as a side - face view resulting in missing information on the other side) in a single input source image, the model is difficult to accurately infer the facial features of the occluded area. For example, when hair covers part of the cheek and forehead, the model may wrongly fill in the details of these occluded areas, resulting in a deviation between the facial contour of the reconstructed portrait and the actual situation. In this embodiment, a feature adaptive fusion method that supports any number of input images is designed, which can extract the missing features from other input images. For example, if the eyes are closed in one image and open in another image, the eye information of the two is fused to generate a reasonable open - eye expression.

[0068] A video portrait semantic compression method based on key - point features in this embodiment decouples the expression, head pose, and shape features of a portrait video into low - dimensional semantic parameters, and transmits the parameters from the encoding end to the decoding end, realizing high - fidelity portrait reconstruction and extremely low - bitrate compression, as Figure 1 shown, including an encoding end and a decoding end;

[0069] The encoding end obtains the portrait video to be compressed, generates a FLAME model of the portrait head based on the portrait video to be compressed, and extracts the dynamic features of the FLAME model of the portrait head; randomly selects several frames of images from the portrait video to be compressed as transmitted images, and transmits the transmitted images and the dynamic features of the FLAME model of the portrait head to the decoding end;

[0070] The decoding end receives the transmitted image and the dynamic features of the FLAME model of the human head, extracts the three-plane features of the transmitted image that can be used for rendering based on the transmitted image; performs dynamic query and weighted fusion on the three-plane features of different transmitted images that can be used for rendering to obtain the normalized three-plane features of each frame of the transmitted image; performs ray sampling on each frame of the transmitted image, extracts the static features of each frame of the transmitted image based on the normalized three-plane features of each frame of the transmitted image, and combines the dynamic features of the FLAME model of the human head to obtain a reconstructed high-fidelity human portrait video.

[0071] Step 1: The encoding end obtains the human portrait video to be compressed and generates a FLAME model of the human head based on the human portrait video to be compressed;

[0072] Step 1.1: Obtain the human portrait video to be compressed, and use multiple pre-trained face capture models to parallelly extract the FLAME parameters of each frame of the image;

[0073] Among them, the FLAME parameters of each frame of the image include expression parameter ψ, head pose parameter θ, and shape parameter β;

[0074] In this embodiment, the pre-trained EMOCA model is used to extract the expression parameter ψ and the head pose parameter θ of each frame of the image, and the pre-trained MICA model is used to extract the shape parameter β of each frame of the image;

[0075] Step 1.2: Optimize the FLAME parameters of each frame of the image based on the multi-frame joint optimization method, and optimize the smoothness of the FLAME parameters of adjacent two frames of the image by constructing a time-series related loss function; the time-series related loss function includes a shape consistency loss for constraining the shape parameter β and a pose continuity loss for constraining the head pose parameter θ;

[0076] Step 1.3: Use the FLAMEDense model to decode the optimized FLAME parameters, decode the optimized FLAME parameters into a number of 3D vertex coordinates, and perform spatial alignment on the number of 3D vertex coordinates using a vertex scaling factor and camera parameters to obtain a FLAME model of the human head; among them, the camera parameters include a focal length parameter and a principal point parameter;

[0077] Step 1.4: Use a differentiable mesh renderer to render the number of 3D vertex coordinates into an Alpha image with a transparent channel, and project the number of 3D vertex coordinates onto a 2D plane.

[0078] Step 2: The encoding end extracts the dynamic features of the FLAME model of the human head;

[0079] Step 2.1: Assign learnable feature weights to each 3D vertex of the FLAME model of the human head;

[0080] Feature weights are used to calculate the features of each vertex of the FLAME model of the human head. When the expression parameter ψ changes, the position coordinates p of vertex i in the FLAME model of the human head i will be dynamically adjusted according to the expression change. For example, when the corners of the mouth move up to simulate a smile, the point cloud shape changes with the expression, but the semantic topology remains unchanged. For example, the vertices representing the eyes always belong to the eye region;

[0081] Step 2.2: Use the NeRF sampling strategy to sample the FLAME model of the human head to obtain a number of dynamic sampling points, and obtain the feature weights and position coordinates of K nearest neighbor points for each dynamic sampling point; for the dynamic sampling point s of the FLAME model of the human head, retrieve K nearest neighbor points of the dynamic sampling point s in the entire FLAME model of the human head, and obtain the feature weight f of the k-th nearest neighbor point among the K nearest neighbor points of the dynamic sampling point s k and the position coordinates p of the k-th nearest neighbor point k ;

[0082] In this embodiment, the number K of nearest neighbor points of the dynamic sampling point is set to 8, and the feature weights and position coordinates of 8 nearest neighbor points of each dynamic sampling point are obtained to avoid the rigid distortion of non-FLAME regions (such as hair) caused by local sampling; as the expression parameter of the FLAME model changes, the positions of the 8 nearest neighbor points of each dynamic sampling point will also change dynamically, thereby generating the dynamic expression feature field of the FLAME model of the human head.

[0083] Step 2.3: Regress the feature weights of the K nearest neighbor points of each dynamic sampling point to extract the dynamic features of each dynamic sampling point;

[0084] Perform linear regression on the feature weights of the K nearest neighbor points of the dynamic sampling point s, perform frequency position encoding on the position coordinates of the K nearest neighbor points of the dynamic sampling point s, and perform weighted fusion on the feature weights of the K nearest neighbor points obtained by linear regression of the dynamic sampling point s and the position coordinates of the K nearest neighbor points with frequency position encoding according to the position weights of the K nearest neighbor points of the dynamic sampling point s to generate the dynamic feature f of the dynamic sampling point s exp as shown in the following formula:

[0085]

[0086] where, Lp(·) is the linear regression function, Fpos(·) is the frequency position encoding function, and w k is the position weight of the k-th nearest neighbor point of the dynamic sampling point s;

[0087] Step 3: Randomly select several frames of images from the human portrait video to be compressed as transmission images, and transmit the transmission images and the dynamic features of the FLAME model of the human head to the decoding end;

[0088] S1: The decoding end receives the transmitted image and the dynamic features of the FLAME model of the human head, and extracts the three-plane features of the transmitted image that can be used for rendering;

[0089] S1.1: Use the ViT network to perform global coordinate transformation and feature extraction on each frame of the transmitted image to obtain the global geometric feature map of each frame of the transmitted image;

[0090] For any frame of the transmitted image F k , use the ViT network to divide the transmitted image F k into non-overlapping blocks NOP of a fixed size, and use linear projection to convert each non-overlapping block NOP of a fixed size into a sequence of embedding vectors PE; perform global coordinate transformation on the sequence of embedding vectors PE, add position encoding to each non-overlapping block NOP of a fixed size in the sequence of embedding vectors PE, and retain the spatial position information of each non-overlapping block NOP of a fixed size to obtain a sequence of embedding vectors with position encoding PE_SEQ; use multiple SegFormer networks to perform feature extraction on the sequence of embedding vectors with position encoding PE_SEQ to obtain the feature sequence SFB_FEAT of the transmitted image F k , and perform multi-scale feature extraction on the feature sequence SFB_FEAT of the transmitted image F k to obtain the multi-resolution feature map MR_FEATS of the transmitted image F k ; perform feature recombination on the multi-resolution feature map MR_FEATS of the transmitted image F k and rearrange it into a spatial feature map as the global geometric feature map GG_FEAT of the transmitted image F k ;

[0091] S1.2: Use the VGG network to perform layer-by-layer feature extraction on each frame of the transmitted image to obtain the high-dimensional feature map of each frame of the transmitted image;

[0092] For any frame of the transmitted image F k , use the VGG network to perform layer-by-layer feature extraction on the transmitted image F k to obtain the feature map CLF after multiple convolutional layers; capture the high-frequency texture information of the feature map CLF through the local receptive field to obtain the feature map LRFF after being processed by the local receptive field; then adjust the resolution of the feature map LRFF to obtain the high-dimensional feature map MRF of the transmitted image F k with the same spatial resolution as the global geometric feature map GG_FEAT of the transmitted image F k ; Each convolutional layer in the VGG network increases the degree of abstraction of the features and gradually reduces the spatial size of the feature map. Each convolutional layer captures the transmitted image F through the local receptive field kThe high-frequency texture information of, retaining the transmitted image F k The fine-grained details;

[0093] S1.3: Concatenate the global geometric feature map and the high-dimensional feature map of each frame of the transmitted image, and extract the three-plane features of each frame of the transmitted image;

[0094] For any frame of the transmitted image F k , the global geometric feature map GG_FEAT and the high-dimensional feature map MRF of the transmitted image F k are concatenated along the channel dimension to obtain the fused feature FF; use 3 lightweight 3*3 convolutional layers to compress the channel number of the fused feature FF to obtain the intermediate feature map IF, and then use a 3*3 convolutional layer to map the intermediate feature map IF to a three-plane representation on three orthogonal planes XY, XZ, and YZ to obtain the three-plane feature TPF of the transmitted image F k ; the three lightweight 3*3 convolutional layers gradually compress the channel number of the fused feature FF, keep the spatial dimension of the fused feature FF, and realize feature compression and non-linear expression of the fused feature FF through the ReLU activation function;

[0095] S1.4: Normalize the three-plane features of each frame of the transmitted image to obtain the three-plane features of each frame of the transmitted image available for rendering;

[0096] For any frame of the transmitted image F k 's three-plane feature TPF, associate the three orthogonal planes XY, XZ, and YZ to a predefined world coordinate system for spatial alignment to obtain the spatially aligned three-plane feature ATPF; normalize the spatially aligned three-plane feature ATPF to obtain the normalized three-plane feature NTPF; based on the compact 3D representation of the normalized three-plane feature NTPF, use the rendering model to obtain the three-plane feature RTPF of the transmitted image F k available for rendering;

[0097] In this embodiment, associate the three orthogonal planes XY, XZ, and YZ to the face standard coordinate system to ensure the consistency of the view transformation during subsequent rendering; normalize the spatially aligned three-plane feature ATPF to prevent numerical instability of the three-plane feature ATPF; based on the compact 3D representation of the normalized three-plane feature NTPF, use the NeRF rendering model to obtain the three-plane features of each frame of the image available for rendering, as Figure 2 shown;

[0098] S2: The decoding end performs dynamic query and weighted fusion on the three-plane features of different images available for rendering to obtain the normalized three-plane features of each frame of the image;

[0099] S2.1: Obtain the shape parameters of the three - plane features available for rendering in each frame of the transmitted image, perform a shape transformation on the three - plane features available for rendering in each frame of the transmitted image, and obtain the new three - plane features of each frame of the transmitted image;

[0100] Obtaining the shape parameters of the three - plane features available for rendering in each frame of the transmitted image includes: the batch size B, the number of feature groups N, the three - plane index T, the number of channels C, the feature map height H, and the feature map width W of the three - plane features RTPF available for rendering in each frame of the transmitted image;

[0101] Use the rearrange function to convert the shape of the three - plane features RTPF available for rendering in each frame of the transmitted image from B*N*T*C*H*W to (B*T*H*W)*N*C, and obtain the new three - plane features XTPF of each frame of the transmitted image;

[0102] S2.2: Introduce a set of learnable query planes, and through multi - layer non - linear transformation, map the new three - plane features XTPF of each frame of the transmitted image to the query space and the key space to obtain the query matrix and the key matrix of the three - plane features of different images;

[0103] S2.3: Calculate the correlation of the new three - plane features of different transmitted images and generate the attention weights of each frame of the transmitted image;

[0104] For any frame of the transmitted image F k and the transmitted image F j , calculate the correlation dots k between the query matrix of the new three - plane features of the transmitted image F j and the key matrix of the new three - plane features of the transmitted image F k,j , as shown in the following formula:

[0105]

[0106] where, Q k is the query matrix of the k - th frame of the transmitted image F k , K j is the key matrix of the j - th frame of the transmitted image F j , and dim is the dimension of the new three - plane features of the transmitted image;

[0107] Use the Softmax function to calculate the attention weight attn k of the transmitted image F k , as shown in the following formula:

[0108] attn k = softmax(dots k,j )

[0109] S2.4: Dynamically query and weighted fuse the three-plane features from different images to obtain the normalized three-plane features of each frame of the transmitted image;

[0110] For the transmitted image F k and the transmitted image F j perform weighted fusion on the attention weights to obtain the fused three-plane feature P of the transmitted image F k as shown in the following formula:

[0111]

[0112] where N is the number of input images, attn i is the weight of the k-th frame image F k and attn j is the weight of the j-th frame image F j and E(·) is the transmitted image encoder;

[0113] Convert the shape of the fused three-plane feature P of the transmitted image F k from (B*T*H*W)*N*C to B*N*T*C*H*W as the normalized three-plane feature GTPF of the image F k ;

[0114] S3: At the decoding end, perform ray sampling on each frame of the transmitted image, extract the static features of each frame of the transmitted image based on the normalized three-plane features of each frame of the transmitted image, and combine with the dynamic features of the FLAME model of the human portrait head to obtain a reconstructed high-fidelity human portrait video;

[0115] S3.1: Calculate the camera origin and ray direction according to the camera parameters of the human portrait video to be compressed, align the rays to the canonical space through a transformation matrix, filter the valid ray segments from the aligned rays using the axis-aligned bounding box of the FLAME model of the human portrait head, and uniformly sample a number of three-plane ray sampling points on the valid ray segments;

[0116] In this embodiment, the number of three-plane ray sampling points sampled on the valid ray segments is set to 48;

[0117] S3.2: Project each three-plane ray sampling point onto the normalized three-plane feature GTPF of each frame of the transmitted image, and extract the static features of each frame of the transmitted image through bilinear interpolation;

[0118] For any three-plane ray sampling point projected onto any frame of the transmitted image, project the three-plane ray sampling point onto the normalized three-plane GTPF of the transmitted image, and extract the features of the three-plane ray sampling point on each plane through bilinear interpolation. Among them, the projection point coordinates of the three-plane ray sampling point on the XY plane are (x, y), and the sampling feature of the three-plane ray sampling point is f xy; The projection point coordinates of the three-plane light sampling point on the XZ plane are (x, z), and the sampling feature of the three-plane light sampling point is f xz , the projection point coordinates of the three-plane light sampling point on the YZ plane are (y, z), and the sampling feature of the three-plane light sampling point is f yz ; Concatenate the sampling features of the three-plane light sampling points on the three planes to obtain the static feature f of the transmitted image canonical ;

[0119] S3.3: Concatenate the dynamic feature of the FLAME model of the human head with the static feature of the transmitted image to obtain a fused feature;

[0120] The dynamic feature f of the FLAME model of the human head exp and the static feature f of the transmitted image canonical are concatenated to obtain the fused feature f fused , as shown in the following formula:

[0121] f fused = Concat(f canonical , f exp )

[0122] where Concat(·) is the concatenation operation;

[0123] S3.4: Based on the fused feature f fused and the light direction, use a multi-layer perceptron MLP to extract the density σ and color RGB of each three-plane light sampling point, and use the ray_marching function to accumulate the extraction results to obtain a reconstructed high-fidelity human portrait video

[0124] Traditional video encoders (such as H.264 / H.265) rely on pixel-domain transformation and motion compensation, and are prone to problems such as face blurring and texture distortion at low bitrates (<10 kbps); although the GAN-based generative compression scheme can improve subjective quality, it has problems of parameter redundancy and poor identity consistency. The video portrait reconstruction and compression technology based on key-point features proposed in this embodiment separates identity-related parameters (such as three-dimensional deformation, lighting coefficients) from dynamic expression parameters (such as facial action coding), and the objective quality at a bitrate of 5 kbps is significantly better than that of traditional encoders:

[0125] In this embodiment, the comparison of evaluation metrics between the video portrait semantic compression method based on key-point features and HEVC at the same bitrate is as Figure 3As shown, where (a) is the PSNR comparison graph and (b) is the SSIM comparison graph. In terms of PSNR comparison, the present embodiment at 5 kbps is comparable to the HEVC preset at 18 kbps. The video portrait semantic compression method based on key point features provided by the embodiment has significant advantages in terms of perceptual quality, with SSIM ≥ 0.92, reaching the level of HEVC at 35 kbps. When the bit rate is low (<10 kbps), in terms of bandwidth efficiency, at the same subjective quality, the bit rate requirement is only 1 / 3 - 1 / 4 of the traditional scheme, which is very suitable for narrowband communication scenarios.

[0126] The video portrait semantic compression method based on key point features provided by the present embodiment can be deployed in scenarios such as remote video conferencing and low-bandwidth emergency rescue. While ensuring the visual quality of the portrait, it reduces the transmission bandwidth requirement by 60% - 75%, providing core support for ultra-low bit rate video services under 5G / 6G networks.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A video portrait semantic compression method based on key point features, characterized in that: It includes an encoding end and a decoding end; The encoding end obtains a portrait video to be compressed, generates a FLAME model of the portrait head based on the portrait video to be compressed, and extracts the dynamic features of the FLAME model of the portrait head; Randomly select several frames of images from the portrait video to be compressed as transmission images, and transmit the transmission images and the dynamic features of the FLAME model of the portrait head to the decoding end; The decoding end receives the transmission images and the dynamic features of the FLAME model of the portrait head, and extracts the three-plane features of the transmission images that can be used for rendering based on the transmission images; Perform dynamic query and weighted fusion on the three-plane features of different transmission images that can be used for rendering to obtain the normalized three-plane features of each frame of transmission image; Perform ray sampling on each frame of transmission image, extract the static features of each frame of transmission image based on the normalized three-plane features of each frame of transmission image, and combine the dynamic features of the FLAME model of the portrait head to obtain a reconstructed high-fidelity portrait video.

2. The video portrait semantic compression method based on key point features according to claim 1, wherein: The encoding end includes: Step 1: The encoding end obtains a portrait video to be compressed and generates a FLAME model of the portrait head based on the portrait video to be compressed; Step 2: The encoding end extracts the dynamic features of the FLAME model of the portrait head; Step 3: Randomly select several frames of images from the portrait video to be compressed as transmission images, and transmit the transmission images and the dynamic features of the FLAME model of the portrait head to the decoding end.

3. The video portrait semantic compression method based on key point features according to claim 2, characterized in that: The said Step 1 includes: Step 1.1: Obtain a portrait video to be compressed, and use multiple pre-trained face capture models to parallelly extract the FLAME parameters of each frame of image; Among them, the FLAME parameters of each frame of image include expression parameter ψ, head pose parameter θ, and shape parameter β; Step 1.2: Optimize the FLAME parameters of each frame of image based on the multi-frame joint optimization method, and optimize the smoothness of the FLAME parameters of adjacent two frames of images by constructing a time-series related loss function; the time-series related loss function includes a shape consistency loss for constraining the shape parameter β and a pose continuity loss for constraining the head pose parameter θ; Step 1.3: Use the FLAMEDense model to decode the optimized FLAME parameters, decode the optimized FLAME parameters into several 3D vertex coordinates, and perform spatial alignment on the several 3D vertex coordinates using a vertex scaling factor and camera parameters to obtain a FLAME model of the portrait head; among them, the camera parameters include a focal length parameter and a principal point parameter; Step 1.4: Use a differentiable mesh renderer to render the several 3D vertex coordinates into an Alpha image with an alpha channel, and project the several 3D vertex coordinates onto a 2D plane.

4. A method for semantic compression of video portraits based on key-point features according to claim 3, characterized in that: The said Step 2 includes: Step 2.1: Assign learnable feature weights to each 3D vertex of the FLAME model of the portrait head; Step 2.2: Use the NeRF sampling strategy to sample the FLAME model of the human head to obtain a number of dynamic sampling points, and acquire the feature weights and position coordinates of the K nearest neighbor points for each dynamic sampling point; for the dynamic sampling point s of the FLAME model of the human head, retrieve the K nearest neighbor points of the dynamic sampling point s in the entire FLAME model of the human head, and obtain the feature weight f of the k-th nearest neighbor point among the K nearest neighbor points of the dynamic sampling point s k and the position coordinate p of the k-th nearest neighbor point k ; Step 2.3: Regress the feature weights of the K nearest neighbor points of each dynamic sampling point to extract the dynamic features of each dynamic sampling point.

5. A method for video portrait semantic compression based on key point features according to claim 4, characterized in that: The specific method of the said Step 2.3 is: Perform linear regression on the feature weights of the K nearest neighbors of the dynamic sampling point s, perform frequency position encoding on the position coordinates of the K nearest neighbors of the dynamic sampling point s, and perform weighted fusion on the feature weights of the K nearest neighbors obtained by linear regression of the dynamic sampling point s and the position coordinates of the K nearest neighbors obtained by frequency position encoding of the dynamic sampling point s according to the position weights of the K nearest neighbors of the dynamic sampling point s to generate the dynamic feature f of the dynamic sampling point s exp , as shown in the following formula: Among them, L p (·) is a linear regression function, F pos (·) is a frequency position encoding function, w k is the position weight of the k-th nearest neighbor point of the dynamic sampling point s.

6. The video portrait semantic compression method based on key point features according to claim 5, wherein: The decoding end includes: S1: The decoding end receives the transmission images and the dynamic features of the FLAME model of the portrait head, and extracts the three-plane features of the transmission images that can be used for rendering; S2: The decoding end performs dynamic query and weighted fusion on the three-plane features available for rendering of different images to obtain the normalized three-plane features of each frame of image; S3: The decoding end performs ray sampling on each frame of transmitted image, extracts the static features of each frame of transmitted image based on the normalized three-plane features of each frame of transmitted image, and combines the dynamic features of the FLAME model of the human head to obtain a reconstructed high-fidelity human portrait video.

7. A method for semantic compression of video portraits based on key-point features according to claim 6, characterized in that: The said S1 includes: S1.1: Use the ViT network to perform global coordinate transformation and feature extraction on each frame of transmitted image to obtain the global geometric feature map of each frame of transmitted image; For any frame transmission image F k , use the ViT network to divide the transmission image F k into non-overlapping blocks NOP of a fixed size, and use linear projection to convert each non-overlapping block NOP of a fixed size into a sequence of embedding vectors PE; perform global coordinate transformation on the sequence of embedding vectors PE, add position encoding to each non-overlapping block NOP of a fixed size in the sequence of embedding vectors PE, and retain the spatial position information of each non-overlapping block NOP of a fixed size to obtain a sequence of embedding vectors PE_SEQ with position encoding; use multiple SegFormer networks to perform feature extraction on the sequence of embedding vectors PE_SEQ with position encoding to obtain the feature sequence SFB_FEAT of the transmission image F k , and perform multi-scale feature extraction on the feature sequence SFB_FEAT of the transmission image F k to obtain the multi-resolution feature map MR_FEATS of the transmission image F k ; perform feature recombination on the multi-resolution feature map MR_FEATS of the transmission image F k and rearrange it into a spatial feature map as the global geometric feature map GG_FEAT of the transmission image F k . S1.2: Use the VGG network to perform layer-by-layer feature extraction on each frame of transmitted image to obtain the high-dimensional feature map of each frame of transmitted image; For any frame transmission image F k , use the VGG network to perform layer-by-layer feature extraction on the transmission image F k to obtain the feature map CLF after multiple convolutional layers; capture the high-frequency texture information of the feature map CLF through the local receptive field to obtain the feature map LRFF after local receptive field processing; then adjust the resolution of the feature map LRFF to obtain the high-dimensional feature map MRF of the transmission image F k with the same spatial resolution as the global geometric feature map GG_FEAT of the transmission image F k ; each convolutional layer in the VGG network increases the degree of abstraction of the features while gradually reducing the spatial size of the feature map. Each convolutional layer captures the high-frequency texture information of the transmission image F k and retains the fine-grained details of the transmission image F k . S1.3: Concatenate the global geometric feature map and the high-dimensional feature map of each frame of transmitted image, and extract the three-plane features of each frame of transmitted image; For any frame transmission image F k , the global geometric feature map GG_FEAT and the high-dimensional feature map MRF of the transmission image F k are concatenated along the channel dimension to obtain the fused feature FF; three lightweight 3×3 convolutional layers are used to compress the channel number of the fused feature FF to obtain the intermediate feature map IF, and then a 3×3 convolutional layer is used to map the intermediate feature map IF into a three-plane representation on three orthogonal planes XY, XZ, and YZ to obtain the three-plane feature TPF of the transmission image F k . S1.4: Perform normalization processing on the three-plane features of each frame of transmitted image to obtain the three-plane features available for rendering of each frame of transmitted image; For any frame transfer image F k For the three-plane feature TPF of k , the three orthogonal planes XY, XZ, and YZ are associated with a predefined world coordinate system for spatial alignment to obtain the spatially aligned three-plane feature ATPF; the spatially aligned three-plane feature ATPF is normalized to obtain the normalized three-plane feature NTPF; based on the compact 3D representation of the normalized three-plane feature NTPF, the transfer image F is obtained using a rendering model k The three-plane feature RTPF that can be used for rendering.

8. A method for semantic compression of video portraits based on key point features according to claim 7, characterized in that: The said S2 includes: S2.1: Obtain the shape parameters of the three-plane features available for rendering of each frame of transmitted image, perform shape transformation on the three-plane features available for rendering of each frame of transmitted image to obtain the new three-plane features of each frame of transmitted image; Obtaining the shape parameters of the three-plane features available for rendering of each frame of transmitted image includes: the batch size B, the number of feature groups N, the three-plane index T, the number of channels C, the feature map height H, and the feature map width W of the three-plane features RTPF available for rendering of each frame of transmitted image; Use the rearrange function to convert the shape of the three-plane features RTPF available for rendering of each frame of transmitted image from B*N*T*C*H*W to (B*T*H*W)*N*C to obtain the new three-plane features XTPF of each frame of transmitted image; S2.2: Introduce a set of learnable query planes, and through multi-layer non-linear transformation, map the new three-plane features XTPF of each frame of transmitted image to the query space and the key space to obtain the query matrix and the key matrix of the three-plane features of different images; S2.3: Calculate the correlation of the new three-plane features of different transmitted images, and generate the attention weights of each frame of transmitted image; For any frame transmission image F k and the transmission image F j , calculate the correlation dots k between the query matrix of the new tri-plane feature of the transmission image F j and the key matrix of the new tri-plane feature of the transmission image F k,j , as shown in the following formula: Among them, Q k is the query matrix of the k-th frame transmitted image F k K j is the key matrix of the j-th frame transmitted image F j and dim is the new three-plane feature dimension of the transmitted image; Calculate the attention weight attn of the transmitted image F using the Softmax function k as follows: k as shown in the following formula: attn k = softmax(dots k,j ) S2.4: Perform dynamic query and weighted fusion on the three-plane features from different images to obtain the normalized three-plane features of each frame of transmitted image; For the transmitted image F k and the transmitted image F j The attention weights are weighted and fused to obtain the transmitted image F k The fused three-plane feature P is as shown in the following formula: where N is the number of input images, attn i is the weight of the k-th frame image F k and attn j is the weight of the j-th frame image F j , and E(·) is the transmitted image encoder; Transfer the transmitted image F k The shape of the fused three-plane feature P is converted from (B*T*H*W)*N*C to B*N*T*C*H*W and used as the image F k The normalized three-plane feature GTPF 9. A method for semantic compression of video portraits based on key point features according to claim 8, characterized in that: The said S3 includes: S3.1: Calculate the camera origin and the ray direction according to the camera parameters of the human portrait video to be compressed, align the rays to the canonical space through the transformation matrix, use the axis-aligned bounding box of the FLAME model of the human head to filter the valid ray segments from the aligned rays, and uniformly sample a number of three-plane ray sampling points on the valid ray segments; S3.2: Project each three-plane ray sampling point onto the normalized three-plane features GTPF of each frame of transmitted image, and extract the static features of each frame of transmitted image through bilinear interpolation; For any three-plane ray sampling point projected onto an arbitrary frame transmission image, project the three-plane ray sampling point onto the normalized three-plane GTPF of the transmission image, and extract the features of the three-plane ray sampling point on each plane through bilinear interpolation. Among them, the projection point coordinates of the three-plane ray sampling point on the XY plane are (x, y), and the sampling feature of the three-plane ray sampling point is f xy ; the projection point coordinates of the three-plane ray sampling point on the XZ plane are (x, z), and the sampling feature of the three-plane ray sampling point is f xz , the projection point coordinates of the three-plane ray sampling point on the YZ plane are (y, z), and the sampling feature of the three-plane ray sampling point is f yz ; splice the sampling features of the three-plane ray sampling points on the three planes to obtain the static feature f of the transmission image canonical ; S3.3: Concatenate the dynamic features of the FLAME model of the human head and the static features of the transmitted image to obtain the fused features; Concatenate the dynamic feature f of the FLAME model of the human head exp with the static feature f of the transmitted image canonical to obtain the fused feature f fused , as shown in the following formula: f fused = Concat(f canonical , f exp ) Among them, Concat(·) is the concatenation operation; S3.4: Based on the fused feature f fused and the light direction, use a multi-layer perceptron (MLP) to extract the density σ and color RGB of each three-plane light sampling point, and use the ray_marching function to accumulate the extraction results to obtain a reconstructed high-fidelity portrait video.