Human body video coding, decoding and communication method and model optimization method

By using pre-trained 3D mannequin to generate a compact and configurable human semantics collection, the problems of insufficient compression efficiency and no interaction in the prior art are solved, and efficient video compression and enhanced interaction are achieved.

CN120034653APending Publication Date: 2025-05-23ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510107799.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art still has room for improvement in the compression efficiency of encoded bitstreams of human videos, and the current compressed bitstream does not support interaction with the original signal of the human body, limiting the interactive function of video communication.

Method used

By inputting the frames of the human video to be encoded into the pretrained 3D human body model, a compact and configurable human semantics collection is generated, compressed to generate an encoded bitstream, and interact with the human body signal by configuring the value of the human body semantics at the decoding end.

Benefits of technology

It improves the compression efficiency of the encoded bitstream, enhances the interactivity of video communication, and supports interaction with the original signals of the human body.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034653A_ABST
    Figure CN120034653A_ABST
Patent Text Reader

Abstract

The invention provides a human body video coding, decoding and communication method and a model optimization method, and the method comprises the steps: inputting a frame in a to-be-coded human body video into a pre-trained 3D human body model OSX, and obtaining a first human body semantic set which is used for simplifying the description of a human body signal; a second human body semantic set is determined from the first human body semantic set, and the second human body semantic set is used for describing signals related to motion in the human body signals, so that the decoding end interacts with the human body signals by configuring values of human body semantics in the second human body semantic set; compressing the second human body semantic set to obtain a first encoding bit stream, and compressing a key reference frame in the to-be-encoded human body video to obtain a second encoding bit stream; and determining the first coding bit stream and the second coding bit stream as a target coding result of the to-be-coded human body video. Compact and configurable human body semantics are used for encoding human body signals, the compression efficiency is improved, and interactivity is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video technology, and in particular to a method for encoding, decoding, communicating, and optimizing a model of a human body video. Background Art

[0002] In recent years, with the advent of the "short video era", media content for humans has exploded. This has not only changed the way we obtain information and entertainment, but also brought new technical challenges and opportunities. Although many advances have been made in improving the compression and storage efficiency of human video content, the compression efficiency of the encoded bitstream still needs to be improved. In addition, the current compressed bitstream does not support interaction with the original human signal, which limits the interactive function of video communication. In response to this technical problem, the relevant technology has not yet proposed an effective solution. Summary of the invention

[0003] The embodiments of the present application provide human body video encoding, decoding, communication methods and model optimization methods to solve one or more of the above-mentioned technical problems.

[0004] In a first aspect, an embodiment of the present application provides a method for encoding a human body video, comprising: inputting a frame in a human body video to be encoded into a pre-trained 3D human body model OSX to obtain a first human body semantic set, wherein the first human body semantic set is used to simplify the description of a human body signal in the human body video to be encoded; determining a second human body semantic set from the first human body semantic set, wherein the second human body semantic set is used to describe motion-related signals in the human body signal, so as to interact with the human body signal at the decoding end by configuring the value of the human body semantics in the second human body semantic set; compressing the second human body semantic set to obtain a first encoded bit stream, and compressing the key reference frame in the human body video to be encoded to obtain a second encoded bit stream; determining the first encoded bit stream and the second encoded bit stream as the target encoding result of the human body video to be encoded.

[0005] In a second aspect, an embodiment of the present application provides a method for decoding a human body video, comprising: after receiving a first coded bit stream and a second coded bit stream determined by the above-mentioned human body video encoding method, decoding and reconstructing based on the first coded bit stream to obtain the second human body semantic set, and decoding and reconstructing based on the second coded bit stream to obtain the key reference frame; inputting the key reference frame into the pre-trained 3D human body model OSX, and using the parameter extraction module of the 3D human body model OSX to obtain a third human body semantic set in the first human body semantic set except the second human body semantic set; inputting the second human body semantic set and the third human body semantic set into the decoder of the 3D human body model OSX to reconstruct a 3D human body mesh of a subsequent frame in the human body video to be encoded, and inputting the key reference frame into the decoder to reconstruct the 3D human body mesh of the key reference frame; performing motion estimation on the 3D human body mesh of the subsequent frame and the 3D human body mesh of the key reference frame; and generating a human body video based on the key reference frame and the motion estimation result.

[0006] In a third aspect, an embodiment of the present application provides a communication method for a human body video, comprising: after receiving a first coded bit stream and a second coded bit stream determined by the above-mentioned human body video encoding method, decoding and reconstructing based on the first coded bit stream to obtain the second human body semantic set, and decoding and reconstructing based on the second coded bit stream to obtain the key reference frame; inputting the key reference frame into the pre-trained 3D human body model OSX, and using the parameter extraction module of the 3D human body model OSX to obtain a third human body semantic set in the first human body semantic set except the second human body semantic set; modifying the value of the human body semantics in the second human body semantic set and / or the value of the human body semantics in the third human body semantic set according to the user's personalized communication needs, to obtain a reconfigured second human body semantic set and a reconfigured third semantic set; using the above-mentioned human body video decoding method to generate a changed human body video.

[0007] In a fourth aspect, an embodiment of the present application provides a model optimization method, wherein the model is set at a decoding end, and the model includes a first network and a second network. The first network and the second network are used to implement the above-mentioned human body video decoding method to perform motion estimation and generate human body video, including: using an end-to-end strategy to jointly train the first network and the second network.

[0008] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the methods described above when executing the computer program.

[0009] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the methods described above is implemented.

[0010] In a seventh aspect, an embodiment of the present application provides a computer program product, including computer instructions, which implement any of the methods described above when executed by a processor.

[0011] Compared with the related art, this application has the following advantages:

[0012] The human body signal is represented by the human body semantics output by the pre-trained 3D human body model OSX, and the output human body semantics are optimized to generate compact human body semantics. In addition, since the compact human body semantics have clear physical meanings and are configurable at the decoding end, the compact and configurable human body semantics finally generated encode the human body signal, which improves the compression efficiency and enhances the interactivity, thereby solving the technical problem that the compression efficiency of the encoded bit stream in the related technology needs to be improved. In addition, the current compressed bit stream does not support interaction with the original human body signal, which limits the interactive function of video communication.

[0013] With semantically hierarchical representations in the encoded bitstream, these semantically meaningful representations are explicitly evolved into a high-dimensional grid representation with the assistance of a 3D human body model, which promotes dynamic perception and further improves the quality of the decoded video required for reconstruction.

[0014] By modifying different semantic level representations, the reconstruction of the 3D human mesh can be flexibly edited so that the head posture and body posture of the 3D mesh can be changed to achieve personalized representation, thereby realizing interactive human video communication.

[0015] The first network and the second network are jointly trained using an end-to-end strategy, so that the trained first network can improve the efficiency of grid-based dense motion estimation, thereby promoting dynamic perception, and the trained second network can improve the efficiency of generating human body frames, thereby improving the quality of decoded video.

[0016] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.

[0018] Figure 1 It shows a schematic diagram of a video compression framework in the related art provided in the embodiments of the present application that follows a prediction-transformation architecture;

[0019] Figure 2 A basic framework diagram of an end-to-end video compression depth model in the related art provided in an embodiment of the present application is shown;

[0020] Figure 3 A basic framework schematic diagram of a deep learning video generation and compression scheme based on a first-order motion model in the related art provided in an embodiment of the present application is shown;

[0021] Figure 4 A basic framework schematic diagram of a deep learning video generation compression scheme based on compact feature representation in the related art provided in an embodiment of the present application is shown;

[0022] Figure 5 A schematic diagram of a human body video encoding method provided in an embodiment of the present application is shown;

[0023] Figure 6 A structural block diagram of a human body video encoding device provided in an embodiment of the present application is shown;

[0024] Figure 7 A schematic diagram of a human body video decoding method provided in an embodiment of the present application is shown;

[0025] Figure 8 A structural block diagram of a human body video decoding device provided in an embodiment of the present application is shown;

[0026] Fig. 9 A schematic diagram of a human body video communication method provided in an embodiment of the present application is shown;

[0027] Fig.10 A structural block diagram of a human body video communication device provided in an embodiment of the present application is shown;

[0028] Fig.11 A schematic diagram of a model optimization method provided in an embodiment of the present application is shown;

[0029] Fig.12 A structural block diagram of a model optimization device provided in an embodiment of the present application is shown;

[0030] Fig.13 A block diagram of an IHVC framework structure provided in an embodiment of the present application is shown; and

[0031] Fig.14 A block diagram of an electronic device used to implement an embodiment of the present application is shown. DETAILED DESCRIPTION

[0032] In the following, only some exemplary embodiments are briefly described. As those skilled in the art will appreciate, the described embodiments may be modified in various ways without departing from the concept or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature and not restrictive.

[0033] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and they all belong to the protection scope of the embodiments of the present application.

[0034] A related art prior to this application is that traditional video compression standards, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), have been highly developed and achieved good compression performance. These standards use a block-based hybrid video coding framework to exploit spatial redundancy, temporal redundancy, and information entropy redundancy in the video. In this framework, the video compression encoder generates a bitstream based on the input current frame, and the decoder reconstructs the video frame based on the received bitstream. Figure 1 The classical video compression framework shown follows the prediction-transform architecture. Specifically, the input frame x t The video is divided into a series of blocks of the same size, i.e., square areas (e.g., blocks of 8×8 pixels). The encoding process of the traditional video compression algorithm at the encoder end mainly includes the following steps: Step 1, motion estimation: estimate the current frame x t Reconstructed frame with the previous The motion between blocks is obtained by moving the motion vector v corresponding to each block. t Step 2: Motion compensation: By using the motion vector v defined in step 1 t , copy the corresponding pixels in the previous reconstructed frame to the current frame to get the predicted frame Then calculate the original frame x t With predicted frame The residual r t ,Right now Step 3: Transform and quantize: The residual r obtained in step 2 t Quantized to Use linear transformation (e.g., Discrete Cosine Transform (DCT)) before quantization to obtain better compression performance. Step 4: Inverse transformation: Use the quantization result in step 3 Perform inverse transform to obtain the reconstructed residual Step 5: Entropy coding: Entropy coding is used to convert the motion vector v in step 1 into t and the quantization result in step 3 Encode into a bit stream and send to the decoder. Step 6, frame reconstruction: By reconstructing the predicted frame in step 2 and the reconstructed residual in step 4 Add to get the reconstructed frame Right now The reconstructed frame will be used for motion estimation of the t+1th frame in step 1. For the decoder, according to the bit stream provided by the encoder in step 5, motion compensation in step 2, inverse quantization in step 4, and then frame reconstruction in step 6 are performed to obtain the reconstructed frame

[0035] Another related technology before this application is: With the rapid development of deep learning, many deep learning-based algorithms have been introduced to replace or enhance video coding tools. Regarding the joint optimization of the entire image / video compression framework, rather than designing a specific module, end-to-end image / video compression algorithms have emerged. For example, the end-to-end video coding scheme DVC, which jointly optimizes all components of video compression. In addition, content adaptation and error propagation perception problems are considered, and an online encoder update scheme is developed to improve video compression performance. However, these traditional or learning-based methods target general natural scenes without special consideration of human content such as faces, bodies, or other parts. Figure 2 We present the basic framework of the first end-to-end deep model for video compression, which jointly optimizes all components of video compression, such as motion estimation, motion compression, and residual compression. Specifically, we use learning-based optical flow estimation to obtain motion information and reconstruct the current frame, and then use two autoencoder-style neural networks to compress the corresponding motion and residual information. All modules are jointly learned through a single loss function, and they work together to make a trade-off between reducing the number of compressed bits and improving the quality of the decoded video. Figure 2 The new end-to-end deep framework presented in Figure 1There is a one-to-one correspondence between the traditional video compression frameworks shown in . These relationships and their differences are briefly summarized as follows: Step 1, motion estimation and compression: Use a convolutional neural network (CNN) model to estimate optical flow, which is regarded as motion information v t Instead of encoding the raw optical flow values ​​directly, a motion vector (MV) encoder-decoder network is used to compress and decode the optical flow values, where the quantized motion representation is denoted as Then the corresponding reconstructed motion information can be decoded through the MV decoder network Step 2: Motion compensation: A motion compensation network is designed to obtain the predicted frame based on the optical flow obtained in step 1. Step 3: Transformation, quantization, and inverse transformation: A highly nonlinear residual encoder-decoder network is used to replace the linear transformation. is nonlinearly mapped to represent y t , then y t Quantified as In order to build an end-to-end training scheme, a quantization method is used. The quantized representation is fed into the residual decoder network to obtain the reconstructed residual Step 4: Entropy coding: During the test phase, the quantized motion representation in step 1 and the residual in step 3 is expressed as is encoded into bits and sent to the decoder. During the training phase, CNN is used to obtain The probability distribution of each symbol in . Step 5: Frame reconstruction: This is the same as the traditional method and will not be described in detail here.

[0036] Another related technology before this application is: With the emergence of deep generative models, especially generative adversarial networks (GANs), face video compression has achieved significant performance improvements. In the video-to-video synthesis task, a novel scheme Face-vidtovid is proposed, which uses a compact three-dimensional key point representation to drive the generative model to render the target frame. In addition, some technologies have also proposed VSBNet, which uses adversarial learning to reconstruct the original frame from the landmarks. In addition, some technologies have proposed an end-to-end talking head video compression framework based on compact feature learning (Compact Feature Transformation and Embedding, CFTE), which is cleverly designed to achieve efficient face video compression and is suitable for ultra-low bandwidth scenarios. The CFTE scheme uses compact feature representation to compensate for time evolution and reconstruct the target face video frame in an end-to-end manner. In addition, it can be combined with a video coding framework with rate-distortion target supervision. Although these algorithms have achieved the reconstruction of frames with only a few facial parameters through the powerful rendering capabilities of deep generative models, some head postures and facial expression movements cannot be accurately rendered compared to the original dynamic video. Figure 3 The basic framework of the deep learning video generation compression scheme based on the first-order motion model is presented. There are also technologies that propose a first-order motion model (FOMM) to deform the reference source frame to follow the motion in the driving video. This method adopts an encoder-decoder architecture and combines a motion transfer component: Step 1. Use equivariant loss to learn a key point extractor without explicit labels. Through the key point extractor, two sets of ten learned key points for the source frame and the driving frame are calculated. The learned key points are converted from the channel 64×64 size feature map through a Gaussian mapping function, and each corresponding key point can represent the feature information of different channels. It should be noted that each key point is a (x, y) point that can represent the most important information in the feature map. Step 2. The dense motion network uses these key points and the source frame to generate a dense motion field and occlusion map. Step 3. The encoder encodes the source frame through traditional image / video compression methods (such as High Efficiency Video Coding (HEVC) / VVC or Joint Photographic Experts Group (JPEG) / BPG). Here, VVC is used to compress the source frame. Step 4: The generated feature map is warped using a dense motion field (via a differentiable grid sampling operation) and then multiplied with the occlusion map. The decoder generates an image from the warped map.

[0037] Another related technology before this application is Figure 4 The basic framework of another deep learning video generation compression scheme based on compact feature representation is presented, which follows an encoder-decoder architecture. On the encoding side, the compression framework consists of three modules: an encoder for compressing key frames, a feature extractor for extracting compact human features from other intermediate frames, and a feature encoding module for compressing inter-frame prediction residuals of compact human features. First, the key frames representing human textures are compressed by the VVC encoder. Each subsequent intermediate frame is represented by a compact feature matrix of size 1×4×4 through the compact feature extractor. It should be noted that the size of the compact feature matrix is ​​not fixed, and the number of feature parameters can be increased or decreased based on the specific needs of bitrate consumption. Subsequently, these extracted features are processed by inter-frame prediction and quantization, and finally the residuals are entropy encoded to generate the final bitstream. On the decoding side, the compression framework also contains three main modules, including a decoding module for reconstructing key frames, a module for reconstructing compact features through entropy decoding and compensation, and a module for generating the final video using the reconstructed features and the decoded key frames. More specifically, in the process of generating the final video, the key frames decoded from the VVC bitstream can be further represented in the form of features through compact feature extraction. Subsequently, the relevant sparse motion fields are calculated based on the features of key frames and intermediate frames, and then pixel-level dense motion maps and occlusion maps are generated. Finally, based on the deep generative model, the decoded key frames, pixel-level dense motion maps, and occlusion maps represented by implicit motion fields are used to generate the final video with accurate appearance, posture, and expression.

[0038] Although the traditional or learning-based end-to-end video compression methods in the above-mentioned related technologies can achieve relatively efficient compression performance in moving human videos, there are some defects in directly applying common compression algorithms to human video compression systems with ultra-low bit rates and enhanced interactivity:

[0039] 1. Traditional compression algorithms compress each video frame by using block-based motion estimation, discrete cosine transform (DCT), etc., so it is still difficult to further reduce the coding bits. As a result, such algorithms are not suitable for ultra-low bit rate human video compression scenarios.

[0040] 2. These traditional or learning-based end-to-end video compression methods focus on general natural scenes without specifically considering human motion information. In particular, features extracted from moving humans are described by feature structure changes with strong priors, such as key points or skeletons, which can greatly help reconstruct higher quality videos.

[0041] 3. Although these generative compression algorithms (such as FOMM or Face_vidtovid) achieve frame reconstruction with a small number of parameters through the powerful rendering capabilities of deep generative models, some human pose movements are still not accurately rendered compared to moving human videos. That is, most algorithms based on 2D face representation (i.e., 2D landmarks and 2D key points) perform poorly in terms of photo realism, or fail to solve the identity preservation problem, or fail to fully transfer the driving pose.

[0042] 4. For these existing 2D generation compression algorithms, semantic information cannot be included in the compressed bitstream to control human motion posture. This obvious defect greatly limits the application of digital human communication in the metaverse.

[0043] In view of this, the embodiments of the present application provide a method for encoding human body videos to solve all or part of the above technical problems. The application scenarios of this method can be ultra-low bandwidth and enhanced interactive human communications in the metaverse (for example, medical care and health, education and training, film and television and entertainment, security and monitoring, transportation and autonomous driving, etc.) and virtual uploaders of live e-commerce. In other words, the human body video encoding method provided in the embodiments of the present application can be applied to related technologies such as human body video compression encoding and storage. Figure 5 As shown, the encoding method of the human body video may include:

[0044] S502, inputting frames in the human body video to be encoded into a pre-trained 3D human body model OSX to obtain a first human body semantic set, wherein the first human body semantic set is used to simplify the description of human body signals in the human body video to be encoded.

[0045] It should be noted that the above human body video is mainly a type of video that focuses on capturing, recording and analyzing human body movements and appearance. The frames in the above human body video to be encoded may include reference key frames and subsequent frames.

[0046] The above 3D human body model is a digital human body simulation, which uses a three-dimensional coordinate system and geometric shapes such as polygons and curved surfaces to construct a virtual human body with a sense of reality. Based on this, the present embodiment uses a pre-trained 3D human body model OSX to concisely represent high-dimensional human body signals with human body semantic information. For example, the above first human body semantic set can be a 3D independent body joint δ body ∈R 21×3 (i.e. the positions of the 21 key joints in 3D space), body shape δ shape ∈R 10 , 3D global translation δ trans ∈R 3 , 3D global rotation δ rot ∈R 3 and the bounding box δ for localizationloc ∈R 3 Compared with the implicit or physically meaningless representations in the related art, the embodiments of the present application provide clear human semantics, which can further achieve interactive compression.

[0047] In order to improve compression efficiency through more economical representation, the embodiment of the present application also proposes S504, which determines a second human body semantic set from the first human body semantic set, wherein the second human body semantic set is used to describe motion-related signals in the human body signal, so as to interact with the human body signal by configuring the values ​​of the human body semantics in the second human body semantic set at the decoding end.

[0048] Optionally, the human body semantics in the second human body semantic set are highly decoupled human body semantics, and configuring one of the human body semantics does not affect the values ​​of other human body semantics. The human body semantics in the second human body semantic set at least include the following human body semantics: human body posture parameters, human body translation parameters, human body rotation parameters, and human body position parameters, etc. Among them, the human body semantics can be multi-dimensional, and the dimensions of different human body semantics can be determined by transmission bandwidth, reconstructed video quality, etc.

[0049] Exemplarily, the total dimension of the second human semantic set may be 40. The human posture parameters may be 21-dimensional, the human translation parameters may be 3-dimensional, the human rotation parameters may be 3-dimensional, and the human position parameters may be 4-dimensional.

[0050] S506: compress the second human body semantic set to obtain a first coded bit stream, and compress the key reference frame in the human body video to be coded to obtain a second coded bit stream.

[0051] Optionally, in an embodiment of the present application, compressing the second human semantic set to obtain the first coded bit stream may be performed on the second human semantic set by inter-frame prediction, quantization, and entropy coding to obtain the first coded bit stream. The key reference frame in the human body video to be encoded is compressed to obtain the second coded bit stream by inputting the key reference frame in the human body video to be encoded into the VVC codec for compression, further providing a texture reference for signal synthesis.

[0052] S508: Determine the first encoded bit stream and the second encoded bit stream as target encoding results of the human body video to be encoded.

[0053] Through the above steps S502 to S508, the frames in the human body video to be encoded are input into the pre-trained 3D human body model OSX to obtain a first human body semantic set, wherein the first human body semantic set is used to simplify the description of the human body signal in the human body video to be encoded; a second human body semantic set is determined from the first human body semantic set, wherein the second human body semantic set is used to describe the motion-related signals in the human body signal, so as to interact with the human body signal at the decoding end by configuring the values ​​of the human body semantics in the second human body semantic set; the second human body semantic set is compressed to obtain a first encoded bit stream, and the key reference frames in the human body video to be encoded are compressed to obtain a second encoded bit stream; the first encoded bit stream and the second encoded bit stream are determined as the target encoding result of the human body video to be encoded. That is to say, in the embodiment of the present application, the human body signal is represented by the human body semantics output by the pre-trained 3D human body model OSX, and the output human body semantics are optimized to generate compact human body semantics. In addition, since the compact human body semantics have clear physical meanings and are configurable at the decoding end, the compact and configurable human body semantics finally generated encode the human body signal, which improves the compression efficiency and enhances the interactivity, thereby solving the technical problem that the compression efficiency of the encoded bit stream in the related technology needs to be improved, and the current compressed bit stream does not support interaction with the original human body signal, which limits the interactive function of video communication.

[0054] Considering that human frames in the same sequence (i.e., a series of consecutive image frames from the same video clip) share matching body shapes δ shape , so these morphological coefficients can be directly extracted from the reconstructed key reference frame without separate signal transmission. In addition, by analyzing the independent body joints δ body From the physical meaning of , we can find that only the last 7 joints are related to the head posture and hand movements. Therefore, the last 42 dimensions of the 63-dimensional vector are used as the posture action coefficient δ body ∈R 10×3Signal transmission is performed, and the remaining first 42 dimensional parameters can be obtained from the reconstructed key reference frame. For other human semantics related to translation, rotation and positioning, they can be retained due to their compact representation and interactive functions. Therefore, the embodiment of the present application proposes that the above-mentioned S504 may include: S5401, determining the human body signal shared by the human body frames in the same sequence, and the human body semantics corresponding to the shared human body signal in the first human body semantic set, to obtain the first target human body semantics; S5402, determining the human body semantics that are not related to motion in the first human body semantic set, to obtain the second target human body semantics; S5403, filtering out the first target human body semantics and the second target human body semantics from the first human body semantic set, to obtain the second human body semantic set. That is, the first human body semantic set is further optimized to screen out compact and interactive human body semantics.

[0055] Exemplarily, it is assumed that the first human body semantic set includes: 3D independent body joints δ body ∈R 21×3 , body size shape ∈R 10 , 3D global translation δ trans ∈R 3 , 3D global rotation δ rot ∈R 3 and the bounding box δ for localization loc ∈R 3 , due to the body size δ shape ∈R 10 The body shape δ can be expressed as the human semantics corresponding to the shared human signal. shape ∈R 10 Determined as the first target human semantics, 3D independent body joint δ body ∈R 21×3 Only the human semantics δ related to head posture and hand movements are included in pose Determined as the second target human semantics, then the final transmitted second human semantics set is δ sem ={δ pose ,δ trans ,δ rot ,δ loc}.

[0056] In order to achieve efficient compression of the second human body semantic set, the embodiment of the present application also proposes that the above-mentioned S506 may include: S5601, performing inter-frame prediction between the human body semantics of the current frame and the human body semantics of the previously reconstructed frame to obtain the residual generated by the inter-frame prediction; S5602, compressing the residual generated by the inter-frame prediction based on context-based arithmetic coding to obtain the first encoded bit stream.

[0057] It should be noted that the above context arithmetic coding combines the compression methods of arithmetic coding and context model. Among them, arithmetic coding is a lossless data compression algorithm that achieves compression by encoding the entire data stream into a single real number. The context model predicts the probability of the current symbol by analyzing the context of each symbol in the data (that is, its previous and subsequent symbols).

[0058] For example, assuming that the current frame is a person in front of the background with the left arm lowered, human body semantics 1 (for example, the coordinates of the left arm) is extracted, and the previously reconstructed frame is a person starting to raise the left arm, human body semantics 2 (for example, the new coordinates of the left arm) is extracted, human body semantics 1 and human body semantics 2 can be inter-frame predicted to obtain the residual generated by the inter-frame prediction, and then the residual generated by the inter-frame prediction is compressed based on the arithmetic coding of the above context, which significantly reduces the bandwidth requirements for storage and transmission while ensuring the video quality.

[0059] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0060] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0061] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a human body video encoding device. Figure 6 The structure block diagram of a human body video encoding device according to an embodiment of the present application is shown, and the device may include:

[0062] A first determination module 62 is used to input a frame in the human body video to be encoded into a pre-trained 3D human body model OSX to obtain a first human body semantic set, wherein the first human body semantic set is used to simplify the description of human body signals in the human body video to be encoded;

[0063] A second determination module 64 is used to determine a second human body semantic set from the first human body semantic set, wherein the second human body semantic set is used to describe a signal related to motion in the human body signal, so that the decoding end interacts with the human body signal by configuring the value of the human body semantics in the second human body semantic set;

[0064] A compression module 66, configured to compress the second human body semantic set to obtain a first coded bit stream, and to compress a key reference frame in the human body video to be coded to obtain a second coded bit stream;

[0065] The third determination module 68 is used to determine the first encoded bit stream and the second encoded bit stream as the target encoding result of the human body video to be encoded.

[0066] pass Figure 6 The device shown inputs frames in a human body video to be encoded into a pre-trained 3D human body model OSX to obtain a first human body semantic set, wherein the first human body semantic set is used to simplify the description of human body signals in the human body video to be encoded; determines a second human body semantic set from the first human body semantic set, wherein the second human body semantic set is used to describe motion-related signals in the human body signal, so as to interact with the human body signal at the decoding end by configuring the values ​​of the human body semantics in the second human body semantic set; compresses the second human body semantic set to obtain a first encoded bit stream, and compresses the key reference frames in the human body video to be encoded to obtain a second encoded bit stream; and determines the first encoded bit stream and the second encoded bit stream as the target encoding result of the human body video to be encoded. That is to say, in the embodiment of the present application, the human body signal is represented by the human body semantics output by the pre-trained 3D human body model OSX, and the output human body semantics are optimized to generate compact human body semantics. In addition, since the compact human body semantics have clear physical meanings and are configurable at the decoding end, the compact and configurable human body semantics finally generated encode the human body signal, which improves the compression efficiency and enhances the interactivity, thereby solving the technical problem that the compression efficiency of the encoded bit stream in the related technology needs to be improved, and the current compressed bit stream does not support interaction with the original human body signal, which limits the interactive function of video communication.

[0067] In an optional embodiment, the second determination module 64 includes: a first determination unit, used to determine the human body signals shared by the human body frames in the same sequence, and the human body semantics corresponding to the shared human body signals in the first human body semantic set, when the human body video to be encoded is a human body video composed of human body frames in the same sequence, to obtain the first target human body semantics; a second determination unit, used to determine the human body semantics that are not related to motion in the first human body semantic set, to obtain the second target human body semantics; a third determination unit, used to filter out the first target human body semantics and the second target human body semantics from the first human body semantic set, to obtain the second human body semantic set.

[0068] The compression module 66 includes: a fourth determination unit, which is used to perform inter-frame prediction between the human body semantics of the current frame and the human body semantics of the previously reconstructed frame to obtain the residual generated by the inter-frame prediction; a fifth determination unit, which is used to compress the residual generated by the inter-frame prediction based on context-based arithmetic coding to obtain the first encoded bit stream.

[0069] Optionally, the second human body semantic set includes at least the following human body semantics: human body posture parameters, human body translation parameters, human body rotation parameters and human body position parameters. The human body semantics are multi-dimensional, and the dimensions of different human body semantics are determined by the transmission bandwidth.

[0070] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0071] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a method for decoding a human body video. Figure 7 The method for encoding a human body video according to an embodiment of the present application is shown, comprising:

[0072] S702, after receiving the first coded bit stream and the second coded bit stream determined by the above-mentioned human body video coding method, decoding and reconstructing are performed based on the first coded bit stream to obtain the second human body semantic set, and decoding and reconstructing are performed based on the second coded bit stream to obtain the key reference frame.

[0073] Optionally, the second human body semantic set can be obtained by decoding and reconstructing the first coded bit stream and performing entropy decoding and compensation on the first coded bit stream to obtain the second human body semantic set. The key reference frame can be obtained by decoding and reconstructing the second coded bit stream and inputting the second coded bit stream into a VVC codec for decoding to obtain the above key reference frame.

[0074] S704, inputting the key reference frame into the pre-trained 3D human body model OSX, and using the parameter extraction module of the 3D human body model OSX to obtain a third human body semantic set in the first human body semantic set except the second human body semantic set.

[0075] For example, it is assumed that the first human body semantic set includes: 3D independent body joints δ body ∈R 21×3 , body size shape ∈R 10 , 3D global translation δ trans ∈R 3 , 3D global rotation δ rot ∈R 3 and the bounding box δ for localization loc∈R 3 , the second human body semantic set δ sem ={δ pose , δ trans , δ rot , δ loc}, then the third human body semantic set may include: body shape δ shape and the first 42 dimensions of δ body .

[0076] S706, input the second human body semantic set and the third human body semantic set into the decoder of the 3D human body model OSX to reconstruct the 3D human body mesh of the subsequent frames in the to-be-encoded human body video, and input the key reference frame into the decoder to reconstruct the 3D human body mesh of the key reference frame.

[0077] S708, perform motion estimation on the 3D human body mesh of the subsequent frames and the 3D human body mesh of the key reference frame.

[0078] S710, generate a human body video based on the key reference frame and the motion estimation result.

[0079] Through the above steps S702 - S710, perform decoding and reconstruction based on the first encoded bitstream to obtain the second human body semantic set, and perform decoding and reconstruction based on the second encoded bitstream to obtain the key reference frame; input the key reference frame into the pre-trained 3D human body model OSX, and use the parameter extraction module of the 3D human body model OSX to obtain the third human body semantic set in the first human body semantic set except the second human body semantic set; input the second human body semantic set and the third human body semantic set into the decoder of the 3D human body model OSX to reconstruct the 3D human body mesh of the subsequent frames in the to-be-encoded human body video, and input the key reference frame into the decoder to reconstruct the 3D human body mesh of the key reference frame; perform motion estimation on the 3D human body mesh of the subsequent frames and the 3D human body mesh of the key reference frame; generate a human body video based on the key reference frame and the motion estimation result. That is to say, in the embodiments of the present application, there is a semantic-level representation in the encoded bitstream, and with the assistance of the 3D human body model, these semantically meaningful representations are explicitly evolved into high-dimensional mesh representations, thereby promoting dynamic perception and further improving the quality of the decoded video required for reconstruction.

[0080] In a possible implementation manner, the above step S706 may include: S7061, use the generation function OSX(·) in the decoder, take the second human body semantic set and the third human body semantic set as the input of the generation function to obtain a generation result; S7062, determine the 3D human body mesh of the subsequent frames through the generation result.

[0081] Since the above-mentioned generation function OSX(·) can combine these semantic information to predict or generate certain structures, the embodiment of the present application uses the generation function OSX(·) to efficiently convert multi-frame human body semantics into a detailed 3D human body mesh model, further improving the quality of human body mesh reconstruction.

[0082] Exemplarily, assuming that the subsequent frame I after reconstruction l The human semantics of (1≤l≤n,l∈Z) is Reconstructed key reference frame The human semantics is First 42 dimensions Then the 3D human body mesh of the subsequent frames It can be determined by the following formula:

[0083]

[0084] 3D human body mesh of the above reconstructed key reference frame In the embodiment of the present application, it can be directly generated through the OSX model.

[0085] Optionally, the above S708 may include: S7081, converting the 3D human body mesh of the subsequent frame to obtain a 2D human body mesh of the subsequent frame, and converting the 3D human body mesh of the key reference frame to obtain a 2D human body mesh of the key reference frame. This step can convert the 3D human body mesh into a 2D human body mesh for motion estimation.

[0086] S7082, using a spatially adaptive normalized SPADE network to convert the 2D human body mesh of the subsequent frame and the 2D human body mesh of the key reference frame into a pixel-level dense motion field, wherein the dense motion field at least includes a dense motion flow and an occlusion map.

[0087] Optionally, a spatially adaptive normalization (SPADE) mechanism (i.e., SPADE(·)) is used as a backbone network to predict these motion fields, since the SPADE network is able to learn the reverse mapping from the 2D body mesh of subsequent frames and the 2D body mesh of the key reference frame. Therefore, the dense motion flow and occlusion map can be used as a guide for signal reconstruction.

[0088] Exemplarily, assume that the 2D human body mesh of the subsequent frame is and the 2D human body mesh of the key reference frame is Then the dense motion flow and occlusion map This can be determined by:

[0089]

[0090] Among them, P1(·) and P2(·) represent two different prediction outputs respectively.

[0091] Considering the powerful reasoning ability of deep generative networks, it is possible to promote realistic signal reconstruction in the generative compression paradigm. The embodiment of the present application adopts the GAN architecture because it has significant advantages in reasoning speed and application deployment compared with other generative models. Under this architecture, S710 may include: S7101, warping the key reference frame with the dense motion flow to obtain a warped result; S7102, performing Hadamard product on the occlusion map and the warped result to generate the human body video.

[0092] That is, the embodiment of the present application adopts a feature distortion strategy to decode the key reference frame Dense motion flow in feature-level domain Then, to improve the reconstruction fidelity, the occlusion map is used To indicate the feature map area with corresponding confidence. The whole process can be described by the following formula:

[0093]

[0094] Among them, f w and ⊙ represent the inverse warp operation and Hadamard product, respectively.

[0095] Optionally, in the embodiment of the present application, the result is generated can be further input into the discriminator module to approximate the original signal I l distribution.

[0096] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0097] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a decoding device for human body video. Figure 8 The structure block diagram of a human body video decoding device according to an embodiment of the present application is shown, and the device may include:

[0098] A first decoding module 82 is used for, after receiving a first coded bit stream and a second coded bit stream determined by a coding method for a human body video, decoding and reconstructing based on the first coded bit stream to obtain the second human body semantic set, and decoding and reconstructing based on the second coded bit stream to obtain the key reference frame;

[0099] A first input module 84 is used to input the key reference frame into the pre-trained 3D human body model OSX, and obtain a third human body semantic set in the first human body semantic set except the second human body semantic set by using a parameter extraction module of the 3D human body model OSX;

[0100] A first reconstruction module 86 is used to input the second human body semantic set and the third human body semantic set into the decoder of the 3D human body model OSX to reconstruct the 3D human body mesh of the subsequent frame in the human body video to be encoded, and input the key reference frame into the decoder to reconstruct the 3D human body mesh of the key reference frame;

[0101] An estimation module 88, configured to perform motion estimation on the 3D human body mesh of the subsequent frame and the 3D human body mesh of the key reference frame;

[0102] The first generating module 90 is used to generate a human body video based on the key reference frame and the motion estimation result.

[0103] pass Fig. 9 The device shown in the embodiment decodes and reconstructs the first coded bit stream to obtain the second human semantic set, and decodes and reconstructs the second coded bit stream to obtain the key reference frame; inputs the key reference frame into the pre-trained 3D human model OSX, and uses the parameter extraction module of the 3D human model OSX to obtain the third human semantic set in the first human semantic set except the second human semantic set; inputs the second human semantic set and the third human semantic set into the decoder of the 3D human model OSX to reconstruct the 3D human mesh of the subsequent frame in the human video to be encoded, and inputs the key reference frame into the decoder to reconstruct the 3D human mesh of the key reference frame; performs motion estimation on the 3D human mesh of the subsequent frame and the 3D human mesh of the key reference frame; generates a human video based on the key reference frame and the motion estimation result. That is to say, the embodiment of the present application has a semantically hierarchical representation in the coded bit stream, and these semantically meaningful representations are explicitly evolved into high-dimensional grid representations with the assistance of the 3D human model, thereby promoting dynamic perception and further improving the quality of the decoded video required for reconstruction.

[0104] In one possible implementation, the first reconstruction module 86 includes: an input unit, used to use the generation function OSX(·) in the decoder, and use the second human body semantic set and the third human body semantic set as inputs of the generation function to obtain a generation result; a reconstruction unit 88, used to determine the 3D human body mesh of the subsequent frame through the generation result.

[0105] The estimation module 88 includes: a first processing unit, used to convert the 3D human body mesh of the subsequent frame to obtain a 2D human body mesh of the subsequent frame, and to convert the 3D human body mesh of the key reference frame to obtain a 2D human body mesh of the key reference frame; a conversion unit, used to use a spatially adaptive normalized SPADE network to convert the 2D human body mesh of the subsequent frame and the 2D human body mesh of the key reference frame into a pixel-level dense motion field, wherein the dense motion field at least includes a dense motion flow and an occlusion map.

[0106] The first generating module 90 includes: a second processing unit, which is used to warp the key reference frame and the dense motion flow to obtain a warped result; and a generating unit, which is used to perform Hadamard product on the occlusion map and the warped result to generate the human body video.

[0107] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0108] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a method for communicating human body video. Fig. 9 The present invention shows a method for communicating a human body video according to an embodiment of the present invention, comprising:

[0109] S902, after receiving the first coded bit stream and the second coded bit stream determined by the above-mentioned human body video coding method, decoding and reconstructing are performed based on the first coded bit stream to obtain the second human body semantic set, and decoding and reconstructing are performed based on the second coded bit stream to obtain the key reference frame.

[0110] S904, input the key reference frame into the pre-trained 3D human body model OSX, and use the parameter extraction module of the 3D human body model OSX to obtain a third human body semantic set in the first human body semantic set except the second human body semantic set.

[0111] Optionally, the specific implementation details of the above steps S902 and S904 can be found in the relevant description above and will not be repeated here.

[0112] S906, modifying the values ​​of the human body semantics in the second human body semantic set and / or the values ​​of the human body semantics in the third human body semantic set according to the user's personalized communication needs, to obtain a reconfigured second human body semantic set and a reconfigured third semantic set.

[0113] Optionally, the above-mentioned user personalized communication needs may be the user's editing needs for the original human body signal, for example, changing the body posture of the original human body signal from A to B, the head posture from A to B, and the camera position from A to B, etc. Since the parameters to be modified in these communication needs have clear human body semantics, the relevant human body semantics can be independently controlled to achieve immersive interaction and personalized reconstruction.

[0114] For example, assuming that in a video call scenario, the user wants to move his entire body in the virtual environment, such as walking forward or backward, then the human body translation parameter in the second human body semantic set can be modified, that is, the 3D global translation δ trans ∈R 3 For another example, if a user wants to make a more exaggerated posture in a virtual environment, the human posture parameters in the second human semantic set can be modified, that is, δ pose For another example, if the user wants to make the virtual image taller, shorter, thinner, etc. in the virtual environment, the body shape parameters in the third human semantic set can be modified, that is, The value of .

[0115] S908: Generate a changed human body video using the human body video decoding method.

[0116] Through the above steps S902 to S906, after receiving the first coded bit stream and the second coded bit stream determined by the above human body video encoding method, decoding and reconstruction are performed based on the first coded bit stream to obtain the second human body semantic set, and decoding and reconstruction are performed based on the second coded bit stream to obtain the key reference frame; the key reference frame is input into the pre-trained 3D human body model OSX, and the parameter extraction module of the 3D human body model OSX is used to obtain the third human body semantic set in the first human body semantic set except the second human body semantic set; according to the user's personalized communication needs, the value of the human body semantics in the second human body semantic set and / or the value of the human body semantics in the third human body semantic set are modified to obtain a reconfigured second human body semantic set and a reconfigured third semantic set; the above human body video decoding method is used to generate a changed human body video. That is to say, the embodiment of the present application can flexibly edit the reconstruction of the 3D human body mesh by modifying different semantic level representations, for example, changing the angle and displacement values, so that the head posture and body posture of the 3D mesh are changed to achieve personalized representation, thereby realizing interactive human body video communication.

[0117] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a communication device for human body video. Fig.10The structure block diagram of a human body video communication device according to an embodiment of the present application is shown, and the device may include:

[0118] The second decoding module 102 is used to decode and reconstruct based on the first coded bit stream to obtain the second human body semantic set after receiving the first coded bit stream and the second coded bit stream determined by the above-mentioned human body video encoding method, and decode and reconstruct based on the second coded bit stream to obtain the key reference frame.

[0119] The second input module 104 is used to input the key reference frame into the pre-trained 3D human body model OSX, and use the parameter extraction module of the 3D human body model OSX to obtain a third human body semantic set in the first human body semantic set except the second human body semantic set.

[0120] The modification module 106 is used to modify the values ​​of the human body semantics in the second human body semantic set and / or the values ​​of the human body semantics in the third human body semantic set according to the user's personalized communication needs, so as to obtain a reconfigured second human body semantic set and a reconfigured third semantic set.

[0121] The second generating module 108 is used to generate a changed human body video using the above human body video decoding method.

[0122] pass Fig.10 The device shown, after receiving the first coded bit stream and the second coded bit stream determined by the above-mentioned human body video encoding method, decodes and reconstructs based on the first coded bit stream to obtain the second human body semantic set, and decodes and reconstructs based on the second coded bit stream to obtain the key reference frame; inputs the key reference frame into the pre-trained 3D human body model OSX, and uses the parameter extraction module of the 3D human body model OSX to obtain the third human body semantic set in the first human body semantic set except the second human body semantic set; according to the user's personalized communication needs, the value of the human body semantics in the second human body semantic set and / or the value of the human body semantics in the third human body semantic set are modified to obtain the reconfigured second human body semantic set and the reconfigured third semantic set; and the above-mentioned human body video decoding method is used to generate the changed human body video. That is to say, the embodiment of the present application can flexibly edit the reconstruction of the 3D human body mesh by modifying different semantic level representations, for example, changing the angle and displacement values, so that the head posture and body posture of the 3D mesh are changed to achieve personalized representation, thereby realizing interactive human body video communication.

[0123] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0124] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a model optimization method. Fig.11 The model optimization method of an embodiment of the present application is shown. The model is set at the decoding end. The model includes a first network and a second network. The first network and the second network are used to realize motion estimation and generate human body video in the decoding of the above human body video, including:

[0125] S1102: Use an end-to-end strategy to jointly train the first network and the second network.

[0126] Optionally, in an embodiment of the present application, the first network may be a SPADE network, and the second network may be a GAN network.

[0127] Through the above step S1102, the first network and the second network are jointly trained using an end-to-end strategy, so that the trained first network can improve the efficiency of grid-based dense motion estimation, thereby promoting dynamic perception, and the trained second network can provide human body frame generation efficiency, thereby improving the quality of decoded video.

[0128] In a possible implementation, S1102 may include:

[0129] S11021, determining a training loss, wherein the training loss is a loss determined by weighting the perceptual loss and the adversarial loss by a preset weight.

[0130] Optionally, in an embodiment of the present application, the preset weight includes a first weight and a second weight, and the first weight is greater than the second weight.

[0131] For example, assume the perceptual loss is The adversarial loss is The first weight is λ per , the second weight is λ adv , then the total training loss Determined by the following formula:

[0132]

[0133] Optionally, the above λ per and λ adv can be set to 100 and 1 respectively. In the embodiment of the present application, λ per and λ adv The specific value of can be randomly set according to different task requirements.

[0134] S11022, optimizing the model through the training loss.

[0135] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a model optimization device. Fig.12 The structure block diagram of the model optimization device of one embodiment of the present application is shown. The model is set at the decoding end. The model includes a first network and a second network. The first network and the second network are used to implement motion estimation and generate human body video in the above-mentioned human body video decoding method, including:

[0136] The training module 122 is used to jointly train the first network and the second network using an end-to-end strategy.

[0137] In one possible implementation, the training module 122 is further used to determine a training loss, wherein the training loss is a loss determined by weighting the perceptual loss and the adversarial loss by a preset weight; and the model is optimized by the training loss.

[0138] pass Fig.12 The device shown uses an end-to-end strategy to jointly train the first network and the second network, so that the trained first network can improve the efficiency of grid-based dense motion estimation, thereby promoting dynamic perception, and the trained second network can provide human body frame generation efficiency, thereby improving the quality of decoded video.

[0139] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0140] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides an interactive human video communication (IHVC) framework, such as Fig.13 As shown in the figure, the framework can achieve low-bandwidth and enhanced interactive human video communication. In particular, at the encoder side, the key reference frame representing the human texture is compressed using the latest VVC codec, which can provide texture reference for signal synthesis. Subsequently, the subsequent frames are input into the pre-trained 3D human model OSX, which is optimized to further generate human semantics (for example, 21-dimensional human posture parameters, 3-dimensional human translation parameters, 3-dimensional human rotation parameters, and 4-dimensional human position parameters) for representing compact and interactive human bodies. The specific dimensions of the human semantics mentioned here are just some examples, and their specific values ​​may depend on the bandwidth. Finally, these highly decoupled human semantics are further inter-frame predicted, quantized, and entropy encoded to generate a transmission bitstream.

[0141] After receiving the coded bit stream, the decoder proposed in the embodiment of the present application will perform mesh editing / reconstruction and signal synthesis to achieve personalized interaction. First, the key reference frame is decoded by the VVC codec, and further projected into the pre-trained 3D human model OSX, and the human semantics corresponding to the key reference frame is output. In addition, the compact semantics of the subsequent frames are obtained by entropy decoding and compensation. Subsequently, the corresponding semantics of the key reference frame and the subsequent frame are input into the preset 3D human model to reconstruct the human mesh, thereby generating a dense motion field at the pixel level (i.e., dense motion flow and occlusion map). Here, by modifying different semantic hierarchical representations, the reconstruction of the 3D human mesh can be flexibly edited, so that the head posture and body posture of the 3D mesh are changed to achieve personalized characterization. Finally, with the decoded key reference frame and the dense motion field at the pixel level, relying on the powerful reasoning ability of the deep generation model, the human video can be reconstructed with high quality.

[0142] The IHVC framework proposed in the embodiments of the present application has several desirable advantages, including compact representations for ultra-low bitrates and semantically meaningful representations for enhanced interactivity. First, the scheme exploits prior knowledge of human body signals, so that 40-dimensional semantic parameters are sufficient to characterize the nonlinear dynamics and complex motions of human body signals, thereby facilitating ultra-low bitrate human video communication. In addition, these highly decoupled representations have clear semantic meanings in terms of body posture, head posture, and camera position, and can be independently controlled to achieve immersive interaction and personalized reconstruction.

[0143] In summary, the embodiments of the present application can control human body movements in compressed bitstreams, which can be further applied to ultra-low bandwidth and enhanced interactive human communications in the metaverse and virtual uploaders of live e-commerce. Due to the highly decoupled parameters, we can redirect the virtually driven human body grid by simply changing the values ​​of angles and displacements. Finally, it is input into the motion estimation module and the frame generation module together with the keyframe grid to reconstruct the facial image with new posture and position, reflecting the adjusted parameters. The embodiment of the present application provides the first generative compression framework that can efficiently encode human body signals with ultra-compact and configurable human body semantics. In this way, the bitstream with these interactive human body semantics can be conveniently manipulated to reconstruct the head posture and body posture movement of the human body signal at the decoding end to achieve personalized communication. In addition, the embodiment of the present application also provides a grid-based motion estimation scheme and a GAN-based human video generation scheme, which can explicitly evolve these semantically meaningful representations into high-dimensional grid representations with the assistance of a 3D human body model, thereby promoting dynamic perception and improving the quality of decoded video.

[0144] Fig.14 FIG. 1 is a block diagram of an electronic device used to implement an embodiment of the present application. Fig.14As shown, the electronic device includes: a memory 1401 and a processor 1402. The memory 1401 stores a computer program that can be run on the processor 1402. When the processor 1402 executes the computer program, the method in the above embodiment is implemented. The number of the memory 1401 and the processor 1402 can be one or more.

[0145] The electronic device also includes:

[0146] The communication interface 1403 is used to communicate with external devices and perform data exchange transmission.

[0147] If the memory 1401, the processor 1402 and the communication interface 1403 are implemented independently, the memory 1401, the processor 1402 and the communication interface 1403 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.14 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0148] Optionally, in a specific implementation, if the memory 1401, the processor 1402 and the communication interface 1403 are integrated on a chip, the memory 1401, the processor 1402 and the communication interface 1403 can communicate with each other through an internal interface.

[0149] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which implements the method provided in the embodiment of the present application when the program is executed by a processor.

[0150] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the method provided by the embodiment of the present application.

[0151] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory, wherein the input interface, the output interface, the processor and the memory are connected via an internal connection path, and the processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.

[0152] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor supporting the Advanced RISC Machines (ARM) architecture.

[0153] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of exemplary but not limiting description, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM).

[0154] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.

[0155] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0156] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0157] Any process or method described in the flow chart or otherwise described herein can be understood as a module, fragment or portion of a code representing one or more executable instructions for implementing the steps of a specific logical function or process. And the scope of the preferred embodiment of the present application includes other implementations, in which the functions may not be performed in the order shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved.

[0158] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or used in combination with these instruction execution systems, devices or apparatuses.

[0159] It should be understood that the various parts of the present application can be implemented with hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented with software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium, and when the program is executed, it includes one of the steps of the method embodiment or a combination thereof.

[0160] In addition, each functional unit in each embodiment of the present application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a disk or an optical disk, etc.

[0161] The above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various changes or substitutions within the technical scope recorded in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A method for encoding a human body video, comprising: Input a frame in a human body video to be encoded into a pre-trained 3D human body model OSX to obtain a first human body semantic set, wherein the first human body semantic set is used to simplify the description of human body signals in the human body video to be encoded; Determine a second human body semantic set from the first human body semantic set, wherein the second human body semantic set is used to describe a signal related to motion in the human body signal, so that the decoding end interacts with the human body signal by configuring the value of the human body semantics in the second human body semantic set; Compressing the second human body semantic set to obtain a first coded bit stream, and compressing the key reference frame in the human body video to be coded to obtain a second coded bit stream; The first coded bit stream and the second coded bit stream are determined as target coding results of the human body video to be encoded.

2. The method according to claim 1, wherein: The human body video to be encoded is a human body video composed of human body frames in the same sequence, and determining a second human body semantic set from the first human body semantic set comprises: Determine the human body signals shared by the human body frames in the same sequence, and the human body semantics corresponding to the shared human body signals in the first human body semantic set, to obtain first target human body semantics; Determine the human body semantics that are not related to the motion in the first human body semantics set to obtain second target human body semantics; The first target human semantics and the second target human semantics are filtered out from the first human semantics set to obtain the second human semantics set.

3. The method according to claim 1, wherein: Compressing the second human semantic set to obtain a first coded bit stream includes: Perform inter-frame prediction between the human body semantics of the current frame and the human body semantics of the previously reconstructed frame to obtain the residual generated by the inter-frame prediction; Based on context arithmetic coding, the residual generated by the inter-frame prediction is compressed to obtain the first encoded bit stream.

4. The method according to any one of claims 1 to 3, wherein: The second human body semantic set includes at least the following human body semantics: human body posture parameters, human body translation parameters, human body rotation parameters and human body position parameters. The human body semantics are multi-dimensional, and the dimensions of different human body semantics are determined by the transmission bandwidth.

5. A method for decoding a human body video, comprising: After receiving the first coded bit stream and the second coded bit stream determined by any one of claims 1 to 4, decoding and reconstructing based on the first coded bit stream to obtain the second human semantic set, and decoding and reconstructing based on the second coded bit stream to obtain the key reference frame; Inputting the key reference frame into the pre-trained 3D human body model OSX, and using a parameter extraction module of the 3D human body model OSX to obtain a third human body semantic set in the first human body semantic set except the second human body semantic set; Inputting the second human body semantic set and the third human body semantic set into the decoder of the 3D human body model OSX to reconstruct the 3D human body mesh of the subsequent frames in the human body video to be encoded, and inputting the key reference frame into the decoder to reconstruct the 3D human body mesh of the key reference frame; Performing motion estimation on the 3D human body mesh of the subsequent frame and the 3D human body mesh of the key reference frame; A human body video is generated based on the key reference frame and the motion estimation result.

6. The method according to claim 5, wherein: Inputting the second human body semantic set and the third human body semantic set into the decoder of the 3D human body model OSX to reconstruct the 3D human body mesh of the subsequent frame, including: Using the generating function OSX(·) in the decoder, taking the second human body semantic set and the third human body semantic set as inputs of the generating function, and obtaining a generating result; The 3D human body mesh of the subsequent frame is determined according to the generated result.

7. The method according to claim 6, wherein: Performing motion estimation on the 3D human body mesh of the subsequent frame and the 3D human body mesh of the key reference frame, comprising: Converting the 3D human body mesh of the subsequent frame to obtain a 2D human body mesh of the subsequent frame, and converting the 3D human body mesh of the key reference frame to obtain a 2D human body mesh of the key reference frame; The 2D human body mesh of the subsequent frame and the 2D human body mesh of the key reference frame are converted into a pixel-level dense motion field using a spatially adaptive normalized SPADE network, wherein the dense motion field includes at least a dense motion flow and an occlusion map.

8. The method according to claim 7, wherein: Generating a human body video based on the key reference frame and the motion estimation result, including: Warping the key reference frame and the dense motion flow to obtain a warped result; A Hadamard product is performed on the occlusion map and the distortion result to generate the human body video.

9. A method for communicating human body video, comprising: After receiving the first coded bit stream and the second coded bit stream determined by any one of claims 1 to 4, decoding and reconstructing based on the first coded bit stream to obtain the second human semantic set, and decoding and reconstructing based on the second coded bit stream to obtain the key reference frame; Inputting the key reference frame into the pre-trained 3D human body model OSX, and using a parameter extraction module of the 3D human body model OSX to obtain a third human body semantic set in the first human body semantic set except the second human body semantic set; According to the personalized communication requirements of the user, modify the values ​​of the human body semantics in the second human body semantic set and / or the values ​​of the human body semantics in the third human body semantic set to obtain a reconfigured second human body semantic set and a reconfigured third semantic set; The method according to any one of claims 5 to 8 is used to generate a changed human body video.

10. A model optimization method, wherein the model is set at a decoding end, the model includes a first network and a second network, the first network and the second network are respectively used to implement motion estimation and generate human body video in the method according to any one of claims 5 to 8, comprising: The first network and the second network are jointly trained using an end-to-end strategy.

11. The method according to claim 10, wherein: Using an end-to-end strategy, jointly training the first network and the second network includes: Determine a training loss, wherein the training loss is a loss determined by weighting the perceptual loss and the adversarial loss by a preset weight; The model is optimized by the training loss.

12. An electronic device comprising a memory, a processor and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer program product, comprising computer instructions, wherein when the computer instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Video compression method based on compact representation of human body features

    CN119011840A

  • Video processing method and device, related equipment, storage medium and computer program product

    CN119182914A

  • Extracting information from images

    US20210082136A1

  • Method and apparatus for face video compression

    US20240251098A1