Apparatus and method of encoding and decoding for facial video
By employing compact rotation information representations bounded within -1 to 1, the inefficiencies in GFVC decoding are addressed, enhancing encoding and decoding efficiency and reducing computational complexity.
Patent Information
- Application Number
- PCT/CN2024/072160
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-17
AI Technical Summary
Existing facial video compression methods, such as Generative Face Video Compression (GFVC), face challenges in efficiently encoding and decoding rotation information due to the computational complexity and inefficiency of current rotation representations, particularly rotation matrices and Euler Angles, which are not compact and require complex trigonometric calculations, impacting decoding time and efficiency.
The use of alternative rotation information representations, including Normalized Euler Angles, Sinus-cosinus Euler Angles, Euler Vector, Normalized Euler Vector, Euler Axis and Angle, Euler Axis and Sinus-cosinus Angle, Quaternion, Rodriguez Vector, Modified Rodriguez Parameters, and Bounded Modified Rodriguez Parameters, which are bounded within the range of -1 to 1, reducing computational complexity and bitstream size.
These representations enhance the efficiency of facial video encoding and decoding by minimizing computational requirements and bitstream size, improving decoding speed and reducing delay, while maintaining reconstruction quality through Generative Facial Video models.
Smart Images

Figure CN2024072160_17072025_PF_FP_ABST
Abstract
Description
APPARATUS AND METHOD OF ENCODING AND DECODING FOR FACIAL VIDEOTECHNICAL FIELD
[0001] The present disclosure generally relates to encoding and decoding technology, and in particular to a decoding method for facial video, an encoding method for facial video, an apparatus, and a computer readable media.BACKGROUND
[0002] Generative Face Video Compression (GFVC) is an ultra-low rate face video compression method proposed in Joint Video Experts Team (JVET) . GFVC utilizes base pictures coded using traditional image or video compression methods, such as Versatile Video Coding (VVC) , together with additional parameters describing variety of facial representations to generate images of the input video sequence or set of pictures.
[0003] Facial representations typically used in the context of GFVC can be contained in a Supplemental Enhancement Information (SEI) message in form of matrices or vectors, which are 2D matrices with size equal to one for one of the dimensions. Each matrix represents a specific type of information, including also motion features such as transformations, e.g. translation, rotation or covariance, which can be applied to other types of facial representations.
[0004] Rotation information represented by matrix elements in GFV SEI message can be utilized by the GFV decoder to apply rotation to a whole set of specific facial representation (e.g. all head 2D or 3D keypoints or landmarks) or, alternatively, to only a selected subset of specific facial representation, such as keypoints representing single eye or mouth region. In the latter case, multiple matrices representing rotations, each being applied to a separate sub-group of facial representation are encoded in the GFV SEI message.
[0005] In the context of applying rotation operation in GFVC, two popular rotation representations are commonly used, namely the rotation matrix representation and the Euler Angles representation.SUMMARY
[0006] Accordingly, the present disclosure aims to provide a decoding method for facial video, an encoding method for facial video, an apparatus, and a computer readable media.
[0007] A technical scheme adopted by the present disclosure is to provide decoding method for facial video. The method includes: decoding a base picture; decoding a facial representation with regard to the base picture; decoding a rotation information of the facial representation, wherein a representation of the rotation information comprises at least one of: a Normalized Euler Angles representation, a Sinus-cosinus Euler Angles representation, an Euler Vector representation, a Normalized Euler Vector representation, an Euler Axis and Angle representation, an Euler Axis and Normalized Angle representation, an Euler Axis and Sinus-cosinus Angle representation, a Quaternion representation, a Rodriguez Vector representation, a Modified Rodriguez Parameters representation, and a Bounded Modified Rodriguez Parameters representation; and generating an output picture based on the base picture and the facial representation through a pre-defined Generative Facial Video (GFV) model.
[0008] Another technical scheme adopted by the present disclosure is to provide a decoding method for facial video. The method includes: decoding a base picture; decoding a facial representation with regard to the base picture, and decoding a rotation information of the facial representation, wherein each element of a representation of the rotation information is normalized within a range of -1 to 1 by an encoding apparatus; and generating an output picture based on the base picture and the facial representation by employing a pre-defined Generative Facial Video (GFV) model.
[0009] Another technical scheme adopted by the present disclosure is to provide a decoding method for facial video. The method includes: decoding a base picture; decoding a facial representation with regard to the base picture, and decoding a rotation information of the facial representation; and generating an output picture based on the base picture and the facial representation by employing a pre-defined Generative Facial Video (GFV) model; wherein the decoding the rotation information of the facial representation comprises: decoding trigonometric values related to elements of a representation of the rotation information, wherein the trigonometric values are received from an encoding apparatus.
[0010] Another technical scheme adopted by the present disclosure is to provide an encoding method for facial video. The method includes: compressing a base picture; generating a facial representation of a subsequent picture with regard to the base picture through a pre-defined Generative Facial Video (GFV) model; and encoding the base picture and the facial representation into a coded bitstream; wherein a representation of the rotation information of the facial representation comprises at least one of : a Normalized Euler Angles representation, a Sinus-cosinus Euler Angles representation, an Euler Vector representation, a Normalized Euler Vector representation, an Euler Axis and Angle representation, an Euler Axis and Normalized Angle representation, an Euler Axis and Sinus-cosinus Angle representation, a Quaternion representation, a Rodriguez Vector representation, a Modified Rodriguez Parameters representation, and a Bounded Modified Rodriguez Parameters representation.
[0011] Another technical scheme adopted by the present disclosure is to provide an apparatus. The apparatus includes a processor and a memory. The memory is configured to store executable instructions that, when executed by the processor, cause the processor to perform at least one of the foregoing methods.
[0012] Another technical scheme adopted by the present disclosure is to provide a computer readable media storing executable instructions that, when executed by a processor, cause the processor to perform at least one of the foregoing methods.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to clearly explain the technical solutions in the embodiments of the present disclosure, the drawings used in the description of the embodiments will be briefly described below. Obviously, the drawings in the following description are merely some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings may also be obtained based on these drawings without any creative work.
[0014] FIG. 1 shows an exemplary diagram illustrating how a Generative Face Video Compression system function.
[0015] FIG. 2 shows an exemplary diagram illustrating how a Generative Face Video codec function.
[0016] FIG. 3 is a flowchart of a decoding process of matrix elements according to an embodiment of the present disclosure.
[0017] FIG. 4 is a flowchart of a decoding process of matrix elements according to another embodiment of the present disclosure.
[0018] FIG. 5 is a flowchart of a decoding method for facial video according to an embodiment of the present disclosure.
[0019] FIG. 6 is a flowchart of a decoding method for facial video according to another embodiment of the present disclosure.
[0020] FIG. 7 is a flowchart of a decoding method for facial video according to yet another embodiment of the present disclosure.
[0021] FIG. 8 is a flowchart of an encoding method for facial video according to an embodiment of the present disclosure.
[0022] FIG. 9 is a schematic diagram of an apparatus according to an embodiment of the present disclosure.DETAILED DESCRIPTION
[0023] The disclosure will now be described in detail with reference to the accompanying drawings and examples. Apparently, the described embodiments are only a part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by a person of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] FIG. 1 shows an exemplary diagram illustrating how a Generative Face Video Compression system functions. As illustrated, the GFVC system 120 may include multiple modules such as a base picture video codec 121 and a generative face video codec 122. The base picture video codec 121 may compress input based picture 111 of the input face video (the input sequence 110) using traditional image or video compression methods. The generative face video codec 122 may model the input face video with a small number of transmitted symbols and decoded input base pictures, and output the decoded input face video, that is, the decoded sequence 130 (e.g., decoded pictures 131-133) . In one embodiment, the input base picture 111 may be utilized to provide essential texture reference, and the sequential pictures (e.g., the input pictures 112 and 113) may be utilized to provide facial representations with regard to the base picture 111.
[0025] In GFVC, most of the approaches use a traditional video codec, such as the Versatile Video Coding (VVC) or High-Efficiency Video Coding (HEVC) to compress the first picture (or a set of key pictures) of an input video sequence as the base picture to provide essential texture reference and efficiently reconstruct the remaining pictures by providing the decoded base pictures to the generative face video codec. It should be appreciated that the present disclosure is not limited to these video codecs, any other approach for compressing the base picture can also be adopted.
[0026] FIG. 2 shows an exemplary diagram illustrating how a Generative Face Video codec function. As shown in FIG. 2, generative face video (GFV) codec 220 may include multiple modules such as an analysis model 221 that models each image in the input face video (e.g., input sequence 210 including input picture 211, 212 and 213) with predefined facial representation (e.g. semantic parameters, motion features, temporal motion features, 2D keypoints, 2D landmarks, 3D keypoints, facial semantics and other formats) , a feature encoder 222 that encodes facial representation obtained from analysis model 221 into small number of transmitted symbols represented in a coded bitstream, a feature decoder 223 that decodes the transmitted symbols from the coded bitstream into the facial representation and a generator network 225 that uses the facial representation of each image in the input face video and at least one decoded base picture (e.g., the base picture 211) to output the decoded input face video (e.g., decoded sequence 230 including decoded picture 231, 232 and 233) . In addition, the generative face video codec 220 may also include a translator 224 that translates the facial representation obtained from feature decoder 223 into a form that is accepted by the generator network 225, which is used in the case of facial representations produced by analysis model 221 that is not paired with the generator network 225. Modules 221 and 222 form a GFV encoder, while modules 223, 224 and 225 form a GFV decoder.
[0027] In GFVC, the analysis model 221 is typically a deep artificial neural network or other image or video processing algorithm that extracts facial representation from input pictures and outputs it in form of features that can be further encoded by feature encoder 222. Similarly, the generator network 225 is typically a deep artificial neural network, e.g. Generative Adversarial Network [Goodfellow I, Pouget-Abadie J, Mirza M, et al. “Generative adversarial nets, ” Advances in neural information processing systems, vol. 27, 2014] , which uses the decoded facial representation obtained from feature decoder 223 (optionally translated into representation accepted by the generator network by the translator 224) and a base picture to reconstruct high quality face pictures to form the decoded sequence 230. The utilization of deep artificial neural networks results from the fact that they provide significantly higher capability to infer and synthesize the reconstructed pictures, resulting in much better reconstruction quality compared to methods without deep artificial neural networks. It should be appreciated that the present disclosure is not limited to the foregoing deep artificial neural network, any other deep artificial neural network or video processing algorithm for extracting facial representations can also be adopted.
[0028] Typical facial representations for GFVC algorithms are: 2D landmarks, 2D keypoints, region matrix, 3D keypoints, compact feature matrix, motion features or facial semantics. Because face images exhibit strong statistical regularities, they can be economically characterized with the above representations, leading to reduced coding bit-rate and improved coding efficiency. One challenge related to possibility to use different facial representations by GFVC algorithms is a requirement of a paired encoder and decoder, in particular, the analysis model and the generator network of the GFV codec. If the facial representations do not match between encoder and decoder, then the reconstruction of face video cannot be successfully realized. A common solution to that problem is utilization of the translator module in the decoder of the GFV codec which can be applied in case of such representations mismatch. Different implementations of the translator 224 have been discussed in JVET and will not be introduced herein.
[0029] Transmission of the coded bitstream (output of the GFV encoder) representing facial representation for the purpose of GFVC can be done by means of a dedicated GFV Supplemental Enhancement Information (SEI) message. One possible realization of the transmission method is to include GFV SEI message in each picture unit (PU) . The decoded picture of a PU may be a base picture (i.e., a decoded output picture that may be used as a texture reference by a generator network to generate a novel face picture) , a dummy picture (e.g., a picture unit of the minimum allowed resolution that contains only SEI messages, including GFV SEI messages) , or a primary coded picture that can be fused by GFV decoder to improve background texture and facial details. When the current picture is not a base picture, the GFV SEI message may be used to generate a novel face picture based on the previously decoded base picture, the facial parameters conveyed by the GFV SEI message, and, optionally, the primary coded picture. It should be appreciated that the present disclosure is not limited to the above-mentioned SEI message, other syntax and semantic elements may be also utilized for same purpose.
[0030] Facial representations used in the context of GFVC may be contained in the GFV SEI message in form of matrices or vectors, which are 2D matrices with size equal to one for one of the dimensions. Each matrix represents a specific type of information, including also motion features such as transformations, e.g. translation, rotation or covariance, which can be applied to other types of facial representations. The type of information represented by the matrix is encoded by a matrix type syntax element. Each matrix type signaled in the GFV SEI message is an index to a predefined table of all available matrix types, that can be used to obtain the matrix size, i.e. number of matrix rows and columns related to number of dimensions of each matrix or vector and, consequently, number of elements encoded to represent that matrix in the coded bitstream. Additionally, each matrix element in the GFV SEI message, which is in general a floating point value, is represented as a set of syntax elements: integer part (typically encoded using Exponential Golomb code with order k equal to 0) , decimal part (typically encoded with number of bits defined by precision signaled by a dedicated syntax element) and, in case of non-zero matrix elements, a sign (typically encoded as 1-bit flag) . A possible realization of a decoding process of matrix elements from GFV SEI is illustrated in Fig. 3. Specifically, a coded bitstream is received, and it is determined whether it contains a matrix type syntax element. If it does not contain a matrix type syntax, other syntax elements are decoded. If it contains a matrix type syntax, the matrix type is firstly decoded and the matrix size can be derived from the matrix type. For each element in the matrix, its integer part and decimal part may be sequentially decoded. If either the integer part or the decimal part is non-zero, the sign flag of the element is decoded. In this way, all elements in the matrix can be decoded.
[0031] Rotation information represented by matrix elements in GFV SEI message can be utilized by the GFV decoder to apply rotation to a whole set of specific facial representation (e.g. all head 2D or 3D keypoints or landmarks) or, alternatively, to only a selected subset of specific facial representation, such as keypoints representing single eye or mouth region. In the latter case, multiple matrices representing rotations, each being applied to a separate sub-group of facial representation are encoded in the GFV SEI message.
[0032] In the context of applying rotation operation in GFVC, different rotation representations result in varying computational complexity of decoding process and require specific set of arithmetic operations available in the GFVC decoder. For example, a conventional rotation representation in form of rotation matrix requires a simple matrix-vector multiplication, at the cost of non-compact representation of rotation parameters. On contrary, other representations, such as Euler Angles, may require additional conversion to rotation matrix form before applying the rotation operation. Additionally, some of the conversions to rotation matrix require utilization of trigonometric functions which, depending on hardware implementation of the decoder, may be especially expensive to compute. Trigonometric functions are usually calculated by summing a specific series of values, resulting in multiple processor cycles per single trigonometric operation -depending on hardware platform, might be even 200 cycles per operation. Examples of such methods are successive approximation methods, e.g., CORDIC, or Taylor expansion. Alternatively, trigonometric operations may be implemented by means of lookup table with pre-calculated values of each trigonometric function and represented angle. This allows to avoid the complex calculations at the expense of a trade-off between approximation accuracy and size of memory required to store the lookup table. In addition, there is a delay in accessing the memory counted in multiple processor cycles, as the lookup tables are typically loaded to RAM, not cache memory. Minimal or compact representation size of rotation information representation has a direct impact on improving coding efficiency of GFVC system by decreasing size of coded bitstream transmitted between GFV encoder and GFV decoder. Utilization of a specific representation of rotation information encoded in the coded bitstream can also impact complexity of applying rotation operation in the GFV decoder, influencing ability of the decoder to perform specific arithmetic operations, overall decoding time and decoding delay, which might be a critical requirement for multiple applications, including e.g. teleconferencing.
[0033] As for the rotation matrix representation, rotation information is represented as full rotation matrix of size 2x2 for 2D rotation or size 3x3 for 3D rotation. In case of 3D rotations, rotation matrix is formed by a triad of unit vectors u, v, w, called a basis (see Eq. 1) . Specifying the coordinates of vectors of this basis in its current, i.e. rotated position, in terms of the reference, i.e. non-rotated coordinate axes, describes the rotation completely. Each of the three unit vectors that form the rotated basis consist of 3 coordinates, yielding a total of 9 parameters per rotation matrix. In case of 2D rotations, there are two unit vectors forming the basis, each consisting of 2 coordinates, yielding 4 parameters per rotation matrix. Not all of the elements of the rotation matrix are independent and, according to Euler's rotation theorem, the rotation matrix has only three degrees of freedom for 3D rotation and two degrees of freedom for 2D rotation. Consequently, the rotation matrix representation of rotation is severely over-represented with 4 values required to describe 1 degree of freedom for 2D rotation and 9 values required to describe 3 degrees of freedom for 3D rotation. On the other hand, rotation matrix allows for straightforward application of rotation operation to a given 2D or 3D point by a simple matrix vector multiplication (see Eq. 2 for rotation of 3D point p) . It also provides a simple means of combining successive rotations by a product of rotation matrices. In addition, values of the rotation matrix are always within -1 and 1 range and there are no trigonometric functions required to perform the rotation operation given a rotation matrix.
[0034] As for the Euler Angles representation, Rotation information is represented as Euler Angles vector of size 3 for 3D rotation or scalar value for 2D rotation. The definition of 3D Euler Angles is not unique and one can find many different conventions in the literature, that depend on the axes about which the rotations are carried out, and their sequence (rotations are not commutative) . Euler Angles representation is a minimal form to represent both 2D and 3D rotations, since it requires 1 (γ) or 3 (α, β, γ) values respectively. However, in general, values of Euler Angles are unbounded, and, with restriction to -180° to 180° rotations, the values are bounded to a range from -π to π. In order to apply a rotation represented in form of Euler Angles to a 2D or 3D point, a conversion of the Euler Angles into rotation matrix is required. The exact conversion method depends on the convention of Euler Angles being used, and e.g. in case of 3D rotation, right-hand system and x-convention, where the rotations are about the x-, y-and z-axes with angles α, β and γ, the final rotation matrix is defined as in Eq. 3, where the individual matrices combined to form the rotation matrix are given in Eq. 4. As a result, the conversion depends heavily on the usage of sinus and cosinus trigonometric functions. A=AzAyAx (Eq. 3)
[0035] As illustrated above, existing representations of rotation information for GFVC consist of the rotation matrix representation and the Euler Angles representation. Each of the two representations presents various set of issues and limitations in terms of application in compression or decoding on a hardware platform with limited computational resources, e.g., a decoded apparatus.
[0036] In related technologies, rotation matrix is extensively used to represent rotation information for both 2D and 3D rotations, which comes down to specifying a rotation using a 2×2 or 3×3 matrix, respectively. However, according to Euler's rotation theorem, the rotation matrix has only two degrees of freedom for a 2D rotation and three degrees of a freedom for 3D rotation, which is significantly less compared to 4 values for 1 degree of freedom in case of 2D rotation and 9 values for 3 degrees of freedom in case of 3D rotation. Consequently, as not all of the elements of the rotation matrix are independent, this representation of rotation information is not well suitable for applications in which size of the representation is a critical element, including especially applications related to data compression, of which GFVC is a representative. This finding suggest that it is desirable to replace rotation matrix representation with a more compact representation of rotation information, ideally, without sacrificing beneficial properties of this representation.
[0037] The Euler Angles representation may provide much more compact representation of the rotation information in comparison to the rotation matrix representation. However, Euler Angles representation has a drawback of having values that are not bounded within -1 and 1 range, which is an issue for efficient coding of rotation representation values commonly used in GFV SEI messages. As the integer part of these values is encoded using Exponential Golomb code with order k equal to 0, limiting its dynamic range is especially important to achieve minimal size of the coded bitstream produced by the encoder. Consequently, a property of values being bounded within -1 and 1 range is desirable, as the integer part for such values may be represented with no more than a single bit, improving the compression rate. Besides, Euler Angles has a drawback of requiring an additional conversion to rotation matrix form before applying the rotation operation. In addition, this conversion depends on a specific set of arithmetic operations that need to be available in the GFVC decoder, namely trigonometric functions which, depending on hardware implementation of the decoder, may be especially expensive to compute. As a result, the compact representation of Euler Angles representation comes at the expense of additional computational complexity which, in case of some decoder implementation may turn out to be prohibitive for real-time applications.
[0038] The present disclosure aims to find a representation of rotation information which can be both compact and not require complex computational functions when applying the rotation operation.
[0039] FIG. 4 is a flowchart of a decoding process of matrix elements according to another embodiment of the present disclosure. Embodiments of the present disclosure modify the matrix decoding process of rotation information illustrated in FIG. 4 by introducing a new matrix type, defining matrix size for the new matrix type and modifying matrix element decoding process for elements of the new matrix type.
[0040] FIG. 5 is a flowchart of a decoding method for facial video according to an embodiment of the present disclosure. The method includes operations described in blocks 302 to 308.
[0041] In block 302, a base picture is decoded.
[0042] The base picture may be compressed in an encoding apparatus using traditional image or video compression methods, such as VVC, and be transmitted in a bitstream to a decoding apparatus which performs the decoding method. The base picture is decoded such that essential texture reference derived from the base picture may also be acquired.
[0043] In block 304, a facial representation with regard to the base picture is decoded.
[0044] The facial representation may include at least one type of 2D landmarks, 2D keypoints, 3D key points, facial semantics, motion features and the like. The facial representation may represent facial features with regard to the base picture, and provide motion parameters for generating the output video sequence. In one embodiment, the facial representation may be acquired from one or more subsequent face pictures of the base picture through an analysis model (e.g., the analysis model 221 shown in FIG. 2) . The facial representation may be transmitted, together or separately with the base picture, from the encoding apparatus.
[0045] In block 306, a rotation information of the facial representation is decoded. The rotation information of the facial representation includes at least one of a certain selection of representations.
[0046] As mentioned above, the rotation information represents rotation, and may be part of the facial representation. The rotation information may be utilized, together with the base picture, to generate the output sequence, i.e., the reconstruction of the input sequence. An embodiment of the present disclosure provides a certain selection of rotation representations, which are compact and do not require complex computational functions.
[0047] In some embodiments, the rotation information of the facial representation may be one of: Normalized Euler Angles representation, Sinus-cosinus Euler Angles representation, Euler Vector representation, Normalized Euler Vector representation, Euler Axis and Angle representation, Euler Axis and Normalized Angle representation, Euler Axis and Sinus-cosinus Angle representation, Quaternion representation, Rodriguez Vector representation, Modified Rodriguez Parameters representation and Bounded Modified Rodriguez Parameters representation. Detailed definition and explanation of each representation will be further described below.
[0048] In one embodiment, the normalized Euler angels representation is applied. The Normalized Euler Angles representation is a modification of the Euler Angles representation, where the original rotation angles, i.e. γ for 2D rotation or (α, β, γ) for 3D rotation, are normalized by a predefined scaling factor value θmax, resulting in γ' for 2D rotation or (α', β', γ') for 3D rotation (see Eq. 5) . The predefined scaling factor value θmax corresponds to a maximum absolute value of rotation angle that is possible to be represented. As a result, values of elements of the Normalized Euler Angles representation are bounded within -1 to 1 range. In the context of GFVC, a possible choice for the predefined scaling factor value θmax is π. {α′, β′, γ′} , α′=α / θmax, β′=β / θmax, γ′=γ / θmax (Eq. 5)
[0049] The matrix size for this representation is 3x1 or 1x3, with all variants being equivalent. Since values of elements of the Normalized Euler Angles representation are bounded within the range of -1 to 1, the coded bitstream (e.g., encoded using Exponential Golomb code with order k equal to 0) produced by an encoder corresponding to these elements can be reduced compared to traditional Euler Angles representation. E. g. value “0” can be represented using Exponential Golomb code with order k equal to 0 as “0” (1 bit) , value “1” as “10” (2 bits) , value “2” as “110” (3 bits) , value “3” as “1110” (4 bits) , etc. As a result, by restricting the values within -1 to 1 range one can achieve reduction of the maximum code length to maximum 2 bits for each integer part for every matrix element, and, assuming the integer part of the bounded representation is represented as a 1-bit flag, to 1 bit only. In some embodiments, the scaling factor value θmax may be determined in the encoding process and transmitted to the decoder. Alternatively or additionally, the scaling factor value θmax may be determined in advance and specified in both the encoder and the decoder, and therefore no transmission is needed.
[0050] In one embodiment, the Sinus-cosinus Euler Angles representation is applied. The Sinus-cosinus Euler Angles representation is a modification of the Euler Angles representation, where each of the original rotation angles, i.e. γ for 2D rotation or (α, β, γ) for 3D rotation, is represented by a pair of sinus and cosinus values of that angle, resulting in scγ for 2D rotation or (scα, scβ, scγ) for 3D rotation (see Eq. 6) . In order to allow this transformation of rotation angles to be a fully invertible, the maximum absolute value of rotation angle that is possible to be represented should be restricted to 2π, which is typically fulfilled in the context of GFVC application. {scα, scβ, scγ} , scα= (sin α, cos α) , scβ= (sin β, cos β) , scγ= (sin γ, cos γ) (Eq. 6)
[0051] The matrix size for this representation is 1x2 or 2x1 in case of 2D rotation or 3x2, 2x3, 6x1 or 1x6 in case of 3D rotation, with all variants being equivalent. Despite this representation is not a minimal one in terms of number of elements as it requires 2 elements for 2D rotation and 6 elements for 3D rotation respectively, it is still considerably smaller compared to rotation matrix, while being similarly complex in terms of computations when applying a rotation operation. It is advantageous in that, at least in part, values of elements of the sinus-sosinus Euler Angles representation are bounded within the range of -1 to 1, and there is no requirement for availability of trigonometric functions in the decoding process as the values of trigonometric functions may be computed in the encoding process and transmitted to the decoder.
[0052] In one embodiment, the Euler Vector representation is applied. The Euler Vector representation is a rotation representation much related to the Euler Axis and Angle representation, which results from multiplying the rotation axis vector e by the scalar rotation angle θ of the former representation. As a result, a 3D rotation r is determined from Euler Vector by normalizing the Euler Vector to have a unit norm (i.e. unit vector) which defines a 3D rotation axis e and the Euler Vector norm which defines a rotation angle θ around the rotation axis (see Eq. 7) . In case of 2D rotation, similarly to Euler Axis and Angle representation, Euler Vector reduces to scalar angle only, which is equivalent to Euler Angles representation. Consequently, Euler Vector is a minimal form to represent both 2D and 3D rotations, since it requires 1 value to represent 2D rotation or 3 values for 3D rotation. The values of Euler Vector are, with restriction to -180° to 180° rotations, bounded to a range from -πto π. In order to apply a 3D rotation represented in form of Euler Vector to a 3D point, the typical methodology is to convert the Euler Vector into 3D rotation matrix requires usage of square root and division functions (see Eq. 8) required to recover θ and e from Euler Vector and, in the next step, sinus and cosinus trigonometric functions, each applied once to the rotation angle.
[0053] The matrix size for this representation is 3x1 or 1x3, with all variants being equivalent. In comparison to the conventional Euler Angles representation, the Euler Vector representation requires less trigonometric computation, and therefore the computation complexity of the decoding process may be reduced.
[0054] In one embodiment, the Normalized Euler Vector representation is applied. The Normalized Euler Vector representation is a modification of the Euler Vector representation described above, where the original rotation angle θ is normalized by a predefined scaling factor value θmax, resulting in rotation angle θ′ and the final representation defined as in Eq. 9. The predefined scaling factor value θmax corresponds to a maximum absolute value of rotation angle that is possible to be represented. Consequently, the representation values are bounded within -1 to 1 range. In the context of GFVC, a possible choice for the predefined scaling factor value θmax is π.
[0055] The matrix size for this representation is 3x1 or 1x3, with all variants being equivalent. Since values of elements of the Normalized Euler Vector representation are bounded within the range of -1 to 1, the coded bitstream (e.g., encoded using Exponential Golomb code with order k equal to 0) produced by an encoder corresponding to these elements can be reduced. In some embodiments, the scaling factor value θmax may be determined in the encoding process and transmitted to the decoder. Alternatively or additionally, the scaling factor value θmax may be determined in advance and specified in both the encoder and the decoder, and therefore no transmission is needed.
[0056] In one embodiment, the Euler Axis and Angle representation is applied. The Euler Axis and Angle representation can be directly derived from Euler's rotation theorem which states that any rotation can be expressed as a single rotation about some axis. Consequently, in order to specify a 3D rotation r, one requires to provide a 3D rotation axis e and a scalar θ defining rotation angle around that axis (see Eq. 10) . In case of 2D rotation, Euler Axis and Angle representation reduces to scalar angle only, which is the same as in Euler Angles representation. As a result, Euler Axis and Angle representation is a compact representation, which requires 4 values to represent a 3D rotation, with rotation axis being a unit vector with values bounded to -1 to 1 range. The values of scalar rotation angle are, with restriction to -180° to 180° rotations, bounded to a range from -π to π. To apply a rotation represented in form of Euler Axis and Angle to a 3D point, a conversion into rotation matrix is required as presented in Eq. 11. In turn, the conversion depends on the usage of sinus and cosinus trigonometric functions, each applied once to the rotation angle, which requires limited trigonometric computation in comparison with the conventional Euler Angles representation.
[0057] The matrix size for this representation is 4x1, 1x4 or 2x2, with all variants being equivalent.
[0058] In one embodiment, the Euler Axis and Normalized Angle representation is applied. The Euler Axis and Normalized Angle representation is a modification of the Euler Axis and Angle representation described above, where the original rotation angle θ is normalized by a predefined scaling factor value θmax, resulting in rotation angle θ′ and the final representation defined as in Eq. 12. The predefined scaling factor value θmax corresponds to a maximum absolute value of rotation angle that is possible to be represented. Consequently, all the representation values are bounded within -1 to 1 range. In the context of GFVC, a possible choice for the predefined scaling factor value θmax is π.
[0059] The matrix size for this representation is 4x1, 1x4 or 2x2, with all variants being equivalent. Since values of elements of the Euler Axis and Normalized Angle representation are bounded within the range of -1 to 1, the coded bitstream (e.g., encoded using Exponential Golomb code with order k equal to 0) produced by an encoder corresponding to these elements can be reduced. In some embodiments, the scaling factor value θmax may be determined in the encoding process and transmitted to the decoder. Alternatively or additionally, the scaling factor value θmax may be determined in advance and specified in both the encoder and the decoder, and therefore no transmission is needed.
[0060] In one embodiment, the Euler Axis and Sinus-cosinus Angle representation is applied. The Euler Axis and Sinus-cosinus Angle representation is a modification of the Euler Axis and Angle representation described above, where the original rotation angle θ is replaced by a pair of sinus and cosinus values of that angle, resulting in the final representation defined as in Eq. 13. In order to allow this transformation of rotation angle to be a fully invertible, the maximum absolute value of rotation angle that is possible to be represented should be restricted to 2π, which is typically fulfilled in the context of GFVC application. Important features of this representation are the following: the representation values are bounded within -1 to 1 range; there is no requirement for availability of trigonometric functions in the GFV decoder as the values of trigonometric functions are explicitly coded.
[0061] The matrix size for this representation is 5x1 or 1x5, with both variants being equivalent. Despite this representation is not a minimal one in terms of number of elements as it requires 5 elements, it is still considerably smaller compared to rotation matrix representation, while being similarly complex in terms of computations when applying a rotation operation. It is advantageous in that, at least in part, values of elements of the Euler Axis and Sinus-cosinus Angle representation are bounded within the range of -1 to 1, and there is no requirement for availability of trigonometric functions in the decoding process as the values of trigonometric functions may be computed in the encoding process and transmitted to the decoder.
[0062] In one embodiment, the Quaternion representation is applied. Quaternion representation specifies a 3D rotation as a 4-vector that consists of a 3D rotation axis e multiplied by sinus of rotation angle θ divided by 2 and a scalar value of cosinus of rotation angle θ divided by 2, as presented in Eq. 14. Quaternion is a compact representation, which requires 4 values to represent a 3D rotation with values always bounded within -1 and 1 range. Combining subsequent rotations is very easy for Quaternion representation and comes down to quaternion multiplication (see Eq. 15) . In order to apply a 3D rotation represented in form of quaternion, one needs to convert quaternion into rotation matrix, as shown in Eq. 16. The conversion does not depend on any function that is computationally expensive. Similarly to rotation matrices, quaternions must sometimes be re-normalized due to rounding errors to make sure that they still correspond to a valid rotations. However, the computational cost of such operation for a quaternion is much less than for normalizing a 3D rotation matrix.
[0063] The matrix size for this representation is 4x1, 1x4 or 2x2, with all variants being equivalent.
[0064] In one embodiment, the Rodriguez Vector representation is applied. Rodriguez Vector represents 3D rotation as a rotation axis e, which is a 3D unit vector, multiplied by a scalar value of tangent of rotation angle θ divided by 2 (see Eq. 17) . In turn, Rodriguez Vector representation is a minimal form to represent 3D rotation, since it requires 3 values to represent a 3D rotation. This representation, however, has a discontinuity for every rotation angle which is an integer multiply of π, making the values range or the representation completely unbounded. This holds even with a restriction to -180° to 180° rotations. On the other hand, Rodriguez Vectors can be easily combined to represent successive rotations, as presented in Eq. 18 for the case of rotation g1 followed by g2. Conversion to 3D rotation matrix, which is required to apply the rotation to a 3D point, is also relatively cheap, and depends on a division function applied once to compute a matrix multiplier (see Eq. 19) .
[0065] The matrix size for this representation is 3x1 or 1x3, with all variants being equivalent.
[0066] In one embodiment, the Modified Rodriguez Parameters representation is applied. The Modified Rodriguez Parameters (MRP) representation is closely related to Rodriguez Vector, however, presents some unique features which make it easier to apply in many scenarios. The key difference between the two representations is the multiplier value which, in case of MRP is a tangent of quarter of the rotation angle θ (see Eq. 20) . This subtle difference allows to mitigate the problem of tangent function discontinuity if the rotation angle is restricted to -360° to 360° range, and for rotations restricted to -180° to 180° range MRP values are also bounded within -1 and 1 range. MRP is also a minimal form to represent 3D rotation and requires 3 values to represent a 3D rotation. Compared to Rodriguez Vector representation, combining subsequent rotations is more complex for MRP, however, still remains moderately easy, as presented in Eq. 21. Lastly, in order to apply rotation to a 3D point, MRP requires conversion to a 3D rotation matrix following Eq. 22, which is also relatively cheap in terms of computational complexity and depends on a division function applied once to compute a matrix multiplier.
[0067] The matrix size for this representation is 3x1 or 1x3, with all variants being equivalent.
[0068] In one embodiment, the Bounded Modified Rodriguez Parameters representation is applied. The Bounded Modified Rodriguez Parameters representation is a modification of the Modified Rodriguez Parameters representation described above, where the original rotation angle θ is bounded within the [-π, π] range. As a result, the maximum absolute value of the rotation angle that is possible to be represented is restricted to π, which is typically fulfilled in the context of GFVC application. An important feature of this representation is that the representation values (px, py, pz) (see Eq. 20) are bounded within -1 to 1 range.
[0069] The matrix size for this representation is 3x1 or 1x3, with all variants being equivalent. Since values of elements of the Bounded Modified Rodriguez Parameters representation are bounded within the range of -1 to 1, the coded bitstream (e.g., encoded using Exponential Golomb code with order k equal to 0) produced by an encoder corresponding to these elements can be reduced.
[0070] In block 308, an output picture is generated based on the base picture and the facial representation through a pre-defined Generative Facial Video model.
[0071] The base picture and the facial representation acquired in the above-described operations may be input to a pre-defined Generative Facial Video model for reconstruction of the input sequence. The pre-defined Generative Facial Video model is typically a deep artificial neural network, e.g., Generative Adversarial Network. The pre-defined Generative Facial Video model may be trained to generated picture with a reference texture, which is provided by the base picture, and a facial representation with regard to the base picture. The utilization of deep artificial neural networks results from the fact that they provide significantly higher capability to infer and synthesize the reconstructed pictures, resulting in much better reconstruction quality compared to methods without deep artificial neural networks.
[0072] In some embodiments, elements of the rotation information are normalized. Accordingly, before the generation operation of the output picture, these elements should be denormalized. For example, if the rotation information is presented as the Normalized Euler Angles representation, the elements α’, β’, and γ’ of the Normalized Euler Angles representation may be denormalized with the scaling factor. If the rotation information is presented as the Normalized Euler Vector representation, the elements rx’, ry’, and rz’ of the Normalized Euler Vector representation may be denormalized with the scaling factor. If the rotation information is presented as the Euler Axis and Normalized Angle representation, the elements θ’ of the Normalized Euler Vector representation may be denormalized with the scaling factor.
[0073] In some embodiments, the rotation information contains trigonometric values, which may be calculated and transmitted by an encoder of the rotation information. For example, If the rotation information is represented as the Sinus-cosinus Euler Angles representation, values of the elements sin α, cos α, sin β, cos β, sin γ, and cos γ of the Sinus-cosinus Euler Angles representation are calculated and transmitted by an encoding apparatus. If the rotation information is represented as the Euler Axis and Sinus-cosinus Angle representation, values of the elements sin θ and cos θ of the Euler Axis and Sinus-cosinus Angle representation are calculated and transmitted by an encoding apparatus.
[0074] The decoding method of the above embodiment includes: decoding a base picture; decoding a facial representation with regard to the base picture; decoding a rotation information of the facial representation; and generating an output picture based on the base picture and the facial representation through a pre-defined Generative Facial Video model. The rotation information may be one of Normalized Euler Angles representation, Sinus-cosinus Euler Angles representation, Euler Vector representation, Normalized Euler Vector representation, Euler Axis and Angle representation, Euler Axis and Normalized Angle representation, Euler Axis and Sinus-cosinus Angle representation, Quaternion representation, Rodriguez Vector representation, Modified Rodriguez Parameters representation, and Bounded Modified Rodriguez Parameters representation. The rotation representation provided in the method has minimal or compact size and low requirement of computational capacity, which may improve the transmission efficiency of the GFVC encoding and decoding process. Additional or alternatively, the rotation representation provided in the method eliminate or reduce the need of complex computation, which may reduce overall decoding time and decoding delay.
[0075] In some embodiments, the rotation information is stored in a Supplemental Enhancement Information (SEI) received from an encoding apparatus. Additionally or alternatively, the rotation information may be stored in other types of message.
[0076] In some embodiments, as shown in FIG. 4, the operation of decoding the rotation information may include: obtaining a matrix type of the rotation information based on the SEI or other message carrying the rotation information; obtaining a matrix size of the rotation information based on the matrix type to acquire a number of dimensions of the rotation information; and decoding each element of the rotation information one by one based on the number of dimensions of the rotation information.
[0077] In some embodiments, the operation of decoding an element of the rotation information may include: decoding an integer part; decoding a decimal part and decoding a sign flag when either the decoded integer part or the decoded decimal part is non-zero. In one embodiment, the value of the element is bounded within a range of -1 to 1, and correspondingly the integer part is represented as 1-bit flag instead of Exponential Golomb code with order k equal to 0, which may make the representation more compact.
[0078] FIG. 6 is a flowchart of a decoding method for facial video according to another embodiment of the present disclosure. The method includes operations described in blocks 402 to 408. For the sake of simplicity and clarity, operations similar to those of the method shown in FIG. 5 will not be elaborated in detail.
[0079] In block 402, a base picture is decoded.
[0080] In block 404, a facial representation with regard to the base picture is decoded.
[0081] In block 406, a rotation information of the facial representation is decoded. Each element of the representation of the rotation information is normalized within the range of -1 to 1.
[0082] Specifically, each element of the representation of the rotation information may be normalized by an encoder, and the normalized values, instead of the original values, may be transmitted for the decoding process.
[0083] In some embodiments, the method may further include: determining a scaling factor and denormalizing each element of the representation of the rotation information with the scaling factor. The scaling factor is used by the encoder in normalization of each element of the representation of the rotation information. The scaling factor may be transmitted from the encoder. Additionally or alternatively, the scaling factor may be pre-defined and stored both in the encoding and decoding apparatus such that no transmission is required.
[0084] In the present embodiment, the representation of the rotation information may include one of Normalized Euler Angles representation, Normalized Euler Vector representation, and Euler Axis and Normalized Angle representation. The detailed description of these representations are given in the foregoing embodiment.
[0085] In the context of GFVC, the scaling factor may be π, as rotation angle of most rotations in GFVC is limited in the range of -π to π. It should be understand that other value can also be utilized based on actual requirement of the encoding and decoding processes.
[0086] In block 408, an output picture is generated based on the base picture and the facial representation through a pre-defined Generative Facial Video model.
[0087] In the present embodiment, the method includes: decoding a base picture; decoding a facial representation with regard to the base picture; decoding a rotation information of the facial representation; and generating an output picture based on the base picture and the facial representation by employing a pre-defined Generative Facial Video model. In the present embodiment, each element of the rotation information representation is normalized within a range of -1 to 1 by an encoding apparatus. Therefore, when the elements of the rotation information representation are encoded, decoded (e.g., using Exponential Golomb code with order k equal to 0) and / or transmitted, bit resources may be saved.
[0088] In some embodiments, the rotation information is stored in a Supplemental Enhancement Information (SEI) received from an encoding apparatus. Additionally or alternatively, the rotation information may be stored in other types of message.
[0089] In some embodiments, as shown in FIG. 4, the operation of decoding the rotation information may include: obtaining a matrix type of the rotation information based on the SEI or other message carrying the rotation information; obtaining a matrix size of the rotation information based on the matrix type to acquire a number of dimensions of the rotation information; and decoding each element of the rotation information one by one based on the number of dimensions of the rotation information.
[0090] In some embodiments, the operation of decoding an element of the rotation information may include: decoding an integer part; decoding a decimal part and decoding a sign flag when either the decoded integer part or the decoded decimal part is non-zero. In one embodiment, the integer part is represented as 1-bit flag instead of Exponential Golomb code with order k equal to 0, which may make the representation more compact.
[0091] FIG. 7 is a flowchart of a decoding method for facial video according to yet another embodiment of the present disclosure. The method includes operations described in blocks 502 to 508. For the sake of simplicity and clarity, operations similar to those of the method shown in FIG. 5 will not be elaborated in detail.
[0092] In block 502, a base picture is decoded.
[0093] In block 504, a facial representation with regard to the base picture is decoded.
[0094] In block 506, a rotation information of the facial representation is decoded. Trigonometric values related to elements of the representation of the rotation information is also decoded.
[0095] In the present embodiment, trigonometric values related to elements of the rotation information may be calculated and transmitted by an encoder. Thus, computational burden of the decoding process / apparatus may be reduced.
[0096] In one embodiment, the facial representation includes an Euler Angles representation, elements of which include α, β, and γ. The trigonometric values related to the elements of the Euler Angles representation include sin α, cos α, sin β, cos β, sin γ, and cos γ. As described above, these trigonometric values, in the decoding process, may be decoded rather than calculated. The original Euler Angles representation provided with trigonometric values of each elements may also be considered as the Sinus-Cosinus Euler Angles representation.
[0097] In one embodiment, the facial representation includes an Euler Axis and Angle representation, elements of which include ex, ey, ez, and θ. The trigonometric values related to the elements of the Euler Axis and Angle representation include sin θ and cos θ. As described above, these trigonometric values, in the decoding process, may be decoded rather than calculated. The original Euler Axis and Angle representation provided with trigonometric values of each elements may also be considered as the Euler Axis and Sinus-cosinus Angle representation.
[0098] In block 508, an output picture is generated based on the base picture and the facial representation through a pre-defined Generative Facial Video model.
[0099] The decoding method of the above embodiment includes: decoding a base picture; decoding a facial representation with regard to the base picture; decoding a rotation information of the facial representation; and generating an output picture based on the base picture and the facial representation through a pre-defined Generative Facial Video model. The trigonometric values related to elements of the rotation information representation may be encoded and transmitted by an encoder, and thus may be decoded directly in the decoding process. Therefore, the need of complex computation in the decoding process may be eliminated or reduced, which may reduce overall decoding time and decoding delay.
[0100] In some embodiments, the rotation information is stored in a Supplemental Enhancement Information (SEI) received from an encoding apparatus. Additionally or alternatively, the rotation information may be stored in other types of message.
[0101] In some embodiments, as shown in FIG. 4, the operation of decoding the rotation information may include: obtaining a matrix type of the rotation information based on the SEI or other message carrying the rotation information; obtaining a matrix size of the rotation information based on the matrix type to acquire a number of dimensions of the rotation information; and decoding each element of the rotation information one by one based on the number of dimensions of the rotation information.
[0102] In some embodiments, the operation of decoding an element of the rotation information may include: decoding an integer part; decoding a decimal part and decoding a sign flag when either the decoded integer part or the decoded decimal part is non-zero. In one embodiment, the value of the element is bounded within a range of -1 to 1, and correspondingly the integer part is represented as 1-bit flag instead of Exponential Golomb code with order k equal to 0, which may make the representation more compact.
[0103] FIG. 8 is a flowchart of an encoding method for facial video according to an embodiment of the present disclosure. The method may include operations described in blocks 602 to 606.
[0104] In block 602, a base picture is compressed.
[0105] As described in embodiments related to FIGS. 1 and 2, a based picture may be compressed in a GFVC encoder to provide essential textures.
[0106] In block 604, a facial representation of a subsequent picture is generated.
[0107] As described in embodiments related to FIGS. 1 and 2, a facial representation of a subsequent picture may be generated through a deep learning neural network. The subsequent picture may represent a target picture that should be generated in the decoding process. The facial representation may be acquired by analyzing the base picture and a sequential input picture. In other embodiment, the facial representation may be provided directly without analysis of the sequential input picture by, for example, directly providing a transformation relation with regard to the base picture.
[0108] In block 606, the base picture and the facial representation are encoded into a coded bitstream. The representation of the rotation information includes at least one of Normalized Euler Angles representation, Sinus-cosinus Euler Angles representation, Euler Vector representation, Normalized Euler Vector representation, Euler Axis and Angle representation, Euler Axis and Normalized Angle representation, Euler Axis and Sinus-cosinus Angle representation, Quaternion representation, Rodriguez Vector representation, Modified Rodriguez Parameters representation; and Bounded Modified Rodriguez Parameters representation.
[0109] The rotation information representation adopted by the present embodiment may be one of Normalized Euler Angles representation, Sinus-cosinus Euler Angles representation, Euler Vector representation, Normalized Euler Vector representation, Euler Axis and Angle representation, Euler Axis and Normalized Angle representation, Euler Axis and Sinus-cosinus Angle representation, Quaternion representation, Rodriguez Vector representation, Modified Rodriguez Parameters representation; and Bounded Modified Rodriguez Parameters representation. The rotation representation provided in the method has minimal or compact size and low requirement of computational resources, which may improve the transmission efficiency of the GFVC encoding and decoding process. Additional or alternatively, the rotation representations provided in the method eliminate or reduce the need of complex computation, which may reduce overall decoding time and decoding delay.
[0110] In some embodiments, the rotation information representation may include one of the Normalized Euler Angles representation, the Normalized Euler Vector representation, and the Euler Axis and Normalized Angle representation. The encoding method may further include: normalizing elements of the rotation information representation of the facial representation with a scaling factor before encoding the facial representations. In one embodiment, the encoding method may further include: encoding the scaling factor in the bitstream. In an example, the scaling factor may be π which may satisfy most case in the context of GFVC.
[0111] In some embodiments, the rotation information representation may include one of the Sinus-cosinus Euler Angles representation and the Euler Axis and Sinus-cosinus Angle representation. The encoding method may further include: determining values of elements of the representation of the rotation information which require trigonometric calculation before encoding the facial representation.
[0112] In some embodiments, the rotation information representation of the facial representation is encoded in a Supplemental Enhancement Information. The SEI includes at least a matrix type of the rotation information. In the SEI, elements of the rotation information representation are represented by an integer part, a decimal part and a sign flag. In one embodiment, if value of an element of the rotation information representation is bounded within a range of -1 to 1, the integer part may be encoded as a 1-bit flag.
[0113] FIG. 9 conceptually illustrates an apparatus 700 with which some embodiments of the invention are implemented. The apparatus 700 may be a computer (e.g., a desktop computer, personal computer, tablet computer, etc. ) , phone, PDA, or any other sort of electronic device. Such an apparatus includes various types of computer readable media and interfaces for various other types of computer readable media. The apparatus 700 includes a processor 702 and a memory 704. The memory 704 is configured to store executable instructions that, when executed by the processor, cause the processor to perform any one of the foregoing decoding or encoding methods.
[0114] The processor 702 may be a single processor or a multi-core processor in different embodiments. In some embodiments, the processor may include a GPU, NPU or DSP which may offload various computations or complement the image processing provided by the processor 702.
[0115] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a machine-readable or computer-readable medium (alternatively referred to as computer-readable storage media, machine-readable media, or machine-readable storage media) . Some examples of such computer-readable media include RAM, ROM, read-only compact discs (CD-ROM) , recordable compact discs (CD-R) , rewritable compact discs (CD-RW) , read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM) , a variety of recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc. ) , flash memory (e.g., SD cards, mini-SD cards, micro-SD cards, etc. ) , magnetic and / or solid state hard drives, read-only and recordable discs, ultra density optical discs, any other optical or magnetic media, and floppy disks. The computer-readable media may store a computer program that is executable by at least one processing unit and includes sets of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter.
[0116] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, some embodiments are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) . In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In addition, some embodiments execute software stored in programmable logic devices (PLDs) , ROM, or RAM devices.
[0117] As used in this specification and any claims of this application, the terms “computer” , “server” , “processor” , and “memory” all refer to electronic or other technological devices. These terms exclude people or groups of people. As used in this specification and any claims of this application, the terms “computer readable medium, ” “computer readable media, ” and “machine readable medium” are entirely restricted to tangible, physical objects that store information in a form that is readable by a computer. These terms exclude any wireless signals, wired download signals, and any other ephemeral signals.
[0118] The present disclosure further provides a computer readable media which is configured to store executable instructions. When the instructions are executed by a processor, the processor may perform any one of the foregoing methods and processes. Many of the above-described features and applications are implemented as software processes that are specified as a set of instructions recorded on a computer readable storage medium (also referred to as computer readable medium) . When these instructions are executed by one or more computational or processing unit (s) (e.g., one or more processors, cores of processors, or other processing units) , they cause the processing unit (s) to perform the actions indicated in the instructions. Examples of computer readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard drives, erasable programmable read only memories (EPROMs) , electrically erasable programmable read-only memories (EEPROMs) , etc. The computer readable media does not include carrier waves and electronic signals passing wirelessly or over wired connections.
[0119] In this specification, the term “software” is meant to include firmware residing in read-only memory or applications stored in magnetic storage which can be read into memory for processing by a processor. Also, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while remaining distinct software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of separate programs that together implement a software invention described here is within the scope of the invention. In some embodiments, the software programs, when installed to operate on one or more electronic systems, define one or more specific machine implementations that execute and perform the operations of the software programs.
[0120] While the disclosure has been described with reference to numerous specific details, one of ordinary skill in the art will recognize that the invention can be embodied in other specific forms without departing from the spirit of the invention. In addition, a number of the figures conceptually illustrate processes and methods. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in one continuous series of operations, and different specific operations may be performed in different embodiments. Furthermore, the process could be implemented using several sub-processes, or as part of a larger macro process.
[0121] The foregoing is merely embodiments of the present disclosure, and is not intended to limit the scope of the disclosure. Any transformation of equivalent structure or equivalent process which uses the specification and the accompanying drawings of the present disclosure, or directly or indirectly application in other related technical fields, are likewise included within the scope of the protection of the present disclosure.
Claims
1.A decoding method for facial video, comprising:decoding a base picture;decoding a facial representation with regard to the base picture;decoding a rotation information of the facial representation, wherein a representation of the rotation information comprises at least one of: a Normalized Euler Angles representation, a Sinus-cosinus Euler Angles representation, an Euler Vector representation, a Normalized Euler Vector representation, an Euler Axis and Angle representation, an Euler Axis and Normalized Angle representation, an Euler Axis and Sinus-cosinus Angle representation, a Quaternion representation, a Rodriguez Vector representation, a Modified Rodriguez Parameters representation, and a Bounded Modified Rodriguez Parameters representation; andgenerating an output picture based on the base picture and the facial representation through a pre-defined Generative Facial Video (GFV) model.2.The method of claim 1,wherein the representation of the rotation information comprises the Normalized Euler Angles representation, wherein elements of the Normalized Euler Angles representation comprise α’, β’, and γ’;wherein the elements α’, β’, and γ’ are normalized values of elements α, β, and γ of an Euler Angles representation.3.The method of claim 2, before the generating the output picture, further comprising:acquiring a scaling factor, wherein the scaling factor is used in normalization of the elements α, β, and γ of the Euler Angles representation; anddenormalizing the elements α’, β’, and γ’ of the Normalized Euler Angles representation with the scaling factor.4.The method of claim 1,wherein the representation of the rotation information comprises the Sinus-cosinus Euler Angles representation, wherein elements of the Sinus-cosinus Euler Angles representation comprise sin α, cos α, sin β, cos β, sin γ, and cos γ;wherein α, β, and γ are elements of an Euler Angles representation;wherein values of the elements sin α, cos α, sin β, cos β, sin γ, and cos γ of the Sinus-cosinus Euler Angles representation are calculated and transmitted by an encoding apparatus.5.The method of claim 1,wherein the representation of the rotation information comprises the Normalized Euler Vector representation, wherein elements of the Normalized Euler Vector representation comprise rx’, ry’, and rz’;wherein rx’, ry’, and rz’ are normalized values of elements rx, ry, and rz of an Euler Vector representation.6.The method of claim 5, before the generating the output picture, further comprising:acquiring a scaling factor, wherein the scaling factor is used in normalization of the elements rx, ry, and rz of the Euler Vector representation; anddenormalizing the elements rx’, ry’, and rz’ of the Normalized Euler Vector representation with the scaling factor.7.The method of claim 1,wherein the representation of the rotation information comprises the Euler Axis and Normalized Angle representation, wherein elements of the Euler Axis and Normalized Angle representation comprise ex, ey, ez, and θ’;wherein the elements ex, ey, ez are equivalent elements of an Euler Axis and Angle representation which indicates a unit vector;wherein the element θ’ is a normalized value of an element θ of the Euler Axis and Angle representation.8.The method of claim 7, before the generating the output picture, further comprising:acquiring a scaling factor, wherein the scaling factor is used in normalization of the element θ of the Euler Axis and Angle representation; anddenormalizing the element θ’ of the Euler Axis and Normalized Angle representation with the scaling factor.9.The method of claim 1,wherein the representation of the rotation information comprises the Euler Axis and Sinus-cosinus Angle representation, wherein elements of the Euler Axis and Sinus-cosinus Angle representation comprise ex, ey, ez, sin θ, and cos θ;wherein the elements ex, ey, ez are equivalent elements of an Euler Axis and Angle representation which indicates a unit vector;wherein θ is an equivalent parameter of an element of the Euler Axis and Angle representation which indicates a rotation angle;wherein values of the elements sin θ and cos θ of the Euler Axis and Sinus-cosinus Angle representation are calculated and transmitted by an encoding apparatus.10.The method of claim 1,wherein the representation of the rotation information comprises the Modified Rodriguez Parameters representation, wherein elements of the Modified Rodriguez Parameters comprise px, py, pz, and θ;wherein relation of the elements of the Modified Rodriguez Parameters follows:wheree is a unit vector indicating a rotation axis.11.The method of claim 1,wherein the representation of the rotation information comprises the Bounded Modified Rodriguez Parameters representation, wherein elements of the Modified Rodriguez Parameters comprise px, py, pz, and θ;wherein a value of θ is bounded within a range of –π to π.12.The method of claim 1,wherein the decoding the facial representation comprises decoding a matrix type of the facial representation;wherein the decoding the rotation information of the facial representation comprises: when the matrix type of the facial representation indicates the rotation information is included in the facial representation, decoding the rotation information.13.The method of claim 1,wherein the rotation information is stored in a Supplemental Enhancement Information (SEI) received from an encoding apparatus.14.The method of claim 1, wherein the decoding the rotation information comprises:obtaining a matrix type of the rotation information;obtaining a matrix size of the rotation information based on the matrix type to acquire a number of dimensions of the rotation information; anddecoding each element of the rotation information one by one based on the number of dimensions of the rotation information.15.The method of claim 14, wherein the decoding each element of the rotation information one by one comprises:for each element of the rotation information:decoding an integer part, wherein when value of the element is bounded within a range of -1 to 1, the integer part is represented as a 1-bit flag;decoding a decimal part; anddecoding a sign flag when either the decoded integer part or the decoded decimal part is non-zero.16.A decoding method for facial video, comprising:decoding a base picture;decoding a facial representation with regard to the base picture, and decoding a rotation information of the facial representation, wherein each element of a representation of the rotation information is normalized within a range of -1 to 1 by an encoding apparatus; andgenerating an output picture based on the base picture and the facial representation by employing a pre-defined Generative Facial Video (GFV) model.17.The method of claim 16, before the generating the output picture, further comprising:determining a scaling factor, wherein the scaling factor is used in normalization of the each element of the representation of the rotation information; anddenormalizing the each element of the representation of the rotation information with the scaling factor.18.The method of claim 17,wherein the representation of the rotation information comprises at least one of a Normalized Euler Angles representation, a Normalized Euler Vector representation, and an Euler Axis and Normalized Angle representation.19.The method of claim 17, wherein the scaling factor is π.20.The method of claim 16,wherein the rotation information is stored in a Supplemental Enhancement Information (SEI) received from the encoding apparatus.21.The method of claim 16, wherein the decoding the rotation information comprises:obtaining a matrix type of the rotation information;obtaining a matrix size of the rotation information based on the matrix type to acquire a number of dimensions of the rotation information; anddecoding the each element of the rotation information one by one based on the number of dimensions of the rotation information.22.The method of claim 21, wherein the decoding each element of the rotation information one by one comprises:for each element of the rotation information:decoding an integer part, wherein the integer part is encoded by the encoding apparatus as a 1-bit flag;decoding a decimal part; anddecoding a sign flag when either the decoded integer part or the decoded decimal part is non-zero.23.A decoding method for facial video, comprising:decoding a base picture;decoding a facial representation with regard to the base picture, and decoding a rotation information of the facial representation; andgenerating an output picture based on the base picture and the facial representation by employing a pre-defined Generative Facial Video (GFV) model;wherein the decoding the rotation information of the facial representation comprises: decoding trigonometric values related to elements of a representation of the rotation information, wherein the trigonometric values are received from an encoding apparatus.24.The method of claim 23,wherein the facial representation comprises an Euler Angles representation;wherein the elements of the Euler Angles representation comprise α, β, and γ;wherein the trigonometric values related to the elements of the Euler Angles representation comprises sin α, cos α, sin β, cos β, sin γ, and cos γ.25.The method of claim 23,wherein the facial representation comprises an Euler Axis and Angle representation;wherein the elements of the Euler Angles representation comprise ex, ey, ez, and θ;wherein the trigonometric values related to the elements of the Euler Axis and Angle representation comprises sin θ and cos θ.26.The method of claim 23,wherein the rotation information is stored in a Supplemental Enhancement Information (SEI) received from the encoding apparatus.27.The method of claim 23, wherein the decoding the rotation information comprises:obtaining a matrix type of the rotation information;obtaining a matrix size of the rotation information based on the matrix type to acquire a number of dimensions of the rotation information; anddecoding each element of the rotation information one by one based on the number of dimensions of the rotation information.28.The method of claim 27, wherein the decoding each element of the rotation information one by one comprises:for each element of the rotation information:decoding an integer part, wherein when value of the element is bounded within a range of -1 to 1, the integer part is represented as a 1-bit flag;decoding a decimal part; anddecoding a sign flag when either the decoded integer part or the decoded decimal part is non-zero.29.An encoding method for facial video, comprising:compressing a base picture;generating a facial representation of a subsequent picture with regard to the base picture through a pre-defined Generative Facial Video (GFV) model; andencoding the base picture and the facial representation into a coded bitstream;wherein a representation of the rotation information of the facial representation comprises at least one of : a Normalized Euler Angles representation, a Sinus-cosinus Euler Angles representation, an Euler Vector representation, a Normalized Euler Vector representation, an Euler Axis and Angle representation, an Euler Axis and Normalized Angle representation, an Euler Axis and Sinus-cosinus Angle representation, a Quaternion representation, a Rodriguez Vector representation, a Modified Rodriguez Parameters representation, and a Bounded Modified Rodriguez Parameters representation.30.The method of claim 31,wherein the representation of the rotation information comprises at least one of the Normalized Euler Angles representation, the Normalized Euler Vector representation, and the Euler Axis and Normalized Angle representation;wherein the method further comprises, before the encoding the facial representation:normalizing elements of the representation of the rotation information of the facial representation with a scaling factor.31.The method of claim 30, further comprising:encoding the scaling factor in the bitstream.32.The method of claim 29,wherein the representation of the rotation information comprises at least one of the Sinus-cosinus Euler Angles representation and the Euler Axis and Sinus-cosinus Angle representation;wherein the method further comprising, before the encoding the facial representation:determining values of elements of the representation of the rotation information which require trigonometric calculation.33.The method of claim 29,wherein the encoding the facial representation into a coded bitstream comprises:encoding the representation of the rotation information of the facial representation in a Supplemental Enhancement Information (SEI) , wherein the SEI comprises at least a matrix type of the rotation information.34.The method of claim 33,wherein, in the SEI, elements of the representation of the rotation information are represented by an integer part, a decimal part and a sign flag;wherein the encoding the facial representation further comprises:when value of an element of the representation of the rotation information of the facial representation is bounded within a range of -1 to 1, encoding the integer part as a 1-bit flag.35.A decoding apparatus comprising a processor and a memory,wherein the memory is configured to store executable instructions that, when executed by the processor, cause the processor to perform:decoding a base picture;decoding a facial representation with regard to the base picture;decoding a rotation information of the facial representation, wherein a representation of the rotation information comprises at least one of: a Normalized Euler Angles representation, a Sinus-cosinus Euler Angles representation, an Euler Vector representation, a Normalized Euler Vector representation, an Euler Axis and Angle representation, an Euler Axis and Normalized Angle representation, an Euler Axis and Sinus-cosinus Angle representation, a Quaternion representation, a Rodriguez Vector representation, and a Modified Rodriguez Parameters representation, and a Bounded Modified Rodriguez Parameters representation; andgenerating an output picture based on the base picture and the facial representation through a pre-defined Generative Facial Video (GFV) model.36.The decoding apparatus of claim 35,wherein the representation of the rotation information comprises the Normalized Euler Angles representation, wherein elements of the Normalized Euler Angles representation comprise α’, β’, and γ’;wherein the elements α’, β’, and γ’ are normalized values of elements α, β, and γ of an Euler Angles representation.37.The decoding apparatus of claim 35,wherein the representation of the rotation information comprises the Sinus-cosinus Euler Angles representation, wherein elements of the Sinus-cosinus Euler Angles representation comprise sin α, cos α, sin β, cos β, sin γ, and cos γ;wherein α, β, and γ are elements of an Euler Angles representation;wherein values of the elements sin α, cos α, sin β, cos β, sin γ, and cos γ of the Sinus-cosinus Euler Angles representation are calculated and transmitted by an encoding apparatus.38.The decoding apparatus of claim 35,wherein the representation of the rotation information comprises the Normalized Euler Vector representation, wherein elements of the Normalized Euler Vector representation comprise rx’, ry’, and rz’;wherein rx’, ry’, and rz’ are normalized values of elements rx, ry, and rz of an Euler Vector representation.39.The decoding apparatus of claim 35,wherein the representation of the rotation information comprises the Euler Axis and Normalized Angle representation, wherein elements of the Euler Axis and Normalized Angle representation comprise ex, ey, ez, and θ’;wherein the elements ex, ey, ez are equivalent elements of an Euler Axis and Angle representation which indicates a unit vector;wherein the element θ’ is a normalized value of an element θ of the Euler Axis and Angle representation.40.The decoding apparatus of claim 35,wherein the representation of the rotation information comprises the Euler Axis and Sinus-cosinus Angle representation, wherein elements of the Euler Axis and Sinus-cosinus Angle representation comprise ex, ey, ez, sin θ, and cos θ;wherein the elements ex, ey, ez are equivalent elements of an Euler Axis and Angle representation which indicates a unit vector;wherein θ is an equivalent parameter of an element of the Euler Axis and Angle representation which indicates a rotation angle;wherein values of the elements sin θ and cos θ of the Euler Axis and Sinus-cosinus Angle representation are calculated and transmitted by an encoding apparatus.41.The decoding apparatus of claim 35,wherein the representation of the rotation information comprises the Bounded Modified Rodriguez Parameters representation, wherein elements of the Modified Rodriguez Parameters comprise px, py, pz, and θ;wherein a value of θ is bounded within a range of –π to π.42.A computer readable media storing executable instructions that, when executed by a processor, cause the processor to perform:decoding a base picture;decoding a facial representation with regard to the base picture;decoding a rotation information of the facial representation, wherein a representation of the rotation information comprises at least one of: a Normalized Euler Angles representation, a Sinus-cosinus Euler Angles representation, an Euler Vector representation, a Normalized Euler Vector representation, an Euler Axis and Angle representation, an Euler Axis and Normalized Angle representation, an Euler Axis and Sinus-cosinus Angle representation, a Quaternion representation, a Rodriguez Vector representation, and a Modified Rodriguez Parameters representation, and a Bounded Modified Rodriguez Parameters representation; andgenerating an output picture based on the base picture and the facial representation through a pre-defined Generative Facial Video (GFV) model.43.The computer readable media of claim 42,wherein the representation of the rotation information comprises the Normalized Euler Angles representation, wherein elements of the Normalized Euler Angles representation comprise α’, β’, and γ’;wherein the elements α’, β’, and γ’ are normalized values of elements α, β, and γ of an Euler Angles representation.44.The computer readable media of claim 42,wherein the representation of the rotation information comprises the Sinus-cosinus Euler Angles representation, wherein elements of the Sinus-cosinus Euler Angles representation comprise sin α, cos α, sin β, cos β, sin γ, and cos γ;wherein α, β, and γ are elements of an Euler Angles representation;wherein values of the elements sin α, cos α, sin β, cos β, sin γ, and cos γ of the Sinus-cosinus Euler Angles representation are calculated and transmitted by an encoding apparatus.45.The computer readable media of claim 42,wherein the representation of the rotation information comprises the Normalized Euler Vector representation, wherein elements of the Normalized Euler Vector representation comprise rx’, ry’, and rz’;wherein rx’, ry’, and rz’ are normalized values of elements rx, ry, and rz of an Euler Vector representation.46.The computer readable media of claim 42,wherein the representation of the rotation information comprises the Euler Axis and Normalized Angle representation, wherein elements of the Euler Axis and Normalized Angle representation comprise ex, ey, ez, and θ’;wherein the elements ex, ey, ez are equivalent elements of an Euler Axis and Angle representation which indicates a unit vector;wherein the element θ’ is a normalized value of an element θ of the Euler Axis and Angle representation.47.The computer readable media of claim 42,wherein the representation of the rotation information comprises the Euler Axis and Sinus-cosinus Angle representation, wherein elements of the Euler Axis and Sinus-cosinus Angle representation comprise ex, ey, ez, sin θ, and cos θ;wherein the elements ex, ey, ez are equivalent elements of an Euler Axis and Angle representation which indicates a unit vector;wherein θ is an equivalent parameter of an element of the Euler Axis and Angle representation which indicates a rotation angle;wherein values of the elements sin θ and cos θ of the Euler Axis and Sinus-cosinus Angle representation are calculated and transmitted by an encoding apparatus.48.The computer readable media of claim 42,wherein the representation of the rotation information comprises the Bounded Modified Rodriguez Parameters representation, wherein elements of the Modified Rodriguez Parameters comprise px, py, pz, and θ;wherein a value of θ is bounded within a range of –π to π.
Citation Information
Patent Citations
Video encoding and decoding method and device, electronic equipment and storage medium
CN112449197A
Method and apparatus for subband encoding and decoding
EP1294196A2
Image and video processing apparatuses and methods
EP3711297A1
Video encoding method and apparatus with syntax element signaling of employed projection layout and associated video decoding method and apparatus
US20180332305A1