Video coding method, video decoding method and device
By extracting and encoding pupil position information in generative face video coding, the problem of inaccurate eye movement representation in existing technologies is solved, and high-quality video reconstruction at extremely low bit rates is achieved.
Patent Information
- Application Number
- CN202410430531.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-10-24
AI Technical Summary
Existing video coding standards struggle to achieve high-quality face video coding at extremely low bit rates, especially when representing eye movements with insufficient accuracy.
During the video encoding process, it is determined whether the current video frame is the base image. If it is not the base image, feature extraction is performed to obtain pupil position information, which is then encoded into generative face video coding to improve the accuracy of eye movement representation.
By adding the encoding of pupil position information, the accuracy of eye movement representation in generative face video coding at extremely low bit rates is improved, achieving higher quality video reconstruction.
Smart Images

Figure CN120835169A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Some embodiments of the present application relate to the technical field of video coding. More specifically, it relates to a video encoding method and a video decoding method and apparatus. BACKGROUND
[0002] In today's digital age, the generation and transmission of video content have become an important way of information exchange. With the rapid development of social media, online education and remote work, the demand for high-quality video coding technology is growing. However, although existing video coding standards such as High Efficiency Video Coding (H.265 / HEVC) and Versatile Video Coding (H.266 / VVC) perform well in high bit rate environments, in very low bit rate environments, these standards often have difficulty achieving satisfactory video quality and compression efficiency. In particular, in scenarios such as mobile devices, Internet of Things and emergency communications, the importance of low bit rate video coding is increasingly highlighted.
[0003] In the context of the increasing importance of low bit rate video coding, the research and standardization of generative face video coding are particularly urgent. Generative models such as Generative Adversarial Network (GAN) and Variational AutoEncoder (VAE) provide new possibilities for generating high-quality face videos at very low bit rates. These models can learn facial features from limited data and generate realistic video frames, thereby significantly reducing the required bit rate while maintaining video quality. Currently, when describing the movement of the eyes, generative face video coding only represents the eye opening and closing state and degree through a matrix. However, simply representing the eye opening and closing state and degree is not sufficient to represent accurate eye movement. SUMMARY
[0004] Exemplary embodiments of the present application provide a video encoding method, a video decoding method and apparatus for improving the accuracy of representing eye movement in generative face video coding.
[0005] Some embodiments of the present application provide technical solutions as follows:
[0006] In a first aspect, some embodiments of the present application provide a video encoding method, comprising:
[0007] determining whether the current video frame is a base image;
[0008] In a case where the current video frame is a base image, the current video frame is encoded based on a preset video encoding standard to obtain encoded data of the current video frame.
[0009] In a case where the current video frame is not a base image, image features corresponding to the current video frame are extracted to obtain the image features corresponding to the current video frame, and the image features corresponding to the current video frame are encoded to obtain encoded data of the current video frame; the image features corresponding to the current video frame include pupil position information of the current video frame.
[0010] In a second aspect, some embodiments of the present application provide an image decoding method, comprising:
[0011] obtaining encoded data of a current video frame;
[0012] determining whether the current video frame is a base image according to the encoded data of the current video frame;
[0013] In a case where the current video frame is a base image, the encoded data of the current video frame is decoded based on a preset video decoding standard to obtain a reconstructed current video frame.
[0014] In a case where the current video frame is not a base image, the encoded data of the current video frame is decoded to obtain image features corresponding to the current video frame, and the current video is reconstructed according to the image features corresponding to the current video frame to obtain a reconstructed current video frame; the image features corresponding to the current video frame include pupil position information of the current video frame.
[0015] In a third aspect, some embodiments of the present application provide an image encoding device, comprising:
[0016] a memory configured to store a computer program;
[0017] a processor configured to, when the computer program is invoked, cause the video encoding device to implement the image encoding method of the first aspect.
[0018] In a fourth aspect, some embodiments of the present application provide an image decoding device, comprising:
[0019] a memory configured to store a computer program;
[0020] a processor configured to, when the computer program is invoked, cause the video decoding device to implement the image encoding method of the first aspect.
[0021] In a fifth aspect, some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a computing device, the computing device implements the method described in the first aspect or the second aspect.
[0022] In a sixth aspect, some embodiments of the present application provide a computer program product, which, when executed on a computer, enables the computer to implement the method described in the first aspect or the second aspect.
[0023] As can be seen from the above technical solution, the video encoding method provided in the embodiment of the present application first determines whether the current video frame is a base image, and if the current video frame is a base image, encodes the current video frame based on a preset video encoding standard to obtain the encoded data of the current video frame; if the current video frame is not a base image, feature extraction is performed on the current video frame to obtain image features corresponding to the current video frame, and then the image features corresponding to the current video frame are encoded to obtain the encoded data of the current video frame. Since the image features corresponding to the current video frame obtained by feature extraction on the current video frame in the embodiment of the present application include pupil position information of the current video frame, the decoding end can improve the accuracy of representing eye movements, thereby more accurately reconstructing the current video frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the implementation methods of some embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0025] Figure 1 The following is a structural block diagram of a generative face video encoding and decoding system in some embodiments of the present application;
[0026] Figure 2 A schematic diagram of a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0027] Figure 3 A schematic diagram of a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0028] Figure 4 A schematic diagram of a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0029] Figure 5 A schematic diagram of a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0030] Figure 6 A diagram illustrating a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0031] Figure 7 A diagram illustrating a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0032] Figure 8 A diagram illustrating a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0033] Figure 9 A diagram illustrating a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0034] Figure 10 A diagram illustrating a coordinate system representing pupil position information in some embodiments of the present application is shown;
[0035] Figure 11 A diagram illustrating a coordinate system representing pupil position information in some embodiments of the present application is shown; DETAILED DESCRIPTION
[0036] For the purpose of clarity, the present application will be described with reference to exemplary embodiments described in the following detailed description. It should be appreciated that the detailed description is only for the purpose of exemplary embodiments and is not intended to limit the present application, as described in the general description above.
[0037] It should be noted that the brief description of terms in the present application is only for the convenience of understanding the following described embodiments, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.
[0038] The terms "comprising" and "having" and any variations thereof are intended to cover but not limited to inclusive, for example, a product or device containing a series of components does not necessarily limit to all components clearly listed, but can include other components not clearly listed or inherent to these products or devices.
[0039] The description of "some implementations", "some embodiments" and the like in the specification are meant to indicate that the described implementations or embodiments can include a particular feature, structure or characteristic, but may not necessarily every implementation or embodiment. In addition, such phrases do not necessarily refer to the same implementation. In addition, when a particular feature, structure or characteristic is described in connection with an implementation, it is considered within the knowledge of those skilled in the art to implement such feature, structure or characteristic in connection with other implementations, whether explicitly described herein or not.
[0040] Under the promotion of AI-Generated Content (AIGC), the Joint Video Experts Team (JVET) of ISO / IEC SC 29 and ITU-T SG16 is committed to researching the standardization of realizing generative facial video coding in an extremely low bit rate environment.
[0041] Referring to Figure 1 , a structural block diagram of a generative facial video coding system is shown. As Figure 1 shown, the generative facial video coding system mainly consists of two parts. The first part is a basic image coding module 11, and the second part is a generative facial video coding module 12. Figure 1
[0042] The basic image coding module 11 includes a basic image encoder 111 and a basic image decoder 112. The basic image encoder 111 is responsible for encoding the basic image carrying the facial basic texture information to generate the code stream corresponding to the basic image. The basic image decoder 112 is responsible for decoding the code stream corresponding to the basic image to obtain the reconstructed basic image. The basic image encoder 111 and the basic image decoder 112 respectively adopt standard coding and decoding technologies for coding and decoding. For example: H.265 / HEVC, H.266 / VVC, etc.
[0043] The generative facial video coding module 12 includes an analysis model 121, a feature encoding module 122, a feature decoding module 123, and a generative model 124. The analysis model 121 is used to extract the motion information of the subsequent image of the basic image to obtain the motion information of the subsequent image. The feature encoding module 122 is used to encode the motion information extracted by the analysis model 121 to generate the code stream corresponding to the subsequent image. The feature decoding module 123 is used to decode the code stream corresponding to the subsequent image to obtain the motion information of the subsequent image. The generative model 124 is used to reconstruct the subsequent image according to the motion information and the facial basic texture information in the basic image.
[0044] Compared with the coding and decoding basis of H.265 / HEVC, H.266 / VVC, etc., the generative facial video coding technology can provide higher quality facial image reconstruction at an extremely low bit rate.
[0045] The following describes the algorithms for implementing the analysis model and the generative model in the generative facial video coding module.
[0046] The facial video coding algorithm for implementing the above-mentioned generative facial video coding module in the current JVET test is shown in Table 1:
[0047] Table 1
[0048] Respective algorithm Face representation FOMM 2D key points + affine transformation matrix FV2V 3D key points + head rotation / translation matrix CFTE Compact feature matrix
[0049] FOMM takes as input a base image and a subsequent image. An unsupervised keypoint detector extracts sparse 2D keypoints of the face and a local affine transformation relative to an abstract reference frame, forming a first-order motion representation. A dense motion network uses the above motion representation to generate a dense optical flow and an occlusion map from the subsequent image to the base image. A generator uses the base image and the output of the dense motion network to render the subsequent image.
[0050] FV2V takes as input a base image and a subsequent image. An appearance feature extractor outputs appearance features of the base image, and a canonical keypoint detector outputs 3D canonical keypoints. For each image, a keypoint perturbation due to head pose and expression is estimated by an estimator network, resulting in a head rotation / translation matrix and sparse 3D keypoints. A motion field estimation network uses the 3D keypoints of the base image and the 3D keypoints of the subsequent image to obtain a dense optical flow from D to S. A generator uses the appearance features of the base image and the output of the dense motion network to render the subsequent image.
[0051] CFTE takes as input a base image and a subsequent image. Each image is passed through a compact feature extractor to obtain a compact feature matrix, and a dense motion estimation network uses the difference between the compact feature matrix of the base image and the compact feature matrix of the subsequent image and the base image features to generate a dense optical flow and an occlusion map from the subsequent image to the base image. A generator uses the base image and the output of the dense motion network to render the subsequent image.
[0052] At the latest JVET meeting, experts focused on the standardization of generative face video coding, particularly on the proposal content of Supplemental Enhancement Information (SEI). As an important coding tool, SEI can provide additional information to improve the efficiency and quality of video coding. The experts agreed to add the SEI proposal content of generative face video to the proposal (Technologies under Consideration, TuC) of future Versatile Supplemental Enhancement Information (VSEI) - ITU-T H.274 | ISO / IEC 23002-7, which demonstrates forward-looking thinking on the development of future video coding technologies, especially the potential and application prospects of generative face video coding in the context of very low bit rates. The following describes the SEI information of generative face video.
[0053] Referring to Table 2, the duration range of the SEI message range for the generative face video is shown:
[0054] Table 2 SEI message range
[0055] SEI information Continuity range … … Generative face video SEI information related image
[0056] That is, the range in which the SEI message for the generative face video is effective is only the video frame related to the SEI message, not the entire video frame in the video.
[0057] The syntax elements in the SEI message for the generative face video are shown in Table 3 below:
[0058] Table 3 Syntax elements in the SEI message for the generative face video
[0059]
[0060]
[0061]
[0062]
[0063]
[0064] The semantics of each syntax element shown in Table 3 are described below: Figure 3
[0065] "gfv_id": contains an identification number for identifying facial feature information, which can be used to indicate a neural network as GenerativeNN(). The value of gfv_id is in the range of 0 to 2 32 -2. Values of 256 to 511, and 2 31 to 2 32 -2 of gfv_id are reserved for future use by ITU-T | ISO / IEC. When the value of "gfv_id" is 256 to 511, and 231 to 232-2, the decoder should ignore the SEI message.
[0066] "gfv_base_pic_flag" is equal to 1, it specifies that the current decoded picture is a base picture. When the value of "gfv_base_pic_flag" is equal to 0, it specifies that the current decoded picture is not a base picture. The constraint for the value of "gfv_base_pic_flag" is that when the SEI message of the generative face video is the first GFV SEI message in the current Coded Layer-wise Video Sequence (CLVS) with the specific value of "gfv_id", the value of "gfv_base_pic_flag" is equal to 1.
[0067] "gfv_nn_present_flag" is equal to 1, it specifies that the SEI message contains or indicates a neural network that can be used as TranslatorNN(). When the value of "gfv_nn_present_flag" is equal to 0, it specifies that the SEI message does not contain or indicate a neural network that can be used as TranslatorNN(). If this syntax element is not present, the value of "gfv_nn_present_flag" is inferred to be equal to 0.
[0068] "gfv_nn_base_flag" is equal to 1, it specifies that the indicated TranslatorNN() is a base neural-network post-filter (NNPF). When the value of "gfv_nn_base_flag" is equal to 1, it specifies that the indicated TranslatorNN() is an update relative to the base NNPF.
[0069] "gfv_nn_mode_idc" is equal to 0, it specifies that the neural network information is contained in the Neural-network post-filter characteristics (NNPFC) SEI message and the neural network information is in the ISO / IEC 15938-17 bitstream format. When the value of "gfv_nn_mode_idc" is equal to 1, it specifies that the neural network information is identified by the URI indicated by "nnpfc_uri" and the format of the neural network information is identified by "nnpfc_tag_uri".
[0070] "gfv_nn_reserved_zero_bit_a" is equal to 0.
[0071] “gfv_nn_tag_uri”: contains a tag URI whose syntax and semantics are specified by IETF RFC 4151, used to identify the neural network format and related information used as the base NNPF or updated relative to the base NNPF.
[0072] “gfv_nn_uri”: contains a URI whose syntax and semantics conform to the provisions of IETF Internet Standard 66, used to identify the neural network used as the base NNPF or updated relative to the base NNPF.
[0073] “gfv_nn_payload_byte[i]”: contains the i-th byte of the bitstream conforming to the ISO / IEC 15938-17 standard. All nnpfc_payload_byte[i] contents should be combined to be a complete bitstream conforming to the ISO / IEC 15938-17 standard.
[0074] “gfv_drive_pic_fusion_flag”: when “gfv_drive_pic_fusion_flag” exists. And when the value is 1, it means that the current decoded image (corresponding to a drive image that may be used for fusion) should be input into GenerativeNN(). When the value is 0, it means that the current decoded image should not be input into GenerativeNN().
[0075] “gfv_coordinate_present_flag”: when the value of “gfv_coordinate_present_flag” is 1, it means that the coordinate information of the key point exists. When the value of “gfv_coordinate_present_flag” is 0, it means that the coordinate information of the key point does not exist. If “gfv_matrix_type_idx[i]” is 0 or 1, the value of “gfv_coordinate_present_flag” is required to be 1.
[0076] “gfv_coordinate_precision_factor_minus1”: represents the length (in bits) of “gfv_coordinate_x_abs[i]”, “gfv_coordinate_y_abs[i]”, and “gfv_coordinate_z_abs[i]”, specifically: the value of “gfv_coordinate_precision_factor_minus1” plus 1.
[0077] "gfv num kps minusl": The value of "gfv num kps minusl" plus 1 indicates the number of key points. The value of "gfv num kps minusl" shall be in the range of 0 to 2 10 -1.
[0078] "gfv kp pred flag": When the value of "gfv kp pred flag" is 1, it indicates that the syntax element "gfv coordinate dx abs[i]", the syntax element "gfv coordinate dy abs[i]" and the syntax element "gfv coordinate dz abs[i]" exist, and the syntax element "gfv coordinate dx sign flag[i]", the syntax element "gfv coordinate dy sign flag[i]" and the syntax element "gfv coordinate dz sign flag[i]" can exist. When the value of "gfv kp pred flag" is 0, it indicates that "gfv coordinate x abs[i]", "gfv coordinate y abs[i]" and "gfv coordinate z abs[i]" exist, and the syntax element "gfv coordinate x sign flag[i]", the syntax element "gfv coordinate y sign flag[i]" and the syntax element "gfv coordinate z sign flag[i]" can exist. When the value of "gfv kp pred flag" is 1, the absolute difference value for the reference frame (base image) is the absolute difference value of the i-th key point and the i-1-th key point; the absolute difference value for the non-reference frame (subsequent image) is the absolute difference value of the i-th key point and the i-th key point of the reference frame (base image).
[0079] "gfv coordinate z present flag": Indicates whether the z-axis coordinate information of the key point exists, and the value of "gfv coordinate z present flag" is 1, indicating that the z-axis coordinate information of the key point exists; the value of "gfv coordinate z present flag" is 0, indicating that the z-axis coordinate information of the key point does not exist.
[0080] "gfv coordinate x abs[i]": Indicates the normalized absolute value of the x-axis coordinate of the i-th key point.
[0081] “gfv_coordinate_x_abs[i]” indicates the normalized absolute value of the x-axis coordinate of the i-th key point.
[0082] “gfv_coordinate_y_abs[i]” indicates the normalized absolute value of the y-axis coordinate of the i-th key point.
[0083] “gfv_coordinate_y_sign_flag[i]” indicates the sign of the y-axis coordinate of the i-th key point. If “gfv_coordinate_y_sign_flag[i]” is not present, the value of “gfv_coordinate_y_sign_flag[i]” is inferred to be 0.
[0084] “gfv_coordinate_z_abs[i]” indicates the normalized absolute value of the z-axis coordinate of the i-th key point.
[0085] “gfv_coordinate_z_sign_flag[i]” indicates the sign of the z-axis coordinate of the i-th key point. If “gfv_coordinate_z_sign_flag[i]” is not present, the value of “gfv_coordinate_z_sign_flag[i]” is inferred to be 0.
[0086] “gfv_coordinate_dx_abs[i]” indicates the absolute difference value of the x-axis coordinate of the i-th key point.
[0087] “gfv_coordinate_dx_sign_flag[i]” indicates the sign of the x-axis coordinate difference value of the i-th key point. If “gfv_coordinate_dx_sign_flag[i]” is not present, the value of “gfv_coordinate_dx_sign_flag[i]” is inferred to be 0.
[0088] “gfv_coordinate_dy_abs[i]” indicates the absolute difference value of the y-axis coordinate of the i-th key point.
[0089] “gfv_coordinate_dy_sign_flag[i]” indicates the sign of the y-axis coordinate difference value of the i-th key point. If “gfv_coordinate_dy_sign_flag[i]” is not present, the value of “gfv_coordinate_dy_sign_flag[i]” is inferred to be 0.
[0090] "gfv_coordinate_dz_abs[i]": indicates the absolute difference of the z-axis coordinate of the i-th key point.
[0091] "gfv_coordinate_dz_sign_flag[i]": indicates the sign of the z-axis coordinate difference of the i-th key point. If "gfv_coordinate_dz_sign_flag[i]" is not present, its value is inferred to be 0.
[0092] "gfv_matrix_present_flag": when the value of "gfv_matrix_present_flag" is 1, it indicates that there are matrix parameters in the SEI message. When the value of "gfv_matrix_present_flag" is 0, it indicates that there are no matrix parameters in the SEI message.
[0093] "gfv_matrix_element_precision_factor_minus1": indicates the number of bits of the decimal part of the matrix element, i.e. the length (in bits) of "gfv_matrix_element_dec[i][j][k][m]". Specifically, the value of "gfv_matrix_element_precision_factor_minus1" plus 1.
[0094] "gfv_num_matrix_types_minus1": the value of "gfv_num_matrix_types_minus1" plus 1 indicates the number of matrix types of the signal in the SEI message. The value of "gfv_num_matrix_types_minus1" shall be in the range of 0 to 2 6 -1.
[0095] "gfv_matrix_pred_flag": When the value of "gfv_matrix_pred_flag" is 1, it indicates that the syntax elements "gfv_matrix_element_int[i][j][k][m]", "gfv_matrix_element_dec[i][j][k][m]" and "gfv_matrix_element_sign_flag[i][j][k][m]" can be present. When the value of "gfv_matrix_pred_flag" is 0, it indicates that the syntax elements "gfv_matrix_delta_element_int[i][j][k][m]", "gfv_matrix_delta_element_dec[i][j][k][m]" are present and the syntax elements "gfv_matrix_delta_element_sign_flag[i][j][k][m]" can be present. If this flag is not present, the value of "gfv_matrix_pred_flag" is inferred to be 0. When the value of "gfv_matrix_pred_flag" is 1, the difference is relative to the reference frame (base image) matrix element value.
[0096] "gfv_matrix_type_idx[i]": Indicates the index of the i-th matrix type, as indicated in Table 4. Table 4 details the different "gfv_matrix_type_idx[i]" values and their corresponding matrix types, such as affine transform matrix, covariance matrix, lip shape matrix, etc.
[0097] Table 4 Semantics of gfv_matrix_type_idx
[0098]
[0099] "gfv_num_matrices_equal_to_num_kps_flag[i]": When the value is equal to 1, it indicates that the number of matrices of the i-th matrix type is equal to the value of "gfv_num_kps_minus1" plus 1. When the value of "gfv_num_matrices_equal_to_num_kps_flag[i]" is 0, it indicates that the number of matrices of the i-th matrix type is not equal to the value of "gfv_num_kps_minus1" plus 1.
[0100] "gfv_num_matrices_info[i]": Provides information to derive the number of matrices of the i-th matrix type.
[0101] "gfv_matrix_width_minus1[ i ]": The value of "gfv_matrix_width_minus1[ i ]" plus 1 represents the matrix width of the i-th matrix type.
[0102] "gfv_matrix_height_minus1[ i ]": The value of "gfv_matrix_height_minus1[ i ]" plus 1 represents the matrix height of the i-th matrix type.
[0103] "gfv_matrix_for_3D_space_flag[ i ]": When the value of "gfv_matrix_for_3D_space_flag[ i ]" is 1, it indicates that the i-th matrix type is a matrix defined in three-dimensional space. When the value of "gfv_matrix_for_3D_space_flag[ i ]" is 0, it indicates that the i-th matrix type is a matrix defined in two-dimensional space.
[0104] "gfv_matrix_element_int[ i ][ j ][ k ][ m ]": Indicates the integer part of the matrix element of the k-th row and m-th column of the j-th matrix of the i-th matrix type.
[0105] "gfv_matrix_element_dec[ i ][ j ][ k ][ m ]": Indicates the decimal part of the matrix element of the k-th row and m-th column of the j-th matrix of the i-th matrix type.
[0106] "gfv_matrix_element_sign_flag[ i ][ j ][ k ][ m ]": Indicates the sign of the matrix element of the k-th row and m-th column of the j-th matrix of the i-th matrix type. If "gfv_matrix_element_sign_flag[ i ][ j ][ k ][ m ]" does not exist, it is inferred that the value of "gfv_matrix_element_sign_flag[ i ][ j ][ k ][ m ]" is 0.
[0107] "gfv_matrix_delta_element_int[ i ][ j ][ k ][ m ]": Indicates the integer part of the difference value of the matrix element of the k-th row and m-th column of the j-th matrix of the i-th matrix type.
[0108] "gfv_matrix_delta_element_dec[ i ][ j ][ k ][ m ]": Indicates the decimal part of the difference value of the matrix element of the k-th row and m-th column of the j-th matrix of the i-th matrix type.
[0109] "gfv_matrix_delta_element_sign_flag[i][j][k][m]": indicates the sign of the difference of the matrix element of the k-th row and m-th column of the j-th matrix of the i-th matrix type. If "gfv_matrix_delta_element_sign_flag[i][j][k][m]" is not present, its value is inferred to be 0.
[0110] The descriptor of each syntax element shown in Table 3 represents the entropy decoding algorithm of the corresponding syntax element, and the correspondence between the descriptor and the entropy decoding algorithm is shown in Table 5 as follows:
[0111] Table 5: Entropy decoding algorithm represented by the descriptor
[0112]
[0113] In addition, when the parameter in the descriptor () is n, it indicates that the corresponding syntax element is fixed-length coded; when the parameter in the descriptor () is v, it indicates that the corresponding element is variable-length coded.
[0114] As shown in Table 4 above, the eye matrix represented by "gfv_matrix_type_idx[3]" in the related art is used to represent the eye opening state and degree. However, the eye matrix alone is not sufficient to represent the real and accurate eye movement. In order to solve the above problem, some embodiments of the present application add the pupil position to the eye movement information, so as to more realistically and accurately represent the eye movement. The implementation scheme of adding the pupil position to the eye movement information is described in detail below.
[0115] The coordinate system for representing the pupil position is described below:
[0116] Referring to Figure 2 In some embodiments, as shown in FIG. 6, the position of the pupil can be represented in a rectangular coordinate system with the left corner of the eye as the origin o and the horizontal axis and the vertical axis as the x and y axes. Figure 2 In the embodiment shown in FIG. 6, the left corner of the eye is taken as the origin, and the horizontal axis and the vertical axis are taken as the x and y axes. Therefore, the manner of determining the position of the pupil can be: a perpendicular line is drawn from the center of the pupil to the y axis, and the length of the perpendicular line from the center of the pupil to the y axis is determined as the x-axis coordinate value x1 of the pupil; a perpendicular line is drawn from the center of the pupil to the x axis, and the length of the perpendicular line from the center of the pupil to the x axis is determined as the y-axis coordinate value y1 of the pupil, thereby determining the coordinate value (x1, y1) of the pupil.
[0117] Referring to Figure 3 In some embodiments, as shown in FIG. 7, the position of the pupil can be represented in a rectangular coordinate system with the right corner of the eye as the origin o and the horizontal axis and the vertical axis as the x and y axes. Figure 3In the shown embodiment, the right eye corner is taken as the origin, the horizontal axis and the vertical axis are taken as the x and y axes, and thus the manner of determining the position of the pupil can be: a perpendicular line is drawn from the pupil center to the y axis, and the length of the perpendicular line from the pupil center to the y axis is determined as the x axis coordinate value x1 of the pupil, a perpendicular line is drawn from the pupil center to the x axis, and the length of the perpendicular line from the pupil center to the x axis is determined as the y axis coordinate value y1 of the pupil, thereby determining the coordinate value (x1, y1) of the pupil.
[0118] Referring to Figure 4 In some embodiments, as shown, the position of the pupil can be represented in a rectangular coordinate system with the left eye corner as the origin o, the line connecting the left eye corner and the right eye corner as the x axis, and the perpendicular line of the vertical x axis passing through the origin o as the y axis. Figure 4 In the shown embodiment, the left eye corner is taken as the origin, the line connecting the left eye corner and the right eye corner is taken as the x axis, and the perpendicular line of the vertical x axis is taken as the y axis, and thus the manner of determining the position of the pupil can be: a perpendicular line is drawn from the pupil center to the y axis, and the length of the perpendicular line from the pupil center to the y axis is determined as the x axis coordinate value x1 of the pupil, a perpendicular line is drawn from the pupil center to the x axis, and the length of the perpendicular line from the pupil center to the x axis is determined as the y axis coordinate value y1 of the pupil, thereby determining the coordinate value (x1, y1) of the pupil.
[0119] Referring to Figure 5 In some embodiments, as shown, the position of the pupil can be represented in a rectangular coordinate system with the right eye corner as the origin o, the line connecting the left eye corner and the right eye corner as the x axis, and the perpendicular line of the vertical x axis passing through the origin o as the y axis. Figure 5 In the shown embodiment, the left eye corner is taken as the origin, the line connecting the left eye corner and the right eye corner is taken as the x axis, and the perpendicular line of the vertical x axis is taken as the y axis, and thus the manner of determining the position of the pupil can be: a perpendicular line is drawn from the pupil center to the y axis, and the length of the perpendicular line from the pupil center to the y axis is determined as the x axis coordinate value x1 of the pupil, a perpendicular line is drawn from the pupil center to the x axis, and the length of the perpendicular line from the pupil center to the x axis is determined as the y axis coordinate value y1 of the pupil, thereby determining the coordinate value (x1, y1) of the pupil.
[0120] Referring to Figure 6 In some embodiments, as shown, the position of the pupil can be represented in a polar coordinate system with the left eye corner as the pole o and the horizontal axis as the polar axis ox. Figure 6 In the shown embodiment, the left eye corner is taken as the pole o and the horizontal axis as the polar axis ox, and thus the manner of determining the position of the pupil can be: the distance from the pupil center to the left eye corner is determined as the polar radius r of the pupil, and the angle between the polar radius and the polar axis ox is determined as the polar angle θ of the pupil, thereby determining the coordinate value (r, θ) of the pupil.
[0121] Referring to Figure 7 In some embodiments, as shown, the position of the pupil can be represented in a polar coordinate system with the right eye corner as the pole o and the horizontal axis as the polar axis ox. Figure 7In the embodiment shown, the right corner of the eye is taken as the pole o and the horizontal axis is the polar axis ox. Therefore, the method for determining the position of the pupil can be: determining the distance from the center of the pupil to the right corner of the eye as the polar diameter r of the pupil, and determining the angle between the polar diameter and the polar axis ox as the polar angle θ of the pupil, thereby determining the coordinate value (r, θ) of the pupil.
[0122] Reference Figure 8 As shown, in some embodiments, the position of the pupil can be represented by a polar coordinate system with the left eye corner as the pole o and the line connecting the left eye corner and the right eye corner as the polar axis ox. Figure 8 In the embodiment shown, the left corner of the eye is the pole o, and the line connecting the left corner of the eye and the right corner of the eye is the polar axis ox. Therefore, the method for determining the position of the pupil can be: determining the distance from the center of the pupil to the left corner of the eye as the polar diameter r of the pupil, and determining the angle between the polar diameter and the polar axis ox as the polar angle θ of the pupil, thereby determining the coordinate value (r, θ) of the pupil.
[0123] Reference Figure 9 As shown, in some embodiments, the position of the pupil can be represented by a polar coordinate system with the right eye corner as the pole o and the line connecting the right eye corner and the left eye corner as the polar axis ox. Figure 9 In the embodiment shown, the right corner of the eye is taken as the pole o, and the line connecting the right corner of the eye and the left corner of the eye is the polar axis ox. Therefore, the method for determining the position of the pupil can be: determining the distance from the center of the pupil to the right corner of the eye as the polar diameter r of the pupil, and determining the angle between the polar diameter and the polar axis ox as the polar angle θ of the pupil, thereby determining the coordinate value (r, θ) of the pupil.
[0124] The following describes the quantization scheme for quantifying pupil position information:
[0125] In some embodiments, when using Figure 2 The coordinate system shown (the left corner of the eye is the origin o, the horizontal axis and the vertical axis are the rectangular coordinate system of the x and y axes) or Figure 3 When the coordinate system shown (the right corner of the eye is the origin o, the horizontal axis and the vertical axis are the x-axis and the y-axis are the rectangular coordinate system), the x-axis can be normalized by the projection distance of the line connecting the left corner of the eye and the right corner of the eye to the x-axis, and the quantization step size of the x-axis is {1 / 32, 1 / 16, 1 / 8}; the y-axis can be normalized by the distance between the highest point and the lowest point of both eyes, and the quantization step size of the y-axis is {1 / 16, 1 / 8}.
[0126] In some embodiments, when using Figure 4 The coordinate system shown (left corner of the eye is the origin o, the line connecting the left and right corners of the eye is x, and the perpendicular line of the x-axis passing through the origin o is the rectangular coordinate system of the y-axis) or Figure 5In the illustrated coordinate system (the right eye corner as the origin o, the line connecting the left eye corner and the right eye corner as the x-axis, and the perpendicular line of the x-axis passing through the origin o as the y-axis), the x-axis can be normalized by the distance between the left eye corner and the right eye corner, and the x-axis can be selected with a quantization step of {1 / 32, 1 / 16, 1 / 8}; and the y-axis can be normalized by the maximum distance of the projection of the left eye and the right eye in the y-axis direction, and the y-axis can be selected with a quantization step of {1 / 16, 1 / 8}.
[0127] In some embodiments, when the polar coordinate system is adopted Figure 2 In the illustrated coordinate system (the left eye corner as the origin o, the horizontal axis and the vertical axis as the x-axis and the y-axis), or Figure 3 In the illustrated coordinate system (the right eye corner as the origin o, the horizontal axis and the vertical axis as the x-axis and the y-axis), or Figure 4 In the illustrated coordinate system (the left eye corner as the origin o, the line connecting the left eye corner and the right eye corner as the x-axis, and the perpendicular line of the x-axis passing through the origin o as the y-axis), or Figure 5 In the illustrated coordinate system (the right eye corner as the origin o, the line connecting the left eye corner and the right eye corner as the x-axis, and the perpendicular line of the x-axis passing through the origin o as the y-axis), the x-axis and the y-axis can also be normalized by the line connecting the left eye corner and the right eye corner, and a quantization step of {1 / 16, 1 / 8} can be selected.
[0128] In some embodiments, when the polar coordinate system is adopted Figure 6 In the illustrated coordinate system (the left eye corner as the origin o, the horizontal axis as the polar axis), or Figure 7 In the illustrated coordinate system (the right eye corner as the origin o, the horizontal axis as the polar axis), or Figure 8 In the illustrated coordinate system (the right eye corner as the origin o, the line connecting the left eye corner and the right eye corner as the polar axis), or Figure 9 In the illustrated coordinate system (the right eye corner as the origin o, the line connecting the left eye corner and the right eye corner as the polar axis), the polar axis can be normalized by the distance from the left eye corner to the right eye corner, and a quantization step of {1 / 32, 1 / 16, 1 / 8} can be selected for the polar radius; and the polar angle can be normalized by the included angle between the highest point and the lowest point of the eye, and a quantization step of {1 / 16, 1 / 8} can be selected for the polar angle.
[0129] In addition, when encoding the pupil position information, the first frame of picture needs to encode a specific quantization parameter or an index of the quantization parameter, and if the quantization parameter of the subsequent image does not change, the new quantization parameter can not be transmitted, and if the quantization parameter changes, the difference between the new quantization parameter and the first frame quantization parameter or the index of the new quantization parameter can be transmitted.
[0130] In the process of generating a face video code, the pupil is relatively fixed in each frame of video position, and the absolute position coding of each frame of image is resource-consuming. In order to accurately reproduce the pupil position at the decoding end and not to occupy too much code stream in the coding process, the application specifies a reference point, which is the left or right corner of the eye where the pupil is located. After the reference point is determined, the relative coordinate information of the pupil position and the reference point is determined. Finally, the relative coordinate information of each frame is coded, or the relative coordinate residual information between each frame and the previous frame is coded.
[0131] Regarding the pupil position content, there are two schemes of one set of pupil position data and two sets of pupil position data. Specifically, one set of pupil positions can be shared by two eyes, in which case the same reference point is selected for both eyes. For example, the left eye selects the left corner and the right eye selects the left corner, or the left eye selects the right corner and the right eye selects the right corner. Alternatively, the pupil position can be transmitted separately for each eye, in which case the reference point for the left and right eyes is selected arbitrarily.
[0132] In combination with the above-mentioned coordinate system for representing the pupil position and the quantization scheme for quantizing the pupil position information, the application embodiment provides a scheme for encoding the pupil position information as shown in Table 6:
[0133] Table 6 Pupil position information encoding scheme
[0134]
[0135]
[0136] Table 6 shows all the schemes provided by selecting different coordinate systems to represent the pupil position and different quantization schemes for pupil position information. A unique coordinate system is determined for different eye movements, and a unique quantization parameter is determined for different transmission conditions and different coordinate system selections.
[0137] In combination with the transmission of pupil position information, which can transmit coordinate information or coordinate residual information, and the pupil coordinates of both eyes or the pupil coordinates of two eyes are not shared, the application embodiment provides 8x6x2x2(8 coordinate system selection x 6 quantization parameter selection x coordinate information\coordinate residual information transmission x two eye pupil coordinate sharing\two eye pupil coordinate not sharing) = 192 different pupil position encoding schemes.
[0138] In some embodiments, the pupil position information can be represented by SEI.
[0139] In video coding, SEI technology is a technology for adding additional information to the video code stream. These additional information can include metadata related to video content, such as timestamp, scene information, copyright information, color space information, etc. The use of SEI can provide more functions and optimize video quality.
[0140] It should be noted that in some embodiments, the pupil position information can also be represented by other ways that can add content to the eyes, and the embodiments of the present application do not limit this.
[0141] The following describes the use of syntax elements when the pupil position information is represented by SEI. Referring to Table 7, the syntax elements used when the pupil position information is represented by SEI include:
[0142] Table 7 represents the SEI syntax elements of the pupil position information
[0143]
[0144]
[0145]
[0146]
[0147]
[0148] "gfv_pupil_coordinate_present_flag": "gfv_pupil_coordinate_present_flag" is used to indicate whether the pupil position information exists. When the value of "gfv_pupil_coordinate_present_flag" is 1, it indicates that the current decoded image contains pupil position information; when the value of "gfv_pupil_coordinate_present_flag" is 0, it indicates that the current decoded image does not contain pupil position information. If the syntax element "gfv_pupil_coordinate_present_flag" does not exist, it is inferred that the value of "gfv_pupil_coordinate_present_flag" is 0.
[0149] "gfv_pupil_both_eyes_flag": is used to indicate whether the pupil position information is common for both eyes. When the value of "gfv_pupil_both_eyes_flag" is 1, it means that the pupil position information is common for both eyes, and the pupil position parameters in the SEI message apply to the pupil position of both eyes; when the value of "gfv_pupil_both_eyes_flag" is 0, it means that the pupil position information is for the left eye and the right eye respectively, and the SEI message can contain two sets of pupil position information.
[0150] "gfv_pupil_coordinate_system_idx": is used to indicate the index of the coordinate system used by the pupil position information. The index is used to select the reference coordinate system of the pupil position coordinates from the predefined coordinate system list. The value of "gfv_pupil_coordinate_system_idx" and the corresponding relationship of the coordinate system can be shown in Table 8 as follows:
[0151] Table 8 Corresponding relationship of index value and coordinate system
[0152]
[0153]
[0154] "gfv_pupil_pred_flag": indicates whether the binocular pupil position information is represented by normalized absolute difference or normalized coordinate value. Specifically, when the value of "gfv_pupil_pred_flag" is 1, it means that the binocular pupil position information is represented by normalized absolute difference, there are syntax elements "gfv_pupil_dx_coordinate_abs" and "gfv_pupil_dy_coordinate_abs", and there can be syntax elements "gfv_pupil_dx_coordinate_sign_flag" and "gfv_pupil_dy_coordinate_sign_flag". When the value of "gfv_pupil_pred_flag" is 0, it means that the binocular pupil position information is represented by normalized coordinate value, there are syntax elements "gfv_pupil_x_coordinate" and "gfv_pupil_y_coordinate". When the value of "gfv_pupil_pred_flag" is 1, the absolute difference is the absolute difference between the pupil position of this frame and the pupil position of the reference frame (base image).
[0155] "gfv_pupil_x_coordinate_precision_factor_minus1": specifies the length (in bits) of the syntax element "gfv_pupil_x_coordinate_abs" or "gfv_pupil_x_coordinate_abs". In particular, the value of "gfv_pupil_x_coordinate_precision_factor_minus1" plus 1.
[0156] "gfv_pupil_y_coordinate_precision_factor_minus1": specifies the length (in bits) of the syntax element "gfv_pupil_y_coordinate_abs" or "gfv_pupil_y_coordinate_abs". In particular, the value of "gfv_pupil_y_coordinate_precision_factor_minus1" plus 1.
[0157] "gfv_pupil_dx_coordinate_abs": specifies the normalized absolute difference of the x-axis coordinate when the pupil position information is shared for both eyes.
[0158] "gfv_pupil_dx_coordinate_sign_flag": specifies the sign of the normalized difference of the x-axis coordinate when the pupil position information is shared for both eyes. If the syntax element "gfv_pupil_dx_coordinate_sign_flag" is not present, the value of "gfv_pupil_dx_coordinate_sign_flag" is inferred to be 0.
[0159] "gfv_pupil_dy_coordinate_abs": specifies the normalized absolute difference of the y-axis coordinate when the pupil position information is shared for both eyes.
[0160] "gfv_pupil_dy_coordinate_sign_flag": specifies the sign of the normalized difference of the y-axis coordinate when the pupil position information is shared for both eyes. If the syntax element "gfv_pupil_dy_coordinate_sign_flag" is not present, the value of "gfv_pupil_dy_coordinate_sign_flag" is inferred to be 0.
[0161] "gfv_pupil_x_coordinate": indicates the x-axis normalized coordinate value when the pupil position information is shared for both eyes.
[0162] “gfv_pupil_y_coordinate”: indicates that the pupil position information is the normalized y-axis coordinate value when it is shared by both eyes.
[0163] "gfv_pupil_polar_radius_precision_factor_minus1": indicates the length (in bits) of the syntax element "gfv_pupil_polar_delta_radius_abs" or "gfv_pupil_polar_radius". Specifically, it is the value of "gfv_pupil_polar_radius_precision_factor_minus1" plus 1.
[0164] "gfv_pupil_polar_angle_precision_factor_minus1": indicates the length (in bits) of the syntax element "gfv_pupil_polar_delta_angle_abs" or "gfv_pupil_polar_angle". Specifically, it is the value of "gfv_pupil_polar_angle_precision_factor_minus1" plus 1.
[0165] "gfv_pupil_polar_delta_radius_abs": specifies that the pupil position information is the normalized absolute difference of the polar radius when shared by both eyes.
[0166] gfv_pupil_polar_delta_radius_sign_flag": Specifies the sign of the normalized difference of the polar radius when the pupil position information is shared by both eyes. If the syntax element "gfv_pupil_polar_delta_radius_sign_flag" is not present, the value of "gfv_pupil_polar_delta_radius_sign_flag" is inferred to be 0.
[0167] "gfv_pupil_polar_delta_angle_abs": specifies that the pupil position information is the normalized absolute difference of the polar angles when the pupil position information is shared by both eyes.
[0168] "gfv_pupil_polar_delta_angle_sign_flag": specifies the sign of the normalized difference of the polar angle of the pupil position information when shared by both eyes. If the syntax element "gfv_pupil_polar_delta_angle_sign_flag" is not present, the value of "gfv_pupil_polar_delta_angle_sign_flag" is inferred to be 0.
[0169] "gfv_pupil_polar_radius": indicates the normalized polar radius value of the pupil position information when shared by both eyes.
[0170] "gfv_pupil_polar_angle": indicates the normalized polar angle value of the pupil position information when shared by both eyes.
[0171] "gfv_pupil_left_eye_present_flag": indicates whether the left eye pupil position information is present. If the syntax element "gfv_pupil_left_eye_present_flag" is not present, the value of "gfv_pupil_left_eye_present_flag" is inferred to be 0.
[0172] "gfv_pupil_right_eye_present_flag": indicates whether the right eye pupil position information is present. If the syntax element "gfv_pupil_right_eye_present_flag" is not present, the value of "gfv_pupil_right_eye_present_flag" is inferred to be 0.
[0173] "gfv_pupil_left_eye_coordinate_system_idx": indicates the index of the coordinate system used for the left eye pupil position information. The index is used to select the reference coordinate system of the pupil position from the list of pre-defined coordinate systems. The coordinate system represented by the index can refer to Table 7.
[0174] "gfv_pupil_left_eye_pred_flag": indicates whether the left eye pupil position information is represented by normalized absolute difference value or normalized coordinate value. Specifically, when the value of "gfv_pupil_left_eye_pred_flag" is 1, it indicates that the left eye pupil position information is represented by normalized absolute difference value, the syntax elements "gfv_pupil_left_eye_dx_coordinate_abs" and "gfv_pupil_left_eye_dy_coordinate_abs" exist, and the syntax elements "gfv_pupil_left_eye_dx_coordinate_sign_flag" and "gfv_pupil_left_eye_dy_coordinate_sign_flag" can exist. When the value of "gfv_pupil_left_eye_pred_flag" is 0, it indicates that the left eye pupil position information is represented by normalized coordinate value, the syntax elements "gfv_pupil_left_eye_x_coordinate" and "gfv_pupil_left_eye_y_coordinate" exist. When the value of "gfv_pupil_left_eye_pred_flag" is 1, the absolute difference value is the absolute difference value between the pupil position of this frame and the pupil position of the reference frame (base image).
[0175] "gfv_pupil_left_eye_x_coordinate_precision_factor_minus1": indicates the length (bit) of the syntax element "gfv_pupil_left_eye_dx_coordinate_abs" or "gfv_pupil_left_eye_x_coordinate_abs". Specifically, the value of "gfv_pupil_left_eye_x_coordinate_precision_factor_minus1" plus 1.
[0176] "gfv_pupil_left_eye_y_coordinate_precision_factor_minus1": indicates the length (bit) of the syntax element "gfv_pupil_left_eye_dy_coordinate_abs" or "gfv_pupil_left_eye_y_coordinate_abs". Specifically, the value of "gfv_pupil_left_eye_y_coordinate_precision_factor_minus1" plus 1.
[0177] "gfv_pupil_left_eye_dx_coordinate_abs": specifies the normalized absolute difference value of the x-axis coordinate of the left eye pupil position information.
[0178] "gfv_pupil_left_eye_dx_coordinate_sign_flag": specifies the normalized difference value sign of the x-axis coordinate of the left eye pupil position information. If the syntax element "gfv_pupil_left_eye_dx_coordinate_sign_flag" is not present, the value 0 of "gfv_pupil_left_eye_dx_coordinate_sign_flag" is inferred.
[0179] "gfv_pupil_left_eye_dy_coordinate_abs": specifies the normalized absolute difference value of the y-axis coordinate of the left eye pupil position information.
[0180] "gfv_pupil_left_eye_dy_coordinate_sign_flag": specifies the normalized difference value sign of the y-axis coordinate of the left eye pupil position information. If the syntax element "gfv_pupil_left_eye_dy_coordinate_sign_flag" is not present, it is inferred to be equal to 0.
[0181] "gfv_pupil_left_eye_x_coordinate": indicates the x-axis normalized coordinate value of the left eye pupil position information.
[0182] "gfv_pupil_left_eye_y_coordinate": indicates the y-axis normalized coordinate value of the left eye pupil position information.
[0183] "gfv_pupil_left_eye_polar_radius_precision_factor_minus1": indicates the length (bits) of the syntax element "gfv_pupil_left_eye_polar_delta_angle_abs" or "gfv_pupil_left_eye_polar_angle". Specifically, the value of "gfv_pupil_left_eye_polar_radius_precision_factor_minus1" plus 1.
[0184] "gfv_pupil_left_eye_polar_delta_radius_sign_flag": specifies the sign of the normalized difference value of the polar radius of the left eye pupil position information. If the syntax element "gfv_pupil_left_eye_polar_delta_radius_sign_flag" is not present, the value 0 of "gfv_pupil_left_eye_polar_delta_radius_sign_flag" is inferred.
[0185] "gfv_pupil_left_eye_polar_delta_radius_abs": specifies the normalized absolute difference value of the polar radius of the left eye pupil position information.
[0186] "gfv_pupil_left_eye_polar_delta_radius_sign_flag": specifies the sign of the normalized difference value of the polar radius of the left eye pupil position information. If the syntax element "gfv_pupil_left_eye_polar_delta_radius_sign_flag" is not present, the value 0 of "gfv_pupil_left_eye_polar_delta_radius_sign_flag" is inferred.
[0187] "gfv_pupil_left_eye_polar_delta_angle_abs": specifies the normalized absolute difference value of the polar angle of the left eye pupil position information.
[0188] "gfv_pupil_left_eye_polar_delta_angle_sign_flag": specifies the sign of the normalized difference value of the polar angle of the left eye pupil position information. If the syntax element "gfv_pupil_left_eye_polar_delta_angle_sign_flag" is not present, the value 0 of "gfv_pupil_left_eye_polar_delta_angle_sign_flag" is inferred.
[0189] "gfv_pupil_left_eye_polar_radius": indicates the normalized polar radius value of the left eye pupil position information.
[0190] "gfv_pupil_left_eye_polar_angle": indicates the normalized polar angle value of the left eye pupil position information.
[0191] "gfv_pupil_right_eye_coordinate_system_idx" indicates the index of the coordinate system used for the right eye pupil position information. The index is used to select the reference coordinate system of the pupil position from the predefined list of coordinate systems. The index can refer to Table 7.
[0192] "gfv_pupil_right_eye_pred_flag" indicates whether the right eye pupil position information is represented by normalized absolute difference values or normalized coordinate values. Specifically, when the value of "gfv_pupil_right_eye_pred_flag" is 1, the right eye pupil position information is represented by normalized absolute difference values, the syntax elements "gfv_pupil_right_eye_dx_coordinate_abs" and "gfv_pupil_right_eye_dy_coordinate_abs" exist, and the syntax elements "gfv_pupil_right_eye_dx_coordinate_sign_flag" and "gfv_pupil_right_eye_dy_coordinate_sign_flag" can exist. When the value of "gfv_pupil_right_eye_pred_flag" is 0, the right eye pupil position information is represented by normalized coordinate values, the syntax elements "gfv_pupil_right_eye_x_coordinate" and "gfv_pupil_right_eye_y_coordinate" exist. When the value of "gfv_pupil_right_eye_pred_flag" is 1, the absolute difference values are the absolute difference values between the pupil position of this frame and the pupil position of the reference frame (base image).
[0193] "gfv_pupil_right_eye_x_coordinate_precision_factor_minus1" indicates the length (bits) of the syntax element "gfv_pupil_left_eye_dx_coordinate_abs" or "gfv_pupil_right_eye_x_coordinate_abs". Specifically, the value of "gfv_pupil_right_eye_x_coordinate_precision_factor_minus1" plus 1.
[0194] "gfv_pupil_right_eye_y_coordinate_precision_factor_minus1": specifies the length (in bits) of the syntax element "gfv_pupil_right_eye_dy_coordinate_abs" or "gfv_pupil_right_eye_y_coordinate_abs". Specifically, the value of "gfv_pupil_right_eye_y_coordinate_precision_factor_minus1" plus 1.
[0195] "gfv_pupil_right_eye_dx_coordinate_abs": specifies the normalized absolute difference value of the x-axis coordinate of the right eye pupil position information.
[0196] "gfv_pupil_right_eye_dx_coordinate_sign_flag": specifies the normalized difference value sign of the x-axis coordinate of the right eye pupil position information. If the syntax element "gfv_pupil_right_eye_dx_coordinate_sign_flag" is not present, the value of "gfv_pupil_right_eye_dx_coordinate_sign_flag" is inferred to be 0.
[0197] "gfv_pupil_right_eye_dy_coordinate_abs": specifies the normalized absolute difference value of the y-axis coordinate of the right eye pupil position information.
[0198] "gfv_pupil_right_eye_dy_coordinate_sign_flag": specifies the normalized difference value sign of the y-axis coordinate of the right eye pupil position information. If the syntax element "gfv_pupil_right_eye_dy_coordinate_sign_flag" is not present, the value of "gfv_pupil_right_eye_dy_coordinate_sign_flag" is inferred to be 0.
[0199] "gfv_pupil_right_eye_x_coordinate": indicates the x-axis normalized coordinate value of the right eye pupil position information.
[0200] "gfv_pupil_right_eye_y_coordinate": indicates the y-axis normalized coordinate value of the right eye pupil position information.
[0201] "gfv_pupil_right_eye_polar_radius_precision_factor_minus1": specifies the length (in bits) of the syntax element "gfv_pupil_right_eye_polar_delta_radius_abs" or "gfv_pupil_right_eye_polar_angle". In particular, the value of "gfv_pupil_right_eye_polar_radius_precision_factor_minus1" plus 1.
[0202] "gfv_pupil_right_eye_polar_angle_precision_factor_minus1": specifies the length (in bits) of the syntax element "gfv_pupil_right_eye_polar_delta_angle_abs" or "gfv_pupil_right_eye_polar_angle". In particular, the value of "gfv_pupil_right_eye_polar_angle_precision_factor_minus1" plus 1.
[0203] "gfv_pupil_right_eye_polar_delta_radius_abs": specifies the normalized absolute difference of the polar radius of the right eye pupil position information.
[0204] "gfv_pupil_right_eye_polar_delta_radius_sign_flag": specifies the sign of the normalized difference of the polar radius of the right eye pupil position information. If the syntax element "gfv_pupil_right_eye_polar_delta_radius_sign_flag" is not present, the value of "gfv_pupil_right_eye_polar_delta_radius_sign_flag" is inferred to be 0.
[0205] "gfv_pupil_right_eye_polar_delta_angle_abs": specifies the normalized absolute difference of the polar angle of the right eye pupil position information.
[0206] "gfv_pupil_right_eye_polar_delta_angle_sign_flag": specifies the sign of the normalized difference of the polar angle of the right eye pupil position information. If the syntax element "gfv_pupil_right_eye_polar_delta_angle_sign_flag" is not present, the value of "gfv_pupil_right_eye_polar_delta_angle_sign_flag" is inferred to be 0.
[0207] "gfv_pupil_right_eye_polar_radius": indicates the normalized polar radius value of the right eye pupil position information.
[0208] "gfv_pupil_right_eye_polar_angle": indicates the normalized polar angle value of the right eye pupil position information.
[0209] Based on the above, the embodiments of the present application also provide a video encoding method. Referring to FIG. 10, the video encoding method comprises the following steps S101-S103: Figure 10
[0210] S101, determine whether the current video frame is a base image.
[0211] In step S101, if the current video frame is a base image, the following step S102 is performed:
[0212] S102, encode the current video frame based on a preset video encoding standard to obtain the encoding data of the current video frame.
[0213] For example, the preset video encoding standard can be H.266 / VVC, H.265 / HEVC, etc.
[0214] In step S101, if the current video frame is not a base image, the following steps S103 and S104 are performed:
[0215] S103, feature extraction is performed on the current video frame to obtain the image features corresponding to the current video frame.
[0216] The image features corresponding to the current video frame include the pupil position information of the current video frame.
[0217] S104, encode the image features corresponding to the current video frame to obtain the encoding data of the current video frame.
[0218] The video coding method provided in the embodiments of the present application first determines whether the current video frame is a base image, and in the case that the current video frame is a base image, encodes the current video frame based on a preset video coding standard to obtain the coding data of the current video frame; in the case that the current video frame is not a base image, extracts features of the current video frame to obtain the image features corresponding to the current video frame, and then encodes the image features corresponding to the current video frame to obtain the coding data of the current video frame. Since the image features corresponding to the current video frame extracted by the embodiments of the present application in the case that the current video frame is not a base image include the pupil position information of the current video frame, the decoding end can improve the accuracy of representing eye movement, thereby more accurately reconstructing the current video frame.
[0219] In some embodiments, the feature extraction of the current video frame to obtain the image features corresponding to the current video frame comprises:
[0220] A first rectangular coordinate system is constructed with the left corner of the eye as the origin, the horizontal axis as the x-axis, and the vertical axis as the y-axis;
[0221] The distance from the geometric center of the pupil to the y-axis is obtained to obtain the x-coordinate of the pupil in the first rectangular coordinate system;
[0222] The distance from the geometric center of the pupil to the x-axis is obtained to obtain the y-coordinate of the pupil in the first rectangular coordinate system;
[0223] The pupil position information is generated according to the x-coordinate and the y-coordinate of the pupil in the first rectangular coordinate system.
[0224] That is, the pupil position of the current video frame is represented by the coordinate system shown in Figure 2
[0225] In some embodiments, the feature extraction of the current video frame to obtain the image features corresponding to the current video frame comprises:
[0226] A second rectangular coordinate system is constructed with the right corner of the eye as the origin, the horizontal axis as the x-axis, and the vertical axis as the y-axis;
[0227] The distance from the geometric center of the pupil to the y-axis is obtained to obtain the x-coordinate of the pupil in the second rectangular coordinate system;
[0228] The distance from the geometric center of the pupil to the x-axis is obtained to obtain the y-coordinate of the pupil in the second rectangular coordinate system;
[0229] The pupil position information is generated according to the x-coordinate and the y-coordinate of the pupil in the second rectangular coordinate system.
[0230] i.e., by Figure 3 The pupil position of the current video frame is represented by a coordinate system as shown in the figure.
[0231] In some embodiments, the feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises:
[0232] A third rectangular coordinate system is constructed with the left corner of the eye as the origin, the line connecting the left corner and the right corner of the eye as the x-axis, and the axis passing through the origin and perpendicular to the x-axis as the y-axis;
[0233] The distance from the geometric center of the pupil to the y-axis is obtained to obtain the x-coordinate of the pupil in the third rectangular coordinate system;
[0234] The distance from the geometric center of the pupil to the x-axis is obtained to obtain the y-coordinate of the pupil in the third rectangular coordinate system;
[0235] According to the x-coordinate and the y-coordinate of the pupil in the third rectangular coordinate system, the pupil position information is generated.
[0236] i.e., by Figure 4 The pupil position of the current video frame is represented by a coordinate system as shown in the figure.
[0237] In some embodiments, the feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises:
[0238] A fourth rectangular coordinate system is constructed with the right corner of the eye as the origin, the line connecting the right corner and the left corner of the eye as the x-axis, and the axis passing through the origin and perpendicular to the x-axis as the y-axis;
[0239] The distance from the geometric center of the pupil to the y-axis is obtained to obtain the x-coordinate of the pupil in the fourth rectangular coordinate system;
[0240] The distance from the geometric center of the pupil to the x-axis is obtained to obtain the y-coordinate of the pupil in the fourth rectangular coordinate system;
[0241] According to the x-coordinate and the y-coordinate of the pupil in the fourth rectangular coordinate system, the pupil position information is generated.
[0242] i.e., by Figure 5 The pupil position of the current video frame is represented by a coordinate system as shown in the figure.
[0243] In some embodiments, the feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises:
[0244] A first polar coordinate system is constructed with the left corner of the eye as a pole and a horizontal axis as a polar axis;
[0245] A distance from the geometric center of the pupil to the pole is obtained to obtain a polar radius of the pupil in the first polar coordinate system;
[0246] An angle between a line connecting the geometric center of the pupil and the pole and the polar axis is obtained to obtain a polar angle of the pupil in the first polar coordinate system;
[0247] The pupil position information is generated according to the polar radius and the polar angle of the pupil in the first polar coordinate system.
[0248] That is, the pupil position of the current video frame is represented by a coordinate system as shown in Figure 6
[0249] In some embodiments, the feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises:
[0250] A second polar coordinate system is constructed with the right corner of the eye as a pole and a horizontal axis as a polar axis;
[0251] A distance from the geometric center of the pupil to the pole is obtained to obtain a polar radius of the pupil in the second polar coordinate system;
[0252] An angle between a line connecting the geometric center of the pupil and the pole and the polar axis is obtained to obtain a polar angle of the pupil in the second polar coordinate system;
[0253] The pupil position information is generated according to the polar radius and the polar angle of the pupil in the second polar coordinate system.
[0254] That is, the pupil position of the current video frame is represented by a coordinate system as shown in Figure 7
[0255] In some embodiments, the feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises:
[0256] A third polar coordinate system is constructed with the left corner of the eye as a pole and a line connecting the left corner and the right corner of the eye as a polar axis;
[0257] A distance from the geometric center of the pupil to the pole is obtained to obtain a polar radius of the pupil in the third polar coordinate system;
[0258] An angle between a line connecting the geometric center of the pupil and the pole and the polar axis is obtained to obtain a polar angle of the pupil in the third polar coordinate system;
[0259] According to the polar radius and the polar angle of the pupil in the third polar coordinate system, the pupil position information is generated.
[0260] That is, by Figure 8 The coordinate system shown in the figure represents the pupil position of the current video frame.
[0261] In some embodiments, the feature extraction on the current video frame is to obtain image features corresponding to the current video frame, including:
[0262] A fourth polar coordinate system is constructed with the right corner of the eye as the polar point and the line connecting the right corner and the left corner of the eye as the polar axis;
[0263] The distance from the geometric center of the pupil to the polar point is obtained to obtain the polar radius of the pupil in the fourth polar coordinate system;
[0264] The angle between the line connecting the geometric center of the pupil and the polar point and the polar axis is obtained to obtain the polar angle of the pupil in the fourth polar coordinate system;
[0265] According to the polar radius and the polar angle of the pupil in the fourth polar coordinate system, the pupil position information is generated.
[0266] That is, by Figure 9 The coordinate system shown in the figure represents the pupil position of the current video frame.
[0267] In some embodiments, the encoding of the image features corresponding to the current video frame is to obtain the encoding data of the current video frame, including:
[0268] The length of the projection of the eye on the x-axis is normalized to the x-axis, and the x-coordinate of the pupil is quantized with a quantization step of 1 / 32 or 1 / 16 or 1 / 8.
[0269] The length of the projection of the eye on the y-axis is normalized to the y-axis, and the y-coordinate of the pupil is quantized with a quantization step of 1 / 16 or 1 / 8.
[0270] In some embodiments, the encoding of the image features corresponding to the current video frame is to obtain the encoding data of the current video frame, including:
[0271] The length of the line connecting the left corner to the right corner of the eye is normalized to the x-axis, and the x-coordinate of the pupil is quantized with a quantization step of 1 / 32 or 1 / 16 or 1 / 8.
[0272] The length of the projection of the eye on the y-axis is normalized to the y-axis, and the y-coordinate of the pupil is quantized with a quantization step of 1 / 16 or 1 / 8.
[0273] In some embodiments, the image features corresponding to the current video frame are encoded to obtain the encoded data of the current video frame, including:
[0274] The x-axis is normalized by the length of the line from the left corner of the eye to the right corner of the eye, and the x-coordinate of the pupil is quantized with a quantization step of 1 / 16 or 1 / 8.
[0275] The y-axis is normalized by the length of the line from the left corner of the eye to the right corner of the eye, and the y-coordinate of the pupil is quantized with a quantization step of 1 / 16 or 1 / 8.
[0276] In some embodiments, the image features corresponding to the current video frame are encoded to obtain the encoded data of the current video frame, including:
[0277] The polar axis is normalized by the length of the line from the left corner of the eye to the right corner of the eye, and the polar radius of the pupil is quantized with a quantization step of 1 / 32 or 1 / 16 or 1 / 8.
[0278] The polar angle is normalized by the angle between the line from the highest point of the eye to the polar point and the line from the lowest point of the eye to the polar point, and the polar angle of the pupil is quantized with a quantization step of 1 / 16 or 1 / 8.
[0279] In some embodiments, the image features corresponding to the current video frame are encoded to obtain the encoded data of the current video frame, including:
[0280] Supplemental enhancement information (SEI) is generated according to the pupil position information.
[0281] The supplemental enhancement information is entropy encoded to obtain the encoded data corresponding to the pupil position information.
[0282] In some embodiments, the supplemental enhancement information is generated according to the pupil position information, including:
[0283] A first identification syntax element for indicating whether the pupil position information exists or not is added in the supplemental enhancement information, and the value of the first identification syntax element is set to a first preset value.
[0284] In some embodiments, the first identification syntax element can be "gfv_pupil_coordinate_present_flag".
[0285] In some embodiments, the first preset value is 1.
[0286] In some embodiments, the supplemental enhancement information is generated according to the pupil position information, including:
[0287] A second identification syntax element is added to the supplemental enhancement information to indicate whether the left eye and the right eye share the same pupil position information. When the left eye and the right eye share the same pupil position information, the value of the second identification syntax element is set to a first preset value; when the left eye and the right eye do not share the same pupil position information, the value of the second identification syntax element is set to a second preset value.
[0288] In some embodiments, the second identification syntax element is "gfv_pupil_both_eyes_flag".
[0289] In some embodiments, the first preset value is 1 and the second preset value is 0.
[0290] In some embodiments, generating supplemental enhancement information based on the pupil position information includes:
[0291] When the left eye and the right eye share the same pupil position information, a first coordinate system syntax element for indicating the coordinate system to which the common pupil position information belongs is added to the supplemental enhancement information, and the value of the first coordinate system syntax element is set according to the coordinate system to which the common pupil position information belongs.
[0292] In some embodiments, the first coordinate system syntax element is "gfv_pupil_coordinate_system_idx".
[0293] In some embodiments, setting the value of the first coordinate system syntax element according to the coordinate system to which the common pupil position information belongs includes:
[0294] When the coordinate system of the shared pupil position information is Figure 2 When the coordinate system is shown, the value of the first coordinate system syntax element is set to 0;
[0295] When the coordinate system of the shared pupil position information is Figure 3 When the coordinate system is shown, the value of the first coordinate system syntax element is set to 1;
[0296] When the coordinate system of the shared pupil position information is Figure 4 When the coordinate system is shown, the value of the first coordinate system syntax element is set to 2;
[0297] When the coordinate system of the shared pupil position information is Figure 5 When the coordinate system is shown, the value of the first coordinate system syntax element is set to 3;
[0298] When the coordinate system of the shared pupil position information is Figure 6 When the coordinate system is shown, the value of the first coordinate system syntax element is set to 4;
[0299] when the coordinate system to which the common pupil position information belongs is the coordinate system shown in FIG. 6, setting the value of the first coordinate system syntax element to 5; Figure 7
[0300] when the coordinate system to which the common pupil position information belongs is the coordinate system shown in FIG. 6, setting the value of the first coordinate system syntax element to 5; Figure 8
[0301] when the coordinate system to which the common pupil position information belongs is the coordinate system shown in FIG. 6, setting the value of the first coordinate system syntax element to 5; Figure 9
[0302] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0303] when the left eye and the right eye share the same pupil position information, and the coordinate system to which the common pupil position information belongs is a rectangular coordinate system, adding a first bit depth syntax element for representing the bit depth of the x coordinate and a second bit depth syntax element for representing the bit depth of the y coordinate in the supplemental enhancement information, and setting the value of the first bit depth syntax element according to the bit depth of the x coordinate, and setting the value of the second bit depth syntax element according to the bit depth of the y coordinate.
[0304] In some embodiments, the first bit depth syntax element is “gfv_pupil_x_coordinate_precision_factor_minus1”, and the second bit depth syntax element is “gfv_pupil_y_coordinate_precision_factor_minus1”.
[0305] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0306] when the left eye and the right eye share the same pupil position information, adding a third identification syntax element for representing the representation manner of the common pupil position information in the supplemental enhancement information, and setting the value of the third identification syntax element to a first preset value when the common pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, and setting the value of the third identification syntax element to a second preset value when the common pupil position information is represented by a normalized value.
[0307] In some embodiments, the third identification syntax element is “gfv_pupil_pred_flag”.
[0308] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0309] In the case where the left eye and the right eye share the same pupil position information, the coordinate system to which the shared pupil position information belongs is a rectangular coordinate system, and the shared pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, the first position information syntax element and the second position information syntax element are added in the supplemental enhancement information, the value of the first position information syntax element is set as the normalized absolute difference value of the x-axis coordinate of the shared pupil position information, and the value of the second position information syntax element is set as the normalized absolute difference value of the y-axis coordinate of the shared pupil position information.
[0310] In some embodiments, the first position information syntax element is “gfv_pupil_dx_coordinate_abs”, and the second position information syntax element is “gfv_pupil_dy_coordinate_abs”.
[0311] In some embodiments, the generating supplemental enhancement information according to the pupil position information further includes:
[0312] In the case where the left eye and the right eye share the same pupil position information, the coordinate system to which the shared pupil position information belongs is a rectangular coordinate system, and the shared pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, the first position information syntax element and the second position information syntax element are added in the supplemental enhancement information, the value of the first position information syntax element is set as the normalized absolute difference value of the x-axis coordinate of the shared pupil position information, and the value of the second position information syntax element is set as the normalized absolute difference value of the y-axis coordinate of the shared pupil position information.
[0313] In some embodiments, the first position information syntax element is “gfv_pupil_dx_coordinate_abs”, and the second position information syntax element is “gfv_pupil_dy_coordinate_abs”.
[0314] In some embodiments, the generating supplemental enhancement information according to the pupil position information further includes:
[0315] In the case where the left eye and the right eye share the same pupil position information, the coordinate system to which the shared pupil position information belongs is a rectangular coordinate system, and the shared pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, the first position information syntax element and the second position information syntax element are added in the supplemental enhancement information, the value of the first position information syntax element is set as the normalized absolute difference value of the x-axis coordinate of the shared pupil position information, and the value of the second position information syntax element is set as the normalized absolute difference value of the y-axis coordinate of the shared pupil position information.
[0316] In some embodiments, the third position information syntax element is "gfv_pupil_x_coordinate", and the fourth position information syntax element is "gfv_pupil_y_coordinate".
[0317] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0318] In the case that the left eye and the right eye share the same pupil position information, and share a coordinate system to which the pupil position information belongs, and the coordinate system is a polar coordinate system, a third depth syntax element for representing a bit depth of a polar radius and a fourth depth syntax element for representing a bit depth of a polar angle are added in the supplemental enhancement information, and a value of the third depth syntax element is set according to the bit depth of the polar radius, and a value of the fourth depth syntax element is set according to the bit depth of the polar angle.
[0319] In some embodiments, the third depth syntax element is "gfv_pupil_polar_radius_precision_factor_minus", and the fourth depth syntax element is "gfv_pupil_polar_angle_precision_factor_minus1".
[0320] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0321] In the case that the left eye and the right eye share the same pupil position information, share a coordinate system to which the pupil position information belongs, and the coordinate system is a polar coordinate system, and the shared pupil position information is represented by a normalized absolute difference value from the pupil position information of the base image, a fifth position information syntax element and a sixth position information syntax element are added in the supplemental enhancement information, a value of the fifth position information syntax element is set as a normalized absolute difference value of a polar radius of the shared pupil position information, and a value of the sixth position information syntax element is set as a normalized absolute difference value of a polar angle of the shared pupil position information.
[0322] In some embodiments, the fifth position information syntax element is "gfv_pupil_polar_delta_radius_abs", and the sixth position information syntax element is "gfv_pupil_polar_delta_angle_abs".
[0323] In some embodiments, the generating the supplemental enhancement information according to the pupil position information further comprises:
[0324] The third sign syntax element and / or the fourth sign syntax element are added in the supplemental enhancement information, and a value of the third sign syntax element is set according to a sign of a normalized difference value of a polar radius of the common pupil position information, and a value of the fourth sign syntax element is set according to a sign of a normalized difference value of a polar angle of the common pupil position information.
[0325] In some embodiments, the third sign syntax element is “gfv_pupil_polar_delta_radius_sign_flag”, and the fourth sign syntax element is “gfv_pupil_polar_delta_angle_sign_flag”.
[0326] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0327] In a case where the left eye and the right eye share the same pupil position information, a coordinate system to which the pupil position information belongs is a polar coordinate system, and a normalized value is used to represent a case where the common pupil position information is shared, a seventh position information syntax element and an eighth position information syntax element are added in the supplemental enhancement information, a value of the seventh position information syntax element is set as a normalized value of a polar radius of the common pupil position information, and a value of the eighth position information syntax element is set as a normalized value of a polar angle of the common pupil position information.
[0328] In some embodiments, the seventh position information syntax element is “gfv_pupil_x_coordinate”, and the eighth position information syntax element is “gfv_pupil_y_coordinate”.
[0329] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0330] In a case where the left eye and the right eye do not share the same pupil position information, a fourth identification syntax element for indicating whether the left eye pupil position information exists and a fifth identification syntax element for indicating whether the right eye pupil position information exists are added in the supplemental enhancement information, and a value of the fourth identification syntax element is set according to whether the left eye pupil position information exists, and a value of the fifth identification syntax element is set according to whether the right eye pupil position information exists.
[0331] In some embodiments, the fourth identification syntax element is “gfv_pupil_left_eye_present_flag”, and the fourth identification syntax element is “gfv_pupil_right_eye_present_flag”.
[0332] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0333] In the case that the left eye pupil position information exists, a second coordinate system syntax element for representing a coordinate system to which the left eye pupil position information belongs is added in the supplemental enhancement information, and a value of the second coordinate system syntax element is set according to a coordinate system constructed when the left eye pupil position information is acquired.
[0334] In some embodiments, the value of the second coordinate system syntax element is “gfv_pupil_left_eye_coordinate_system_idx”.
[0335] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0336] In the case that the left eye pupil position information exists and a coordinate system to which the left eye pupil position information belongs is a rectangular coordinate system, a fifth bit depth syntax element for representing a bit depth of an x coordinate and a sixth bit depth syntax element for representing a bit depth of a y coordinate are added in the supplemental enhancement information, a value of the fifth bit depth syntax element is set according to the bit depth of the x coordinate, and a value of the sixth bit depth syntax element is set according to the bit depth of the y coordinate.
[0337] In some embodiments, the fifth bit depth syntax element is “gfv_pupil_left_eye_x_coordinate_precision_factor_minus1”, and the sixth bit depth syntax element is “gfv_pupil_left_eye_y_coordinate_precision_factor_minus1”.
[0338] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0339] In the case that the left eye pupil position information exists, a sixth identification syntax element for representing a representation manner of the left eye pupil position information is added in the supplemental enhancement information, and in the case that the left eye pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, a value of the sixth identification syntax element is set as a first preset value; in the case that the left eye pupil position information is represented by a normalized value, the value of the sixth identification syntax element is set as a second preset value.
[0340] In some embodiments, the sixth identification syntax element is “gfv_pupil_left_eye_pred_flag”.
[0341] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0342] In the case that the left eye pupil position information exists, the coordinate system to which the left eye pupil position information belongs is a rectangular coordinate system, and the left eye pupil position information is represented by a normalized absolute difference value with the pupil position information of the base image, the ninth position information syntax element and the tenth position information syntax element are added in the supplemental enhancement information, the value of the ninth position information syntax element is set as the normalized absolute difference value of the x-axis coordinate of the left eye pupil position information, and the value of the tenth position information syntax element is set as the normalized absolute difference value of the y-axis coordinate of the left eye pupil position information.
[0343] In some embodiments, the ninth position information syntax element is “gfv_pupil_left_eye_dx_coordinate_abs”, and the tenth position information syntax element is “gfv_pupil_left_eye_dy_coordinate_abs”.
[0344] In some embodiments, the generating the supplemental enhancement information according to the pupil position information further comprises:
[0345] The fifth sign syntax element and / or the sixth sign syntax element are added in the supplemental enhancement information, the value of the fifth sign syntax element is set according to the sign of the normalized difference value of the x-axis coordinate of the left eye pupil position information, and the value of the sixth sign syntax element is set according to the sign of the normalized difference value of the y-axis coordinate of the left eye pupil position information.
[0346] In some embodiments, the fifth sign syntax element is “gfv_pupil_left_eye_dx_coordinate_sign_flag”, and the fifth sign syntax element is “gfv_pupil_left_eye_dy_coordinate_sign_flag”.
[0347] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0348] In the case that the left eye pupil position information exists, the coordinate system to which the left eye pupil position information belongs is a rectangular coordinate system, and the left eye pupil position information is represented by a normalized value, the eleventh position information syntax element and the twelfth position information syntax element are added in the supplemental enhancement information, the value of the eleventh position information syntax element is set as the normalized value of the x-axis coordinate of the left eye pupil position information, and the value of the twelfth position information syntax element is set as the normalized value of the y-axis coordinate of the left eye pupil position information.
[0349] In some embodiments, the eleventh position information is "gfv_pupil_left_eye_x_coordinate", and the twelfth position information is "gfv_pupil_left_eye_y_coordinate".
[0350] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0351] In a case where the left eye pupil position information exists and a coordinate system to which the left eye pupil position information belongs is a polar coordinate system, a seventh depth syntax element for representing a bit depth of a polar radius and an eighth depth syntax element for representing a bit depth of a polar angle are added in the supplemental enhancement information, and a value of the seventh depth syntax element is set according to the bit depth of the polar radius, and a value of the eighth depth syntax element is set according to the bit depth of the polar angle.
[0352] In some embodiments, the seventh depth syntax element is "gfv_pupil_left_eye_polar_radius_precision_factor_minus1", and the eighth depth syntax element is "gfv_pupil_left_eye_polar_angle_precision_factor_minus1".
[0353] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0354] In a case where the left eye pupil position information exists, a coordinate system to which the left eye pupil position information belongs is a polar coordinate system, and the left eye pupil position information is represented by a normalized absolute difference value from the pupil position information of the base image, a thirteenth position information syntax element and a fourteenth position information syntax element are added in the supplemental enhancement information, a value of the thirteenth position information syntax element is set as a normalized absolute difference value of the polar radius of the left eye pupil position information, and a value of the fourteenth position information syntax element is set as a normalized absolute difference value of the polar angle of the left eye pupil position information.
[0355] In some embodiments, the thirteenth position information syntax element is "gfv_pupil_left_eye_polar_delta_radius_abs", and the fourteenth position information syntax element is "gfv_pupil_left_eye_polar_delta_angle_abs.
[0356] In some embodiments, the generating the supplemental enhancement information according to the pupil position information further comprises:
[0357] The seventh and eighth sign syntax elements are added in the supplemental enhancement information, and a value of the seventh sign syntax element is set according to a sign of a normalized difference value of a polar radius of the left eye pupil position information, and a value of the eighth sign syntax element is set according to a sign of a normalized difference value of a polar angle of the left eye pupil position information.
[0358] In some embodiments, the seventh sign syntax element is “gfv_pupil_left_eye_polar_delta_radius_sign_flag”, and the eighth sign syntax element is “gfv_pupil_left_eye_polar_delta_angle_sign_flag”.
[0359] In some embodiments, the generating supplemental enhancement information according to the pupil position information comprises:
[0360] In a case where the left eye pupil position information exists, a coordinate system to which the left eye pupil position information belongs is a polar coordinate system, and the left eye pupil position information is represented by a normalized value, a fifteenth position information syntax element and a sixteenth position information syntax element are added in the supplemental enhancement information, a value of the fifteenth position information syntax element is set as a normalized value of a polar radius of the left eye pupil position information, and a value of the sixteenth position information syntax element is set as a normalized value of a polar angle of the left eye pupil position information.
[0361] In some embodiments, the fifteenth position information syntax element is “gfv_pupil_left_eye_polar_radius”, and the sixteenth position information syntax element is “gfv_pupil_left_eye_polar_angle”.
[0362] In some embodiments, the generating supplemental enhancement information according to the pupil position information comprises:
[0363] In a case where the right eye pupil position information exists, a third coordinate system syntax element for representing a coordinate system to which the right eye pupil position information belongs is added in the supplemental enhancement information, and a value of the third coordinate system syntax element is set according to a coordinate system constructed when the right eye pupil position information is acquired.
[0364] In some embodiments, the third coordinate system syntax element is “gfv_pupil_right_eye_coordinate_system_idx”.
[0365] In some embodiments, the generating supplemental enhancement information according to the pupil position information comprises:
[0366] In a case where the right eye pupil position information exists and a coordinate system to which the right eye pupil position information belongs is a rectangular coordinate system, a ninth depth syntax element for representing a bit depth of an x coordinate and a tenth depth syntax element for representing a bit depth of a y coordinate are added in the supplemental enhancement information, and a value of the ninth depth syntax element is set according to the bit depth of the x coordinate, and a value of the tenth depth syntax element is set according to the bit depth of the y coordinate.
[0367] In some embodiments, the value of the ninth depth syntax element is “gfv_pupil_right_eye_x_coordinate_precision_factor_minus1”, and the value of the tenth depth syntax element is “gfv_pupil_right_eye_y_coordinate_precision_factor_minus1”.
[0368] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0369] In a case where the right eye pupil position information exists, a seventh identification syntax element for representing a representation manner of the right eye pupil position information is added in the supplemental enhancement information, and in a case where the right eye pupil position information is represented by a normalized absolute difference value from pupil position information of a base image, a value of the seventh identification syntax element is set to a first preset value; and in a case where the right eye pupil position information is represented by a normalized value, the value of the seventh identification syntax element is set to a second preset value.
[0370] In some embodiments, the seventh identification syntax element is “gfv_pupil_right_eye_pred_flag”.
[0371] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0372] In a case where the right eye pupil position information exists, a coordinate system to which the right eye pupil position information belongs is a rectangular coordinate system, and the right eye pupil position information is represented by a normalized absolute difference value from pupil position information of a base image, a seventeenth position information syntax element and an eighteenth position information syntax element are added in the supplemental enhancement information, and a value of the seventeenth position information syntax element is set to a normalized absolute difference value of an x-axis coordinate of the right eye pupil position information, and a value of the eighteenth position information syntax element is set to a normalized absolute difference value of a y-axis coordinate of the right eye pupil position information.
[0373] In some embodiments, the seventeenth position information syntax element is "gfv_pupil_right_eye_dx_coordinate_abs", and the eighteenth position information syntax element is "gfv_pupil_right_eye_dy_coordinate_abs".
[0374] In some embodiments, the generating the supplemental enhancement information according to the pupil position information further comprises:
[0375] In the supplemental enhancement information, a ninth sign syntax element and / or a tenth sign syntax element are added, and a value of the ninth sign syntax element is set according to a sign of a normalized difference value of an x-axis coordinate of the right eye pupil position information, and a value of the tenth sign syntax element is set according to a sign of a normalized difference value of a y-axis coordinate of the right eye pupil position information.
[0376] In some embodiments, the ninth sign syntax element is "gfv_pupil_right_eye_dx_coordinate_sign_flag", and the tenth sign syntax element is "gfv_pupil_right_eye_dy_coordinate_sign_flag".
[0377] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0378] In a case where the right eye pupil position information exists, a coordinate system to which the right eye pupil position information belongs is a rectangular coordinate system, and the right eye pupil position information is represented by a normalized value, in the supplemental enhancement information, a nineteenth position information syntax element and a twentieth position information syntax element are added, and a value of the nineteenth position information syntax element is set as a normalized value of an x-axis coordinate of the right eye pupil position information, and a value of the twentieth position information syntax element is set as a normalized value of a y-axis coordinate of the right eye pupil position information.
[0379] In some embodiments, the nineteenth position information syntax element is "gfv_pupil_right_eye_x_coordinate", and the twentieth position information syntax element is "gfv_pupil_right_eye_y_coordinate".
[0380] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0381] In a case where the right eye pupil position information exists and a coordinate system to which the right eye pupil position information belongs is a polar coordinate system, a eleventh depth syntax element for representing a bit depth of a polar radius and a twelfth depth syntax element for representing a bit depth of a polar angle are added in the supplemental enhancement information, and a value of the eleventh depth syntax element is set according to the bit depth of the polar radius, and a value of the twelfth depth syntax element is set according to the bit depth of the polar angle.
[0382] In some embodiments, the eleventh depth syntax element is “gfv_pupil_right_eye_polar_radius_precision_factor_minus1”, and the twelfth depth syntax element is “gfv_pupil_right_eye_polar_angle_precision_factor_minus1”.
[0383] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0384] In a case where the right eye pupil position information exists, a coordinate system to which the right eye pupil position information belongs is a polar coordinate system, and the pupil position information is represented by a normalized absolute difference value from the pupil position information of the base image, a twenty-first position information syntax element and a twenty-second position information syntax element are added in the supplemental enhancement information, a value of the twenty-first position information syntax element is set as a normalized absolute difference value of a polar radius of the right eye pupil position information, and a value of the twenty-second position information syntax element is set as a normalized absolute difference value of a polar angle of the right eye pupil position information.
[0385] In some embodiments, the twenty-first position information syntax element is “gfv_pupil_right_eye_polar_delta_radius_abs”, and the twenty-second position information syntax element is “gfv_pupil_right_eye_polar_delta_angle_abs”.
[0386] In some embodiments, the generating the supplemental enhancement information according to the pupil position information further comprises:
[0387] A eleventh sign syntax element and / or a twelfth sign syntax element are added in the supplemental enhancement information, a value of the eleventh sign syntax element is set according to a sign of the normalized difference value of the polar radius of the right eye pupil position information, and a value of the twelfth sign syntax element is set according to a sign of the normalized difference value of the polar angle of the right eye pupil position information.
[0388] In some embodiments, the eleventh syntax element is gfv_pupil_right_eye_polar_delta_radius_sign_flag, and the twelfth syntax element is gfv_pupil_right_eye_polar_delta_angle_sign_flag.
[0389] In some embodiments, the generating the supplemental enhancement information according to the pupil position information comprises:
[0390] In the case that the right eye pupil position information exists, the coordinate system to which the right eye pupil position information belongs is a polar coordinate system, and the right eye pupil position information is represented by a normalized value, the twenty-third position information syntax element and the twenty-fourth position information syntax element are added in the supplemental enhancement information, the value of the twenty-third position information syntax element is set as the normalized value of the polar radius of the right eye pupil position information, and the value of the twenty-fourth position information syntax element is set as the normalized value of the polar angle of the right eye pupil position information.
[0391] In some embodiments, the twenty-third position information syntax element is gfv_pupil_right_eye_polar_radius, and the twenty-fourth position information syntax element is gfv_pupil_right_eye_polar_angle.
[0392] The embodiments of the present application can also provide an image decoding method. As shown in FIG. 11, the image decoding method comprises the following steps: Figure 11
[0393] S111, obtaining the encoding data of a current video frame.
[0394] S112, determining whether the current video frame is a base image according to the encoding data of the current video frame.
[0395] In some embodiments, whether the current video frame is a base image can be determined by the value of a preset position in the encoding data of the current video frame. For example, if the value of the preset position is 1, it is determined that the current video frame is a base image, and if the value of the preset position is 0, it is determined that the current video frame is not a base image.
[0396] In the step S112, if it is determined that the current video frame is a base image, the following step S113 is performed:
[0397] S113, decoding the encoding data of the current video frame based on a preset video decoding standard to obtain a reconstructed current video frame.
[0398] The preset video decoding standard can be H.266 / VVC, H.265 / HEVC, etc.
[0399] In step S112, if it is determined that the current video frame is a base image, steps S114 and S115 are performed.
[0400] In step S114, the encoded data of the current video frame is decoded to obtain the image features corresponding to the current video frame.
[0401] The image features corresponding to the current video frame include the pupil position information of the current video frame.
[0402] In step S115, the current video is reconstructed according to the image features corresponding to the current video frame to obtain the reconstructed current video frame.
[0403] The implementation of decoding the encoded data of the current video frame to obtain the image features of the current video frame in step S114 corresponds to the implementation of encoding the image features of the current video frame to generate the encoded data of the current video frame. Those skilled in the art should understand that decoding the encoded data of the current video frame can obtain the image features including the pupil position information of the current video frame, for example, the second identification syntax element can be decoded to determine whether the left eye and the right eye share the same set of pupil position information. For another example, the value of the first coordinate system syntax element can be decoded to determine the coordinate system to which the shared pupil position information belongs. Without avoiding repetition, the above-mentioned implementations are not repeated here.
[0404] In some embodiments, some embodiments of the present application provide a video encoding apparatus, comprising:
[0405] a memory configured to store a computer program;
[0406] a processor configured to, when the computer program is invoked, cause the video encoding apparatus to implement the video encoding method of any of the above-mentioned embodiments.
[0407] In some embodiments, some embodiments of the present application provide a video decoding apparatus, comprising:
[0408] a memory configured to store a computer program;
[0409] a processor configured to, when the computer program is invoked, cause the video decoding apparatus to implement the video decoding method of any of the above-mentioned embodiments.
[0410] In some embodiments, some embodiments of the present application provide a computer readable storage medium, having stored thereon a computer program, which, when executed by a computing device, causes the computing device to implement the video encoding method or the video decoding method according to any of the above embodiments.
[0411] In some embodiments, some embodiments of the present application provide a computer program product, which, when running on a computer, causes the computer to implement the video encoding method or the video decoding method according to any of the above embodiments.
[0412] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0413] In order to facilitate explanation, the above description has been made in combination with specific embodiments. However, the above exemplary discussion is not intended to exhaust or limit the embodiments to the specific forms disclosed above. Various modifications and variations can be derived according to the above teachings. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A method of video coding, the method comprising: The method comprises the following steps: determining whether the current video frame is a base image; in the case that the current video frame is a base image, encoding the current video frame based on a preset video coding standard to obtain the encoding data of the current video frame; in the case that the current video frame is not a base image, performing feature extraction on the current video frame to obtain the image features corresponding to the current video frame, and encoding the image features corresponding to the current video frame to obtain the encoding data of the current video frame; wherein the image features corresponding to the current video frame include the pupil position information of the current video frame.
2. The method of claim 1, wherein, The feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises: constructing a first rectangular coordinate system with the left corner of the eye as the origin, the horizontal axis as the x-axis, and the vertical axis as the y-axis; obtaining the distance from the geometric center of the pupil to the y-axis to obtain the x-coordinate of the pupil in the first rectangular coordinate system; obtaining the distance from the geometric center of the pupil to the x-axis to obtain the y-coordinate of the pupil in the first rectangular coordinate system; generating the pupil position information according to the x-coordinate and the y-coordinate of the pupil in the first rectangular coordinate system.
3. The method of claim 1, wherein, The feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises: constructing a second rectangular coordinate system with the right corner of the eye as the origin, the horizontal axis as the x-axis, and the vertical axis as the y-axis; obtaining the distance from the geometric center of the pupil to the y-axis to obtain the x-coordinate of the pupil in the second rectangular coordinate system; obtaining the distance from the geometric center of the pupil to the x-axis to obtain the y-coordinate of the pupil in the second rectangular coordinate system; generating the pupil position information according to the x-coordinate and the y-coordinate of the pupil in the second rectangular coordinate system.
4. The method of claim 1, wherein, The feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises: constructing a third rectangular coordinate system with the left corner of the eye as the origin, the line connecting the left corner and the right corner of the eye as the x-axis, and the axis passing through the origin and perpendicular to the x-axis as the y-axis; obtaining the distance from the geometric center of the pupil to the y-axis to obtain the x-coordinate of the pupil in the third rectangular coordinate system; obtaining the distance from the geometric center of the pupil to the x-axis to obtain the y-coordinate of the pupil in the third rectangular coordinate system; generating the pupil position information according to the x-coordinate and the y-coordinate of the pupil in the third rectangular coordinate system.
5. The method of claim 1, wherein, The feature extraction on the current video frame to obtain the image features corresponding to the current video frame comprises: constructing a fourth rectangular coordinate system with the right corner of the eye as the origin, the line connecting the right corner and the left corner of the eye as the x-axis, and the axis passing through the origin and perpendicular to the x-axis as the y-axis; obtaining the distance from the geometric center of the pupil to the y-axis to obtain the x-coordinate of the pupil in the fourth rectangular coordinate system; obtaining the distance from the geometric center of the pupil to the x-axis to obtain the y-coordinate of the pupil in the fourth rectangular coordinate system; generating the pupil position information according to the x-coordinate and the y-coordinate of the pupil in the fourth rectangular coordinate system.
6. The method of claim 1, wherein, The feature extraction on the current video frame is performed to obtain image features corresponding to the current video frame, and the feature extraction includes: A first polar coordinate system is constructed with the left corner of the eye as a pole and a horizontal axis as a polar axis; A distance from the geometric center of the pupil to the pole is obtained to obtain a polar radius of the pupil in the first polar coordinate system; An angle between a line connecting the geometric center of the pupil and the pole and the polar axis is obtained to obtain a polar angle of the pupil in the first polar coordinate system; The pupil position information is generated according to the polar radius and the polar angle of the pupil in the first polar coordinate system.
7. The method of claim 1, wherein, The feature extraction on the current video frame is performed to obtain image features corresponding to the current video frame, and the feature extraction includes: A second polar coordinate system is constructed with the right corner of the eye as a pole and a horizontal axis as a polar axis; A distance from the geometric center of the pupil to the pole is obtained to obtain a polar radius of the pupil in the second polar coordinate system; An angle between a line connecting the geometric center of the pupil and the pole and the polar axis is obtained to obtain a polar angle of the pupil in the second polar coordinate system; The pupil position information is generated according to the polar radius and the polar angle of the pupil in the second polar coordinate system.
8. The method of claim 1, wherein, The feature extraction on the current video frame is performed to obtain image features corresponding to the current video frame, and the feature extraction includes: A third polar coordinate system is constructed with the left corner of the eye as a pole and a line connecting the left corner and the right corner of the eye as a polar axis; A distance from the geometric center of the pupil to the pole is obtained to obtain a polar radius of the pupil in the third polar coordinate system; An angle between a line connecting the geometric center of the pupil and the pole and the polar axis is obtained to obtain a polar angle of the pupil in the third polar coordinate system; The pupil position information is generated according to the polar radius and the polar angle of the pupil in the third polar coordinate system.
9. The method of claim 1, wherein, The feature extraction on the current video frame is performed to obtain image features corresponding to the current video frame, and the feature extraction includes: A fourth polar coordinate system is constructed with the right corner of the eye as a pole and a line connecting the right corner and the left corner of the eye as a polar axis; A distance from the geometric center of the pupil to the pole is obtained to obtain a polar radius of the pupil in the fourth polar coordinate system; An angle between a line connecting the geometric center of the pupil and the pole and the polar axis is obtained to obtain a polar angle of the pupil in the fourth polar coordinate system; The pupil position information is generated according to the polar radius and the polar angle of the pupil in the fourth polar coordinate system.
10. The method of claim 2 or 3, wherein, The feature extraction on the current video frame is performed to obtain image features corresponding to the current video frame, and the feature extraction includes: The x-axis is normalized according to the length of the projection of the eye on the x-axis, and the x-coordinate of the pupil is quantized with a quantization step of 1 / 32 or 1 / 16 or 1 / 8; The y-axis is normalized according to the length of the projection of the eye on the y-axis, and the y-coordinate of the pupil is quantized with a quantization step of 1 / 16 or 1 / 8.
11. The method of claim 4 or 5, wherein, The feature extraction on the current video frame is performed to obtain image features corresponding to the current video frame, and the feature extraction includes: Normalizing the x-axis with respect to the length of the line connecting the left corner of the eye to the right corner of the eye, and quantizing the x-coordinate of the pupil with a quantization step of 1 / 32 or 1 / 16 or 1 / 8; Normalizing the y-axis with respect to the length of the projection of the eye on the y-axis, and quantizing the y-coordinate of the pupil with a quantization step of 1 / 16 or 1 / 8.
12. The method according to any one of claims 2 to 5, characterized in that, The encoding of the image features corresponding to the current video frame to obtain the encoded data of the current video frame comprises: Normalizing the x-axis with respect to the length of the line connecting the left corner of the eye to the right corner of the eye, and quantizing the x-coordinate of the pupil with a quantization step of 1 / 16 or 1 / 8; Normalizing the y-axis with respect to the length of the line connecting the left corner of the eye to the right corner of the eye, and quantizing the y-coordinate of the pupil with a quantization step of 1 / 16 or 1 / 8.
13. The method according to any one of claims 6 to 9, characterized in that, The encoding of the image features corresponding to the current video frame to obtain the encoded data of the current video frame comprises: Normalizing the polar axis with respect to the length of the line connecting the left corner of the eye to the right corner of the eye, and quantizing the polar radius of the pupil with a quantization step of 1 / 32 or 1 / 16 or 1 / 8; Normalizing the polar angle with respect to the angle between the line connecting the highest point of the eye and the polar point and the line connecting the lowest point of the eye and the polar point, and quantizing the polar angle of the pupil with a quantization step of 1 / 16 or 1 / 8.
14. The method of claim 1, wherein, The encoding of the image features corresponding to the current video frame to obtain the encoded data of the current video frame comprises: Generating supplemental enhancement information SEI according to the pupil position information; Entropy encoding the supplemental enhancement information to obtain the encoded data corresponding to the pupil position information.
15. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: Adding a first identification syntax element for indicating whether the pupil position information exists or not in the supplemental enhancement information, and setting the value of the first identification syntax element to a first preset value.
16. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: Adding a second identification syntax element for indicating whether the left eye and the right eye share the same pupil position information in the supplemental enhancement information, and setting the value of the second identification syntax element to a first preset value in the case that the left eye and the right eye share the same pupil position information, and setting the value of the second identification syntax element to a second preset value in the case that the left eye and the right eye do not share the same pupil position information.
17. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: In the case that the left eye and the right eye share the same pupil position information, adding a first coordinate system syntax element for indicating the coordinate system to which the shared pupil position information belongs in the supplemental enhancement information, and setting the value of the first coordinate system syntax element according to the coordinate system to which the shared pupil position information belongs.
18. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye and the right eye share the same pupil position information and a coordinate system to which the pupil position information belongs is a rectangular coordinate system, a first depth syntax element for representing a bit depth of an x coordinate and a second depth syntax element for representing a bit depth of a y coordinate are added in the supplemental enhancement information, and a value of the first depth syntax element is set according to the bit depth of the x coordinate, and a value of the second depth syntax element is set according to the bit depth of the y coordinate.
19. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye and the right eye share the same pupil position information, a third identification syntax element for representing a representation manner of the shared pupil position information is added in the supplemental enhancement information, and in a case where the shared pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, a value of the third identification syntax element is set as a first preset value; and in a case where the shared pupil position information is represented by a normalized value, the value of the third identification syntax element is set as a second preset value.
20. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye and the right eye share the same pupil position information, a coordinate system to which the shared pupil position information belongs is a rectangular coordinate system, and the shared pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, a first position information syntax element and a second position information syntax element are added in the supplemental enhancement information, and a value of the first position information syntax element is set as a normalized absolute difference value of an x axis coordinate of the shared pupil position information, and a value of the second position information syntax element is set as a normalized absolute difference value of a y axis coordinate of the shared pupil position information.
21. The method of claim 20, wherein, The generating the supplemental enhancement information according to the pupil position information further comprises: A first sign syntax element and / or a second sign syntax element are added in the supplemental enhancement information, and a value of the first sign syntax element is set according to a sign of the normalized difference value of the x axis coordinate of the shared pupil position information, and a value of the second sign syntax element is set according to a sign of the normalized difference value of the y axis coordinate of the shared pupil position information.
22. The method of claim 14, wherein, In a case where the left eye and the right eye share the same pupil position information, a coordinate system to which the shared pupil position information belongs is a rectangular coordinate system, and the shared pupil position information is represented by a normalized value, a third position information syntax element and a fourth position information syntax element are added in the supplemental enhancement information, and a value of the third position information syntax element is set as a normalized value of an x axis coordinate of the shared pupil position information, and a value of the fourth position information syntax element is set as a normalized value of a y axis coordinate of the shared pupil position information. The generating the supplemental enhancement information according to the pupil position information comprises:
23. The method of claim 14, wherein, In the case that the left eye and the right eye share the same pupil position information and the coordinate system to which the pupil position information belongs is a polar coordinate system, a third bit depth syntax element for representing the bit depth of the polar radius and a fourth bit depth syntax element for representing the bit depth of the polar angle are added in the supplemental enhancement information, and the value of the third bit depth syntax element is set according to the bit depth of the polar radius, and the value of the fourth bit depth syntax element is set according to the bit depth of the polar angle.
24. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: In the case that the left eye and the right eye share the same pupil position information, the coordinate system to which the pupil position information belongs is a polar coordinate system, and the shared pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, a fifth position information syntax element and a sixth position information syntax element are added in the supplemental enhancement information, the value of the fifth position information syntax element is set as the normalized absolute difference value of the polar radius of the shared pupil position information, and the value of the sixth position information syntax element is set as the normalized absolute difference value of the polar angle of the shared pupil position information.
25. The method of claim 24, wherein, The generating of the supplemental enhancement information according to the pupil position information further comprises: A third sign syntax element and / or a fourth sign syntax element are added in the supplemental enhancement information, the value of the third sign syntax element is set according to the sign of the normalized difference value of the polar radius of the shared pupil position information, and the value of the fourth sign syntax element is set according to the sign of the normalized difference value of the polar angle of the shared pupil position information.
26. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: In the case that the left eye and the right eye share the same pupil position information, the coordinate system to which the pupil position information belongs is a polar coordinate system, and the shared pupil position information is represented by a normalized value, a seventh position information syntax element and an eighth position information syntax element are added in the supplemental enhancement information, the value of the seventh position information syntax element is set as the normalized value of the polar radius of the shared pupil position information, and the value of the eighth position information syntax element is set as the normalized value of the polar angle of the shared pupil position information.
27. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: In the case that the left eye and the right eye do not share the same pupil position information, a fourth identification syntax element for representing whether the left eye pupil position information exists and a fifth identification syntax element for representing whether the right eye pupil position information exists are added in the supplemental enhancement information, the value of the fourth identification syntax element is set according to whether the left eye pupil position information exists, and the value of the fifth identification syntax element is set according to whether the right eye pupil position information exists.
28. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: In the case that the left eye pupil position information exists, a second coordinate system syntax element for representing the coordinate system to which the left eye pupil position information belongs is added in the supplemental enhancement information, and the value of the second coordinate system syntax element is set according to the coordinate system constructed when the left eye pupil position information is acquired.
29. The method of claim 14, wherein, The generating of the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye pupil position information exists and a coordinate system to which the left eye pupil position information belongs is a rectangular coordinate system, a fifth depth syntax element for representing a bit depth of an x coordinate and a sixth depth syntax element for representing a bit depth of a y coordinate are added in the supplemental enhancement information, and a value of the fifth depth syntax element is set according to the bit depth of the x coordinate, and a value of the sixth depth syntax element is set according to the bit depth of the y coordinate.
30. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye pupil position information exists, a sixth identification syntax element for representing a representation manner of the left eye pupil position information is added in the supplemental enhancement information, and in a case where the left eye pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, a value of the sixth identification syntax element is set as a first preset value; in the case where the left eye pupil position information is represented by the normalized value, the value of the sixth identification syntax element is set as a second preset value.
31. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye pupil position information exists, a coordinate system to which the left eye pupil position information belongs is a rectangular coordinate system, and the left eye pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, a ninth position information syntax element and a tenth position information syntax element are added in the supplemental enhancement information, a value of the ninth position information syntax element is set as a normalized absolute difference value of an x axis coordinate of the left eye pupil position information, and a value of the tenth position information syntax element is set as a normalized absolute difference value of a y axis coordinate of the left eye pupil position information.
32. The method of claim 31, wherein, The generating the supplemental enhancement information according to the pupil position information further comprises: A fifth sign syntax element and / or a sixth sign syntax element are added in the supplemental enhancement information, a value of the fifth sign syntax element is set according to a sign of a normalized difference value of an x axis coordinate of the left eye pupil position information, and a value of the sixth sign syntax element is set according to a sign of a normalized difference value of a y axis coordinate of the left eye pupil position information.
33. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye pupil position information exists, a coordinate system to which the left eye pupil position information belongs is a rectangular coordinate system, and the left eye pupil position information is represented by a normalized value, an eleventh position information syntax element and a twelfth position information syntax element are added in the supplemental enhancement information, a value of the eleventh position information syntax element is set as a normalized value of an x axis coordinate of the left eye pupil position information, and a value of the twelfth position information syntax element is set as a normalized value of a y axis coordinate of the left eye pupil position information.
34. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye pupil position information exists and a coordinate system to which the left eye pupil position information belongs is a polar coordinate system, a seventh bit depth syntax element for representing a bit depth of a polar radius and an eighth bit depth syntax element for representing a bit depth of a polar angle are added in the supplemental enhancement information, and a value of the seventh bit depth syntax element is set according to the bit depth of the polar radius, and a value of the eighth bit depth syntax element is set according to the bit depth of the polar angle.
35. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye pupil position information exists, a coordinate system to which the left eye pupil position information belongs is a polar coordinate system, and the left eye pupil position information is represented by a normalized absolute difference value from the pupil position information of the base image, a thirteenth position information syntax element and a fourteenth position information syntax element are added in the supplemental enhancement information, a value of the thirteenth position information syntax element is set as a normalized absolute difference value of the polar radius of the left eye pupil position information, and a value of the fourteenth position information syntax element is set as a normalized absolute difference value of the polar angle of the left eye pupil position information.
36. The method of claim 35, wherein, The generating the supplemental enhancement information according to the pupil position information further comprises: A seventh sign syntax element and / or an eighth sign syntax element are added in the supplemental enhancement information, a value of the seventh sign syntax element is set according to a sign of the normalized difference value of the polar radius of the left eye pupil position information, and a value of the eighth sign syntax element is set according to a sign of the normalized difference value of the polar angle of the left eye pupil position information.
37. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the left eye pupil position information exists, a coordinate system to which the left eye pupil position information belongs is a polar coordinate system, and the left eye pupil position information is represented by a normalized value, a fifteenth position information syntax element and a sixteenth position information syntax element are added in the supplemental enhancement information, a value of the fifteenth position information syntax element is set as a normalized value of the polar radius of the left eye pupil position information, and a value of the sixteenth position information syntax element is set as a normalized value of the polar angle of the left eye pupil position information.
38. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the right eye pupil position information exists, a third coordinate system syntax element for representing a coordinate system to which the right eye pupil position information belongs is added in the supplemental enhancement information, and a value of the third coordinate system syntax element is set according to a coordinate system constructed when the right eye pupil position information is acquired.
39. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In a case where the right eye pupil position information exists and a coordinate system to which the right eye pupil position information belongs is a rectangular coordinate system, a ninth bit depth syntax element for representing a bit depth of an x coordinate and a tenth bit depth syntax element for representing a bit depth of a y coordinate are added in the supplemental enhancement information, a value of the ninth bit depth syntax element is set according to the bit depth of the x coordinate, and a value of the tenth bit depth syntax element is set according to the bit depth of the y coordinate.
40. The method of claim 14, wherein, The generating the supplemental enhancement information according to the pupil position information comprises: In the case that the right eye pupil position information exists, a seventh identification syntax element for representing a right eye pupil position information representation mode is added in the supplemental enhancement information, and in the case that the right eye pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, a value of the seventh identification syntax element is set as a first preset value; in the case that the right eye pupil position information is represented by a normalized value, the value of the seventh identification syntax element is set as a second preset value.
41. The method of claim 14, wherein, The generating supplemental enhancement information according to the pupil position information comprises: In the case that the right eye pupil position information exists, a coordinate system to which the right eye pupil position information belongs is a rectangular coordinate system, and the right eye pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, a seventeenth position information syntax element and an eighteenth position information syntax element are added in the supplemental enhancement information, a value of the seventeenth position information syntax element is set as a normalized absolute difference value of an x-axis coordinate of the right eye pupil position information, and a value of the eighteenth position information syntax element is set as a normalized absolute difference value of a y-axis coordinate of the right eye pupil position information.
42. The method of claim 41, wherein, The generating supplemental enhancement information according to the pupil position information further comprises: A ninth sign syntax element and / or a tenth sign syntax element are added in the supplemental enhancement information, a value of the ninth sign syntax element is set according to a sign of a normalized difference value of an x-axis coordinate of the right eye pupil position information, and a value of the tenth sign syntax element is set according to a sign of a normalized difference value of a y-axis coordinate of the right eye pupil position information.
43. The method of claim 14, wherein, The generating supplemental enhancement information according to the pupil position information comprises: In the case that the right eye pupil position information exists, a coordinate system to which the right eye pupil position information belongs is a rectangular coordinate system, and the right eye pupil position information is represented by a normalized value, a nineteenth position information syntax element and a twentieth position information syntax element are added in the supplemental enhancement information, a value of the nineteenth position information syntax element is set as a normalized value of an x-axis coordinate of the right eye pupil position information, and a value of the twentieth position information syntax element is set as a normalized value of a y-axis coordinate of the right eye pupil position information.
44. The method of claim 14, wherein, The generating supplemental enhancement information according to the pupil position information comprises: In the case that the right eye pupil position information exists, and a coordinate system to which the right eye pupil position information belongs is a polar coordinate system, an eleventh bit depth syntax element for representing a bit depth of a polar radius and a twelfth bit depth syntax element for representing a bit depth of a polar angle are added in the supplemental enhancement information, a value of the eleventh bit depth syntax element is set according to the bit depth of the polar radius, and a value of the twelfth bit depth syntax element is set according to the bit depth of the polar angle.
45. The method of claim 14, wherein, The generating supplemental enhancement information according to the pupil position information comprises: In the case that the right eye pupil position information exists, the coordinate system to which the right eye pupil position information belongs is a polar coordinate system, and the right eye pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, the twenty-first position information syntax element and the twenty-second position information syntax element are added in the supplemental enhancement information, the value of the twenty-first position information syntax element is set as the normalized absolute difference value of the polar radius of the right eye pupil position information, and the value of the twenty-second position information syntax element is set as the normalized absolute difference value of the polar angle of the right eye pupil position information.
46. The method of claim 45, wherein, The generating supplemental enhancement information according to the pupil position information further includes: The eleventh sign syntax element and / or the twelfth sign syntax element are added in the supplemental enhancement information, the value of the eleventh sign syntax element is set according to the sign of the normalized difference value of the polar radius of the right eye pupil position information, and the value of the twelfth sign syntax element is set according to the sign of the normalized difference value of the polar angle of the right eye pupil position information.
47. The method of claim 14, wherein, The generating supplemental enhancement information according to the pupil position information includes: In the case that the right eye pupil position information exists, the coordinate system to which the right eye pupil position information belongs is a polar coordinate system, and the right eye pupil position information is represented by a normalized absolute difference value of the pupil position information of the base image, the twenty-first position information syntax element and the twenty-second position information syntax element are added in the supplemental enhancement information, the value of the twenty-first position information syntax element is set as the normalized absolute difference value of the polar radius of the right eye pupil position information, and the value of the twenty-second position information syntax element is set as the normalized absolute difference value of the polar angle of the right eye pupil position information.
48. An image decoding method, comprising: It includes: Obtaining the encoding data of the current video frame; Determining whether the current video frame is a base image according to the encoding data of the current video frame; In the case that the current video frame is a base image, decoding the encoding data of the current video frame based on a preset video decoding standard to obtain a reconstructed current video frame; In the case that the current video frame is not a base image, decoding the encoding data of the current video frame, obtaining the image features corresponding to the current video frame, and performing reconstruction of the current video according to the image features corresponding to the current video frame to obtain a reconstructed current video frame; wherein the image features corresponding to the current video frame include pupil position information of the current video frame.
49. An apparatus for video encoding, the apparatus comprising: It includes: The memory is configured to store a computer program; The processor is configured to cause the video encoding device to implement the video encoding method in any one of claims 1-47 when the computer program is invoked.
50. An apparatus for video decoding, the apparatus comprising: It includes: The memory is configured to store a computer program; The processor is configured to cause the video decoding device to implement the video decoding method in claim 48 when the computer program is invoked.