Constraint Flag Indication in Video Bitstream
By using general_non_packed_constraint_flag signaling technology in the video decoder, it is solved by determining whether the VR video code stream contains fisheye omnidirectional video information, which is a problem that the decoder in the prior art cannot correctly present the fisheye omnidirectional video image, and the effect of reducing transmission bandwidth and decoding complexity is achieved.
Patent Information
- Application Number
- CN201980018932.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-21
- Filing Date
- 2019-02-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2039-02-19
AI Technical Summary
The prior art is difficult to effectively process the encoding and decoding of fisheye video information in virtual reality (VR) videos, resulting in the decoder being unable to correctly present the fisheye omnidirectional video image, affecting the user experience.
By using general_non_packed_constraint_flag signaling technology in the video decoder, it is determined whether the code stream contains fisheye omnidirectional video information. If not included, set flag to 1, and prohibit the code stream from including omnidirectional fisheye SEI messages; if possible, set flag to 0, allowing the decoder to search for omnidirectional fisheye SEI messages.
It realizes that while ensuring the video resolution and quality are unchanged, the transmission bandwidth and decoding complexity are reduced, the performance of the VR video system is improved, and unnecessary user experience problems are avoided.
Smart Images

Figure CN112262581B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 645,922, filed on Mar. 21, 2018, entitled "Signaling Of Omnidirectional Fisheye Video Property In A Video Bitstream" by Ye - Kui Wang, the entire content of which is incorporated herein by reference. Technical Field
[0003] The present invention generally relates to video coding, and more particularly to video coding in the context of virtual reality. Background Art
[0004] Virtual reality (VR) can be virtually displayed in a non - physical world, enabling a user to interact with that world. The non - physical world is created by presenting natural and / or synthetic images and sounds related to the movements of an immersed user. With recent advancements in display devices (e.g., head mounted displays (HMDs)) and VR video (commonly also referred to as 360 - degree video or omnidirectional video) production, a high - quality experience can be provided. VR applications include games, training, education, sports videos, online shopping, adult entertainment, and so on. Summary of the Invention
[0005] In one aspect, it relates to an encoding method performed by a video decoder. The method includes: a receiver in the video decoder receives a general_non_packed_constraint_flag; a processor in the video decoder determines that the value of the general_non_packed_constraint_flag is 1, where a value of 1 indicates that there is no supplemental enhancement information (SEI) message for fisheye video information in a coded video sequence (CVS).
[0006] According to the above aspect, in the first implementation of the method, it is determined that the value of the general_non_packed_constraint_flag in the active sequence parameter set (SPS) of the current layer is 1, where the value of 1 indicates that there should be no applicable fisheye video information SEI message in any image of the coded layer-wise video sequence (CLVS) of the current layer.
[0007] The first aspect relates to a method for processing video data. The method includes: determining whether the video decoder can correctly present a decoded image obtained by decoding the bitstream according to an indication attribute of the bitstream, where the bitstream includes an encoded representation of the video data, and the indication attribute includes a specific value indicating that the bitstream does not contain any fisheye omnidirectional video images; processing the bitstream according to the determination.
[0008] The second aspect relates to an encoding method performed by a video encoder. The method includes: encoding a flag and a representation of video data into a bitstream, where the flag includes a specific value indicating to a video decoder receiving the bitstream that the bitstream does not contain any fisheye omnidirectional video images; transmitting the bitstream to the video decoder.
[0009] The third aspect relates to a decoding method performed by a video decoder. The method includes: receiving an encoded bitstream, where the encoded bitstream includes a flag and a representation of video data, and the flag includes a specific value indicating to the video decoder that the bitstream does not contain any fisheye omnidirectional video images; decoding the encoded bitstream according to the specific value.
[0010] The above methods all contribute to the realization of signaling technology. When the value of general_non_packed_constraint_flag in the active sequence parameter set (SPS) is equal to 1, the bitstream is prohibited from including the omnidirectional fisheye supplemental enhancement information (SEI) message of the image. When the bitstream includes the omnidirectional fisheye SEI message of the image, the video encoder sets the value of general_non_packed_constraint_flag in the active SPS to 0. Therefore, the video decoder knows that any bitstream containing general_non_packed_constraint_flag equal to 1 does not contain any fisheye omnidirectional video images. When the video decoder receives a bitstream containing general_non_packed_constraint_flag equal to 0, the video decoder knows that the bitstream may contain an omnidirectional fisheye SEI message corresponding to one or more fisheye omnidirectional video images.
[0011] According to the first, second, or third aspect, in the first implementation manner of the method, the specific value is 1(1).
[0012] According to the first, second, or third aspect or any of the above implementation manners of the first, second, or third aspect, in the second implementation manner of the method, the flag is general_non_packed_constraint_flag.
[0013] According to the first, second, or third aspect or any of the above implementation manners of the first, second, or third aspect, in the third implementation manner of the method, the fisheye omnidirectional video image is an image captured by a fisheye camera.
[0014] According to the first, second, or third aspect or any of the above implementation manners of the first, second, or third aspect, in the fourth implementation manner of the method, the fisheye omnidirectional video image is an image including multiple active regions.
[0015] According to the first, second, or third aspect or any of the above implementation manners of the first, second, or third aspect, in the fifth implementation manner of the method, the flag is set in the sequence parameter set (SPS) of the bitstream.
[0016] The fourth aspect relates to a method for processing video data. The method includes: determining whether a video decoder can correctly present a decoded image obtained by decoding the bitstream according to an indication attribute of the bitstream, where the bitstream includes an encoded representation of the video data, and the indication attribute includes a specific value indicating that the bitstream may include one or more fisheye omnidirectional video images; processing the bitstream according to the determination.
[0017] The fifth aspect relates to an encoding method performed by a video encoder. The method includes: encoding a flag and a representation of video data into a bitstream, where the flag includes a specific value indicating to a video decoder receiving the bitstream that the bitstream may include one or more fisheye omnidirectional video images; transmitting the bitstream to the video decoder.
[0018] The sixth aspect relates to a decoding method performed by a video decoder. The method includes: receiving an encoded bitstream, where the encoded bitstream includes a flag and a representation of video data, and the flag includes a specific value indicating to the video decoder that the bitstream may include one or more fisheye omnidirectional video images; decoding the encoded bitstream according to the specific value.
[0019] The above methods all contribute to implementing signaling technology. When the value of general_non_packed_constraint_flag in the active SPS is equal to 1, it is prohibited for the bitstream to include an omnidirectional fisheye SEI message of an image. When the bitstream includes an omnidirectional fisheye SEI message of an image, the video encoder sets the value of general_non_packed_constraint_flag in the active SPS to be equal to 0. Therefore, the video decoder knows that any bitstream containing general_non_packed_constraint_flag equal to 1 does not contain any fisheye omnidirectional video images. When the video decoder receives a bitstream containing general_non_packed_constraint_flag equal to 0, the video decoder knows that the bitstream may include an omnidirectional fisheye SEI message corresponding to one or more fisheye omnidirectional video images.
[0020] According to the fourth, fifth, or sixth aspect, in a first implementation manner of the method, the specific value is 0(0).
[0021] According to the fourth, fifth, or sixth aspect or any of the above implementation manners of the fourth, fifth, or sixth aspect, in a second implementation manner of the method, the flag is general_non_packed_constraint_flag.
[0022] According to the fourth, fifth or sixth aspect or any of the above implementation manners of the fourth, fifth or sixth aspect, in the third implementation manner of the method, the omnidirectional fisheye video image is an image captured by a fisheye camera.
[0023] According to the fourth, fifth or sixth aspect or any of the above implementation manners of the fourth, fifth or sixth aspect, in the fourth implementation manner of the method, the omnidirectional fisheye video image is an image including a plurality of active regions.
[0024] According to the fourth, fifth or sixth aspect or any of the above implementation forms of the fourth, fifth or sixth aspect, in the fifth implementation form of the method, the flag is set in the sequence parameter set (SPS) of the bitstream.
[0025] The seventh aspect relates to an encoding device. The encoding device includes: a receiver for receiving an image for encoding or receiving a bitstream for decoding; a transmitter coupled to the receiver, wherein the transmitter is configured to transmit the bitstream to a decoder or transmit a decoded image to a display; a memory coupled to at least one of the receiver and the transmitter, wherein the memory is configured to store instructions; a processor coupled to the memory, wherein the processor is configured to execute the instructions stored in the memory to perform the method according to any one of the above aspects or implementation manners.
[0026] The encoding device helps to implement signaling technology. When the value of general_non_packed_constraint_flag in the active SPS is equal to 1, the bitstream is prohibited from including an omnidirectional fisheye SEI message of an image. When the bitstream includes an omnidirectional fisheye SEI message of an image, the video encoder sets the value of general_non_packed_constraint_flag in the active SPS to be equal to 0. Therefore, the video decoder knows that any bitstream including general_non_packed_constraint_flag equal to 1 does not include any omnidirectional fisheye video images. When the video decoder receives a bitstream including general_non_packed_constraint_flag equal to 0, the video decoder knows that the bitstream may include an omnidirectional fisheye SEI message corresponding to one or more omnidirectional fisheye video images.
[0027] According to the third aspect, in the first implementation manner of the device, the device further includes: a display for displaying an image.
[0028] The eighth aspect relates to a system. The system includes an encoder and a decoder communicatively coupled to the encoder. The encoder or the decoder includes an encoding device according to any one of the above aspects or implementations.
[0029] The system helps to implement signaling technology. When the value of general_non_packed_constraint_flag in the active SPS is equal to 1, the bitstream is prohibited from including an omnidirectional fisheye SEI message of an image. When the bitstream includes an omnidirectional fisheye SEI message of an image, the video encoder sets the value of general_non_packed_constraint_flag in the active SPS to 0. Thus, the video decoder knows that any bitstream containing general_non_packed_constraint_flag equal to 1 does not contain any fisheye omnidirectional video images. When the video decoder receives a bitstream containing general_non_packed_constraint_flag equal to 0, the video decoder knows that the bitstream may contain an omnidirectional fisheye SEI message corresponding to one or more fisheye omnidirectional video images.
[0030] The fifth aspect relates to an encoding component. The encoding component includes: a receiving component for receiving an image for encoding or receiving a bitstream for decoding; a transmission component coupled to the receiving component, wherein the transmission component is configured to transmit the bitstream to a decoder or transmit a decoded image to a display; a storage component coupled to at least one of the receiving component or the transmission component, wherein the storage component is configured to store instructions; and a processing component coupled to the storage component, wherein the processing component is configured to execute the instructions stored in the storage component to perform the method according to any one of the above aspects or implementations.
[0031] The encoding component helps to implement signaling technology. When the value of general_non_packed_constraint_flag in the active SPS is equal to 1, the bitstream is prohibited from including the omnidirectional fisheye SEI message of the image. When the bitstream includes the omnidirectional fisheye SEI message of the image, the video encoder sets the value of general_non_packed_constraint_flag in the active SPS to be equal to 0. Therefore, the video decoder knows that any bitstream containing general_non_packed_constraint_flag equal to 1 does not contain any fisheye omnidirectional video images. When the video decoder receives a bitstream containing general_non_packed_constraint_flag equal to 0, the video decoder knows that the bitstream may contain the omnidirectional fisheye SEI message corresponding to one or more fisheye omnidirectional video images.
[0032] The features disclosed herein can be used to improve the indication of omnidirectional fisheye video attributes in a video bitstream. The improved indication enhances the performance of the VR video system. For example, when ensuring that the resolution / quality of the video part presented to the user is the same, compared with the traditional VR video system, it reduces the transmission bandwidth and / or reduces the decoding complexity.
[0033] For clarity, any of the above embodiments can be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present invention.
[0034] These and other features will be more clearly understood from the following detailed description in conjunction with the drawings and the claims. Description of the Drawings
[0035] To more fully understand the present invention, reference is now made to the following brief description taken in conjunction with the drawings and the detailed description, in which like reference numerals represent like components.
[0036] Figure 1 Schematic diagram of an exemplary system for VR video encoding.
[0037] Figure 2 Flowchart of an exemplary method for encoding a VR image bitstream.
[0038] Figure 3 Flowchart of an exemplary method for encoding a video signal.
[0039] Figure 4 Schematic diagram of an exemplary encoding and decoding (codec) system for video encoding.
[0040] Figure 5 Schematic diagram of an exemplary video encoder.
[0041] Figure 6 Schematic diagram of an exemplary video decoder.
[0042] Figure 7 Schematic diagram of an example of a bitstream structure.
[0043] Figure 8 An embodiment of a method for processing video data.
[0044] Figure 9 An embodiment of an encoding method performed by a video encoder.
[0045] Figure 10 An embodiment of an encoding method performed by a video decoder.
[0046] Figure 11 An embodiment of a method for processing video data.
[0047] Figure 12 An embodiment of an encoding method performed by a video encoder.
[0048] Figure 13 An embodiment of an encoding method performed by a video decoder.
[0049] Figure 14 Schematic diagram of an exemplary video encoding device.
[0050] Figure 15 Schematic diagram of an embodiment of an encoding component. Detailed implementation
[0051] First, it should be understood that although the following provides illustrative implementations of one or more embodiments, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0052] Video coding standards include International Telecommunications Union Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding plus Depth (MVC+D), as well as three-dimension (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).
[0053] One characteristic that VR video is significantly different from standard video is that VR usually only displays a part of the entire video area represented by the video image, and this part corresponds to the current field of view (FOV) (i.e., the area that the user currently sees), while in standard video applications, usually the entire video area is displayed. FOV is sometimes also called the viewing angle. This characteristic can improve the performance of VR video systems through methods such as perspective-based projection mapping or perspective-based video coding. Compared with traditional VR video systems, improving performance can mean reducing the transmission bandwidth or reducing the decoding complexity or achieving both while ensuring that the resolution / quality of the video part presented to the user is the same.
[0054] The VR system can also use fisheye omnidirectional video instead of projected omnidirectional video (e.g., projecting a spherical video stream onto a rectangular sub-image video stream, as described below). In a VR system including fisheye omnidirectional video, the video is captured by a fisheye camera group, which includes a plurality of independent fisheye cameras. These fisheye cameras point in different directions and, ideally, cover all viewing directions around the camera group. The circular video images captured by the fisheye cameras are not stitched and projected on the encoder side, but are directly carried on a two-dimensional (2D) rectangular image at each time point. Other steps for video encoding, storage, transmission, and presentation are similar to those for projected omnidirectional video.
[0055] MPEG has recently developed the Omnidirectional Media Format (OMAF) standard. It is expected to be published as ISO / IEC International Standard 23090 Part 2. OMAF specifically describes the omnidirectional media format for encoding, storing, distributing, and presenting omnidirectional media including video, images, audio, and timed text. In an OMAF player, the user views the inner surface of the sphere from the center of the sphere outward. OMAF supports both projected omnidirectional video and fisheye omnidirectional video.
[0056] The indication of omnidirectional video metadata in the video bitstream is discussed. "HEVC Additional Supplemental Enhancement Information (Draft 4)" published by J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K. Wang (editors) in the output file JCTVC-AC1005 of the Joint Collaborative Team on Video Coding (JCT-VC) on October 24, 2017, specifically describes the latest amendment to HEVC. This HEVC amendment includes the specific description of some Supplemental Enhancement Information (SEI) messages, which are used to indicate the omnidirectional video metadata required for the correct presentation of omnidirectional video. Omnidirectional video is usually also called 360-degree video or VR video. Five types of SEI messages are specifically described in JCTVC-AC1005 to indicate omnidirectional video metadata: equirectangular projection SEI message, cube projection SEI message, sphere rotation SEI message, region wrapping SEI message, and omnidirectional view SEI message.
[0057] In JCTVC-AC1005, the semantics of the syntax element general_non_packed_constraint_flag are specifically described as follows:
[0058] When general_non_packed_constraint_flag equals 1, it indicates that there are no frame packing arrangement SEI messages, segmented rectangle frame packing arrangement SEI messages, equidistant cylindrical projection SEI messages, or cube projection SEI messages in the coded video sequence (CVS). When general_non_packed_constraint_flag equals 0, it indicates that there may or may not be frame packing arrangement SEI messages, segmented rectangle frame packing arrangement SEI messages, equidistant cylindrical projection SEI messages, or cube projection SEI messages in the CVS.
[0059] Note 2: The decoder may ignore the value of general_non_packed_constraint_flag because the decoding process does not depend on the presence or absence of frame packing arrangement SEI messages, segmented rectangle frame packing arrangement SEI messages, equidistant cylindrical projection SEI messages, or cube projection SEI messages in the CVS and their interpretations.
[0060] The above semantics of general_non_packed_constraint_flag ensure that when general_non_packed_constraint_flag equals 1, the CVS is part of a "standard" video bitstream that does not employ any frame packing arrangement scheme or any omnidirectional video projection scheme. Therefore, a "standard" decoder that does not support special post-decoder rendering operations (such as frame unpacking or the inverse operation of omnidirectional video projection) will be able to correctly render the video. Since the syntax element general_non_packed_constraint_flag is carried in a special part of the parameter set, this information is useful to the system; and this special parameter set part can usually be accessed by system functions that perform content selection and session negotiation.
[0061] The amendment draft of HEVC is specifically described in "Additional Supplemental Enhancement Information for HEVC (Draft 1)" published in JCT-VC output file JCTVC-AD1005 in March 2018 by J. Boyce, H.-M. Oh, G. J. Sullivan, A. Tourapis, and Y.-K. Wang (editors). JCTVC-AD1005 includes the specific description of the omnidirectional fisheye SEI message. The presence of the omnidirectional fisheye SEI message in the coded layer-wise video sequence (CLVS) indicates that each coded video picture in the CLVS is a fisheye omnidirectional video picture, and the fisheye omnidirectional video picture contains multiple active regions captured by a fisheye camera. The receiver can use the information of the fisheye omnidirectional video carried in the omnidirectional fisheye SEI message to correctly present the fisheye omnidirectional video. It is stipulated that the omnidirectional fisheye SEI message is applicable to the CLVS containing the SEI message (also called the current CLVS); when the omnidirectional fisheye SEI message exists in the CLVS, the omnidirectional fisheye SEI message should exist in the first access unit of the CLVS, and the omnidirectional fisheye SEL message may exist in other access units of the CLVS.
[0062] For convenience, the syntax and semantics of the omnidirectional fisheye SEI message specifically described in JCTVC-AD1005 are reproduced below.
[0063] Omnidirectional Fisheye SEI Message Syntax
[0064] omfy_view_dimension_idc indicates the alignment and viewing direction of the fisheye lens, as described below:
[0065] - omfy_view_dimension_idc being equal to 0 means that omfy_num_active_areas is equal to 2; the values of omfy_camera_centre_azimuth, omfy_camera_centre_elevation, omfy_camera_centre_tilt, omfy_camera_centre_offset_x, omfy_camera_centre_offset_y, and omfy_camera_centre_offset_z are such that the active regions have aligned optical axes and face in opposite directions; the sum of the values of omfy_field_of_view is greater than or equal to 360×2 16 .
[0066] - When omfy_view_dimension_idc equals 1, it means omfy_num_active_areas equals 2; the values of omfy_camera_centre_azimuth, omfy_camera_centre_elevation, omfy_camera_centre_tilt, omfy_camera_centre_offset_x, omfy_camera_centre_offset_y, and omfy_camera_centre_offset_z make the active areas have parallel optical axes, where these optical axes are orthogonal to the lines intersecting the camera centre point; the camera corresponding to i equal to 0 is in the left view.
[0067] - When omfy_view_dimension_idc equals 2, it means omfy_num_active_areas equals 2; the values of omfy_camera_centre_azimuth, omfy_camera_centre_elevation, omfy_camera_centre_tilt, omfy_camera_centre_offset_x, omfy_camera_centre_offset_y, and omfy_camera_centre_offset_z make the active areas have parallel optical axes, where these optical axes are orthogonal to the lines intersecting the camera centre point; the camera corresponding to i equal to 0 is in the right view.
[0068] - When omfy_view_dimension_idc equals 7, it means there are no additional constraint conditions for the syntax element values in the omnidirectional fisheye SEI message.
[0069] - Reserve the values of omfy_view_dimension_idc in the range from 3 to 6 (including 3 and 6) for future use by ITU-T or ISO / IEC. When the value of omfy_view_dimension_idc is in the range from 3 to 6 (including 3 and 6), the decoder shall ignore this value.
[0070] omfy_reserved_zero_5bits shall equal 0 in the bitstream to conform to the version in this specification. Reserve other values of omfy_reserved_zero_5bits for future use by ITU-T or ISO / IEC. The decoder shall ignore the value of omfy_reserved_zero_5bits.
[0071] omfy_num_active_areas_minus1 plus 1 represents the number of active areas in the coded picture. The value of omfy_num_active_areas_minus1 shall be in the range of 0 to 3 (including 0 and 3). Values of omfy_num_active_areas_minus1 greater than 3 are reserved for future use by ITU-T or ISO / IEC. When omfy_num_active_areas_minus1 in the omnidirectional fisheye SEI message is greater than 3, the decoder shall ignore the omnidirectional fisheye SEI message.
[0072] omfy_circular_region_centre_x[i] and omfy_circular_region_centre_y[i] represent the horizontal and vertical coordinates of the centre of the circular region respectively, in units of 2 to 16 luma picture elements, where the circular region contains the i-th active area in the coded picture. The values of omfy_circular_region_centre_x[i] and omfy_circular_region_centre_y[i] shall be in the range of 0 to 65536×2 16 −1 (i.e., 4294967295) (including 0 and 4294967295).
[0073] omfy_rect_region_top[i], omfy_rect_region_left[i], omfy_rect_region_width[i] and omfy_rect_region_height[i] represent the coordinates of the top left corner, width and height of the i-th rectangular region, where the i-th rectangular region contains the i-th active area. These values are in units of luma picture elements.
[0074] omfy_circular_region_radius[i] represents the radius of the circular region containing the i-th active area, in units of 2 to 16 luma picture elements, where the radius is defined as the length from the centre of the circular region represented by omfy_circular_region_centre_x[i] and omfy_circular_region_centre_y[i] to the outermost pixel boundary of the circular region, which corresponds to the maximum field of view of the i-th fisheye lens, represented by omfy_field_of_view[i]. The value of omfy_circular_region_radius[i] shall be in the range of 0 to 65536×2 16within the range of -1 (i.e., 4294967295) including 0 and 4294967295.
[0075] The i-th active region is defined as the intersection of the i-th rectangular region and the i-th circular region. Among them, the i-th rectangular region is represented by omfy_rect_region_top[i], omfy_rect_region_left[i], omfy_rect_region_width[i], and omfy_rect_region_height[i], and the i-th circular region is represented by omfy_circular_region_centre_x[i], omfy_circular_region_centre_y[i], and omfy_circular_region_radius[i].
[0076] omfy_scene_radius[i] represents the radius of the circular region within the i-th active region, in units of 2 to 16 luminance pixels. Among them, the region represented by omfy_circular_region_centre_x[i], omfy_circular_region_centre_y[i], and omfy_scene_radius[i] does not include occlusions such as the camera body. The value of omfy_scene_radius[i] should be less than or equal to omfy_circular_region_radius[i], and should be within the range of 0 to 65536×2 16 -1 (i.e., 4294967295) including 0 and 4294967295. The closed region is the region recommended by the encoder for stitching.
[0077] omfy_camera_centre_azimuth[i] and omfy_camera_centre_elevation[i] represent the spherical coordinates corresponding to the center of the circular region, in units of 2 to 16 degrees. Among them, the circular region contains the i-th active region in the trimmed output image. The value of omfy_camera_centre_azimuth[i] should be within the range of -180×2 16 (i.e., -11796480) to 180×2 16 -1 (i.e., 11796479) including -11796480 and 11796479, while the value of omfy_camera_centre_elevation[i] should be within the range of -90×2 16 (i.e., 5898240) to 90×2 16within the range of (i.e., 5898240), including 5898240 and 5898240.
[0078] omfy_camera_centre_tilt[i] represents the tilt angle of the i-th active region in the cropped output image, in units of 2 to 16 degrees. The value of omfy_camera_centre_tilt[i] should be within the range of -180 × 2 16 (i.e., -11796480) to 180 × 2 16 -1 (i.e., 11796479), including -11796480 and 11796479.
[0079] omfy_camera_centre_offset_x[i], omfy_camera_centre_offset_y[i], and omfy_camera_centre_offset_z[i] represent the XYZ offset values of the focal center of the fisheye camera corresponding to the i-th active region from the origin of the focal center of the entire fisheye camera structure, in units of 2 to 16 millimeters. The values of omfy_camera_centre_offset_x[i], omfy_camera_centre_offset_y[i], and omfy_camera_centre_offset_z[i] should be within the range of 0 to 65536 × 2 16 -1 (i.e., 4294967295), including 0 and 4294967295.
[0080] omfy_field_of_view[i] represents the spherical field coverage of the i-th active region in the encoded image, in units of 2 to 16 degrees. The value of omfy_field_of_view[i] should be within the range of 0 to 360 × 2 16 including 0 and 360 × 2 16 )
[0081] omfy_num_polynomial_coeffs[i] represents the number of polynomial coefficients corresponding to the i-th active region. The value of omfy_num_polynomial_coeffs[i] should be within the range of 0 to 8, including 0 and 8. Values of omfy_num_polynomial_coeffs[i] greater than 8 are reserved for future use by ITU-T or ISO / IEC. When omfy_num_polynomial_coeffs[i] in the omnidirectional fisheye SEI message is greater than 8, the decoder should ignore the omnidirectional fisheye SEI message.
[0082] omfy_polynomial_coeff[i][j] represents the j-th polynomial coefficient value in the curve function, in units of 2 to 24, where the curve function maps the normalized distance of the luminance pixel points from the center of the i-th circular area to the angular value of the spherical coordinates of the normal vector of the i-th image plane. The value of omfypolynomial_coeff[i][j] should be in the range of -128×2 24 (i.e., 2147483648) to 128×2 24 -1 (i.e., 2147483647) (including -2147483648 and 2147483648).
[0083] Currently, when the value of general_non_packed_constraint_flag in the active sequence parameter set (SPS) is equal to 1, the bitstream can include the omnidirectional fisheye SEI message of the image. When the bitstream is sent to a "standard" decoder (i.e., a video decoder that cannot correctly process and / or render fisheye video images), the standard decoder cannot correctly decode or render the fisheye video bitstream. This will result in a poor user experience.
[0084] This document discloses signaling techniques and / or methods that prohibit the bitstream from including the omnidirectional fisheye SEI message (also known as the fisheye video information SEI message) of the image when the value of general_non_packed_constraint_flag in the active SPS is equal to 1. When the bitstream includes the omnidirectional fisheye SEI message of the image, the video encoder sets the value of general_non_packed_constraint_flag in the active SPS to 0. Therefore, the video decoder knows that any bitstream containing general_non_packed_constraint_flag equal to 1 does not contain any fisheye omnidirectional video images. When the video decoder receives a bitstream containing general_non_packed_constraint_flag equal to 0, the video decoder knows that the bitstream may contain an omnidirectional fisheye SEI message corresponding to one or more fisheye omnidirectional video images.
[0085] Figure 1Schematic diagram of an exemplary system 100 for VR video encoding. System 100 includes an omnidirectional camera 101, a VR encoding device 104 including an encoder 103, a decoder 107, and a rendering device 109. The omnidirectional camera 101 includes an array of camera devices. Each camera device points in a different angle, such that the omnidirectional camera 101 can capture multiple-direction video streams of the surrounding environment from multiple angles. For example, the omnidirectional camera 101 can capture the environmental video as a sphere, where the omnidirectional camera 101 is located at the center of the sphere. As used herein, the terms "sphere" and "spherical video" refer to both a geometric sphere and a sub-part of a geometric sphere, such as a spherical segment, a spherical dome, a spherical frustum, etc. For example, the omnidirectional camera 101 can capture 180-degree video to cover half of the environment, such that the production staff can stay behind the omnidirectional camera 101. The omnidirectional camera 101 can also capture video within 360 degrees (or any subset within 360 degrees). However, a part of the floor under the omnidirectional camera 101 may be ignored, resulting in the video not being a complete sphere. Therefore, the term "sphere" used herein is a general term used for clear discussion and should not be considered as a geometric limitation.
[0086] Forward the video captured by the omnidirectional camera 101 to the VR encoding device 104. The VR encoding device 104 can be a computing system including dedicated VR encoding software. The VR encoding device 104 can include an encoder 103 (also referred to as a video encoder). In some examples, the encoder 103 can also be included in a computer system independent of the VR encoding device 104. The VR encoding device 104 is used to convert multiple-direction video streams into a video stream including multiple-direction video streams, covering the entire area recorded from all relevant angles. This conversion can be referred to as image stitching. For example, the frames in each simultaneously captured video stream can be stitched together to generate a single spherical image. Secondly, a spherical video stream can be generated from the spherical image. For the sake of clear discussion, it should be noted that the terms "frame" and "picture / image" can be used interchangeably herein unless otherwise specified.
[0087] Next, the spherical video stream can be forwarded to the encoder 103 for compression. The encoder 103 is a device and / or program capable of converting information from one format to another for standardization, acceleration, and / or compression. The standardization encoder 103 is used to encode rectangular images and / or square images. Therefore, the encoder 103 is used to map each spherical image in the spherical video stream to a plurality of rectangular sub-images. Then, these sub-images can be carried in separate sub-image video streams. Therefore, each sub-image video stream shows a stream of images that change over time, and each sub-image video stream records a part of the spherical video stream. Finally, the encoder 103 can encode each sub-image video stream to compress the video stream into a manageable file size. The encoding process is discussed in further detail below. Generally speaking, the encoder 103 divides each frame in each sub-image video stream into pixel blocks, compresses the pixel blocks through inter-frame prediction and / or intra-frame prediction to generate encoded blocks including predicted blocks and residual blocks, transforms the residual blocks for further compression, and filters these blocks using various filters. The compressed blocks and the corresponding syntax are stored in the bitstream in the International Standardization Organization base media file format (ISOBMFF) and / or the omnidirectional media format (OMAF), etc.
[0088] The VR encoding device 104 can store the encoded bitstream in a memory on the local and / or server, so as to communicate with the decoder 107 on demand. These data can be forwarded via the network 105. The network 105 can include the Internet, a mobile communication network (e.g., a data network based on long term evolution (LTE)), or other data communication data systems.
[0089] The decoder 107 (also known as the video decoder) is a device located at the user's position and is used to perform the reverse process of the encoding process to reconstruct the sub-image video stream according to the encoded bitstream. The decoder 107 also combines the sub-image video streams to reconstruct the spherical video stream. Then, the spherical video stream or a sub-part thereof can be forwarded to the rendering device 109. The rendering device 109 is a device for displaying the spherical video stream to the user. For example, the rendering device 109 may include an HMD fixed to the user's head and covering the user's eyes. The rendering device may include a screen, a camera, a motion sensor, a speaker, etc. corresponding to each eye, and may communicate with the decoder 107 through a wireless and / or wired connection. The rendering device 109 may display a part of the spherical video stream to the user. The shown part is captured according to the FOV and / or viewing angle of the rendering device. For example, the rendering device 109 may change the position of the FOV according to the user's head movement through a motion tracking sensor. In this way, the user can see different parts of the spherical video stream according to the head movement. In addition, the rendering device 109 may adjust the FOV corresponding to each eye according to the user's interpupillary distance (IPD), thereby creating an impression of a three-dimensional space.
[0090] Figure 2 FIG. 200 is a flowchart of an exemplary method 200. The method 200 is used to encode a VR image bitstream into a plurality of sub-image bitstreams using components in the system 100, etc. In step 201, a multi-directional camera group such as the multi-directional camera 101 is used to capture a plurality of directional video streams. The plurality of directional video streams include environmental views at different angles. For example, the plurality of directional video streams may include videos captured around the camera at 360 degrees, 180 degrees, 240 degrees, etc. in the horizontal plane, and the multi-directional video streams may also include videos captured around the camera at 360 degrees, 180 degrees, 240 degrees, etc. in the vertical plane. The generated videos include information sufficient to cover the spherical area around the camera within a certain period of time.
[0091] In step 203, the plurality of directional video streams are synchronized in the time domain. Specifically, each directional video stream includes a series of images captured at a corresponding angle. Ensure that the frames captured at the same time domain position in each directional video stream can be uniformly processed, so that the plurality of directional video streams are synchronized. Then, the frames in the directional video streams can be stitched together in the spatial domain to generate a spherical video stream. Therefore, each frame in the spherical video stream contains data obtained from the frames of all directional video streams that occur at the same time position.
[0092] In step 205, the spherical video stream is mapped to a rectangular sub-image video stream. This process can also be referred to as projecting the spherical video stream onto the rectangular sub-image video stream. As described above, encoders and decoders are typically designed to encode rectangular frames and / or square frames. Therefore, mapping the spherical video stream to the rectangular sub-image video stream generates video streams that can be encoded and decoded by non-VR specific encoders and decoders. It should be noted that steps 203 and 205 involve VR video processing and can thus be performed by dedicated VR hardware, software, or a combination thereof.
[0093] In step 207, the rectangular sub-image video stream is forwarded to an encoder such as encoder 103. Then, the encoder encodes the sub-image video stream into a sub-image bitstream in the format of a corresponding media file. Specifically, the encoder can treat each sub-image video stream as a video signal. The encoder can encode each frame in each sub-image video stream through inter-frame prediction, intra-frame prediction, etc. This encoding and the corresponding decoding, as well as the encoders and decoders, will be discussed in detail below in conjunction with Figures 3 to 15 This encoding and the corresponding decoding, as well as the encoders and decoders, will be discussed in detail below in conjunction with
[0094] In step 209, the sub-image bitstream is sent as a track to the decoder. In some examples, tracks in the same representation are transmitted such that all sub-image bitstreams are transmitted at the same quality. A drawback of this approach is that less interesting regions in the final VR video stream are transmitted at the same resolution as all other regions. Viewpoint-based coding can be used to improve compression by this method. In viewpoint-based coding, for tracks containing sub-image bitstreams with user FOV data, a high-quality representation is selected to be transmitted at a higher resolution. For tracks containing sub-image bitstreams with regions outside the user FOV, a low-quality representation is selected to be transmitted at a gradually decreasing resolution. In some examples, certain regions can even be completely ignored. For example, if the user decides to change the FOV to include regions adjacent to the FOV, these regions can be transmitted at a slightly reduced quality. Since regions further away from the FOV are increasingly unlikely to enter the FOV and thus increasingly unlikely to be presented to the user, these regions can be transmitted at a gradually decreasing quality. A track can include a relatively short (e.g., about 3 seconds) video segment, so the representation selected for a particular region of the video can change over time according to changes in the FOV. In this way, the quality can vary with changes in the user FOV. Viewpoint-based coding can significantly reduce the file size of the tracks sent to the user without significantly degrading the visual quality, since the regions with reduced quality are less likely to be seen by the user.
[0095] In step 211, a decoder such as decoder 107 receives a track containing a sub-image bitstream. Then, the decoder can decode the sub-image bitstream into a sub-image video stream for display. The decoding process involves the inverse of the encoding process (e.g., performing inter-frame prediction and intra-frame prediction), which is discussed in detail below in conjunction with Figures 3 to 10 this.
[0096] In step 213, the decoder can merge the sub-image video streams into a spherical video stream for display to the user. Specifically, the decoder can use a so-called lightweight merging algorithm to select frames that occur at the same display time from each sub-image video stream and merge these frames together at the positions and / or angles corresponding to the respective sub-image video streams. The decoder can also use filters to smooth the edges between sub-image video streams, remove artifacts, etc. Then, the decoder can forward the spherical video stream to a rendering device (e.g., rendering device 109).
[0097] In step 215, the rendering device renders the viewpoint of the spherical video stream for display to the user. As described above, regions outside the FOV in the spherical video stream may not be rendered at each time point. Thus, in viewpoint-based coding, the low-quality representation is effectively ignored, and therefore, the reduction in viewing quality has a negligible impact on the user experience while reducing the file size.
[0098] Figure 3 Flowchart of an exemplary method 300 for encoding a video signal including a sub-image video stream. For example, method 300 may include receiving a plurality of sub-image video streams obtained according to step 205 in step 200. In method 300, each sub-image video stream is input as a video signal. In method 300, steps 301 to 317 are performed on each sub-image video stream to perform steps 207 to 211 in method 200. Therefore, the output video signal from method 300 includes the decoded sub-image video streams, which can be merged and displayed according to steps 213 and 215 in method 200.
[0099] In method 300, an encoder encodes a video signal including a sub-image video stream and the like. During the encoding process, the video signal is compressed by using various mechanisms to reduce the video file size. With a smaller file, the compressed video file can be transmitted to the user while reducing the associated bandwidth overhead. Then, a decoder decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is typically the inverse of the encoding process, such that the video signal reconstructed by the decoder is consistent with the video signal on the encoder side.
[0100] In step 301, the video signal is input into the encoder. For example, the video signal may be an uncompressed video file stored in a memory. Alternatively, the video file may be captured by a video capture device (e.g., a camera) and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component contains a series of image frames. When viewed in sequence, these image frames give the visual effect of motion. These frames contain pixels represented by light, referred to herein as luminance components (or luminance pixel points), and also contain pixels represented by color, referred to as chrominance components (or color pixel points).
[0101] In step 303, the video is segmented into blocks. The segmentation includes subdividing the pixels in each frame into square and / or rectangular blocks for compression. For example, in HEVC (also known as H.265 and MPEG-H Part 2), a frame can first be divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). These CTUs contain both luminance pixel points and chrominance pixel points. The CTUs can be divided into blocks using a coding tree, and then the blocks can be repeatedly subdivided until a configuration that supports further encoding is achieved. For example, the luminance component of the frame can be subdivided until each block contains relatively uniform illumination values. Additionally, the chrominance component of the frame can be subdivided until each block contains relatively uniform color values. Therefore, the segmentation mechanism varies according to the content of the video frame.
[0102] In step 305, various compression mechanisms are used to compress the image blocks segmented in step 303. For example, inter-frame prediction and / or intra-frame prediction can be employed. Inter-frame prediction takes advantage of the fact that objects in common scenarios often appear in consecutive frames. Thus, blocks that describe objects in a reference frame do not need to be repeatedly described in adjacent frames. Specifically, an object (e.g., a table) may remain in a fixed position in multiple frames. Therefore, the table is described once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism can be used to match objects in multiple frames. Additionally, due to reasons such as object movement or camera movement, moving objects can be represented in multiple frames. In a particular example, a video can show a car moving across the screen in multiple frames. Motion vectors can be used to describe this movement, or the absence of such motion. A motion vector is a two-dimensional vector that provides the offset from the coordinates of an object in one frame to the coordinates of that object in a reference frame. Thus, inter-frame prediction can encode the image blocks in the current frame as a set of motion vectors that indicate the offsets of the image blocks in the current frame from the corresponding blocks in the reference frame.
[0103] Intra-frame prediction is used to encode blocks in a common frame. Intra-frame prediction takes advantage of the fact that luminance components and chrominance components tend to cluster in a frame. For example, a patch of green in a part of a tree often neighbors several similar patches of green. Intra-frame prediction uses a variety of directional prediction modes (e.g., there are 33 in HEVC), a planar mode, and a direct current (DC) mode. These directional modes indicate that the pixel points of the current block are similar / same to the pixel points of adjacent blocks along the corresponding direction. The planar mode indicates that a series of blocks on a row / column (e.g., a plane) are interpolated based on adjacent blocks at the row edge. The planar mode actually represents that light / color smoothly transitions between rows / columns by using a relatively constant slope with numerical changes. The DC mode is used for boundary smoothing and indicates that the block is similar / same to the average of the pixel points of all adjacent blocks that are related to the angular direction of the directional prediction mode. Thus, intra-frame prediction blocks can represent image blocks as various relationship prediction mode values rather than actual values. Additionally, inter-frame prediction blocks can represent image blocks as motion vector values rather than actual values. In either case, the prediction blocks may not accurately represent the image blocks in some situations. Any differences are stored in residual blocks. These residual blocks can be transformed to further compress the file.
[0104] In step 307, various filtering techniques can be used. In HEVC, filters are used according to the in-loop filtering scheme. The block-based prediction discussed above may generate a blocky image on the decoder side. In addition, the block-based prediction scheme can encode blocks and then reconstruct the encoded blocks for subsequent use as reference blocks. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to blocks / frames. These filters mitigate block artifacts so that the encoded file can be accurately reconstructed. In addition, these filters mitigate artifacts in the reconstructed reference blocks, making it less likely for additional artifacts to be generated in subsequent blocks encoded based on the reconstructed reference blocks.
[0105] Once the video signal is segmented, compressed, and filtered, at step 309, the resulting data is encoded into a bitstream. The bitstream includes the data discussed above and any signaling data (e.g., syntax) needed to support proper video signal reconstruction on the decoder side. For example, this data can include segmentation data, prediction data, residual blocks, and various flags that provide encoding instructions to the decoder. The bitstream can be stored in a memory for transmission to the decoder as a track and / or track fragment of an ISOBMFF upon request. The bitstream can also be broadcast and / or multicast to multiple decoders. The generation of the bitstream is an iterative process. Thus, steps 301, 303, 305, 307, and 309 can occur continuously and / or simultaneously in multiple frames and blocks. Figure 3 The order shown is for clarity and ease of discussion purposes and is not intended to limit the video encoding process to a specific order.
[0106] In step 311, the decoder receives the bitstream and starts the decoding process. For example, the decoder may adopt an entropy decoding scheme to convert the bitstream into corresponding syntax data and video data. In step 311, the decoder uses the syntax data in the bitstream to determine the segmentation of the frame. The segmentation should match the result of the block segmentation at step 303. Now, the entropy encoding / decoding adopted in step 311 is described. The encoder makes many choices during the compression process. For example, it selects a block segmentation scheme from several possible choices according to the spatial localization of the values in one (or more) input images. Indicating the exact choice may use a large number of bits. The "bit" used herein is a binary value, which is a variable (e.g., a bit value that may vary according to the context). Entropy encoding enables the encoder to discard any options that are clearly not suitable for a particular situation, leaving a set of available options. Then, a codeword is assigned to each available option. The length of the codeword depends on the number of available options (e.g., one bit corresponds to two options, two bits correspond to three to four options, and so on). Then, the encoder encodes the codewords of the selected options. This scheme reduces the size of the codewords because the codewords are as large as the expected codewords, thus uniquely indicating a selection from a small subset of available options rather than uniquely indicating a selection from a potentially large set of all possible options. Then, the decoder decodes the selection by determining the set of available options in a manner similar to the encoder. By determining the set of available options, the decoder can read the codewords and determine the choices made by the encoder.
[0107] In step 313, the decoder performs block decoding. Specifically, the decoder performs an inverse transform to generate a residual block. Then, the decoder uses the residual block and the corresponding prediction block to reconstruct the image block according to the segmentation. The prediction block may include both the intra-prediction block and the inter-prediction block generated by the encoder at step 305. Then, the reconstructed image block is placed in the frame of the reconstructed video signal according to the segmentation data determined at step 311. The syntax for step 313 may also be indicated in the bitstream by the entropy encoding discussed above.
[0108] In step 315, the frame of the reconstructed video signal is filtered in a manner similar to step 307 on the encoder side. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove block artifacts. Once the frame is filtered, the video signal can be forwarded for merging at step 317 and then output to a display such as an HMD for viewing by the end user.
[0109] Figure 4FIG. 0 is a schematic diagram of an exemplary encoding and decoding (codec) system 400 for video encoding. Specifically, the codec system 400 provides functions to support encoding and decoding of a sub-image video stream according to methods 200 and 300. In addition, the codec system 400 can be used to implement the encoder 103 and / or the decoder 107 in the system 100.
[0110] The codec system 400 is generally applicable to describe components used in both encoders and decoders. The codec system 400 receives a video signal (e.g., including a sub-image video stream) and segments the video signal, as described in connection with steps 301 and 303 in the operation method 300, to produce a segmented video signal 301. Then, when acting as an encoder, the codec system 400 compresses the segmented video signal 401 into an encoded bitstream, as described in connection with steps 305, 307, and 309 in method 300. When acting as a decoder, the codec system 400 generates an output video signal from the bitstream, as described in connection with steps 311, 313, 315, and 317 in the operation method 300. The codec system 400 includes a general encoder control component 411, a transform scaling and quantization component 413, an intra-frame estimation component 415, an intra-frame prediction component 417, a motion compensation component 419, a motion estimation component 421, a scaling and inverse transform component 429, a filtering control analysis component 427, an in-loop filtering component 425, a decoded image buffer component 423, and a header formatting and context adaptive binary arithmetic coding (CABAC) component 431. These components are coupled as shown. In Figure 4 FIG., the black lines represent the motion of the data to be encoded / decoded, and the dashed lines represent the motion of the control data that controls the operation of other components. All components in the codec system 400 can be present in the encoder. The decoder can include a subset of the components in the codec system 400. For example, the decoder can include an intra-frame prediction component 417, a motion compensation component 419, a scaling and inverse transform component 429, an in-loop filtering component 425, and a decoded image buffer component 423. These components are now described.
[0111] The split video signal 401 is a captured video sequence that is segmented into pixel blocks by an encoding tree. The encoding tree uses various partitioning modes to subdivide the pixel blocks into smaller pixel blocks. These blocks can then be further subdivided into smaller blocks. The resulting blocks can be referred to as nodes on the encoding tree. A larger parent node is partitioned into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / encoding tree. In some cases, the partitioned blocks are called coding units (CUs). For example, a CU can be a sub - part of a CTU and contain a luminance block, a red difference chroma (Cr) block, a blue difference chroma (Cb) block, and syntax instructions corresponding to the CU. The partitioning modes can include a binary tree (BT), a triple tree (TT), and a quad tree (QT) for splitting a node into two, three, or four child nodes of different shapes, depending on the partitioning mode used. The split video signal 401 is forwarded to a general encoder control component 411, a transform - scaling and quantization component 413, an intra - prediction component 415, a filtering control analysis component 427, and a motion - estimation component 421 for compression.
[0112] The general encoder control component 411 is used to determine whether to encode an image in a video sequence into a bitstream according to application constraints. For example, the general encoder control component 411 manages the optimization of the bitrate / bitstream size and the reconstructed quality. These decisions can be made based on storage space / bandwidth availability and image resolution requests. The general encoder control component 411 also manages buffer utilization according to the transmission speed to alleviate buffer underflow and overflow problems. To solve these problems, the general encoder control component 411 manages the segmentation, prediction, and filtering performed by other components. For example, the general encoder control component 411 can dynamically increase the compression complexity to increase the resolution and bandwidth utilization, or reduce the compression complexity to decrease the resolution and bandwidth utilization. Thus, the general encoder control component 411 controls other components in the codec system 400 to balance the video signal reconstruction quality and the bitrate. The general encoder control component 411 generates control data that is used to control the operations of other components. The control data is also forwarded to a header formatting and CABAC component 431 to be encoded into the bitstream, thereby indicating the parameters for the decoder to perform decoding.
[0113] The split video signal 401 is also sent to the motion estimation component 421 and the motion compensation component 419 for inter-frame prediction. The frames or stripes in the split video signal 401 can be divided into multiple video blocks. The motion estimation component 421 and the motion compensation component 419 perform inter-frame predictive coding on the received video blocks with reference to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 400 can perform multiple encoding processes to select a suitable encoding mode for each block in the video data, and so on.
[0114] The motion estimation component 421 and the motion compensation component 419 can be highly integrated, but are described separately for conceptual purposes. Motion estimation performed by the motion estimation component 421 is a process of generating motion vectors, where the motion vectors are used to estimate the motion of video blocks. For example, the motion vectors can indicate the displacement of the coding object relative to the prediction block. The prediction block is a block found to highly match the block to be coded in terms of pixel differences. The prediction block can also be referred to as a reference block. Such pixel differences can be determined by the sum of absolute difference (SAD), the sum of squared difference (SSD), or other difference metrics. HEVC uses several coding objects, including CTU, coding tree block (CTB), and CU. For example, a CTU can be divided into CTBs, and then the CTBs can be divided into CUs to be included in the CU. A CU can be encoded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing the transform residual data of the CU. The motion estimation component 421 uses rate-distortion analysis as part of the rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 421 can determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and can select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance the quality of video reconstruction (e.g., the amount of data loss caused by compression) and the encoding efficiency (e.g., the size of the final encoding).
[0115] In some examples, the codec system 400 may calculate values of sub-integer pixel positions of a reference image stored in the decoded image buffer component 423. For example, the video codec system 400 may interpolate values of quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, the motion estimation component 421 may perform motion search on integer pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. The motion estimation component 421 calculates motion vectors for the PUs of video blocks in an inter-coded slice by comparing the position of the PU with the position of a predicted block in the reference image. The motion estimation component 421 outputs the calculated motion vectors as motion data to the header formatting and CABAC component 431 for encoding and as motion data to the motion compensation component 419.
[0116] The motion compensation performed by the motion compensation component 419 may include obtaining or generating a predicted block according to the motion vector determined by the motion estimation component 421. In some examples, the motion estimation component 421 and the motion compensation component 419 may also be functionally integrated. When receiving the motion vector of the PU of the current video block, the motion compensation component 419 may locate the predicted block pointed to by the motion vector. Then, the pixel values of the predicted block are subtracted from the pixel values of the current video block being encoded, generating a pixel difference, thereby forming a residual video block. Generally, the motion estimation component 421 performs motion estimation on the luminance component, and the motion compensation component 419 uses the motion vector calculated based on the luminance component for both the chrominance component and the luminance component. The predicted block and the residual block are forwarded to the transform scaling and quantization component 413.
[0117] The segmented video signal 401 is also sent to the intra estimation component 415 and the intra prediction component 417. Similar to the motion estimation component 421 and the motion compensation component 419, the intra estimation component 415 and the intra prediction component 417 may be highly integrated, but are separately described for conceptual purposes. The intra estimation component 415 and the intra prediction component 417 perform intra prediction on the current block with reference to each block in the current frame to replace the inter prediction performed between frames by the motion estimation component 421 and the motion compensation component 419 as described above. Specifically, the intra estimation component 415 determines an intra prediction mode for encoding the current block. In some examples, the intra estimation component 415 selects a suitable intra prediction mode from multiple tested intra prediction modes to encode the current block. Then, the selected intra prediction mode is forwarded to the header formatting and CABAC component 431 for encoding.
[0118] For example, the intra prediction component 415 calculates the rate-distortion value by performing rate-distortion analysis on various tested intra prediction modes, and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was coded to produce the coded block, and determines the bit rate (e.g., the number of bits) used to produce the coded block. The intra prediction component 415 calculates the ratio based on the distortion and rate of various coded blocks to determine which intra prediction mode yields the best rate-distortion value for the block. Additionally, the intra prediction component 415 can be used to encode depth blocks in the depth map using the depth modeling mode (DMM) according to rate-distortion optimization (RDO).
[0119] When implemented on the encoder, the intra prediction component 417 can generate a residual block from the prediction block according to the selected intra prediction mode determined by the intra prediction component 415, or when implemented on the decoder, read the residual block from the bitstream. The residual block includes the value difference between the prediction block and the original block, represented as a matrix. Then, the residual block is forwarded to the transform scaling and quantization component 413. The intra prediction component 415 and the intra prediction component 417 can perform operations on the luminance component and the chrominance component.
[0120] The transform scaling and quantization component 413 is used to further compress the residual block. The transform scaling and quantization component 413 performs a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform on the residual block, thereby generating a video block including residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms can also be performed. The transform can convert the residual information from the pixel value domain to the transform domain, e.g., the frequency domain. The transform scaling and quantization component 413 is also used to scale the transform residual information according to frequencies, etc. This scaling involves applying a scaling factor to the residual information in order to quantize different frequency information at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 413 is also used to quantize the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 413 can perform a scan on the matrix including the quantized transform coefficients. The quantized transform coefficients are forwarded to the header formatting and CABAC component 431 for encoding into the bitstream.
[0121] The scaling and inverse transform component 429 performs operations opposite to those of the transform scaling and quantization component 413 to support motion estimation. The scaling and inverse transform component 429 performs inverse scaling, inverse transform, and / or inverse quantization to reconstruct the residual block in the pixel domain, e.g., for subsequent use as a reference block. This reference block can become the prediction block for another current block. The motion estimation component 421 and / or the motion compensation component 419 can calculate the reference block by adding the residual block to the corresponding prediction block for use in motion estimation of subsequent blocks / frames. Filters are applied to the reconstructed reference block to mitigate artifacts generated during the scaling, quantization, and transform processes. These artifacts can make the prediction inaccurate (and generate additional artifacts) when predicting subsequent blocks.
[0122] The filter control analysis component 427 and the in-loop filter component 425 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 429 can be combined with the corresponding prediction block from the intra prediction component 417 and / or the motion compensation component 419 to reconstruct the original image block. Then, filters can be applied to the reconstructed image block. In some examples, filters can be applied to the residual block. Similar to Figure 4 other components in, the filter control analysis component 427 and the in-loop filter component 425 are highly integrated and can be implemented together, but are described separately for conceptual purposes. The filters applied to the reconstructed reference block are applied to specific spatial regions, and these filters include multiple parameters for adjusting how these filters are used. The filter control analysis component 427 analyzes the reconstructed reference block to determine where to use the filters and sets the corresponding parameters. This data is forwarded to the header formatting and CABAC component 431 and used as filter control data for encoding. The in-loop filter component 425 uses these filters according to the filter control data. These filters can include a deblocking filter, a noise suppression filter, a SAO filter, and an adaptive loop filter. These filters can be applied in the spatial domain / pixel domain (e.g., on the reconstructed pixel block) or in the frequency domain according to examples.
[0123] When operating as an encoder, the filtered reconstructed image block, residual block, and / or prediction block are stored in the decoded image buffer component 423 for subsequent use in motion estimation as described above. When operating as a decoder, the decoded image buffer component 423 stores the reconstructed filtered block and forwards the reconstructed filtered block to the display as part of the output video signal. The decoded image buffer component 423 can be any storage device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0124] The header formatting and CABAC component 431 receives data from various components in the codec system 400 and encodes this data into an encoded bitstream for transmission to a decoder. Specifically, the header formatting and CABAC component 431 generates various headers to encode control data (e.g., general control data and filtering control data). In addition, prediction data including intra prediction and motion data, as well as residual data in the form of quantized transform coefficient data, are encoded in the bitstream. The final bitstream includes all the information required by the decoder to reconstruct the original segmented video signal 401. This information may also include an intra prediction mode index table (also referred to as a codeword mapping table), definitions of the coding contexts of various blocks, indications of the most probable intra prediction modes, indications of segmentation information, etc. These data can be encoded through entropy coding. For example, these information can be encoded through context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. After entropy coding, the encoded bitstream can be transmitted to another device (e.g., a video decoder) or archived for subsequent transmission or retrieval.
[0125] As described above, the present invention provides signaling techniques and / or methods that prohibit the bitstream from including an omnidirectional fisheye SEI message of an image when the value of general_non_packed_constraint_flag in the active SPS is equal to 1. When the bitstream includes an omnidirectional fisheye SEI message of an image, the video encoder sets the value of general_non_packed_constraint_flag in the active SPS to be equal to 0. Thus, the video decoder knows that any bitstream containing general_non_packed_constraint_flag equal to 1 does not contain any fisheye omnidirectional video images. When the video decoder receives a bitstream containing general_non_packed_constraint_flag equal to 0, the video decoder knows that the bitstream may contain an omnidirectional fisheye SEI message corresponding to one or more fisheye omnidirectional video images.
[0126] The signaling technology and / or method has at least the following advantages and benefits compared to traditional signaling technologies and methods. Specifically, the disclosed embodiments can be used to avoid undesirable, annoying, or unexpected user experiences in order to enable a better user experience, and to reduce the implementation complexity of a decoder that supports processing projection indication SEI messages (i.e., equirectangular projection SEI messages or cube projection SEI messages) and omnidirectional fisheye SEI messages.
[0127] To implement the signaling technology and / or method disclosed herein, the semantics of the syntax element general_non_packed_constraint_flag are changed as follows.
[0128] When general_non_packed_constraint_flag is equal to 1, it indicates that there is no frame-packing arrangement SEI message, segmented rectangular frame-packing arrangement SEI message, equirectangular projection SEI message, cube projection SEI message, or omnidirectional fisheye SEI message in the CVS. When general_non_packed_constraint_flag is equal to 0, it indicates that there may or may not be one or more frame-packing arrangement SEI messages, segmented rectangular frame-packing arrangement SEI messages, equirectangular projection SEI messages, cube projection SEI messages, or omnidirectional fisheye SEI messages in the CVS.
[0129] Note 2: The decoder can ignore the value of general_non_packed_constraint_flag because the decoding process does not depend on the presence or absence of frame-packing arrangement SEI messages, segmented rectangular frame-packing arrangement SEI messages, equirectangular projection SEI messages, or cube projection SEI messages in the CVS and their interpretations.
[0130] Figure 5 FIG. is a block diagram of an exemplary video encoder 500. The video encoder 500 can encode a sub-image bitstream. The video encoder 500 can be used to implement the encoding function of the codec system 400 and / or perform steps 301, 303, 305, 307, and / or 309 in method 300. Similar to encoder 103, encoder 500 can also be used to perform steps 205 to 209 in method 200. The encoder 500 splits the input video signal (e.g., a sub-image video stream), thereby generating a split video signal 501 that is substantially similar to the split video signal 401. Then, the components in the encoder 500 compress and encode the split video signal 501 into a bitstream.
[0131] Specifically, the segmented video signal 501 is forwarded to the intra prediction component 517 for intra prediction. The intra prediction component 517 can be substantially similar to the intra estimation component 415 and the intra prediction component 417. The segmented video signal 501 is also forwarded to the motion compensation component 521 for inter prediction based on the reference blocks in the decoded picture buffer component 523. The motion compensation component 521 can be substantially similar to the motion estimation component 421 and the motion compensation component 419. The predicted blocks and residual blocks from the intra prediction component 517 and the motion compensation component 521 are forwarded to the transform and quantization component 513 to perform transform and quantization on the residual blocks. The transform and quantization component 513 can be substantially similar to the transform scaling and quantization component 413. The transform quantized residual blocks and the corresponding predicted blocks (along with relevant control data) are forwarded to the entropy coding component 531 for encoding into the bitstream. The entropy coding component 531 can be substantially similar to the header formatting and CABAC component 431.
[0132] The transform quantized residual blocks and / or the corresponding predicted blocks are also forwarded from the transform and quantization component 513 to the inverse transform and quantization component 529 to reconstruct the reference blocks for use by the motion compensation component 521. The inverse transform and quantization component 529 can be substantially similar to the scaling and inverse transform component 429. The in-loop filter in the in-loop filter component 525 is also applied to the residual blocks and / or the reconstructed reference blocks, depending on the example. The in-loop filter component 525 can be substantially similar to the filter control analysis component 427 and the in-loop filter component 425. The in-loop filter component 525 can include multiple filters, as described in connection with the in-loop filter component 425. Then, the filtered blocks are stored in the decoded picture buffer component 523 for use as reference blocks by the motion compensation component 521. The decoded picture buffer component 523 can be substantially similar to the decoded picture buffer component 423.
[0133] The encoder 500 receives a sub-image video stream segmented from a spherical video stream for use with a VR system using view-based coding. As described above, when the sub-image video streams are transmitted to the decoder at different resolutions, artifacts may occur because data is lost during the process of reducing the resolution of the sub-image video stream with poorer quality. This is because both intra-frame prediction and inter-frame prediction encode a block based on the pixel points (pixels) of adjacent blocks. When the reference pixel points cross the sub-image video stream boundary, the reference pixel points may become inaccurate because the data in the adjacent sub-image video streams is lost. To mitigate these problems, the motion compensation component 521 and the intra-frame prediction component 517 in the encoder 500 encode each sub-image video stream to be self-contained. Specifically, the motion compensation component 521 and the intra-frame prediction component 517 are used to only reference the integer pixel point positions within the same sub-image video stream during encoding. Specifically, when encoding the first sub-image video stream, the encoder 500 is prevented from referencing the pixel point positions in other sub-image video streams. This applies to both the intra-frame prediction mode and the inter-frame prediction motion vectors. In addition, the motion compensation component 521 and the intra-frame prediction component 517 may reference the fractional pixel point positions in the first sub-image video stream, provided that the pixels at the reference fractional pixel point positions can be reconstructed only based on the pixel point positions within the first sub-image bitstream (e.g., without referencing any other sub-image bitstreams) through interpolation. In addition, the motion compensation component 521 may generate a motion vector candidate list for the first sub-image bitstream when performing inter-frame prediction. However, when the motion vectors in the candidate list come from blocks in other sub-image bitstreams, the motion compensation component 521 may not include these motion vectors. These restrictions ensure that each sub-image bitstream can be decoded without referencing adjacent sub-image bitstreams, thus avoiding resolution mismatches.
[0134] In addition, video encoding can adopt parallelization. For example, wavefront parallel processing (WPP) can be used to accelerate the video encoding process. WPP allows encoding of the current block (e.g., CTU) as long as the blocks above the current block and the blocks above and to the right of the current block have been decoded. WPP gives the impression similar to a wave, where the first row of blocks is encoded two blocks earlier than the second row of blocks, that is, two blocks earlier than the third row of blocks, and so on. On the decoder side, the sub-image bitstream frames can be regarded as slices, and these slices can be merged to reconstruct the spherical video stream. When there are slices, WPP can be configured not to operate because WPP only operates on the entire frame once (e.g., the frame from the spherical video stream) and does not operate at the slice level. Accordingly, the encoder 500 can disable WPP when encoding the sub-image bitstream. For example, WPP uses entropy_coding_sync_enabled_flag. This flag is included in the PPS syntax of each image. The encoder 500 can set entropy_coding_sync_enabled_flag to 0 to disable WPP for the sub-image video stream.
[0135] In addition, by encoding the sub-image video stream into tracks and ensuring that these tracks have the same display time, the encoder 500 can avoid the problem of timing mismatch between sub-image bitstreams. The encoder 500 can also ensure that each pixel point of the common VR image (e.g., the frame in the spherical video stream) has the same image sequence number, even if these pixel points are segmented into different sub-image bitstreams and / or carried in different tracks.
[0136] Figure 6 FIG. is a block diagram of an exemplary video decoder 600 that can decode the sub-image bitstream. The video decoder 600 can be used to implement the decoding function of the codec system 400 and / or execute steps 311, 313, 315, and / or 317 in the operation method 300. Similar to the decoder 107, the decoder 600 can also be used to execute steps 211 to 213 in the method 200. The decoder 600 receives multiple sub-image bitstreams from the encoder 500 and the like, generates a reconstructed output video signal including the sub-image video stream, merges the sub-image video stream into a spherical video stream, and forwards the spherical video stream through a presentation device for display to the user.
[0137] The entropy decoding component 633 receives the bitstream. The entropy decoding component 633 is used to perform an entropy decoding scheme, e.g., CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 633 can use the header information to provide context for parsing additional data encoded as codewords in the bitstream. The decoded information includes any information required to decode the video signal, e.g., overall control data, filtering control data, segmentation information, motion data, prediction data, and quantized transform coefficients in the residual blocks. The quantized transform coefficients are forwarded to the inverse transform and quantization component 629 to reconstruct the residual blocks. The inverse transform and quantization component 629 can be similar to the inverse transform and quantization component 529.
[0138] The reconstructed residual blocks and / or prediction blocks are forwarded to the intra prediction component 617 to reconstruct the image blocks according to the intra prediction operation. The intra prediction component 617 can be similar to the intra estimation component 415 and the intra prediction component 417. Specifically, the intra prediction component 617 uses the prediction mode to locate the reference blocks in the frame and adds the residual blocks to the above results to reconstruct the intra prediction image blocks. The reconstructed intra prediction image blocks and / or residual blocks and the corresponding inter prediction data are forwarded to the decoded image buffer component 623 via the in-loop filter component 625. The decoded image buffer component 423 and the in-loop filter component 425 can be substantially similar to the decoded image buffer component 423 and the in-loop filter component 425. The in-loop filter component 625 filters the reconstructed image blocks, residual blocks, and / or prediction blocks. This information is stored in the decoded image buffer component 623. The reconstructed image blocks from the decoded image buffer component 623 are forwarded to the motion compensation component 621 for inter prediction. The motion compensation component 621 can be substantially similar to the motion estimation component 421 and the motion compensation component 419. Specifically, the motion compensation component 621 uses the motion vectors of the reference blocks to generate the prediction blocks and adds the residual blocks to the above results to reconstruct the image blocks. The generated reconstructed blocks can also be forwarded to the decoded image buffer component 623 via the in-loop filter component 625. The decoded image buffer component 623 continues to store additional reconstructed image blocks, which can be reconstructed into frames through the segmentation information. These frames can also be placed in sequence. The sequence is output to the display as the reconstructed output video signal.
[0139] Figure 7 An example of the structure of a bitstream 700 for carrying a flag (e.g., general_non_packed_constraint_flag) is shown. The flag is used to indicate to the decoder whether the bitstream 700 includes any fisheye omnidirectional video images. The "flag" used herein can be referred to as an indication attribute.
[0140] As shown, the bitstream 700 includes CLVS 702. Although Figure 7Only one CLVS 702 is shown, but it should be understood that in practical applications, the bitstream 700 may include one or more CLVSs. The CLVS 702 is divided into access units, namely the first access unit 704 and the second access unit 706. Although Figure 7 two access units are shown, it should be understood that in practical applications, the bitstream 700 may include one or more access units. The first access unit 704 includes a plurality of network access layer (NAL) data units, namely the first NAL data unit 708, the second NAL data unit 710, the third NAL data unit 712, and the fourth NAL data unit 714. Although not shown, the second access unit 706 may include similar NAL data units. In addition, although Figure 7 four NAL data units are shown, it should be understood that in practical applications, the bitstream 700 may include one or more access units.
[0141] In one embodiment, the first NAL data unit 708 contains the SPS. As described above, in one embodiment, the SPS includes a flag (e.g., general_non_packed_constraint_flag) or an indication attribute with a value of 1 or 0. In one embodiment, the SPS is set in the first NAL data unit 708. In one embodiment, the SPS may also be set in one or more of other data units (e.g., the second NAL data unit 710, the third NAL data unit 712, and the fourth NAL data unit 714, etc.).
[0142] In Figure 7Among them, the second NAL data unit 710 contains a picture parameter set (PPS), and the third NAL data unit 712 contains slice information. In one embodiment, the PPS and slices can be set in other data units. The fourth NAL data unit 714 may or may not contain an SEI message, depending on the value of the flag in the SPS. For example, when the encoder sets the flag or indication attribute to 1, there is no frame packing arrangement SEI message, segmented rectangular frame packing arrangement SEI message, equidistant cylindrical projection SEI message, cube projection SEI message, or omnidirectional fisheye SEI message in the CLVS 702 (e.g., the associated non-video coding layer (VCL) NAL unit on the base layer of the picture sequence and the coded video sequence (CVS)). Alternatively, when the encoder sets the flag or indication attribute to 1, there may or may not be one or more frame packing arrangement SEI messages, segmented rectangular frame packing arrangement SEI messages, equidistant cylindrical projection SEI messages, cube projection SEI messages, or omnidirectional fisheye SEI messages in the CLVS 702.
[0143] Figure 8 This is an embodiment of a method 800 for processing video data. In one embodiment, the method 800 is executed by a video decoder (e.g., decoder 107). The method 800 can be executed when the video decoder has received an encoded bitstream. Executing the method 800 can ensure that the video decoder can correctly or appropriately present the bitstream.
[0144] In step 802, it is determined whether the video decoder can correctly present the decoded image obtained by decoding the bitstream. The determination is based on the indication attribute of the bitstream, and the bitstream includes the encoded representation of the video data. The indication attribute includes a specific value indicating that the bitstream does not contain any omnidirectional fisheye video images. In one embodiment, the indication attribute is a constraint flag (e.g., general_non_packed_constraint_flag). In one embodiment, the value of the indication attribute is 1 (one). In one embodiment, the indication attribute is set in the SPS of the bitstream. In one embodiment, the indication attribute can be located elsewhere in the bitstream.
[0145] In step 804, the bitstream is processed according to the determination. In one embodiment, the video decoder (e.g., decoder 107) decodes the bitstream in a consistent interpretation manner. In one embodiment, the bitstream is decoded to present the decoded image. The decoded image does not contain any omnidirectional fisheye video images.
[0146] Figure 9 This is an embodiment of an encoding method 900 performed by a video encoder (e.g., encoder 103). Method 900 can be performed when video data is to be encoded into a bitstream and transmitted to a video decoder (e.g., decoder 107). Performing method 900 can ensure that the video decoder can correctly or appropriately render the bitstream.
[0147] In step 902, flags and representations of video data are encoded into the bitstream. The flag includes a specific value that indicates to a video decoder receiving the bitstream that the bitstream does not contain any fisheye omnidirectional video images. In one embodiment, the flag is a constraint flag (e.g., general_non_packed_constraint_flag). In one embodiment, the value of the flag is 1 (one). In one embodiment, the flag is set in the SPS of the bitstream. In step 904, the bitstream is transmitted to the video decoder.
[0148] Figure 10 This is an embodiment of a decoding method 1000 performed by a video decoder (e.g., decoder 107). Method 1000 can be performed when a coded bitstream is received. In one embodiment, the coded bitstream is received from a video encoder (e.g., encoder 103). Performing method 1000 can ensure that the video decoder can correctly or appropriately render the bitstream.
[0149] In step 1002, a coded bitstream is received. The coded bitstream includes flags and representations of video data. The flag includes a specific value that indicates to a video decoder receiving the bitstream that the bitstream does not contain any fisheye omnidirectional video images. In one embodiment, the flag is a constraint flag (e.g., general_non_packed_constraint_flag). In one embodiment, the value of the flag is 1 (one). In one embodiment, the flag is set in the SPS of the bitstream. In step 1004, the coded bitstream is decoded according to the specific value.
[0150] Figure 11 This is an embodiment of a method 1100 for processing video data. In one embodiment, method 1100 is performed by a video decoder (e.g., decoder 107). Method 1100 can be performed when the video decoder has received a coded bitstream. Performing method 1100 can ensure that the video decoder can correctly or appropriately render the bitstream.
[0151] In step 1102, it is determined whether the video decoder can correctly present the decoded image obtained by decoding the bitstream. In one embodiment, the indication attribute includes a specific value indicating that the bitstream may contain one or more fisheye omnidirectional video images. In one embodiment, the indication attribute is a constraint flag (e.g., general_non_packed_constraint_flag). In one embodiment, the value of the indication attribute is 0(0). In one embodiment, the indication attribute is set in the SPS of the bitstream.
[0152] In step 1104, the bitstream is processed according to the determination. In one embodiment, the video decoder (e.g., decoder 107) decodes the bitstream in a consistent interpretation manner. In one embodiment, the bitstream is decoded to present the decoded image. The decoded image may or may not contain any fisheye omnidirectional video images. In one embodiment, the video decoder checks the bitstream for a specific type of SEI message to process the bitstream.
[0153] Figure 12 This is an embodiment of an encoding method 1200 performed by a video encoder (e.g., encoder 103). Method 1200 can be executed when video data is to be encoded into a bitstream and transmitted to a video decoder (e.g., decoder 107). Executing method 1200 can ensure that the video decoder can correctly or appropriately present the bitstream.
[0154] In step 1202, the flags and representations of the video data are encoded into the bitstream. The flags include a specific value indicating to the video decoder receiving the bitstream that the bitstream may contain one or more fisheye omnidirectional video images. In one embodiment, the flag is a constraint flag (e.g., general_non_packed_constraint_flag). In one embodiment, the value of the flag is 0(0). In one embodiment, the flag with the specific value indicates that the video decoder searches for an omnidirectional fisheye SEI message in the bitstream. In one embodiment, the flag is set in the SPS of the bitstream. In step 1204, the bitstream is transmitted to the video decoder.
[0155] Figure 13 This is an embodiment of a decoding method 1300 performed by a video decoder (e.g., decoder 107). Method 1300 can be executed when an encoded bitstream is received. In one embodiment, the encoded bitstream is received from a video encoder (e.g., encoder 103). Executing method 1300 can ensure that the video decoder can correctly or appropriately present the bitstream.
[0156] In step 1302, an encoded bitstream is received. The encoded bitstream includes flags and representations of video data. The flags include a specific value that indicates to a video decoder receiving the bitstream that the bitstream may contain one or more fisheye omnidirectional video images. In one embodiment, the flag is a constraint flag (e.g., general_non_packed_constraint_flag). In one embodiment, the value of the flag is 0(zero). In one embodiment, the flag having the specific value indicates that the video decoder searches for an omnidirectional fisheye SEI message in the bitstream. In one embodiment, the flag is set in the SPS of the bitstream. In step 1304, the encoded bitstream is decoded according to the specific value.
[0157] Alternative methods for solving the problems described herein are provided below.
[0158] In one embodiment, a new syntax element (e.g., a new flag) located in the same parameter set special part as the syntax element general_non_packed_constraint_flag is used to indicate the possible presence of an omnidirectional fisheye SEI message in an encoded video sequence. For example, the new flag being equal to 1 indicates that there is no omnidirectional fisheye SEI message in the CVS, while the new flag being equal to 0 indicates that there may or may not be an omnidirectional fisheye SEI message in the CVS.
[0159] Change the existing semantics of the syntax element general_non_packed_constraint_flag so that it only applies to frame-packing arrangement SEI messages and equidistant cylindrical projection SEI messages or segmented rectangular frame-packing arrangement SEI messages; use a new syntax element (e.g., a new flag) located in the same parameter set special part as the syntax element general_non_packed_constraint_flag to control omnidirectional fisheye SEI messages, equidistant cylindrical projection SEI messages, and cube projection SEI messages. In one embodiment, the semantics of general_non_packed_constraint_flag and the new flag (e.g., named general_non_360packed_constraint_flag) are specifically described as follows:
[0160] general_non_packed_constraint_flag being equal to 1 indicates that there is no frame-packing arrangement SEI message or segmented rectangular frame-packing arrangement SEI message in the CVS. general_non_packed_constraint_flag being equal to 0 indicates that there may or may not be one or more frame-packing arrangement SEI messages or segmented rectangular frame-packing arrangement SEI messages in the CVS.
[0161] When the general_non_360packed_constraint_flag is equal to 1, it indicates that there is no equidistant cylindrical projection SEI message, cube map projection SEI message, or omnidirectional fisheye SEI message in the CVS. When the general_non_360packed_constraint_flag is equal to 0, it means that there may or may not be an equidistant cylindrical projection SEI message, cube projection SEI message, or omnidirectional fisheye SEI message in the CVS.
[0162] By any of these methods, when it is known that the value of the new flag is equal to 0, the system functions for content selection and session negotiation can avoid sending a fisheye video stream to a "standard" decoder that does not support the special fisheye post - decoder rendering operations. Without such avoidance, a poor user experience will be generated.
[0163] The ideas of the present invention have been described above in the context of HEVC. However, these ideas can be applied to any other video codec, including future video codecs, standard or non - standard video codecs. In addition, these ideas can be applied individually or in combination.
[0164] Figure 14 Schematic diagram of the encoding device 1400 provided for the embodiments of the present application. The encoding device 1400 is suitable for performing the methods and processes disclosed herein. The encoding device 1400 includes an input port 1410 for receiving data and a receiving unit (Rx) 1420; a processor, logic unit, or central processing unit (CPU) 1430 for processing the data; a transmission unit (Tx) 1440 and an output port 1450 for transmitting the data; a memory 1460 for storing the data. The encoding device 1400 may further include optical - to - electrical (OE) components and electrical - to - optical (EO) components coupled to the input port 1410, the receiving unit 1420, the transmission unit 1440, and the output port 1450 for the exit or entry of optical or electrical signals.
[0165] The processor 1430 is implemented by hardware and software. The processor 1430 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 1430 communicates with the input port 1410, the receiving unit 1420, the transmitting unit 1440, the output port 1450, and the memory 1460. The processor 1430 includes an encoding module 1470. The encoding module 1470 implements the embodiments disclosed above. Thus, the encoding module 1470 provides a substantial improvement to the functions of the encoding device 1400 and affects the transition of the encoding device 1400 to different states. Alternatively, the encoding module 1470 is implemented as instructions stored in the memory 1460 and executed by the processor 1430.
[0166] The video encoding device 1400 may further include an input and / or output (I / O) device 1480 for communicating data with a user. The I / O device 1480 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 1480 may further include input devices such as a keyboard, a mouse, a trackball, etc., and corresponding interfaces for interacting with these output devices.
[0167] The memory 1460 includes one or more disks, tape drives, or solid state drives, and can be used as an overflow data storage device to store such programs when a program is selected for execution, or to store instructions and data read during the execution of a program. The memory 1460 can be volatile and / or non-volatile, and can be read-only memory (ROM), random-access memory (RAM), ternary content-addressable memory (TCAM), or static random-access memory (SRAM).
[0168] Figure 15Schematic diagram of an embodiment of an encoding component 1500. In the embodiment, the encoding component 1500 is implemented in a video encoding device 1502 (e.g., encoder 103 or decoder 107). The video encoding device 1502 includes a receiving component 1501. The receiving component 1501 is used to receive an image for encoding or receive a bitstream for decoding. The video encoding device 1502 includes a transmission component 1507 coupled to the receiving component 1501. The transmission component 1507 is used to transmit the bitstream to a decoder or transmit the decoded image to a display component (e.g., one of the I / O devices 1480).
[0169] The video encoding device 1502 includes a storage component 1503. The storage component 1503 is coupled to at least one of the receiving component 1501 or the transmission component 1507. The storage component 1503 is used to store instructions. The video encoding device 1502 further includes a processing component 1505. The processing component 1505 is coupled to the storage component 1503. The processing component 1505 is used to execute the instructions stored in the storage component 1503 to perform the methods disclosed herein.
[0170] Although the present invention provides several embodiments, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention should be considered illustrative rather than restrictive, and the present invention is not limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0171] In addition, without departing from the scope of the present invention, the various technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments may be combined or integrated with other systems, components, technologies, or methods. Other changes, substitutions, and alterations will be apparent to those skilled in the art and are all within the spirit and scope disclosed herein.
Claims
1. A decoding method performed by a video decoder, characterized in that, the method includes: a receiver in the video decoder receives general_non_packed_constraint_flag; a processor in the video decoder determines that the value of general_non_packed_constraint_flag is 1, where the value of 1 indicates that there is no supplemental enhancement information (SEI) message of fisheye video information in the coded video sequence (CVS); the receiver receives a track containing a sub-image bitstream, the track is obtained by downsampling the captured sub-image video stream and then performing view-based entropy coding. Among them, for the track containing the sub-image bitstream with the user's current field of view (FOV) data, it is represented with the first quality and transmitted at the first resolution. For the track containing the sub-image bitstream with regions other than the user's FOV data, it is represented with the second quality and transmitted at the second resolution. Among them, the first quality is higher than the second quality, the first resolution is higher than the second resolution, and the farther the region is from the FOV, the lower the quality of the corresponding track. The resolution of the downsampled sub-image video stream is lower than the resolution of the captured sub-image video stream; the codewords used in the entropy coding include the codewords assigned to each available option in a set of available options, the set of available options is determined from the set of all possible options, and the length of the codewords depends on the number of available options.
2. The method according to claim 1, characterized in that, the method further includes: determining that the value of general_non_packed_constraint_flag in the active sequence parameter set (SPS) of the current layer is 1, where the value of 1 indicates that there should be no applicable fisheye video information SEI message in any image in the coded layer-wise video sequence (CLVS) of the current layer.
3. A method for a video decoder to process video data, characterized in that, the method includes: judging whether the video decoder can correctly present the decoded image obtained by decoding the bitstream according to the indication attribute of the bitstream, where the bitstream includes the encoded representation of the video data, and the indication attribute includes a first specific value indicating that the bitstream does not contain any fisheye omnidirectional video images; processing the bitstream according to the judgment; The bitstream includes a track containing a sub-image bitstream, which is obtained by downsampling a captured sub-image video stream and then performing view-based entropy coding. Among them, for the track containing the sub-image bitstream with the user's current field of view (FOV) data, a representation of the first quality is adopted and transmitted at the first resolution. For the track containing the sub-image bitstream with regions other than the user's FOV data, a representation of the second quality is adopted and transmitted at the second resolution. Wherein, the first quality is higher than the second quality, the first resolution is higher than the second resolution, and the farther the region is from the FOV, the lower the quality of the corresponding track. The resolution of the downsampled sub-image video stream is lower than the resolution of the captured sub-image video stream; The codewords used in the entropy coding include codewords assigned to each available option in a set of available options, and the set of available options is determined from the set of all possible options. The length of the codeword depends on the number of available options.
4. The method according to claim 3, characterized in that, the first specific value is 1.
5. The method according to claim 3, characterized in that, the indication attribute is general_non_packed_constraint_flag.
6. The method according to claim 3, characterized in that, the fisheye omnidirectional video image is an image captured by a fisheye camera.
7. The method according to any one of claims 3-6, characterized in that, the fisheye omnidirectional video image is an image including a plurality of active regions.
8. An encoding method performed by a video encoder, characterized in that, the method includes: Encoding the flags and representations of video data into a bitstream, where the flags include a first specific value, indicating to a video decoder receiving the bitstream that the bitstream does not contain any fisheye omnidirectional video images; Transmitting the bitstream to the video decoder; The video data includes a sub-image video stream, and the bitstream further includes a track containing a sub-image bitstream, which is obtained by downsampling a captured sub-image video stream and then performing view-based entropy coding. Among them, for the track containing the sub-image bitstream with the user's current field of view (FOV) data, a representation of the first quality is adopted and sent at the first resolution. For the track containing the sub-image bitstream with regions other than the user's FOV data, a representation of the second quality is adopted and sent at the second resolution. Wherein, the first quality is higher than the second quality, the first resolution is higher than the second resolution, and the farther the region is from the FOV, the lower the quality of the corresponding track. The resolution of the downsampled sub-image video stream is lower than the resolution of the captured sub-image video stream; The codewords used in the entropy coding include codewords assigned to each available option in a set of available options, and the set of available options is determined from the set of all possible options. The length of the codeword depends on the number of available options.
9. The method according to claim 8, characterized in that, the first specific value is 1.
10. The method according to claim 8, characterized in that, The method further includes: Encoding a flag and a representation of video data into a bitstream, where the flag includes a second specific value, indicating to a video decoder receiving the bitstream that the bitstream may include one or more fisheye omnidirectional video images; Transmitting the bitstream to the video decoder.
11. The method according to claim 10, wherein, the flag having the second specific value indicates to the video decoder to search for an omnidirectional fisheye supplemental enhancement information (SEI) message in the bitstream.
12. The method according to claim 11, wherein, the omnidirectional fisheye SEI message indicates that the images in the bitstream are fisheye omnidirectional video images.
13. The method according to any one of claims 10 - 12, wherein, the second specific value is 0.
14. The method according to any one of claims 8 - 12, wherein, the flag is general_non_packed_constraint_flag.
15. The method according to any one of claims 8 - 12, wherein, the flag is set in the sequence parameter set (SPS) of the bitstream.
16. The method according to any one of claims 8 - 12, wherein, the fisheye omnidirectional video image is an image captured by a fisheye camera.
17. The method according to any one of claims 8 - 12, wherein, the fisheye omnidirectional video image is an image including a plurality of active regions.
18. A decoding method performed by a video decoder, wherein, the method includes: Receiving an encoded bitstream, where the encoded bitstream includes a flag and a representation of video data, and the flag includes a first specific value, indicating to the video decoder that the bitstream does not include any fisheye omnidirectional video images; Decoding the encoded bitstream according to the specific value; The video data includes a sub - image video stream, and the encoded bitstream further includes a track including a sub - image bitstream, which is obtained by performing downsampling on the captured sub - image video stream and then performing view - based entropy coding. Among them, for the track including the sub - image bitstream with the user's current field of view (FOV) data, a representation of the first quality is adopted and transmitted at the first resolution, and for the track including the sub - image bitstream with regions other than the user's FOV data, a representation of the second quality is adopted and transmitted at the second resolution, where the first quality is higher than the second quality, the first resolution is higher than the second resolution, and the farther the region is from the FOV, the lower the quality of the corresponding track. The resolution of the downsampled sub - image video stream is lower than the resolution of the captured sub - image video stream; The codewords adopted by the entropy coding include codewords assigned to each available option in a set of available options, and the set of available options is determined from a set of all possible options. The length of the codewords depends on the number of available options.
19. The method according to claim 18, wherein, the first specific value is 1.
20. The method according to claim 18, wherein, the method further comprises: receiving an encoded bitstream, wherein the encoded bitstream includes a flag and a representation of video data, the flag includes a second specific value, and indicates to the video decoder that the bitstream may contain one or more fisheye omnidirectional video images; decoding the encoded bitstream according to the specific value.
21. The method according to claim 20, wherein, the flag having the second specific value indicates that the video decoder searches for an omnidirectional fisheye supplemental enhancement information (SEI) message in the bitstream.
22. The method according to claim 21, wherein, the omnidirectional fisheye SEI message indicates that the image in the bitstream is a fisheye omnidirectional video image.
23. The method according to claim 20, wherein, the second specific value is 0.
24. The method according to any one of claims 18 - 23, wherein, the flag is general_non_packed_constraint_flag.
25. The method according to any one of claims 18 - 23, wherein, the flag is set in the sequence parameter set (SPS) of the bitstream.
26. The method according to any one of claims 18 - 23, wherein, the fisheye omnidirectional video image is an image captured by a fisheye camera.
27. The method according to any one of claims 18 - 23, wherein, the fisheye omnidirectional video image is an image including a plurality of active regions.
28. A method for processing video data, wherein, the method comprises: judging whether a video decoder can correctly present a decoded image obtained by decoding the bitstream according to an indication attribute of the bitstream, wherein the bitstream includes an encoded representation of video data, and the indication attribute includes a second specific value, indicating that the bitstream may contain one or more fisheye omnidirectional video images; processing the bitstream according to the judgment; the video data includes a sub - image video stream, and the bitstream further includes a track containing a sub - image bitstream, the track is obtained by performing downsampling on the captured sub - image video stream and then performing view - based entropy coding. For a track containing a sub - image bitstream with user's current field - of - view (FOV) data, a representation of the first quality is used and transmitted at the first resolution. For a track containing a sub - image bitstream with regions other than the user's FOV data, a representation of the second quality is used and transmitted at the second resolution, wherein the first quality is higher than the second quality, the first resolution is higher than the second resolution, and the farther the region is from the FOV, the lower the quality of the corresponding track, and the resolution of the downsampled sub - image video stream is lower than the resolution of the captured sub - image video stream; The codewords used in the entropy coding include codewords assigned to each available option in a set of available options, the set of available options being determined from a set of all possible options, and the length of the codewords depending on the number of available options.
29. The method according to claim 28, wherein, the second specific value is 0.
30. The method according to claim 28, wherein, the indication attribute is general_non_packed_constraint_flag.
31. The method according to claim 28, wherein, the fisheye omnidirectional video image is an image captured by a fisheye camera.
32. The method according to any one of claims 28 - 31, wherein, the fisheye omnidirectional video image is an image including a plurality of active regions.
33. An encoding device, wherein, the encoding device comprises: a receiver for receiving an image for encoding or receiving a bitstream for decoding; a transmitter coupled to the receiver, wherein the transmitter is configured to transmit the bitstream to a decoder or transmit a decoded image to a display; a memory coupled to at least one of the receiver and the transmitter, wherein the memory is configured to store instructions; a processor coupled to the memory, wherein the processor is configured to execute the instructions stored in the memory to perform the method according to any one of claims 1 to 32.
34. The encoding device according to claim 33, wherein, the encoding device further comprises: a display for displaying an image.
35. A codec system, wherein, the codec system comprises: an encoder; a decoder in communication with the encoder, wherein the encoder or the decoder comprises the encoding device according to claim 33 or 34.