Volume visual media processing method and apparatus
By dividing the 3D scene into multiple view groups and using multi-track encapsulation technology, the problem of rendering multiple views in existing technologies is solved, achieving efficient 3D scene rendering from different viewpoints and improving the user experience.
Patent Information
- Application Number
- CN202080094122.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-15
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2040-04-15
AI Technical Summary
Existing video encoding technologies struggle to efficiently represent and render 3D visual scenes, especially when users want to view 3D scenes from different viewpoints, and cannot effectively support the encoding and decoding of multiple views.
Using group-based encoding and rendering technology, the 3D scene is divided into multiple view groups. Through view group information structure and multi-track encapsulation technology, the encoding and decoding of multiple views are realized, supporting rendering from different viewpoints.
It enables efficient rendering of 3D scenes from different viewpoints, supports encoding and decoding of multiple views, and enhances the immersive experience of the user experience.
Smart Images

Figure CN115039404B_ABST
Abstract
Description
Technical Field
[0001] This patent document relates to volumetric visual media processing and transmission technology. Background Technology
[0002] Video coding uses compression tools to encode two-dimensional video frames into a compressed bitstream representation, which is more efficient for storage or transmission over a network. Traditional video coding techniques that use two-dimensional video frames are sometimes inefficient for representing visual information in three-dimensional scenes. Summary of the Invention
[0003] This patent document specifically describes techniques for encoding and decoding digital video carrying visual information related to volumetric visual media.
[0004] In one example aspect, a method for volumetric visual data processing is disclosed. The method includes: decoding a bitstream by a decoder, the bitstream containing volumetric visual information of a 3D scene represented as one or more atlas sub-bitstreams and one or more coded video sub-bitstreams; reconstructing the 3D scene using the decoding results of the one or more atlas sub-bitstreams and the one or more coded video sub-bitstreams; and rendering a target view of the 3D scene based on a desired viewing position and / or desired viewing orientation.
[0005] In another example, a method for generating a bitstream that includes volumetric visual information is disclosed. The method includes: generating a bitstream containing volumetric visual information of a three-dimensional scene by an encoder using one or more atlas sub-bitstreams and one or more coded video sub-bitstreams; and including information in the bitstream that enables rendering a target view of the three-dimensional scene based on a desired viewing position and / or a desired viewing direction.
[0006] In another example aspect, an apparatus for implementing one or more of the methods described above is disclosed. The apparatus may include a processor configured to implement the described encoding or decoding methods.
[0007] In yet another example, a computer program storage medium is disclosed. This computer program storage medium includes code stored thereon. When executed by a processor, this code causes the processor to implement the described methods.
[0008] These and other aspects are described in this document. Attached Figure Description
[0009] Figure 1 An example processing flow for group-based encoding used in atlas generation is shown.
[0010] Figure 2An example of multi-track encapsulation of a V-PCC bitstream with atlas groups is shown.
[0011] Figure 3 An example of a multitrack encapsulation of a V-PCC bitstream with multiple atlases is shown.
[0012] Figure 4 This is a flowchart of an example volumetric visual media processing method.
[0013] Figure 5 This is a flowchart of an example volumetric visual media processing method.
[0014] Figure 6 This is a block diagram illustrating an example of a volumetric visual media data encoding device according to the present technology.
[0015] Figure 7 This is a block diagram illustrating an example of a volumetric visual media data processing device according to the present technology.
[0016] Figure 8 This is a block diagram of the hardware platform used to implement the volumetric visual media processing method described in this paper. Detailed Implementation
[0017] The use of section headings in this document is solely for readability purposes and does not limit the scope of the embodiments and techniques disclosed in each section to that section. Examples from the H.264 / AVC and H.265 / HEVC, MPEG, and MPEG-DASH standards are used to describe certain features. However, the applicability of the disclosed techniques is not limited to these standards.
[0018] Various syntax elements are disclosed in different sections of this document for point cloud data processing. However, it should be understood that syntax elements with the same name will have the same format and syntax as used in different sections, unless otherwise stated. Furthermore, in different embodiments, different syntax elements and structures described under different section headings may be combined. Moreover, although specific structures are described as embodiments, it should be understood that the order of the various entries of the syntax structure may be changed unless otherwise stated in this document.
[0019] 1. Brief discussion
[0020] Traditionally, the capture, processing, storage, and presentation of digital visual media (such as images and videos) have utilized two-dimensional, frame-based capture of visual scenes. In recent years, there has been a growing interest in extending the user experience to three dimensions. Various industry standards have begun to address issues related to the capture, transmission, and presentation of 3D visual scenes. Notably, a set of technologies uses traditional frame-based (2D) video coding tools to encode 3D visual information by projecting 3D information onto a 2D plane.
[0021] Two noteworthy technologies include the use of video-based pointcloud compression (V-PCC) and the Moving Picture Experts Group (MPEG) Immersive Video (MIV) proposal.
[0022] 1.1 Video-based point cloud compression (V-PCC)
[0023] Video-based point cloud compression (V-PCC) represents the volumetric encoding of point cloud visual information and enables efficient capture, compression, reconstruction, and rendering of point cloud data using MPEG video codecs such as AVC, HEVC, and VVC. A V-PCC bitstream containing a encoded point cloud sequence (CPCS) includes VPCC units carrying sequence parameter set (SPS) data, an atlas information bitstream, a 2D video-coded occupancy map bitstream, a 2D video-coded geometry bitstream, and zero or more 2D video-coded attribute bitstreams. Each V-PCC unit has a V-PCC unit header describing its type and a V-PCC unit payload. The payloads of the occupancy, geometry, and attribute V-PCC units correspond to video data units (e.g., HEVC NAL units) that can be decoded by the video decoder specified in the corresponding occupancy, geometry, and attribute parameter set V-PCC unit.
[0024] 1.2 Carrying of V-PCC in ISOBMFF
[0025] In a V-PCC primary stream, V-PCC units are mapped to individual tracks in an ISOBMFF file based on their type. Multitrack ISOBMFF V-PCC containers contain two types of tracks: V-PCC tracks and V-PCC component tracks. ISOBMFF is a popular file format used for representing multiple tracks of digital video and audio information.
[0026] The V-PCC track carries the volumetric visual information from the V-PCC bitstream, including the patch information sub-bitstream and sequence parameter set. The V-PCC component track is a constrained video scheme track that carries the occupancy map, geometry, and attribute sub-bitstream's 2D video encoded data from the V-PCC bitstream. Based on this layout, the V-PCC ISOBMFF container should include the following:
[0027] The V-PCC track contains a sequence parameter set (in the sample ingress) and a sample carrying payloads of sequence parameter set V-PCC units (unit type VPCC_VPS) and atlas V-PCC units (unit type VPCC_AD). This track also includes track references to other tracks carrying payloads of video compression V-PCC units (i.e., unit types VPCC_OVD, VPCC_GVD, and VPCC_AVD).
[0028] The restricted video scheme track, in which the samples contain access units of the video-coded basic stream of occupied graph data (i.e., payloads of V-PCC units of type VPCC_OVD).
[0029] One or more restricted video scheme tracks, wherein the samples contain access units of the video-coded basic stream (i.e., payloads of V-PCC units of type VPCC_GVD) containing geometric data.
[0030] Zero or more restricted video scheme tracks, wherein the samples contain access units of the video-coded basic stream (i.e., payloads of V-PCC units of type VPCC_AVD) containing attribute data.
[0031] 1.3 MPEG Immersive Video (MIV)
[0032] MPEG is developing an international standard (ISO / IEC 23090-12), namely MPEG Immersive Video (MIV), to support the compression of immersive video content, in which real or virtual 3-D scenes are captured by multiple real or virtual cameras. MIV content supports the playback of three-dimensional (3D) scenes in 6 degrees of freedom (6DoF) within a limited range of viewing positions and orientations.
[0033] While both MIV and V-PCC technologies aim to provide a similar end-user experience of viewing 3D scenes and objects, these solutions employ some different approaches. For example, MIV promises to provide view-based access to 3D volumetric visual data, while V-PCC offers projection-based access. Therefore, MIV promises to provide a more realistic, user-controlled user experience and will offer viewers a more immersive experience. However, it remains beneficial to utilize some of the existing bitstream syntax and file format information available in V-PCC to ensure rapid and compatible adoption of MIV.
[0034] 2. Example problems to consider at the encoder end
[0035] On the encoder side of MIV, the view representation is a 2D sample array with at least a depth / occupancy component and optional texture and solid components, representing a 3D scene projected onto a surface using view parameters. View parameters define the projection used to generate the view representation from the 3D scene, including intrinsic and extrinsic parameters. In this context, the source view refers to the source video material prior to encoding, which corresponds to the format of the view representation. This can be obtained by capturing the 3D scene with a real camera or by projecting it onto a surface using source camera parameters via a virtual camera.
[0036] 2.1 Group-based encoder
[0037] Group-based encoders are top-level encoders for MIV. They split a view into multiple view groups and encode each view group independently using multiple single-group encoders. The source view is distributed across multiple single-group encoders, each with a view optimizer and an atlas builder. The view optimizer labels the source view as a base view or an additional view, while the atlas builder takes the base view and additional view along with their parameters, as well as the output atlas and related parameters, as input.
[0038] MPEG video codecs (such as HEVC (High Efficiency Video Coding) encoders) will be used to encode the texture and depth of the atlas. The resulting attribute and geometric video streams will be multiplexed along with MIV metadata to form the final MIV bitstream.
[0039] 3. Example problems to consider at the decoding end
[0040] The MIV decoder processes the parsing and decoding of the MIV bitstream to output decoded geometric images, texture attribute images, and MIV metadata frame by frame.
[0041] For the rendering portion of the MIV decoder, the MIV rendering engine reconstructs the geometry frame at nominal atlas resolution, then converts samples of the decoded geometry frame, upscaled at nominal atlas resolution, into floating-point depth values in meters. The output of the MIV decoder is a perspective viewport or omnidirectional view based on the desired viewing posture, thus enabling motion parallax cues within a limited space. To this end, the MIV rendering engine performs the reconstruction of the reconstructed view and the projection pixels of the reconstructed view onto the viewport.
[0042] In V-PCC-based representations of 3D scenes, a fixed number of projections of the 3D visual media are represented in the bitstream. For example, six projections corresponding to the six surfaces of a bounding box can be converted into 2D visual images and encoded using conventional video codec techniques. However, V-PCC cannot support the user experience of viewing a 3D scene from different viewpoints rather than viewing a limited number of projections of the 3D scene. Therefore, in this viewpoint-based rendering of volumetric video data, it is currently unknown how to represent such visual data at the bitstream level (e.g., representing the actual scene in bits), at the file level (e.g., organizing media data into logical file groups), or at the system level (e.g., transport and metadata levels) to allow the encoder to construct a bitstream representing the 3D volumetric data so that the renderer at the decoder can parse the bitstream and retrieve the media data based on the user's desired viewpoint.
[0043] Furthermore, it remains unclear how to extend the current organization of V-PCC tracks to accommodate the use of multiple views in MIV. For example, how to map V-PCC tracks to the desired views used to render a 3D scene is unknown. For instance, a MIV implementation could use 10, 40, or even 100 different views, which could be encoded in a bitstream. It is also unknown how to use the track structure to signal the different views so that the decoder or renderer can parse the system layer of the bitstream to locate the desired video or image tracks and render the view from the viewer's desired location or viewpoint.
[0044] Various embodiments are disclosed in this document to address the aforementioned problems. For example, as further described in this document, solutions are provided to encode and decode multiple views in a view group and to use one or more sub-streams for atlases, as further described in this document.
[0045] 3.1 Group-based renderer
[0046] Group-based renderers can render from local patches within each atlas group separately. The renderer process includes a group selection stage, multiple channels, and a merging stage. Each channel runs the compositor using a different atlas group and outputs composite intermediate views. The merging stage combines all intermediate composite views into a final desired viewport, such as a target view that indicates a perspective viewport or omnidirectional view in the desired viewing position and orientation.
[0047] 3.2 Carrying V-PCC data with multiple atlases
[0048] Despite differences in intended applications, input data formats, and rendering, Video-Based Point Cloud Compression (V-PCC) and MPEG Immersive Video (MIV) share the same core tools for representing information in the coding domain: splitting 3D spatial data into 2D patch maps and encoding them as 2D atlas frames. Therefore, a V-PCC base bitstream can contain more than one atlas to carry MIV content.
[0049] To support efficient access, delivery, and rendering of volumetric visual media compressed into MPEG immersive video as defined in ISO / IEC 23090-12 in a 6DOF environment, it is necessary to specify the storage format of V-PCC bitstreams with multiple atlases.
[0050] 3.3 Example File Format
[0051] Generally, embodiments based on the disclosed technology can be used for video data processing. In some embodiments, omnidirectional video data is stored in a file based on the ISO (International Organization for Standardization) Basic Media File Format. The ISO Basic Media File Format, such as constrained scheme information boxes, track reference boxes, and track group boxes, can be operated with reference to the ISO Basic Media File Format of ISO / IEC JTC1 / SC29 / WG11 Moving Picture Experts Group (MPEG) MPEG-4 Part 12.
[0052] All data in the ISO Basic File Format is contained within boxes. An ISO Basic File Format file, represented by an MP4 file, consists of several boxes, each with a type and length and can be considered a data object. A box can contain another box, called a container box. An MP4 file will initially have only one box of type "ftyp," serving as a marker for the file format and containing some information about the file. There will also be only one type of box, "MOOV" (movie box), which is the container box, and its sub-boxes contain the media's metadata information. The media data of the MP4 file is contained within "mdat" type media boxes (media data boxes), which are also container boxes. These container boxes may or may not be available (when the media data references other files), and the structure of the media data includes metadata.
[0053] Timing metadata tracks are a mechanism in the ISO Basic Media File Format (ISOBMFF) that establishes timing metadata associated with a specific sample. Timing metadata is less coupled to media data and is typically "descriptive".
[0054] Each volumetric vision scene can be represented by a unique volumetric vision track. An ISOBMFF file can contain multiple scenes, and therefore multiple volumetric vision tracks can exist within the file.
[0055] As previously mentioned, this document provides several technical solutions to allow the representation of 3D or spatial regions of point cloud data (such as MPEG V-PCC data) in formats compatible with traditional 2D video formats such as MP4 or ISOBMFF. One advantage of the proposed solutions is their ability to reuse traditional 2D video technologies and syntax to implement new functionalities.
[0056] 4. Solution 1
[0057] In some embodiments, a new syntactic structure, referred to as the view group information structure, can be encoded into a bitstream by an encoder and correspondingly decoded by a decoder to render the desired view of the 2D scene to a display. The syntactic structure and some example implementations of the associated encoding and decoding techniques are described herein.
[0058] 4.1 Example 1
[0059] Example view group information structure
[0060] definition
[0061] ViewGroupInfoStruct provides view group information for volumetric visual media such as MIV content captured and processed during the encoding phase, including at least: view group identifier, view group description, number of views, view identifier, and camera parameters for each view.
[0062] grammar
[0063]
[0064] Semantics
[0065] view_group_id provides the identifier for the view group.
[0066] view_group_description is a null-terminated UTF-8 string that provides a textual description of the view group.
[0067] num_views specifies the number of views in a view group.
[0068] view_id provides the identifier for a given view in a view group.
[0069] A basic_view_flag value of 1 indicates that the relevant view is selected as the basic view. A basic_view_flag value of 0 indicates that the relevant view is not selected as the basic view.
[0070] A value of 1 for `camera_parameters_included_flag` indicates that `CameraParametersStruct` exists. A value of 0 for `camera_parameters_included_flag` indicates that `CameraParametersStruct` does not exist.
[0071] Camera parameter structure
[0072] definition
[0073] CameraParametersStruct provides real or virtual camera position and orientation information, which can be used to render V-PCC or MIV content as a perspective or omnidirectional view of the desired viewing position and orientation.
[0074] During the decoding phase, the group-based renderer can use this information to calculate the distance from the view group to the desired pose being composited. The view-weighted compositor can use this information to calculate the distance between the view position and the target viewport position.
[0075] grammar
[0076]
[0077] camera_id provides an identifier for a given real or virtual camera.
[0078] A camera_pos_present value of 1 indicates the presence of camera position parameters. A camera_pos_present value of 0 indicates the absence of camera position parameters.
[0079] A camera_ori_present value of 1 indicates the presence of camera orientation parameters. A camera_ori_present value of 0 indicates the absence of camera orientation parameters.
[0080] A camera_fov_present value of 1 indicates the presence of camera field of view parameters. A camera_fov_present value of 0 indicates the absence of camera field of view parameters.
[0081] A camera_depth_present value of 1 indicates that a camera depth parameter exists. A camera_depth_present value of 0 indicates that a camera depth parameter does not exist.
[0082] `camera_pos_x`, `camera_pos_y`, and `camera_pos_z` indicate the x, y, and z coordinates of the camera position in the global reference coordinate system, respectively, in meters. Values should be in increments of 2. -16 The unit is meters.
[0083] `camera_quat_x`, `camera_quat_y`, and `camera_quat_z` indicate the x, y, and z components of the camera orientation, represented using quaternions. These values should be floating-point values in the range of -1 to 1 (inclusive). These values specify the x, y, and z components used to represent the rotation from the global coordinate axes to the camera's local coordinate axes using quaternions, i.e., `qX`, `qY`, and `qZ`. The fourth component of the quaternion `qW` is calculated as follows:
[0084] qW = sqrt(1-(qX) 2 +qY 2 +qZ 2 ))
[0085] The point (w, x, y, z) represents a rotation about the axis pointed to by the vector (x, y, z) up to an angle 2*cos^{-1}(w)=2*sin^{-1}(sqrt(x^{2}+y^{2}+z^{2})).
[0086] camera_hor_range indicates the horizontal field of view of the view frustum associated with the camera, in radians. This value should be in the range of 0 to 2π.
[0087] camera_ver_range indicates the vertical field of view of the view frustum associated with the camera, in radians. This value should be in the range of 0 to π.
[0088] `camera_near_depth` and `camera_far_depth` indicate the near and far depths (or distances) based on the near and far planes of the view frustum associated with the camera. These values should be in increments of 2. -16 The unit is meters.
[0089] Example of V-PCC parameter track
[0090] V-PCC Parameter Track Example Entry
[0091] Example entry type: "vpcp"
[0092] Container: SampleDescriptionBox
[0093] Mandatory: Yes
[0094] Quantity: There can be one or more sample entry points.
[0095] The V-PCC parameter track should use VPCCParametersSampleEntry, which uses the sample entry type "vpcp" to extend VolumetricVisualSampleEntry.
[0096] The VPCC parameter track sample entry should contain a VPCCConfigurationBox and a VPCCUnitHeaderBox.
[0097] grammar
[0098]
[0099] Semantics
[0100] The VPCCConfigurationBox should contain the V-PCC parameter set of the multi-atlas V-PCC bitstream, that is, the V-PCC unit of vuh_unit_type equal to VPCC_VPS.
[0101] The VPCCConfigurationBox should contain only non-ACLNAL units common to all V-PCC tracks of the multi-atlas V-PCC data (including but not limited to NAL_ASPS, NAL_AAPS, NAL_PREFIX_SEI or NAL_SUFFIX_SEI NAL units), as well as EOB and EOS NAL units (if present).
[0102] For different V-PCC track groups, the VPCCConfigurationBox can contain different values for the NAL_AAPS atlas NAL cells.
[0103] V-PCC Track Grouping
[0104] MIV's group-based encoder can divide the source view into multiple groups. It takes the source camera parameters as input and the number of groups as a preset, and outputs a list of views to be included in each group.
[0105] Grouping forces the atlas builder to output local coherent projections of important regions in the atlas (e.g., those belonging to foreground objects or occluded areas), thereby improving both subjective and objective results, especially for natural content or at high bitrate levels.
[0106] Figure 1 An example of the processing flow for group-based encoding used in atlas generation is described.
[0107] like Figure 1 As shown, during the group encoding phase, each single-group encoder generates metadata with its own index atlas or view. A unique group ID is assigned to each group, and this unique group ID is appended to the atlas parameters of the relevant group. To enable the renderer to correctly interpret the metadata and properly map patches across all views, the merger renumbers the atlas and view IDs for each patch and merges the trimmed graphics. Each base view is carried in the atlas as a single, fully occupied patch (assuming the atlas size is equal to or greater than the base view size), or (if the aforementioned assumption is not met) split into multiple atlases. Additional views are trimmed into multiple patches, which may be carried in the same atlas along with patches of the base view if the atlas is larger or in a separate atlas.
[0108] like Figure 1 As shown, all atlases generated by the atlas builder from the same view group should be grouped together as an atlas group. For group-based rendering, the decoder needs to decode patches within one or more atlas groups corresponding to one or more view groups, from which one or more views of volumetric visual data (e.g., MIV content) are selected for target view rendering.
[0109] The decoder can select one or more views of volumetric visual data for a target view based on one or more view group information, as described in the example view group information structure, where each view group information describes one or more views and each view group information includes camera parameters for one or more views.
[0110] Figure 2 An example of a multitrack encapsulation of a V-PCC bitstream with atlas groups is shown.
[0111] like Figure 2 As shown, before decoding the atlas group, the file parser needs to determine and decapsulate a set of volumetric visual tracks (e.g., V-PCC track groups) corresponding to the atlas group based on the syntax elements of the volumetric visual parameter tracks in the file storage of the bitstream (e.g., VPCCViewGroupsBox of V-PCC parameter tracks); where the set of volumetric visual tracks and volumetric visual parameter tracks carry all the atlas data of the atlas group.
[0112] The file parser can identify volumetric visual parameter tracks based on specific sample ingress types. In the case of V-PCC parameter tracks, the sample ingress type "vpcp" should be used to identify the V-PCC parameter track, and the V-PCC parameter track specifies a constant parameter set and common atlas data for all referenced V-PCC tracks with a specific track reference.
[0113] For the storage of a V-PCC bitstream with multiple atlases, all V-PCC tracks corresponding to all atlases from the same atlas group should be indicated by a track group of type "vptg".
[0114] definition
[0115] TrackGroupTypeBox with track_group_type equal to "vptg" indicates that this V-PCC track belongs to a group of V-PCC tracks corresponding to the atlas group.
[0116] V-PCC tracks belonging to the same atlas group have the same track_group_id value for track_group_type "vptg", and the track_group_id of a track from one atlas group is different from the track_group_id of a track from any other atlas group.
[0117] grammar
[0118] aligned(8)class VPCCTrackGroupBox extends trackGroupTypeBox(′Vptg′){}
[0119] Semantics
[0120] V-PCC tracks with the same track_group_id value in a TrackGroupTypeBox with track_group_type equal to "vptg" belong to the same atlas group. Therefore, the track_group_id in a TrackGroupTypeBox with track_group_type equal to "vptg" is used as the identifier of the atlas group.
[0121] Static view group information box
[0122] definition
[0123] Static view groups used for volumetric visual media (such as MIV content) and their corresponding associated V-PCC track groups should be signaled in the VPCCViewGroupsBox.
[0124] grammar
[0125]
[0126] Semantics
[0127] num_view_groups indicates the number of view groups used for MIV content.
[0128] vpcc_track_group_id identifies the group used for a V-PCC track, which carries all atlas data for a related view group of volumetric visual media such as MIV content.
[0129] Dynamic view group information
[0130] If a V-PCC parameter track has an associated timing metadata track with a sample inlet type of "dyvg", then the source view group defined for the MIV stream carried by the V-PCC parameter track is considered a dynamic view group (i.e., the view group information can change dynamically over time).
[0131] The associated timing metadata track should contain a "cdsc" track reference to the V-PCC parameter track carrying the atlas stream.
[0132] Sample entry
[0133]
[0134] Sample format
[0135] grammar
[0136]
[0137] Semantics
[0138] `num_view_groups` indicates the number of view groups in the sample that are signaling updates. It is not necessarily equal to the total number of available view groups. Only view groups whose source views are being updated exist in the sample.
[0139] ViewGroupInfoStruct() is defined in the preceding sections of Example 1. If camera_parameters_included_flag is set to 0, this indicates that the camera parameters of the view group have been previously signaled in a previous sample or in a previous instance of ViewGroupInfoStruct with the same view_group_id in the sample entry.
[0140] 4.2 Example Implementation 2
[0141] Encapsulation and Signaling in MPEG-DASH
[0142] Each V-PCC component track should be represented as a separate V-PCC component AdaptationSet in the DASH manifest (MPD) file. Each V-PCC track should also be represented as a separate V-PCC atlas AdaptationSet. Additional AdaptationSets for common atlas information are used as the main AdaptationSet for the V-PCC content. If a V-PCC component has multiple layers, a separate AdaptationSet can be used to signal each layer.
[0143] The main AdaptationSet should have its @codecs attribute set to "vpcp", and the atlas AdaptationSet should have its @codecs attribute set to "vpc1". The @codecs attribute of V-PCC component AdaptationSets or Representations (in the absence of a signal to @codecs for the AdaptationSet element) is set based on the corresponding codec used to encode the component.
[0144] The main AdaptationSet should contain a single initialization segment at the adaptation set level. The initialization segment should contain all sequence parameter sets and non-ACL NAL units common to all V-PCC tracks required for initializing the V-PCC decoder (including the V-PCC parameter sets for multi-atlas V-PCC bitstreams, and NAL_ASPS, NAL_AAPS, NAL_PREFIX_SEI, or NAL_SUFFIX_SEI NAL units), as well as EOB and EOS NAL units (if present).
[0145] The Atlas AdaptationSet should contain a single initialization segment at the adaptation set level. The initialization segment should contain all the sequence parameter sets required for decoding the V-PCC tracks, including the V-PCC atlas sequence parameter set and other parameter sets for the component substreams.
[0146] The media segment used for the representation of the main AdaptationSet should contain one or more track segments of the V-PCC parametric track. The media segment used for the representation of the atlas AdaptationSet should contain one or more track segments of the V-PCC track. The media segment used for the representation of the component AdaptationSet should contain one or more track segments of the corresponding component track at the file format level.
[0147] V-PCC Pre-selection
[0148] V-PCC preselection is signaled in MPD using the PreSelection element defined in MPEG-DASH (ISO / IEC 23009-1). This element has a list of IDs with the @preselectionComponents attribute, including the ID of the main AdaptationSet of the point cloud, followed by the ID of the atlas AdaptationSet and the ID of the AdaptationSet corresponding to the point cloud components. The @codecs attribute of PreSelection should be set to "vpcp", indicating that the PreSelection media is a video-based point cloud. PreSelection can be signaled using a PreSelection element within a Period element or a preselection descriptor at the adapter set level.
[0149] V-PCC descriptor
[0150] An EssentialProperty element with the @schemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2019:vpc" is called a VPCC descriptor. At most one VPCC descriptor can exist at the adaptation set level of the main AdaptationSet in a point cloud.
[0151] Table 1 Attributes of VPCC Descriptors
[0152]
[0153]
[0154] VPCCViewGroups descriptor
[0155] To identify static view groups and their corresponding associated V-PCC track groups within the main AdaptationSet of V-PCC content, the VPCCViewGroups descriptor should be used. VPCCViewGroups is an EssentialProperty or SupplementalProperty descriptor with the @schemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2020:vpvg".
[0156] At most one single VPCCViewGroups descriptor should exist at the AdaptationSet level, the representation level in the main AdaptationSet, or the preselected level used for point cloud content.
[0157] The @value attribute of the VPCCViewGroups descriptor should not exist. The VPCCViewGroups descriptor should include the elements and attributes specified in Table 2.
[0158] Table 2 Elements and Attributes of the VPCCViewGroups Descriptor
[0159]
[0160]
[0161]
[0162]
[0163] Dynamic View Group
[0164] When the view group is dynamic, the timing metadata track used to signal the view information of each view group in the representation timeline should be carried in a separate AdaptationSet with a single representation and associated (linked) with the main V-PCC track using the @associationId attribute defined in ISO / IEC 23009-1 [MPEG-DASH], and the @associationType value includes 4CC "vpcm" for the corresponding AdaptationSet or Representation.
[0165] 5. Solution 2
[0166] 5.1 Example 3
[0167] Example view information structure
[0168] definition
[0169] ViewInfoStruct provides view information for MIV content captured and processed during the encoding phase, including at least: the view identifier, the identifier of the view group to which it belongs, the view description, and the view's camera parameters.
[0170] grammar
[0171]
[0172] Semantics
[0173] view_id provides the identifier for the view.
[0174] view_group_id provides the identifier of the view group to which it belongs.
[0175] view_description is a null-terminated UTF-8 string that provides a textual description of the view.
[0176] A basic_view_flag value of 1 indicates that the relevant view is selected as the basic view. A basic_view_flag value of 0 indicates that the relevant view is not selected as the basic view.
[0177] A value of 1 for `camera_parameters_included_flag` indicates that `CameraParametersStruct` exists. A value of 0 for `camera_parameters_included_flag` indicates that `CameraParametersStruct` does not exist.
[0178] CameraParametersStruct() is defined in the previous chapter of Example 1.
[0179] Static view information box
[0180] Figure 3 An example of a multitrack encapsulation of a V-PCC bitstream with multiple atlases is shown.
[0181] For target view rendering, the decoder needs to decode one or more patches in the atlas that correspond to one or more views of the volumetric visual data (e.g., MIV content) that have been selected for target view rendering.
[0182] The decoder can select one or more views of volumetric visual data for a target view based on view information of one or more views, as described in the example view information structure, where each view information describes the camera parameters of the corresponding view.
[0183] like Figure 3 As shown, before decoding one or more atlases, the file parser needs to determine and decapsulate one or more volumetric visual tracks (e.g., V-PCC tracks) corresponding to one or more atlases based on the syntax elements of the volumetric visual parameter tracks (e.g., VPCCViewsBox of V-PCC parameter tracks) in the file storage of the bitstream; wherein the one or more volumetric visual tracks and volumetric visual parameter tracks carry all atlas data of the atlas.
[0184] The file parser can identify volumetric visual parameter tracks based on specific sample ingress types. In the case of V-PCC parameter tracks, the sample ingress type "vpcp" should be used to identify the V-PCC parameter track, and the V-PCC parameter track specifies a constant set of parameters and common atlas data for all referenced V-PCC tracks with a specific track reference.
[0185] definition
[0186] The source view of the MIV content and its corresponding atlas should be signaled in the VPCCViewsBox.
[0187] grammar
[0188] Box type: "vpvw"
[0189] Container: VPCCParametersSampleEntry("vpcp")
[0190] Must: No
[0191] Quantity: zero or one
[0192]
[0193] Semantics
[0194] num_views indicates the number of source views in the MIV content.
[0195] num_vpcc_tracks indicates the number of V-PCC tracks associated with the source view.
[0196] vpcc_track_id identifies a V-PCC track that carries atlas data of the associated source view.
[0197] Dynamic view information
[0198] If a V-PCC parameter track has an associated timing metadata track with a sample inlet type of "dyvw", then the source view defined for the MIV stream carried by the V-PCC parameter track is considered a dynamic view (i.e., the view information can change dynamically over time).
[0199] The associated timing metadata track should contain a "cdsc" track reference to the V-PCC parameter track carrying the atlas stream.
[0200] Sample entry
[0201]
[0202] Sample format
[0203] grammar
[0204]
[0205] Semantics
[0206] `num_views` indicates the number of views in the sample that have signaled notifications. This may not necessarily equal the total number of available views. Only views in the sample whose view information is being updated exist.
[0207] ViewInfoStruct() is defined in the preceding sections of Example 2. If camera_parameters_included_flag is set to 0, this indicates that the view's camera parameters were previously signaled in a previous sample or in a previous instance of ViewInfoStruct with the same view_id in the sample entry.
[0208] 5.2 Example 4
[0209] Examples of encapsulation and signaling in MPEG-DASH
[0210] V-PCC descriptor
[0211] An EssentialProperty element with the @schemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2019:vpc" is called a VPCC descriptor. At most one VPCC descriptor can exist at the adaptation set level of the main AdaptationSet in a point cloud.
[0212] Table 3 Attributes of VPCC Descriptors
[0213]
[0214]
[0215] VPCCViews descriptor
[0216] To identify static views in the main AdaptationSet of V-PCC content and their corresponding associated V-PCC tracks, the VPCCViews descriptor should be used. VPCCViews is an EssentialProperty or SupplementalProperty descriptor with the @schemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2020:vpvw".
[0217] There should be at most one single VPCCViews descriptor for the representation level or point cloud content in the AdaptationSet or the main AdaptationSet.
[0218] The @value attribute of the VPCCViews descriptor should not exist. The VPCCViews descriptor should include the elements and attributes specified in Table 4.
[0219] Table 4 Elements and Attributes of the VPCCViewGroups Descriptor
[0220]
[0221]
[0222]
[0223]
[0224] Dynamic View
[0225] When the view is dynamic, the timing metadata track used to signal each view's information in the representation timeline should be carried in a separate AdaptationSet with a single representation, and associated (linked) with the main V-PCC track using the @associationId attribute defined in ISO / IEC 23009-1 [MPEG-DASH], with the @associationType value including 4CC "vpcm" for the corresponding AdaptationSet or Representation.
[0226] Figure 4 This is a flowchart of an example method 400 for processing volumetric visual media data. As discussed throughout this document, in some embodiments, the volumetric visual media data may include point cloud data. In some embodiments, the volumetric visual media data may represent 3-D objects. 3-D objects may be projected onto a 2-D surface and arranged as video frames. In some embodiments, the volumetric visual data may represent multi-view video data, etc.
[0227] Method 400 can be implemented by an encoder device, as further described in this document. Method 400 includes: at 402, generating a bitstream containing volumetric visual information of a three-dimensional scene by the encoder using one or more atlas sub-bitstreams and one or more coded video sub-bitstreams. Method 400 includes: at 404, adding information to the bitstream that enables rendering a target view of the three-dimensional scene based on the desired viewing position and / or desired viewing direction.
[0228] In some embodiments, generation (402) may include encoding a set of atlases corresponding to a set of views by an encoder, wherein one or more views of the volumetric visual data are selectable from the set of views for rendering the target view. For example, a set of atlases may refer to a set of atlases as a set of atlas substreams in a bitstream.
[0229] In some embodiments, generating (402) includes: encapsulating a set of volumetric visual parameter tracks corresponding to an atlas group in a bitstream-based file storage using syntax elements. In some embodiments, the set of volumetric visual tracks and volumetric visual parameter tracks may be constructed (using the corresponding atlas substream) to carry all atlas data of the atlas group. In some examples, the syntax elements may be implemented using (static or dynamic) view group information boxes. For example, static view groups as described in Section 4.1 or Section 5.1 may be used in such embodiments.
[0230] In some embodiments, generation (402) includes: encapsulating a set of volumetric visual tracks corresponding to the atlas group based on syntax elements of a timing metadata track, the timing metadata track containing specific track references to volumetric visual parameter tracks in file storage of the bitstream, in order to encode the atlas group. Here, the set of volumetric visual tracks and volumetric visual parameter tracks may carry all atlas data of the atlas group. As further described in this document, the specific track references may be used by the decoder during parsing / rendering operations. This generation operation may use dynamic view groups described in this document (e.g., Section 4.1 or Section 5.1).
[0231] In some embodiments, method 400 further includes adding information to the bitstream that identifies the group of volumetric visual tracks based on a specific track group type and a specific track group identifier, wherein each volumetric visual track in the group of volumetric visual tracks contains a specific track reference to a volumetric visual parameter track.
[0232] In some embodiments, method 400 further includes: encoding one or more views of the volumetric visual data for the target view by an encoder based on one or more view group information, wherein each view group information describes one or more views. In some embodiments, each view group information further includes camera parameters for one or more views.
[0233] In some embodiments, method 400 further includes encoding by a decoder one or more atlases corresponding to one or more views of volumetric visual data selected for the target view.
[0234] In some embodiments, information from one or more atlas substreams is encoded by encapsulating one or more volumetric visual tracks corresponding to one or more atlases using syntax elements of volumetric visual parameter tracks in a bitstream-based file storage syntax structure (e.g., view information box syntax structure – static or dynamic); wherein the one or more volumetric visual tracks and volumetric visual parameter tracks carry all atlas data for one or more atlases.
[0235] In some embodiments, information from one or more atlas substreams is encoded by encapsulating one or more volumetric visual tracks corresponding to one or more atlases based on syntax elements of timing metadata tracks (e.g., view information box syntax structures—static or dynamic), the timing metadata tracks containing specific track references to volumetric visual parameter tracks in file storage of the bitstream; wherein the one or more volumetric visual tracks and volumetric visual parameter tracks carry all atlas data for one or more atlases.
[0236] In some embodiments, method 400 includes adding information to the bitstream for the purpose of identifying one or more views of the volumetric visual data for rendering a target view based on view information of one or more views, wherein each view information describes camera parameters of the corresponding view.
[0237] In some embodiments, method 400 includes: including in the bitstream information for identifying volumetric visual parameter tracks based on a specific sample ingress type, wherein the volumetric visual parameter tracks correspond to one or more volumetric visual tracks having a specific track reference, wherein the volumetric visual parameter tracks specify a constant set of parameters and common atlas data for all referenced volumetric visual tracks having a specific track reference.
[0238] In some embodiments, method 400 includes adding information to the bitstream for identifying timing metadata tracks based on a specific sample ingress type, the specific sample ingress type indicating that one or more views of the volumetric visual data selected for rendering a target view are dynamic.
[0239] The encoded video substream may include: one or more video coding elementary streams for geometric data, zero or one video coding elementary stream for occupancy graph data, and zero or more video coding elementary streams for attribute data, wherein the geometric data, occupancy graph data, and attribute data are descriptions of a three-dimensional scene.
[0240] Figure 5 This is a flowchart of an example method 500 for processing volumetric visual media data. Method 500 can be implemented by a decoder. The various terms used to describe the syntax elements in method 500 are similar to the terms used above to describe the syntax elements of encoder-side method 400.
[0241] Method 500 includes: at 502, decoding a bitstream by a decoder, the bitstream containing volumetric visual information of a 3D scene represented as one or more atlas sub-bitstreams and one or more coded video sub-bitstreams. Method 500 includes: at 504, reconstructing the 3D scene using the decoding results of one or more atlas sub-bitstreams and one or more coded video sub-bitstreams.
[0242] Method 500 includes, at 506, rendering a target view of a 3D scene based on the desired viewing position and / or desired viewing direction. In some embodiments, decoding and reconstruction may be performed by a first hardware platform, while rendering may be performed by another hardware platform working in conjunction with the decoding hardware platform. In other words, the first hardware platform may only perform steps 502 and 504 described above to implement the 3D scene reconstruction method. In some embodiments, the decoder may receive the viewer's desired viewing position or desired viewing direction in xyz or polar coordinates. Based on this information, the decoder may use a decoded sub-bitstream of a graph corresponding to the view set used to generate the target view to create a target view aligned with the viewer's position / direction from the decoded sub-bitstream including video information.
[0243] In some embodiments, the reconstruction includes: decoding a set of atlases corresponding to a view group by a decoder, wherein one or more views of the volumetric visual data have been selected from the view group for rendering the target view.
[0244] In some embodiments, decoding includes, prior to decoding the atlas group, decapsulating a set of volumetric visual tracks corresponding to the atlas group by a file parser based on the syntax elements of the volumetric visual parameter tracks in the file storage of the bitstream, wherein the set of volumetric visual tracks and volumetric visual parameter tracks carry all atlas data of the atlas group.
[0245] In some embodiments, decoding includes, prior to decoding the atlas group, the decapsulation of a set of volumetric visual tracks corresponding to the atlas group by a file parser based on the syntax elements of a timing metadata track, which contains specific track references to volumetric visual parameter tracks in the file storage of the bitstream; wherein the set of volumetric visual tracks and volumetric visual parameter tracks carry all atlas data for the atlas group. For example, the dynamic view group structure described in this document can be used during this operation.
[0246] In some embodiments, method 500 further includes: identifying the group of volumetric visual tracks according to a specific track group type and a specific track group identifier, wherein each volumetric visual track in the group of volumetric visual tracks contains a specific track reference to a volumetric visual parameter track.
[0247] In some embodiments, method 500 further includes: selecting one or more views of volumetric visual data for a target view by a decoder based on one or more view group information, wherein each view group information describes one or more views.
[0248] In some embodiments, each view group information also includes camera parameters for one or more views.
[0249] In some embodiments, the method further includes: decoding by a decoder one or more atlases corresponding to one or more views of the volumetric visual data selected for the target view.
[0250] In some embodiments, information from one or more atlas substreams is decoded by decapsulating one or more volumetric visual tracks corresponding to one or more atlases using syntax elements (e.g., ViewInfoBox elements) in a bitstream-based file storage syntax structure; wherein the one or more volumetric visual tracks and volumetric visual parameter tracks carry all atlas data for one or more atlases.
[0251] In some embodiments, information from one or more atlas substreams is decoded by decapsulating one or more volumetric vision tracks corresponding to one or more atlases based on syntax elements of timing metadata tracks, which contain specific track references to volumetric vision parameter tracks in file storage of the bitstream; wherein the one or more volumetric vision tracks and volumetric vision parameter tracks carry all atlas data of one or more atlases.
[0252] In some embodiments, the method further includes: selecting one or more views of the volumetric visual data by a decoder based on view information of one or more views for rendering the target view, wherein each view information describes camera parameters of the corresponding view.
[0253] In some embodiments, method 500 further includes: identifying a volumetric visual parameter track based on a specific sample inlet type, wherein the volumetric visual parameter track corresponds to one or more volumetric visual tracks having a specific track reference, wherein the volumetric visual parameter track specifies a constant parameter set and common atlas data for all referenced volumetric visual tracks having a specific track reference.
[0254] In some embodiments, method 500 further includes: identifying a timing metadata track based on a specific sample ingress type, the specific sample ingress type indicating that one or more views of the volumetric visual data selected for rendering the target view are dynamic.
[0255] In some embodiments, one or more encoded video sub-bitstreams include: one or more video encoded elementary streams for geometric data, zero or one video encoded elementary stream for occupancy graph data, and zero or more video encoded elementary streams for attribute data, wherein the geometric data, occupancy graph data, and attribute data are descriptions of a three-dimensional scene.
[0256] refer to Figures 4 to 5In some embodiments, a map set can refer to a set of map substreams. In some embodiments, a set of volumetric visual tracks used in the above method can represent a set of volumetric visual tracks.
[0257] In some embodiments, in method 400 or 500, the syntax elements of the volume visual parameter track can be the ViewGroupInfoBox syntax structure described in this document.
[0258] Figure 6 This is a block diagram of an example of an encoder for volumetric media data according to the present technology, device 600. Device 600 includes an acquisition module 601 configured to collect 3D scene and volumetric visual media information in the form of point cloud data, multi-view video data, or multi-projection. This module may include an input / output controller circuitry for reading video data from memory or a camera frame buffer. This module may include processor-executable instructions for reading volumetric data. Device 600 includes a bitstream generator module 602 configured to generate a bitstream that is an encoded representation of volumetric visual information according to various techniques described herein (e.g., method 400). This module may be implemented as processor-executable software code. Device 600 also includes a module 603 configured to perform subsequent processing on the bitstream (e.g., metadata insertion, encryption, etc.). The device also includes a storage / transmission module 904 configured to perform storage or network transport layer encoding on video encoded data or media data. Module 604 can implement, for example, the MPEG-DASH technology described in this document, which is used for streaming data over digital communication networks or storing bit streams in a DASH-compatible format.
[0259] Modules 601 to 604 described above can be implemented using dedicated hardware or hardware capable of performing processing in combination with appropriate software. Such hardware or dedicated hardware may include application-specific integrated circuits (ASICs), various other circuits, various processors, etc. When implemented by a processor, the functionality may be provided by a single dedicated processor, a single shared processor, or multiple independent processors, some of which may be shared. Furthermore, a processor should not be construed as referring to hardware capable of executing software, but may implicitly include, but is not limited to, digital signal processor (DSP) hardware, read-only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage devices.
[0260] like Figure 6 The device 600 shown can be a device used in video applications, such as a mobile phone, computer, server, set-top box, portable mobile terminal, digital camera, television broadcasting system equipment, etc.
[0261] Figure 7This is a block diagram illustrating an example of an apparatus 700 according to the present technology. Apparatus 700 includes an acquisition module 701 configured to acquire a bitstream from a network or by reading from a storage device. For example, module 701 can perform parsing and extraction of media files encoded using the MPEG-DASH technology described in this document, and perform decoding from network transport layer data including volumetric visual media data. A system and file parser module 702 can extract various system layer and file layer syntax elements (e.g., atlas sub-bitstreams, group information, etc.) from the received bitstream. A video decoder 703 is configured to decode encoded video sub-bitstreams, which include media data or volumetric media data of a 3D scene, such as point cloud data or multi-view video data. A renderer module 704 is configured to render a target view of the 3D scene based on a desired viewing position or desired viewing direction received from a user via user interface controls.
[0262] Modules 701 to 704 described above can be implemented using dedicated hardware or hardware capable of performing processing in combination with appropriate software. Such hardware or dedicated hardware may include application-specific integrated circuits (ASICs), various other circuits, various processors, etc. When implemented by a processor, the functionality may be provided by a single dedicated processor, a single shared processor, or multiple independent processors, some of which may be shared. Furthermore, a processor should not be construed as referring to hardware capable of executing software, but may implicitly include, but is not limited to, digital signal processor (DSP) hardware, read-only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage devices.
[0263] like Figure 7 The devices shown can be those used in video applications, such as mobile phones, computers, servers, set-top boxes, portable mobile terminals, digital cameras, and television broadcasting system equipment.
[0264] Figure 8 This is a block diagram of an example of device 800, which can be used as a hardware platform for implementing the various encoding and / or decoding functions described herein, including... Figures 6 to 7 The encoder / decoder implementation described herein. Apparatus 800 includes a processor 802 programmed to implement the methods described herein. Apparatus 800 may also include a dedicated hardware circuitry system for performing specific functions, such as bitstream encoding or decoding. Apparatus 800 may also include a memory storing the processor's executable code and / or volume data, as well as other data, including data conforming to the various syntax elements described herein.
[0265] In some embodiments, a 3D point cloud data encoder may be implemented to generate a bitstream representation of a 3D point cloud by encoding 3D spatial information using the syntax and semantics described in this document.
[0266] Volumetric visual media data encoding or decoding devices can be integrated into user devices such as computers, laptops, tablets, or gaming devices.
[0267] The disclosures and other embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuit systems or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or combinations of one or more of the foregoing. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of substances affecting machine-readable propagation signals, or combinations of one or more of the foregoing. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations of one or more of the foregoing. Propagation signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, that are generated to encode information for transmission to a suitable receiver device.
[0268] A computer program (also called a program, software, software application, script, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites and interconnected via a communication network.
[0269] The processes and logic flows described in this document can be executed by one or more programmable processors, which execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuit systems, and the devices can be implemented as dedicated logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0270] For example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magneto-optical, magneto-optical, or optical disc) for storing data, or operatively coupled to receive data from or transfer data to such mass storage devices, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices like EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be complemented or integrated into a dedicated logic circuit system.
[0271] While this patent document contains numerous details, these should not be construed as limiting the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described as functioning in certain combinations, and even initially claimed in this way, in some cases one or more features from a claimed combination may be removed from the combination, and a claimed combination may refer to a sub-combination or a variation of a sub-combination.
[0272] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or that all the operations shown be performed to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0273] Only a few implementations and examples are described, and other implementations, enhancements and variations may be made based on the content described and illustrated in this patent document.
Claims
1. A method for processing volumetric visual data, comprising: The bitstream is decoded by a decoder, the bitstream containing volumetric visual information for a 3D scene, the 3D scene being represented as one or more atlas sub-bitstreams and one or more coded video sub-bitstreams; The three-dimensional scene is reconstructed using the results of decoding the one or more atlas sub-bitstreams and the results of decoding the one or more coded video sub-bitstreams. as well as The target view of the 3D scene is rendered based on the desired viewing position and / or desired viewing direction. The decoding of the bitstream includes: Based on the first syntax element of the volumetric visual parameter track identified by the first sample entry type, one or more volumetric visual tracks corresponding to a map set are decapsulated, wherein the map set includes all map sets generated from the same view set, and one or more views of the volumetric visual data are selected from the view set for the rendering of the target view; and Decode the atlas group corresponding to the same view group; and Each of the one or more volumetric visual tracks is associated with a second syntax element, which has a track group type equal to a specific value indicating that the corresponding volumetric visual track belongs to a set of volumetric visual tracks corresponding to the atlas group.
2. The method of claim 1, wherein the atlas group corresponds to a view group from which volumetric visual data is selected for the rendering of the target view.
3. The method of claim 1, wherein the decapsulation of a set of volumetric visual tracks corresponding to the atlas group is performed before decoding the atlas group, and wherein the set of volumetric visual tracks and the volumetric visual parameter tracks carry all atlas data of the atlas group.
4. The method according to claim 3, further comprising: The set of volumetric vision tracks is identified according to a specific track group type and a specific track group identifier, wherein each volumetric vision track in the set of volumetric vision tracks contains a specific track reference to the volumetric vision parameter track.
5. The method according to claim 2, comprising: The decoder selects one or more views of the volumetric visual data for the target view based on one or more view group information, wherein each view group information describes one or more views.
6. The method of claim 5, wherein each view group information further includes camera parameters for the one or more views.
7. The method of claim 1, wherein the volumetric visual information from the one or more atlas sub-bitstreams is decoded by the decapsulation of the one or more volumetric visual tracks corresponding to the atlas group; wherein the one or more volumetric visual tracks and the volumetric visual parameter tracks carry all atlas data of the atlas group.
8. The method according to claim 2, further comprising: The decoder selects one or more views of the volumetric visual data based on view information for the one or more views, for rendering the target view, wherein each view information describes the camera parameters of the corresponding view.
9. The method according to claim 3 or 7, further comprising: The volumetric visual parameter track is identified based on the first sample inlet type. The volumetric visual parameter track specifies a constant set of parameters and common atlas data for all referenced volumetric visual tracks with a specific track reference.
10. The method of claim 3, further comprising: The timing metadata track is identified based on a second sample entry type, which indicates that one or more views of the volumetric visual data selected for rendering the target view are dynamic.
11. The method of claim 1, wherein the one or more encoded video sub-bitstreams comprise: One or more video-coded elementary streams for geometric data, and Used to occupy zero or one video coded elementary stream of graph data, and Zero or more video-coded elementary streams used for attribute data, The geometric data, the occupancy map data, and the attribute data described the three-dimensional scene.
12. A method for processing volumetric visual data, comprising: An encoder generates a bitstream containing volumetric visual information of a 3D scene by representing the scene using one or more atlas sub-bitstreams and one or more coded video sub-bitstreams. The bitstream includes information that enables the rendering of a target view of the 3D scene based on the desired viewing position and / or desired viewing direction. The generation includes: The encoder encodes a set of atlases corresponding to the view group, and selects one or more views of volumetric visual data from the view group for rendering the target view. The atlas group includes all atlases generated from the view group. Based on the first syntax element of the volumetric visual parameter track identified by the first sample entry type, one or more volumetric visual tracks corresponding to the atlas group are encapsulated, and Each of the one or more volumetric visual tracks is associated with a second syntax element, which has a track group type equal to a specific value indicating that the corresponding volumetric visual track belongs to a set of volumetric visual tracks corresponding to the atlas group.
13. The method of claim 12, wherein the encapsulation is performed for a set of volumetric vision tracks including the one or more volumetric vision tracks, wherein the set of volumetric vision tracks and the volumetric vision parameter tracks carry all atlas data of the atlas group.
14. The method of claim 13, further comprising: The bitstream includes information identifying the set of volumetric vision tracks based on a specific track group type and a specific track group identifier, wherein each volumetric vision track in the set of volumetric vision tracks contains a specific track reference to the volumetric vision parameter track.
15. The method of claim 12, wherein the one or more views of the volumetric visual data for the target view are based on one or more view group information, wherein each view group information describes one or more views.
16. The method of claim 15, wherein each view group information further includes camera parameters for the one or more views.
17. The method of claim 12, wherein the information from the one or more atlas substreams is through The encapsulation performed on the set of volumetric visual tracks corresponding to the atlas group is encoded; wherein the set of volumetric visual tracks and the volumetric visual parameter tracks carry all atlas data of the atlas group.
18. The method of claim 12, further comprising: This includes information for the purpose of identifying the one or more views of the volumetric visual data based on view information of the one or more views for the rendering of the target view, wherein the view information describes the camera parameters of the corresponding view.
19. The method according to claim 13 or 17, comprising: The bitstream includes information for identifying the volumetric visual parameter orbital based on the first sample inlet type. The volumetric visual parameter track specifies a constant set of parameters and common atlas data for all referenced volumetric visual tracks with a specific track reference.
20. The method of claim 12, further comprising: The bitstream includes information for identifying timing metadata tracks based on a second sample ingress type, which indicates that one or more views of the volumetric visual data selected for target view rendering are dynamic.
21. The method of claim 12, wherein the one or more encoded video sub-bitstreams comprise: One or more video-coded elementary streams for geometric data, and Used to occupy zero or one video coded elementary stream of graph data, and Zero or more video-coded elementary streams used for attribute data, The geometric data, the occupancy map data, and the attribute data described the three-dimensional scene.
22. A video processing apparatus comprising a processor configured to implement the method according to any one of claims 1 to 21.
23. A computer-readable medium having code stored thereon, the code encoding instructions to cause a processor to perform the method according to any one or more of claims 1 to 21.
Citation Information
Patent Citations
3D point cloud compression systems for delivery and access of a subset of a compressed 3D point cloud
US20190318488A1
Methods and apparatus for volumetric video transport
WO2020013976A1