Efficient patch rotation in point cloud coding

KR103023167B1Active Publication Date: 2026-09-21HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020247043568
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-01-14
Filing Date
2020-01-13
Publication Date
2026-09-21
Estimated Expiration
2040-01-13

Smart Images

  • Figure 112024146579215-PAT00013_ABST
    Figure 112024146579215-PAT00013_ABST
Patent Text Reader

Abstract

A point cloud coding (PCC) method implemented by a decoder is provided. The method comprises the steps of: a receiver of the decoder receiving a bitstream containing a patch rotation enable flag and atlas information for a two-dimensional (2D) patch; a processor of the decoder determining that the 2D patch can be rotated based on the patch rotation enable flag; the processor rotating the 2D patch; and the processor reconstructing a three-dimensional (3D) image using the atlas information and the rotated two-dimensional patch.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This patent application claims priority to U.S. Provisional Application No. 62 / 792,259, filed by Vladislav Zakharchenko on January 14, 2019, with the title of the invention “Point Cloud Bitstream Structure and Auxiliary Information Differential Coding,” the contents of which are incorporated herein by reference in their entirety.

[0002] The present disclosure generally relates to point cloud coding, and specifically to high-level syntax for point cloud coding. Background Technology

[0003] Point clouds are used in various applications, including the entertainment industry, intelligent automotive navigation, geospatial inspection, three-dimensional (3D) modeling of real-world objects, and visualization. Given the non-uniform sampling geometry of point clouds, a concise representation is useful for the storage and transmission of this data. Compared to other 3D presentations, non-uniform point clouds are more general and applicable to a wider range of sensors and data acquisition strategies. For example, when performing 3D presentations in a virtual reality world or remote rendering in a telepresence environment, the rendering of virtual figures and real-time commands are processed using dense point cloud datasets.

[0004] A first aspect relates to a point cloud coding (PCC) method implemented by a decoder. The method comprises the steps of: a receiver of the decoder receiving a bitstream containing a patch rotation enabled flag and atlas information for a two-dimensional (2D) patch; a processor of the decoder determining that the 2D patch can be rotated based on the patch rotation enabled flag; the processor rotating the 2D patch; and the processor reconstructing a three-dimensional (3D) image using the atlas information and the rotated 2D patch.

[0005] This flexible patch orientation scheme allows patches to be packed more efficiently into a bounding box. Due to this more efficient packing, the area occupied by the patch within the bounding box can be reduced compared to techniques where patch rotation is unavailable or not allowed, which leads to better coding efficiency (e.g., fewer bits are required to code a bitstream).

[0006] In a first embodiment of the method according to this first aspect, the bitstream includes a default patch rotation and a preferred patch rotation, and the 2D patch is rotated according to the default patch rotation or the preferred patch rotation.

[0007] In a second embodiment of the method according to such a first aspect or any prior embodiment of the first aspect, the bitstream includes one of a plurality of available patch rotations, and the 2D patch is rotated according to one of the plurality of available patch rotations.

[0008] In a third implementation of the method according to such a first aspect or any prior implementation of the first aspect, a three-bit flag in the bitstream identifies one of the plurality of available patch rotations.

[0009] In a fourth embodiment of the method according to such a first aspect or any prior embodiment of the first aspect, the bitstream includes a limited rotation enable flag.

[0010] In a fifth embodiment of the method according to such a first aspect or any prior embodiment of the first aspect, the 2D patch is rotated based on an exhaustive orientation mode when the limited rotation flag has a first value, and rotated based on a simple orientation mode when the limited rotation flag has a second value.

[0011] In a sixth embodiment of the method according to such a first aspect or any prior embodiment of the first aspect, the bitstream includes a patch rotation present flag, and the 2D patch is rotated according to a default patch rotation when the patch rotation present flag has a first value.

[0012] In a seventh embodiment of the method according to such a first aspect or any prior embodiment of the first aspect, the bitstream includes a patch rotation presence flag, and the 2D patch is rotated according to a priority patch rotation or one of a plurality of available patch rotations when the patch rotation presence flag has a second value.

[0013] In the eighth embodiment of the method according to such a first aspect or any prior embodiment of the first aspect, when the patch rotation enable flag has a second value, the restricted rotation enable flag has the second value, and the patch rotation existence flag has the second value, the method further includes the step of first rotating the 2D patch according to the patch rotation.

[0014] In a ninth embodiment of the method according to such a first aspect or any prior embodiment of the first aspect, when the patch rotation enable flag has a second value, the restricted rotation enable flag has the second value, and the patch rotation existence flag has the first value, the method further includes the step of rotating the 2D patch according to the default patch rotation.

[0015] In a 10th implementation of a method according to such a first aspect or any prior implementation of the first aspect, when the patch rotation enable flag has a second value, the limited rotation enable flag has a first value, and the patch rotation existence flag has the first value, the method further includes the step of rotating the 2D patch according to the default patch rotation.

[0016] In a first embodiment of the method according to such a first aspect or any prior embodiment of the first aspect, when the patch rotation enable flag has a second value, the limited rotation enable flag has a first value, and the patch rotation existence flag has the second value, the method further includes the step of rotating the 2D patch according to one of a plurality of available patch rotations.

[0017] A second aspect relates to a point cloud coding (PCC) method implemented by an encoder. The method comprises the steps of: a receiver of the encoder acquiring a three-dimensional (3D) image; a processor of the encoder determining a plurality of two-dimensional (2D) projections on the 3D image using a plurality of available patch rotations; the processor selecting one of the plurality of 2D projections; the processor setting a plurality of flags according to the selected one of the plurality of 2D projections; the processor generating a bitstream including the plurality of flags and atlas information to reconstruct the 3D image; and storing the bitstream in the memory of the encoder for transmission toward a decoder.

[0018] This flexible patch orientation scheme allows patches to be packed more efficiently into bounding boxes. Due to this more efficient packing, the area occupied by the patch within the bounding box can be reduced compared to techniques where patch rotation is unavailable or not allowed, which leads to better coding efficiency (e.g., fewer bits are required to code a bitstream).

[0019] In a first embodiment of the method according to this second aspect, the plurality of flags includes a patch rotation enable flag, a restricted rotation enable flag, and a patch rotation presence flag.

[0020] In a second implementation of a method according to such a second aspect or any prior implementation of the second aspect, at least one of the plurality of flags is set to signal the decoder to use a default patch rotation.

[0021] In a third embodiment of the method according to such a second aspect or any prior embodiment of the second aspect, at least one of the plurality of flags is set to signal the decoder to use a priority patch rotation.

[0022] In a fourth embodiment of the method according to such a second aspect or any prior embodiment of the second aspect, at least one of the plurality of flags is set to signal the decoder to use one of the plurality of available patch rotations.

[0023] A third aspect relates to a decoding device, wherein the decoding device comprises: a receiver configured to receive a bitstream including a patch rotation enable flag and atlas information for a two-dimensional (2D) patch; a memory coupled to the receiver and also storing instructions; and a processor coupled to the memory, wherein the processor executes the instructions so that the decoding device determines that the 2D patch can be rotated based on the patch rotation enable flag; rotates the 2D patch; and reconstructs a three-dimensional (3D) image using the atlas information and the rotated 2D patch.

[0024] This flexible patch orientation scheme allows patches to be packed more efficiently into bounding boxes. Due to this more efficient packing, the area occupied by the patch within the bounding box can be reduced compared to techniques where patch rotation is unavailable or not allowed, which leads to better coding efficiency (e.g., fewer bits are required to code a bitstream).

[0025] In a first embodiment of a decoding device according to this third aspect, the decoding device further includes a display configured to display the 3D image.

[0026] A fourth aspect relates to an encoding device, wherein the encoding device comprises: a receiver configured to receive a three-dimensional (3D) image; a memory coupled to the receiver and also storing instructions; and a processor coupled to the memory, wherein the processor implements the instructions so that the encoding device determines a plurality of two-dimensional (2D) projections on the 3D image using a plurality of available patch rotations; selects one of the plurality of 2D projections; sets a plurality of flags according to the selected one of the plurality of 2D projections; generates a bitstream including the plurality of flags and atlas information to reconstruct the 3D image; and stores the bitstream in the memory for transmission toward a decoder.

[0027] This flexible patch orientation scheme allows patches to be packed more efficiently into bounding boxes. Due to this more efficient packing, the area occupied by the patch within the bounding box can be reduced compared to techniques where patch rotation is unavailable or not allowed, which leads to better coding efficiency (e.g., fewer bits are required to code a bitstream).

[0028] In a first embodiment of an encoding device according to such a fourth aspect, the encoding device further comprises a transmitter coupled to the processor, and the transmitter is configured to transmit the bitstream toward the decoder.

[0029] A fifth aspect relates to a coding device, the coding device comprising: a receiver configured to receive a bitstream; a transmitter coupled to the receiver—the transmitter is configured to transmit the bitstream to a decoder or to transmit a decoded volume image to a reconstruction device configured to reconstruct the decoded volume image—; a memory coupled to at least one of the receiver or the transmitter and also storing instructions; and a processor coupled to the memory, the processor being configured to execute instructions stored in the memory to perform any method disclosed herein.

[0030] This flexible patch orientation scheme allows patches to be packed more efficiently into bounding boxes. Due to this more efficient packing, the area occupied by the patch within the bounding box can be reduced compared to techniques where patch rotation is unavailable or not allowed, which leads to better coding efficiency (e.g., fewer bits are required to code a bitstream).

[0031] In a first embodiment of a coding device according to the fifth aspect, the coding device further includes a display configured to display a projected image based on the decoded volume image.

[0032] A sixth aspect relates to a system, wherein the system comprises an encoder; and a decoder communicating with the encoder, and the encoder or the decoder comprises an encoding device or a decoding device or a coding device as described herein.

[0033] This flexible patch orientation scheme allows patches to be packed more efficiently into bounding boxes. Due to this more efficient packing, the area occupied by the patch within the bounding box can be reduced compared to techniques where patch rotation is unavailable or not allowed, which leads to better coding efficiency (e.g., fewer bits are required to code a bitstream).

[0034] A seventh aspect relates to means for coding, wherein the means for coding comprises: receiving means configured to receive a volume image to be encoded, decoded, reconstructed, and receive a bitstream to be projected; a transmission means coupled to the receiving means, wherein the transmission means is configured to transmit the bitstream to a decoder or transmit the decoded image to a display means; a storage means coupled to at least one of the receiving means or the transmission means and also storing a command; and a processing means coupled to the storage means, wherein the processing means is configured to execute a command stored in the storage means to perform any method disclosed herein.

[0035] For clarity, any one of the aforementioned embodiments may be combined with any one or more other aforementioned embodiments to create a new embodiment within the scope of the present disclosure.

[0036] These and other features will be more clearly understood from the following detailed description taken in relation to the attached drawings and claims. Brief explanation of the drawing

[0037] For a more complete understanding of the present disclosure, the following brief description taken in connection with the accompanying drawings and detailed description, in which similar reference numbers indicate similar parts, is now referred to. Figure 1 is a block diagram illustrating an exemplary coding system that can utilize context modeling techniques. FIG. 2 is a block diagram illustrating an exemplary encoder capable of implementing context modeling technology. FIG. 3 is a block diagram illustrating an exemplary decoder capable of implementing context modeling technology. Figure 4 is a representation of a frame bitstream group. Figure 5 is a representation of a three-dimensional (3D) point cloud. Figure 6 is a representation of the 3D point cloud of Figure 5 projected onto a bounding box. Figure 7 is a representation of an occupancy map corresponding to a two-dimensional (2D) projection from the bounding box of Figure 6. Figure 8 is a representation of a geometry map corresponding to a 2D projection from the bounding box of Figure 6. Figure 9 is a representation of an attribute map corresponding to a 2D projection from the bounding box of Figure 6. FIG. 10 is an example of a representation of a patch orientation index. Figure 11 is an example of a patch orientation decoding process. FIG. 12 is an example of a point cloud coding (PCC) method implemented by a decoder. FIG. 13 is an example of a PCC method implemented by an encoder. Figure 14 is a schematic diagram of a coding device. FIG. 15 is a schematic diagram of an embodiment of a means for coding. Specific details for implementing the invention

[0038] While exemplary implementations of one or more embodiments are provided below, it should be understood from the outset that the disclosed system and / or method may be implemented using any number of technologies currently known or existing. The present disclosure shall by no means be limited to the exemplary implementations, drawings, and technologies illustrated below, including the exemplary designs and implementations illustrated and described herein, and may be modified within the scope of the appended claims, together with the full scope of equivalents.

[0039] Video coding standards include ITU-T (International Telecommunications Union Telecommunication Standardization Sector) H.261, ISO (International Organization for Standardization) / IEC (International Electrotechnical Commission) MPEG (Moving Picture Experts Group)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, AVC (Advanced Video Coding) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and HEVC (High Efficiency Video Coding) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extensions such as SVC (Scalable Video Coding), MVC (Multiview Video Coding), MVC D (Multiview Video Coding plus Depth), and 3D AVC (3D-AVC). HEVC includes extensions such as SHVC (Scalable HEVC), MV-HEVC (Multiview HEVC), and 3D-HEVC (3D HEVC).

[0040] A point cloud is a set of data points in 3D space. Each data point contains location (e.g., X, Y, Z), color (e.g., R, G, B or Y, U, V), and other possible attributes such as transparency, reflectivity, and time of acquisition. Generally, each point in a cloud has the same number of attributes attached to it. Point clouds can be used in various applications, such as real-time 3D immersive telepresence, virtual reality (VR) viewing of content with interactive parallax, 3D free-viewpoint sports broadcasting, geographic information systems, cultural heritage, large-scale 3D dynamic map-based autonomous navigation, and automotive applications.

[0041] In 2016, the ISO / IEC MPEG (Moving Picture Experts Group) began developing a new codec standard for point cloud coding for lossless and lossy compressed point cloud data, which offers significant coding efficiency and robustness for network environments. Using this codec standard, point clouds can be manipulated into the form of computer data, stored on various storage media, transmitted and received over existing and future networks, and distributed on existing and future broadcast channels.

[0042] Recently, PCC (Point Cloud coding) work has been classified into three categories: PCC Category 1, PCC Category 2, and PCC Category 3, and a working draft for PCC Category 2 (PCC Cat2) and a working draft for PCC Category 1 and 3 (PCC Cat13) have been developed. The latest working draft (WD) for PCC Cat2 is included in MPEG output document N17534, and the latest WD for PCC Cat13 is included in MPEG output document N17533.

[0043] The key philosophy behind the design of the PCC Cat2 codec in PCC Cat2 WD is to compress the geometry and texture information of dynamic point clouds by leveraging existing video codecs through the compression of point cloud data into a set of different video sequences. Specifically, two video sequences are generated and compressed using the video codec: one representing the geometry information of the point cloud data and the other representing the texture information. Additional metadata for interpreting the two video sequences, namely the occupancy map and auxiliary patch information, is also generated and compressed separately.

[0044] Unfortunately, the existing design of PCC has drawbacks. For example, data units belonging to a single time instance—that is, a single access unit (AU)—are not contiguous in the decoding order. In PCC Cat 2 WD, data units for textures, geometry, auxiliary information, and occupancy maps for each AU are interleaved at the frame group level. That is, geometry data for all frames in the group is together. The same applies to texture data, etc. In PCC Cat 13 WD, data units for geometry and general attributes for each AU are interleaved at the entire PCC bitstream level (e.g., the same as in PCC Cat 2 WD if there is only one frame group with the same length as the entire PCC bitstream). The interleaving of data units belonging to a single AU inherently causes a huge end-to-end delay in the application system's presentation duration, which is at least equal to the length of the frame group.

[0045] Another disadvantage is related to the bitstream format. Since the bitstream format allows the emulation of start code patterns such as 0x0003, it does not work for transmission over MPEG-2 transport streams (TS) where start code emulation prevention is required. For PCC Cat2, start code emulation prevention is currently only available in `group_of_frames_geometry_video_payload()` and `group_of_frames_texture_video_payload()` when HEVC or AVC is used for coding geometry and texture components. For PCC Cat13, start code emulation prevention is not present anywhere in the bitstream.

[0046] In PCC Cat2 WD, some codec information for geometry and texture bitstreams (e.g., which codec, the codec's profile, levels, etc.) is buried deep within multiple instances of the group_of_frames_geometry_video_payload() and group_of_frames_texture_video_payload() structures. Additionally, some information such as profiles and levels that direct auxiliary information and the decoding and point cloud reconstruction functions of the occupancy map component are missing.

[0047] A high-level syntax design is provided that resolves one or more of the aforementioned problems related to point cloud coding. As described more fully below, the present disclosure utilizes a type indicator of a data unit header (also known as a PCC NAL (Network Access Layer) header) to specify the type of content in the payload of a PCC NAL unit. Additionally, the present disclosure utilizes a group of frame header NAL units to carry a group of frame header parameters. A group of frame header NAL units may also be used to signal the profile and level of each geometry or texture bitstream.

[0048] FIG. 1 is a block diagram illustrating an exemplary coding system (10) that may utilize PCC video coding technology. As illustrated in FIG. 1, the coding system (10) includes a source device (12) that provides encoded video data to be later decoded by a destination device (14). In particular, the source device (12) may provide video data to the destination device (14) via a computer-readable medium (16). The source device (12) and the destination device (14) may include any of a wide range of devices, including a desktop computer, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a so-called "smart" pad, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some cases, the source device (12) and the destination device (14) may be equipped for wireless communication.

[0049] A destination device (14) can receive encoded video data to be decoded via a computer-readable medium (16). The computer-readable medium (16) may include any type of medium or device capable of moving the encoded video data from a source device (12) to a destination device (14). In one example, the computer-readable medium (16) may include a communication medium that enables the source device (12) to transmit the encoded video data directly to the destination device (14) in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device (14). The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, switch, base station, or any other equipment that may be useful for facilitating communication from a source device (12) to a destination device (14).

[0050] In some examples, the encoded data may be output from the output interface (24) to a storage device. Similarly, the encoded data may be accessed from the storage device by an input interface. The storage device may include various distributed or local access data storage media, such as a hard drive, a Blu-ray disc, a digital video disc (DVD), a CD-ROM (Compact Disc Read-Only Memories), flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In additional examples, the storage device may correspond to a file server or other intermediate storage device capable of storing the encoded video generated by the source device (12). The destination device (14) may access the stored video data from the storage device via streaming or download. The file server may be any type of server capable of storing the encoded video data and transmitting the encoded video data to the destination device (14). Examples of file servers include web servers (e.g., for websites), FTP (file transfer protocol) servers, NAS (Network Attached Storage) devices, or local disk drives. The destination device (14) may access the encoded video data via any standard data connection, including an internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL (digital subscriber line), cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.

[0051] The technology of the present disclosure is not necessarily limited to wireless applications or settings. This technology may be applied to video coding that supports various multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission such as HTT dynamic adaptive streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the coding system (10) may be configured to support unidirectional or bidirectional video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0052] In the example of FIG. 1, the source device (12) includes a video source (18) configured to provide a volumetric image, a projection device (20), a video encoder (22), and an output interface (24). The destination device (14) includes an input interface (26), a video decoder (28), a reconstruction device (30), and a display device (32). According to the present disclosure, the encoder (22) of the source device (12) and / or the decoder (28) of the destination device (14) may be configured to apply techniques for video coding. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device (12) may receive video data from an external video source, such as an external camera. Likewise, the destination device (14) may interface with an external display device rather than including an integrated display device.

[0053] The coding system (10) illustrated in FIG. 1 is merely one example. Techniques for video coding may be performed by any digital video encoding and / or decoding device. While the techniques of the present disclosure are generally performed by a coding device, the techniques may also be performed by an encoder / decoder generally referred to as a "CODEC." Furthermore, the techniques of the present disclosure may also be performed by a video preprocessor. The encoder and / or decoder may be a graphics processing unit (GPU) or a similar device.

[0054] The source device (12) and the destination device (14) are merely examples of such coding devices in which the source device (12) generates video data coded for transmission to the destination device (14). In some examples, the source device (12) and the destination device (14) may operate in a substantially symmetric manner such that each of the source and destination devices (12, 14) includes video encoding and decoding components. Thus, the coding system (10) can support unidirectional or bidirectional video transmission between video devices (12, 14) for, for example, video streaming, video playback, video broadcasting, or video telephone.

[0055] The video source (18) of the source device (12) may include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed interface for receiving volume images or video from a video content provider. As an additional alternative, the video source (18) may generate volume images or computer graphics-based data as source video or as a combination of live video, archived video, and computer-generated video.

[0056] In some cases, when the video source (18) is a video camera, the source device (12) and the destination device (14) may form a so-called camera phone or video phone. However, as mentioned above, the techniques described in this disclosure may generally be applicable to video coding and may also be applicable to wireless and / or wired applications.

[0057] The projection device (20) is configured to project a volume image onto a planar surface (e.g., a bounding box) as described more fully below. That is, the projection device (20) is configured to convert a three-dimensional (3D) image into a two-dimensional (2D) image.

[0058] In any case, a volume image, a captured video, a pre-captured video, or a computer-generated video may be encoded by an encoder (22). The encoded video information may then be output to a computer-readable medium (16) by an output interface (24).

[0059] The computer-readable medium (16) may include transient media, such as wireless broadcast or wired network transmission, or storage media (i.e., non-transient storage media), such as hard disks, flash drives, compact discs, digital video discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) may receive encoded video data from a source device (12) and provide the encoded video data to a destination device (14), for example, via network transmission. Similarly, a computing device of a media production facility, such as a disk stamping facility, may receive encoded video data from the source device (12) and produce a disk containing the encoded video data. Thus, the computer-readable medium (16) may be understood to include one or more computer-readable media of various forms in various examples.

[0060] The input interface (26) of the destination device (14) receives information from a computer-readable medium (16). The information in the computer-readable medium (16) may include syntax information defined by an encoder (22), which is also used by a decoder (28), and the syntax information includes syntax elements describing the processing and / or characteristics of blocks and / or other coded units, e.g., a group of pictures (GOP).

[0061] The reconstruction device (30) is configured to convert flat images or images back into volume images, as described more fully below. That is, the reconstruction device (30) is configured to convert 2D images or images back into 3D images.

[0062] The display device (32) displays volume images or decoded video data to the user and may include any of various display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED), or other types of display devices.

[0063] The encoder (22) and decoder (28) may operate according to a video coding standard, such as the HEVC (High Efficiency Video Coding) standard currently under development, and may follow the HEVC Test Model (HM). Alternatively, the encoder (22) and decoder (28) may operate according to other proprietary or industry standards, such as the ITU-T (International Telecommunications Union Telecommunication Standardization Sector) H.264 standard referred to as MPEG (Moving Picture Expert Group)-4, Part 10, AVC (Advanced Video Coding), H.265 / HEVC, or extensions of these standards. However, the technology of this disclosure is not limited to any specific coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. Although not illustrated in FIG. 1, in some aspects, the encoder (22) and decoder (28) may be integrated with the audio encoder and decoder, respectively, and may include a suitable multiplexer-demultiplexer (MUX-DEMUX) unit or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams. If applicable, the MUX-DEMUX device may comply with the ITU H.223 multiplexer protocol or other protocols such as UDP (user datagram protocol).

[0064] Each of the encoder (22) and decoder (28) may be implemented as any one of various suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or a combination thereof. When the technology is partially implemented in software, the device may store instructions for the software on a suitable non-transient computer-readable medium and perform the technology of the present disclosure by executing instructions in hardware using one or more processors. Each of the encoder (22) and decoder (28) may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined encoder / decoder (CODEC) in each device. A device including an encoder (22) and / or a decoder (28) may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular phone.

[0065] FIG. 2 is a block diagram illustrating an example of an encoder (22) capable of implementing video coding technology. The encoder (22) may perform intra-coding and inter-coding of video blocks within a video slice. Intra-coding relies on spatial prediction to reduce or eliminate spatial redundancy of video within a given video frame or image. Inter-coding relies on temporal prediction to reduce or eliminate temporal redundancy of video within adjacent frames or images of a video sequence. An intra-mode (I mode) may refer to one of several spatial-based coding modes. An inter-mode, such as uni-prediction (also known as uni-prediction) (P mode) or bi-prediction (also known as bi-prediction) (B mode), may refer to any of several temporal-based coding modes.

[0066] As illustrated in FIG. 2, the encoder (22) receives the current video block within the video frame to be encoded. In the example of FIG. 2, the encoder (22) includes a mode selection unit (40), a reference frame memory (64), a summer (50), a transform processing unit (52), a quantization unit (54), and an entropy coding unit (56). The mode selection unit (40) in turn includes a motion compensation unit (44), a motion estimation unit (42), an intra prediction (also known as intra prediction) unit (46), and a partition unit (48). For video block reconstruction, the encoder (22) also includes an inverse quantization unit (58), an inverse transform unit (60), and a summer (62). A deblocking filter (not shown in FIG. 2) for filtering block boundaries to remove blockiness artifacts from the reconstructed video may also be included. If desired, the deblocking filter generally filters the output of the summer (62). In addition to the deblocking filter, additional filters (in-loop or post-loop) may be used. These filters are not shown for brevity, but if desired, they can filter the output of the summer (50) (as an in-loop filter).

[0067] During the encoding process, the encoder (22) receives a video frame or slice to be coded. The frame or slice may be divided into multiple video blocks. The motion estimation unit (42) and the motion compensation unit (44) perform inter-predictive coding of the received video blocks for one or more blocks in one or more reference frames to provide temporal prediction. Alternatively, the intra-predictive unit (46) may perform intra-predictive coding of the received video blocks for one or more neighboring blocks in the same frame or slice as the block to be coded to provide spatial prediction. The encoder (22) may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0068] Furthermore, the partition unit (48) may partition blocks of video data into sub-blocks based on an evaluation of previous partitioning methods in previous coding passes. For example, the partition unit (48) may initially partition a frame or slice into a largest coding unit (LCU) and then partition each LCU into sub-coding units (sub-CUs) based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit (40) may additionally generate a quad-tree data structure indicating the partitioning of LCUs into sub-CUs. A leaf node CU of the quad-tree may include one or more prediction units (PUs) and one or more transform units (TUs).

[0069] This disclosure uses the term "block" to refer to any of a CU, PU, ​​or TU in the context of HEVC, or uses similar data structures in the context of other standards (e.g., macroblocks and subblocks in H.264 / AVC). A CU includes a coding node, a PU, and a TU associated with the coding node. The size of a CU corresponds to the size of a coding node and is square. The size of a CU may range from 8×8 pixels up to 64×64 pixels or a treeblock size greater than or equal to a larger size. Each CU may include one or more PUs and one or more TUs. Syntax data associated with a CU may describe, for example, partitioning the CU into one or more PUs. The partitioning mode may differ whether the CU is encoded in skip or direct mode, intra-prediction mode, or inter-prediction (also known as inter-prediction) mode. PUs may be partitioned into non-square shapes. Syntax data associated with CUs can also describe partitioning a CU into one or more TUs according to a quadtree, for example. TUs may be square or non-square in shape (e.g., rectangular).

[0070] The mode selection unit (40) may, for example, select one of an intra-coding mode or an inter-coding mode based on an error result, provide the resulting intra-coded or inter-coded block to the summerator (50) to generate residual block data, and provide it to the summerator (62) to reconstruct an encoded block for use as a reference frame. The mode selection unit (40) also provides syntax elements, such as motion vectors, intra-mode indicators, partition information, and other such syntax information, to the entropy coding unit (56).

[0071] The motion estimation unit (42) and the motion compensation unit (44) may be highly integrated, but are exemplified separately for conceptual purposes. Motion estimation performed by the motion estimation unit (42) is a process of generating motion vectors that estimate motion for a video block. For example, the motion vector may represent the displacement of the PU of the video block within the current video frame or image for the prediction block within the reference frame (or other coded unit) for the current block being coded within the current frame (or other coded unit). The prediction block is a block found to closely match the block to be coded in relation to pixel differences that may be determined by the Sum of Absolute Difference (SAD), Sum of the Square Difference (SSD), or other difference metrics. In some examples, the encoder (22) may also calculate values ​​for sub-integer pixel positions of the reference image stored in the reference frame memory (64). For example, the encoder (22) can interpolate the values ​​of a 1 / 4 pixel position, a 1 / 8 pixel position, or other fractional pixel positions of the reference image. Thus, the motion estimation unit (42) can perform motion search for the whole pixel position and the fractional pixel position and output a motion vector with fractional pixel precision.

[0072] The motion estimation unit (42) calculates a motion vector for the PU of the video block in the intercoded slice by comparing the position of the PU with the position of the prediction block of the reference image. The reference image may be selected from the first reference image list (List 0) or the second reference image list (List 1), each of which identifies one or more reference images stored in the reference frame memory (64). The motion estimation unit (42) transmits the calculated motion vector to the entropy encoding unit (56) and the motion compensation unit (44).

[0073] Motion compensation performed by the motion compensation unit (44) may involve fetching or generating a prediction block based on the motion vector determined by the motion estimation unit (42). Again, the motion estimation unit (42) and the motion compensation unit (44) may be functionally integrated in some examples. Upon receiving a motion vector for the PU of the current video block, the motion compensation unit (44) may find the location of the prediction block pointed to by the motion vector in one of the reference image lists. The summer (50) forms a residual video block by subtracting the pixel value of the prediction block from the pixel value of the current video block being coded, thereby forming a pixel difference value as discussed below. Generally, the motion estimation unit (42) performs motion estimation for the lumina component, and the motion compensation unit (44) uses a motion vector calculated based on the lumina component for both the chroma component and the lumina component. The mode selection unit (40) may also generate video blocks and syntax elements associated with the video slice for use by the decoder (28) when decoding the video blocks of the video slice.

[0074] The intra prediction unit (46) can intra-predict the current block as an alternative to the inter-prediction performed by the motion estimation unit (42) and the motion compensation unit (44) as described above. In particular, the intra prediction unit (46) may determine the intra-prediction mode to use for encoding the current block. In some examples, the intra prediction unit (46) may encode the current block using various intra-prediction modes, for example, during separate encoding passes, and the intra prediction unit (46) (or the mode selection unit (40) in some examples) may select an appropriate intra-prediction mode to use from the tested modes.

[0075] For example, the intra prediction unit (46) can calculate a rate distortion value using rate distortion analysis for various tested intra prediction modes and select the intra prediction mode having the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to generate the encoded block, and the bitrate (i.e., number of bits) used to generate the encoded block. The intra prediction unit (46) can calculate a ratio from the distortion and rate for various encoded blocks to determine the intra prediction mode that represents the best rate distortion value for the block.

[0076] Additionally, the intra prediction unit (46) may be configured to code depth blocks of the depth map using a depth modeling mode (DMM). The mode selection unit (40) may determine whether the available DMM mode produces better coding results than the intra prediction mode and other DMM modes, for example, by using rate-distortion optimization (RDO). Data for texture images corresponding to the depth map may be stored in the reference frame memory (64). The motion estimation unit (42) and the motion compensation unit (44) may also be configured to mutually predict depth blocks of the depth map.

[0077] After selecting an intra prediction mode for a block (e.g., one of a conventional intra prediction mode or a DMM mode), the intra prediction unit (46) may provide information indicating the selected intra prediction mode for the block to the entropy coding unit (56). The entropy coding unit (56) may also encode the information indicating the selected intra prediction mode. The encoder (22) may include, in the transmitted bitstream configuration data which may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also called codeword mapping tables), definitions of encoding contexts for various blocks, and indications of the most likely intra prediction mode to be used for each context, the intra prediction mode index table, and the modified intra prediction mode index table.

[0078] The encoder (22) forms a residual video block by subtracting the prediction data from the mode selection unit (40) from the original video block being coded. The summer (50) represents a component that performs this subtraction operation.

[0079] The transform processing unit (52) applies a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. The transform processing unit (52) may perform other transforms conceptually similar to the DCT. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used.

[0080] The transformation processing unit (52) applies the transformation to the residual block to generate a block of residual transformation coefficients. The transformation can transform residual information from a pixel value domain to a transformation domain, such as a frequency domain. The transformation processing unit (52) may transmit the resulting transformation coefficients to the quantization unit (54). The quantization unit (54) quantizes the transformation coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameters. In some examples, the quantization unit (54) may then perform a scan of a matrix containing the quantized transformation coefficients. Alternatively, the entropy encoding unit (56) may perform the scan.

[0081] Following quantization, the entropy coding unit (56) entropies the quantized transformation coefficients. For example, the entropy coding unit (56) may perform context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. In the case of context-based entropy coding, the context may be based on neighboring blocks. Following entropy coding by the entropy coding unit (56), the encoded bitstream may be transmitted to another device (e.g., decoder (28)) or stored for later transmission or retrieval.

[0082] The inverse quantization unit (58) and the inverse transform unit (60) each apply inverse quantization and inverse transform, respectively, to reconstruct a residual block in the pixel domain for later use, for example, as a reference block. The motion compensation unit (44) may also calculate the reference block by adding the residual block to the prediction block of one of the frames in the reference frame memory (64). The motion compensation unit (44) may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values ​​for use in motion estimation. The summer (62) adds the reconstructed residual block to the motion-compensated prediction block generated by the motion compensation unit (44) to generate a reconstructed video block to be stored in the reference frame memory (64). The reconstructed video block may be used as a reference block for intercoding blocks in subsequent video frames by the motion estimation unit (42) and the motion compensation unit (44).

[0083] FIG. 3 is a block diagram illustrating an example of a decoder (28) capable of implementing video coding technology. In the example of FIG. 3, the decoder (28) includes an entropy decoding unit (70), a motion compensation unit (72), an intra prediction unit (74), an inverse quantization unit (76), an inverse transformation unit (78), a reference frame memory (82), and a summer (80). The decoder (28) may perform a decoding pass that is generally opposite to the encoding pass described in relation to the encoder (22) (Fig. 2) in some examples. The motion compensation unit (72) may generate prediction data based on a motion vector received from the entropy decoding unit (70), while the intra prediction unit (74) may generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit (70).

[0084] During the decoding process, the decoder (28) receives an encoded video bitstream from the encoder (22) representing a video block of encoded video slices and associated syntax elements. The entropy decoding unit (70) of the decoder (28) entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra-predict mode indicators, and other syntax elements. The entropy decoding unit (70) passes the motion vectors and other syntax elements to the motion compensation unit (72). The decoder (28) may also receive syntax elements at the video slice level and / or video block level.

[0085] When a video slice is coded as an intra-coded (I) slice, the intra-prediction unit (74) may generate prediction data for a video block of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current frame or image. When a video frame is coded as an inter-coded (e.g., B, P, or GPB) slice, the motion compensation unit (72) generates a prediction block for a video block of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit (70). The prediction block may be generated from one of the reference images within one of the reference image lists. The decoder (28) may configure List 0 and List 1, which are reference frame lists, using a default configuration technique based on the reference images stored in the reference frame memory (82).

[0086] The motion compensation unit (72) determines prediction information for a video block of the current video slice by parsing motion vectors and other syntax elements, and generates a prediction block for the current video block to be decoded using the prediction information. For example, the motion compensation unit (72) uses a portion of the received syntax elements to determine the prediction mode used to code the video block of the video slice (e.g., intra prediction or inter prediction), the inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the reference image lists for the slice, motion vectors for each inter-encoded video block of the slice, the inter prediction state for each inter-encoded video block, and other information for decoding the video block of the current video slice.

[0087] The motion compensation unit (72) may also perform interpolation based on an interpolation filter. The motion compensation unit (72) may calculate an interpolated value for a sub-integer pixel of a reference block using an interpolation filter used by the encoder (22) during the encoding of a video block. In this case, the motion compensation unit (72) may determine the interpolation filter used by the encoder (22) from the received syntax elements and may generate a prediction block using the interpolation filter.

[0088] Data for a texture image corresponding to a depth map can be stored in a reference frame memory (82). The motion compensation unit (72) may also be configured to mutually predict depth blocks of the depth map.

[0089] FIG. 4 is a representation of a frame bitstream group (400). As illustrated, the frame bitstream group (400) includes a first frame group (402) (GOF_0) and a second frame group (404) (GOF_1). For illustrative purposes only, the first frame group (402) and the second frame group (404) are separated from each other by a dotted line. Although two frame groups (402, 404) are illustrated in FIG. 4, it should be understood that in actual applications, any number of frames may be included in the frame bitstream group (400).

[0090] The first frame group (402) and the second frame group (404) are each formed from a collection of access units (406). The access unit (406) is configured to include a frame containing all or part of a compressed image (e.g., a point cloud). The access unit (406) of FIG. 4 may include an atlas frame, or be referred to herein as an atlas frame. In one embodiment, the atlas frame is a frame containing sufficient information to reconstruct a point cloud by mapping together coded components, wherein the components include point geometry, point attributes, occupancy maps, patches, etc.

[0091] A coding technique that allows for flexible patch orientation is disclosed herein. As used herein, patch orientation considers patch rotation, mirroring, and patch swapping (collectively, rotation). By using a flexible patch orientation scheme, patches can be packed more efficiently into a bounding box. Due to more efficient packing, the area occupied by the patch within the bounding box can be reduced compared to techniques where patch rotation is unavailable or not allowed, which leads to better coding efficiency (e.g., fewer bits are required to code a bitstream).

[0092] FIG. 5 is a representation of a point cloud (500). The point cloud (500) is a volumetric representation of space on a regular 3D grid. That is, the point cloud (500) is three-dimensional (3D). As illustrated in FIG. 5, the point cloud (500) contains point cloud content (502) within a 3D space (504). The point cloud content (502) is represented by a set of points (e.g., voxels) within the 3D space (504). A voxel is a volume element representing some numerical quantity, such as a point color in a 3D space, used for the visualization and analysis of 3D data. Thus, a voxel can be thought of as a 3D element corresponding to a pixel in a 2D image.

[0093] Each voxel of the point cloud (500) of FIG. 5 has coordinates (e.g., xyz coordinates) and one or more attributes (e.g., red / green / blue (RGB) color components, reflectance, etc.). On the other hand, while the point cloud content (502) of FIG. 5 depicts a person, the point cloud content (502) may be any other volumetric object or image in an actual application.

[0094] FIG. 6 is a representation of the point cloud (500) of FIG. 5 projected onto a bounding box (600). As illustrated in FIG. 6, the bounding box (600) includes a patch (602) projected onto a two-dimensional (2D) surface or its plane (604). Thus, the patch (602) is a 2D representation of a portion of a 3D image. The patch (602) collectively corresponds to the point cloud content (502) of FIG. 5. Data representation in video-based point cloud coding (V-PCC), also known as point cloud compression, relies on this 3D-to-2D conversion.

[0095] Data representation in V-PCC is described as a set of planar 2D images (e.g., patches (602)) using, for example, the occupancy map (710) of FIG. 7, the geometry map (810) of FIG. 8, and the attribute map (910) of FIG. 9.

[0096] FIG. 7 is a representation of an occupancy map (710) corresponding to a 2D projection (e.g., patch (602)) from the bounding box (600) of FIG. 6. The occupancy map (710) is coded in binary form. For example, 0 (zero) indicates that a portion of the bounding box (700) is not occupied by one of the patches (702). The portions of the bounding box (700) represented by 0 do not participate in the reconstruction of the volume representation (e.g., point cloud content (502)). In contrast, 1 (one) indicates that a portion of the bounding box (700) is occupied by one of the patches (702). The portions of the bounding box (700) represented by 1 participate in the reconstruction of the volume representation (e.g., point cloud content (502)).

[0097] FIG. 8 is a representation of a geometry map (810) corresponding to a 2D projection (e.g., patch (602)) from the bounding box (600) of FIG. 6. The geometry map (810) provides or depicts the contour or topography of each patch (802). That is, the geometry map (810) indicates the distance of each point of the patch (802) from the plane (e.g., plane (604)) of the bounding box (800).

[0098] FIG. 9 is a representation of an attribute map (910) corresponding to a 2D projection (e.g., patch (602)) from the bounding box (600) of FIG. 6. The attribute map (910) provides or depicts the attributes of each point within the patch (902) of the bounding box (900). The attributes of the attribute map (910) may be, for example, the color components of the points. The color components may be based on an RGB color model, a YUV color model, or other known color models.

[0099] Currently, patch projection (e.g., the process of projecting the 3D image (504) of FIG. 5 onto the bounding box (600) of FIG. 6 as a collection of patches (602)) is performed without changing the orientation of the patches. That is, the patches are not rotated or manipulated relative to their original orientation. However, this is a suboptimal solution for efficient packing. To overcome this, the present disclosure provides a technique that allows for flexible patch orientation (e.g., rotation, mirroring, and axis swapping of the patches). By using a flexible patch orientation scheme, patches can be packed more efficiently into the bounding box. Due to more efficient packing, the area occupied by the patches within the bounding box can be reduced compared to techniques where patch rotation is unavailable or not allowed, which leads to better coding efficiency (e.g., fewer bits are required to code the bitstream).

[0100] FIG. 10 is an example of a representation of a patch orientation index (1000). The patch orientation index (1000) includes various predefined patch orientations (1008) for a patch (1002). Different patch orientations (1008) of the patch orientation index (1000) are provided with numbers from #0 to #7 for identification. The initial patch orientation (1008) assigned to #0 may be referred to as the anchor patch orientation. As illustrated, the patch orientation (1008) assigned to #0 is rotated 90 degrees clockwise to obtain the patch orientation (1008) having #1. Likewise, the patch orientation (1008) assigned to #1 is rotated 90 degrees clockwise to obtain the patch orientation (1008) having #2, and the patch orientation (1008) assigned to #2 is rotated 90 degrees clockwise to obtain the patch orientation (1008) having #3.

[0101] Referring still to FIG. 10, the patch orientation (1008) assigned to #0 is mirrored (e.g. flipped with respect to the vertical axis) to obtain the patch orientation (1008) having #4. The patch orientation (1008) assigned to #4 is rotated 90 degrees clockwise to obtain the patch orientation (1008) having #5, the patch orientation (1008) assigned to #5 is rotated 90 degrees clockwise to obtain the patch orientation (1008) having #6, and the patch orientation (1008) assigned to #6 is rotated 90 degrees clockwise to obtain the patch orientation (1008) having #7.

[0102] In one embodiment, a pair of patch orientations (1008) is formed to create a simple orientation. For example, a patch orientation (1008) assigned to #0 and a patch orientation (1008) assigned to #7 form a simple orientation. Similarly, a patch orientation (1008) assigned to #1 and a patch orientation (1008) assigned to #6 form a simple orientation, a patch orientation (1008) assigned to #2 and a patch orientation (1008) assigned to #5 form a simple orientation, and a patch orientation (1008) assigned to #3 and a patch orientation (1008) assigned to #4 form a simple orientation. Thus, the patch orientation index (1000) provides four simple orientations.

[0103] Exhaustive orientation (also known as an exhaustive orientation map) includes all patch orientations (1008) of the patch orientation index (1000). That is, in the embodiment illustrated in FIG. 10, all eight of the patch orientations (1008) (labeled #0 through #7) are included in the exhaustive orientation. Using exhaustive orientation can provide better compression compared to simple orientation, but incurs additional computational costs in the encoder.

[0104] In one embodiment, the patch (1002) may have a default patch orientation (e.g., one of the patch orientations (1008)). For example, the patch (1002) may have a default rotation. Any one of the patch orientations (1008) (e.g., #0 through #7) may be the default rotation. The default patch orientation may be signaled in a sequence parameter set (SPS), a picture parameter set (PPS), or a frame group of the bitstream (e.g., at the patch level).

[0105] In one embodiment, the patch (1002) may have a preferred patch orientation (e.g., one of the patch orientations (1008)). For example, the patch (1002) may have a preferred rotation. Any one of the patch orientations (1008) (e.g., #0 through #7) may be the preferred rotation. The preferred patch orientation may be signaled in the SPS, PPS, or frame group of the bitstream (e.g., patch level).

[0106] In one embodiment, a flag is used to indicate that a patch (e.g., patch (1002)) can be rotated. In one embodiment, the flag is designated as a patch rotation enable flag. In one embodiment, the patch rotation enable flag is a 1-bit flag. When the patch rotation enable flag has a first value (e.g., 0), the patch cannot be rotated. When the patch rotation enable flag has a second value (e.g., 1), the patch can be rotated.

[0107] In one embodiment, a flag is used to indicate that a patch (e.g., patch (1002)) has a simple orientation (e.g., one of two patch orientations (1008)). In one embodiment, the flag is designated as a limited rotation enable flag. In one embodiment, the limited rotation enable flag is a 1-bit flag. When the limited rotation enable flag has a first value (e.g., 0), full orientation may be available for the patch. When the limited rotation enable flag has a second value (e.g., 1), simple orientation may be available for the patch.

[0108] In one embodiment, a flag is used to indicate that a patch (e.g., patch (1002)) has been rotated. In one embodiment, the flag is designated as the patch rotation present flag. In one embodiment, the patch rotation present flag is a 1-bit flag. When the patch rotation present flag has a first value (e.g., 0), the default patch orientation is used. When the patch rotation present flag has a second value (e.g., 1), one of the eight patch orientations (1008) from the priority patch orientation or full orientation is used depending on the value of the restricted rotation enable flag, which is described in more detail below. In one embodiment, the restricted rotation enable flag is a 3-bit flag.

[0109] FIG. 11 is an example of a patch orientation decoding process (1100). The patch orientation decoding process (1100) can be used to decode an encoded bitstream to reconstruct a volume image. In block (1102), a default patch rotation and a priority patch rotation for a patch (e.g., a 2D patch) are obtained from the encoded bitstream. The default patch rotation and the priority patch rotation can be represented by index numbers #0 through #7.

[0110] In block (1104), a patch is processed. In one embodiment, processing of the patch includes obtaining atlas information, the 2D location of the patch, the 3D location of the patch, and extracting the flags described above for the patch. Atlas information is information that can decode the patch for reconstruction and map from a 2D representation to a 3D representation. In one embodiment, atlas information includes a list of patches.

[0111] In block (1106), the value of the patch rotation enable flag is determined. When the patch rotation enable flag has a first value (e.g., 0), the patch cannot be rotated. Therefore, in block (1108), the patch is not rotated. Then, in block (1110), the patch is processed based on atlas information (also known as auxiliary information) to reconstruct the volume image.

[0112] Returning to block (1106), when the patch rotation enable flag has a second value (e.g., 1), the patch can be rotated. In block (1112), the value of the restricted rotation enable flag is determined. When the restricted rotation enable flag has a first value (e.g., 0), the patch is rotated using one of the eight available patch rotations according to the full orientation mode.

[0113] In block (1114), the value of the patch rotation presence flag is determined. When the patch rotation presence flag has a first value (e.g., 0), the default patch orientation is used. Thus, in block (1116), the default rotation is used for the patch. In one embodiment, no additional signaling is required because all patches in the patch group use the same default patch orientation. Then, in block (1110), the patch is processed based on atlas information to reconstruct the volume image.

[0114] Returning to block (1114), when the patch rotation presence flag has a second value (e.g., 1), one of the available orientations (represented by index numbers #0 through #7) is used. In one embodiment, the patch orientation to be used is signaled using a 3-bit flag. Then, in block (1110), the patch is processed based on atlas information to reconstruct the volume image.

[0115] Returning to block (1112), when the restricted rotation enable flag has a second value (e.g., 1), the patch is rotated using one of two patch rotations according to the simple orientation mode. In block (1120), the value of the patch rotation presence flag is determined. When the patch rotation presence flag has a first value (e.g., 0), the default patch orientation is used. Thus, in block (1116), the default rotation is used for the patch. In one embodiment, no additional signaling is required because all patches in the patch group use the same default patch orientation. Then, in block (1110), the patch is processed based on atlas information to reconstruct the volume image.

[0116] When the patch rotation presence flag has a second value (e.g., 1), the priority patch orientation is used. Thus, in block (1122), the priority rotation is used for the patch. In one embodiment, no additional signaling is required because all patches in the patch group use the same priority patch orientation. Then, in block (1110), the patch is processed based on atlas information to reconstruct the volume image.

[0117] FIG. 12 is an example of a point cloud coding (PCC) (1200) method implemented by a decoder (e.g., an entropy decoding unit (70)). The method (1200) may be performed to decode an encoded bitstream to reconstruct a volume image. In block (1202), a bitstream containing a patch rotation enable flag and atlas information for a two-dimensional (2D) patch is received. The atlas information (also known as auxiliary information) contains sufficient information to reconstruct a point cloud by mapping together the coded components, wherein the components include point geometry, point attributes, occupancy maps, patches, etc.

[0118] In block (1204), a decision is made that the 2D patch can be rotated based on the patch rotation enable flag. In one embodiment, the patch can be rotated when the patch rotation enable flag is set to the value of (1).

[0119] In block (1206), a 2D patch is rotated. The 2D patch may be rotated according to a default patch rotation, a priority patch rotation, or one of a plurality of available patch rotations as described herein. That is, in one embodiment, the 2D patch may be rotated according to a simple orientation mode in which the default patch rotation or the priority patch rotation is used, or according to a full orientation mode in which any one of the eight available patch rotations is used.

[0120] In block (1208), a three-dimensional (3D) image is reconstructed using atlas information and a rotated 2D patch. Once reconstructed, the 3D image can be displayed for a user on the display of an electronic device (e.g., a smartphone, tablet, laptop computer, etc.).

[0121] FIG. 13 is an example of a point cloud coding (PCC) (1300) method implemented by an encoder (e.g., an entropy encoding unit (56)). The method (1300) may be performed to encode a volume image into a bitstream for transmission toward a decoder. In block (1302), a three-dimensional (3D) image (e.g., a volume image) is acquired. In block (1304), a plurality of two-dimensional (2D) projections on the 3D image are determined using a plurality of available patch rotations. In one embodiment, the available patch rotation is the patch orientation (1008) shown in FIG. 10. However, other patch rotations may be used.

[0122] In block (1306), one of a plurality of 2D projections is selected. One of the plurality of 2D projections may be selected when the 2D projection results in the most efficient packing of the bounding box compared to other 2D projections. The most efficient packing of the bounding box may result in the use of the smallest amount of area and the least intensive computation by the encoder, etc.

[0123] In block (1308), a plurality of flags are set according to one of the selected plurality of 2D projections. In one embodiment, a patch rotation flag, a restricted rotation enable flag, a patch rotation presence flag, and a flag identifying which of the plurality of available patch rotations is set are each set.

[0124] In block (1310), a bitstream containing a plurality of flags and atlas information is generated to reconstruct a 3D image. In block (1312), the bitstream is stored for transmission toward a decoder. In one embodiment, the bitstream is transmitted toward a decoder.

[0125] FIG. 14 is a schematic diagram of a coding device (1400) (e.g., encoder (22), decoder (28), etc.) according to an embodiment of the present disclosure. The coding device (1400) is suitable for implementing the method and process disclosed herein. The coding device (1400) includes an ingress port (1410) and a receiver unit (Rx) (420) for receiving data; a processor, logic unit, or central processing unit (CPU) (1430) for processing data; a transmitter unit (Tx) (1440) and an egress port (1450) for transmitting data; and a memory (1460) for storing data. The coding device (1400) may also include an optical-to-electrical (OE) component and an electrical-to-optical (EO) component coupled to an entry port (1410), a receiver unit (1420), a transmitter unit (1440), and an exit port (1450) for the outflow or inflow of an optical or electrical signal.

[0126] The processor (1430) is implemented in hardware and software. The processor (1430) may be implemented as one or more CPU chips, cores (e.g., as multi-core processors), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor (1430) communicates with an entry port (1410), a receiver unit (1420), a transmitter unit (1440), an exit port (1450), and a memory (1460). The processor (1430) includes a coding module (1470). The coding module (1470) implements the disclosed embodiments described above. In one embodiment, the coding module (1470) is a reconstruction module configured to project a reconstructed volume image. Thus, the inclusion of the coding module (1470) provides a substantial improvement to the functionality of the coding device (1400) and has the effect of transforming the coding device (1400) into a different state. Alternatively, the coding module (1470) is implemented as an instruction stored in memory (1460) and executed by the processor (1430).

[0127] The coding device (1400) may also include an input and / or output (I / O) device (1480) for communicating data with the user. The I / O device (1480) may include an output device such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device (1480) may also include an input device such as a keyboard, mouse, trackball, etc., and / or a corresponding interface for interacting with these output devices.

[0128] The memory (1460) includes one or more disks, tape drives and solid-state drives and can be used as an overflow data storage device, and stores the program when the program is selected for execution and stores instructions and data read during program execution. The memory (1460) may be volatile and non-volatile and may be ROM (read-only memory), RAM (random-access memory), TCAM (ternary content-addressable memory), and SRAM (static random-access memory).

[0129] FIG. 15 is a schematic diagram of an embodiment of means for coding (1500). In the embodiment, the means for coding (1500) is implemented in a coding device (1502) (e.g., an encoder (22) or a decoder (28)). The coding device (1502) includes a receiving means (1501). The receiving means (1501) is configured to receive an image to be encoded or to receive a bitstream to be decoded. The coding device (1502) includes a transmission means (1507) coupled to the receiving means (1501). The transmission means (1507) is configured to transmit the bitstream to a decoder or to transmit the decoded image to a display means (e.g., one of an I / O device (1480)).

[0130] The coding device (1502) includes a storage means (1503). The storage means (1503) is coupled to at least one of a receiving means (1501) or a transmitting means (1507). The storage means (1503) is configured to store commands. The coding device (1502) also includes a processing means (1505). The processing means (1505) is coupled to the storage means (1503). The processing means (1505) is configured to execute commands stored in the storage means (1503) to perform the method disclosed herein.

[0131] In one embodiment, a syntax suitable for implementing the concept disclosed herein is provided.

[0132] Possible syntax definitions are described below.

[0133]

[0134] Modification of auxiliary information data unit

[0135]

[0136] From the foregoing, it should be recognized that the orientation of a patch may vary depending on the default projection process. The orientation of a patch can be signaled in a simplified manner using a 1-bit flag. A mechanism for switching between the default orientation and the preferred orientation has been introduced.

[0137] While some embodiments have been provided in this disclosure, it is understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of this disclosure. These examples should be considered exemplary and are not limiting, and the intent is not limited to the details given herein. For example, various elements or components may be combined or integrated into other systems, or certain functions may be omitted or not implemented.

[0138] Additionally, the technologies, systems, subsystems, and methods described and illustrated in various embodiments as discrete or separate may be combined with or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure. Other examples of modifications, substitutions, and alterations are identifiable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

Claim 1 A method for decoding a three-dimensional (3D) image, comprising the steps of: receiving a bitstream including at least one flag and atlas information for a two-dimensional (2D) patch, wherein the at least one flag includes a first flag; determining an orientation mode for the 2D patch based on the at least one flag; orienting the 2D patch based on the orientation mode; and reconstructing a 3D image based on the oriented 2D patch and the atlas information, wherein the first flag has a first value indicating that the orientation mode used to orient the 2D patch is selected from a first set of orientation modes; or the first flag has a second value indicating that the orientation mode used to orient the 2D patch is selected from a second set of orientation modes. Claim 2 A method for decoding a 3D image according to claim 1, wherein the first flag is a 1-bit flag. Claim 3 A method for decoding a 3D image according to claim 1, wherein the first orientation mode set includes two orientation modes, and the at least one flag further includes a second flag, and the second flag is a 1-bit flag indicating whether any one of the first orientation mode set is selected as the orientation mode. Claim 4 A method for decoding a 3D image according to claim 1, wherein the second orientation mode set includes eight orientation modes, and the at least one flag further includes a third flag that identifies the orientation mode in the second orientation mode set. Claim 5 A method for decoding a 3D image, wherein the eight orientation modes are each labeled from 0 to 7. Claim 6 A method for decoding a 3D image, wherein, in paragraph 4, the third flag is a 3-bit flag. Claim 7 A method for decoding a 3D image according to claim 1, wherein the first value is 0 and the second value is 1. Claim 8 A method for encoding a three-dimensional (3D) image, comprising: acquiring a 3D image; determining a two-dimensional (2D) projection of the 3D image based on an orientation mode; setting at least one flag for a 2D patch according to the 2D projection, wherein the at least one flag includes a first flag; and generating a bitstream including atlas information for reconstructing the 3D image and the at least one flag, wherein the first flag has a first value indicating that the orientation mode used to orient the 2D patch is selected from a first set of orientation modes; or the first flag has a second value indicating that the orientation mode used to orient the 2D patch is selected from a second set of orientation modes. Claim 9 A method for encoding a 3D image, wherein the first flag is a 1-bit flag, in paragraph 8. Claim 10 A method for encoding a 3D image according to claim 8, wherein the first orientation mode set comprises two orientation modes, and the at least one flag further comprises a second flag, and the second flag is a 1-bit flag indicating whether any one of the first orientation mode set is selected as the orientation mode. Claim 11 A method for encoding a 3D image according to claim 8, wherein the second orientation mode set comprises eight orientation modes, the at least one flag further comprises a third flag, and the third flag identifies the orientation mode in the second orientation mode set. Claim 12 A method for encoding a 3D image, wherein the eight orientation modes are each labeled from 0 to 7. Claim 13 A method for encoding a 3D image, wherein the third flag is a 3-bit flag, in paragraph 11. Claim 14 A method for encoding a 3D image according to claim 8, wherein the first value is 0 and the second value is 1. Claim 15 A decoding device comprising: a memory for storing instructions; and a processor coupled to said memory and configured to execute instructions stored in said memory so that said decoding device executes the method of any one of claims 1 to 7; or a decoding device comprising a processing circuit for performing the method of any one of claims 1 to 7. Claim 16 An encoding device comprising: a memory for storing instructions; and a processor coupled to said memory and configured to execute instructions stored in said memory so that said encoding device executes the method of any one of claims 8 to 14; or an encoding device comprising a processing circuit for performing the method of any one of claims 8 to 14.

Citation Information

Patent Citations

  • Methods and apparatus for reflective symmetry based 3D model compression

    KR1020140098094A