Transfer format for encoding and decoding point clouds
By copying the three-channel texture image to the four-channel image and storing the combined information in the fourth channel, the four-channel image is transmitted to solve the problem of inefficient image structured data transmission bandwidth, and more efficient memory transmission is achieved.
Patent Information
- Application Number
- CN202080060356.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-29
- Filing Date
- 2020-08-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-08-18
AI Technical Summary
The prior art has low bandwidth efficiency when transmitting image structured data, especially when multiple monochrome geometric images and three-channel texture images need to be transmitted, resulting in low memory transmission efficiency.
The four-channel image is transmitted to realize the transmission of image structured data by copying the three channels of the four-channel image and storing the combined information in the fourth channel of the four-channel image.
The transmission bandwidth between memories is optimized, the access requirements of memory cache is reduced, while the 2D structure of data is preserved.
Smart Images

Figure CN114341941B_ABST
Abstract
Description
Technical Field
[0001] At least one of the present embodiments generally relates to the processing of point clouds. In particular, the transmission of image structured data between a processing unit and a memory is presented, where the processing unit and the memory are used to implement an image-based point cloud decoder. Background Art
[0002] This section aims to introduce the reader to aspects of the field that may be related to the aspects of at least one of the embodiments described and / or claimed herein. This discussion is considered to be helpful in providing background information to enable the reader to better understand the aspects of at least one of the embodiments.
[0003] Point clouds can be used for various purposes, such as cultural heritage / buildings, where objects such as statues or buildings are 3D scanned in order to share the spatial configuration of the object without having to transport or access it. Additionally, this is a way to ensure the preservation of knowledge about the object in case it may be damaged; for example, a temple after an earthquake. Such point clouds are typically static, colored, and huge.
[0004] Another use case is for topography and cartography, where the use of 3D representations allows maps to be not limited to a plane and may include relief. Google Maps is now a good example of a 3D map, but uses meshes instead of point clouds. Nevertheless, point clouds may be a suitable data format for 3D maps, and such point clouds are typically static, colored, and huge.
[0005] Point clouds can also be used in the automotive industry and the field of autonomous vehicles. Autonomous vehicles should be able to "sense" their environment in order to make appropriate driving decisions based on the actual situation of their neighbors. Typical sensors such as LIDAR (Light Detection and Ranging) produce dynamic point clouds that are used by the decision-making engine. These point clouds are not intended to be viewed by humans, and they are typically small, not necessarily colored, and dynamic, with a high capture frequency. These point clouds may have other attributes, such as the reflectivity provided by LIDAR, as this attribute provides good information about the material of the sensed object and may help in making decisions.
[0006] Virtual reality and immersive worlds have recently become a hot topic and are foreseen by many as the future of 2D flat video. The basic idea is to immerse the viewer in the surrounding environment, which is different from a standard TV where the viewer can only see the virtual world in front of the viewer. Depending on the degree of freedom of the viewer in the environment, there are several levels of immersion. Point clouds are a good candidate format for distributing virtual reality (VR) worlds.
[0007] In many applications, it is important to be able to distribute dynamic point clouds to end-users (or store them in a server) by consuming only a reasonable amount of bitrate (or storage space for storage applications), while maintaining an acceptable (or preferably very high) quality of experience. The efficient compression of these dynamic point clouds is a key point that makes the distribution chain of many immersive worlds practical.
[0008] In view of the above, at least one embodiment has been designed. Summary of the Invention
[0009] To provide a basic understanding of some aspects of the present disclosure, a simplified overview of at least one embodiment in this embodiment is presented below. This overview is not an extensive review of the embodiment. It is not intended to identify key or decisive elements of the embodiment. The following overview only presents some aspects of at least one embodiment in this embodiment in a simplified form, as a prelude to the more detailed description provided elsewhere in this document.
[0010] According to a general aspect of at least one embodiment, a method is provided, including transmitting at least one three-channel texture image representing a point cloud texture and at least two image structured data for reconstructing the geometry of the point cloud between two processing units or memories, where the transmission includes:
[0011] - Copying three channels of the at least one three-channel texture image into three channels of a four-channel image;
[0012] - Storing combined information in a fourth channel of the four-channel image, where the combined information is obtained by combining the at least two image structured data; and
[0013] - Transmitting the four-channel image.
[0014] Such a method can be used, for example, to reconstruct the point cloud.
[0015] According to an embodiment, one image structured data is a monochromatic geometry image representing the geometry of the point cloud, and the other image structured data is an occupancy map, where the pixel value in the occupancy map indicates whether a block of the texture and the monochromatic geometry image includes at least one orthographic projection point of the point cloud, and where the pixel value of the fourth channel is the product of the deviation value of the pixel of the monochromatic geometry image and the value of the co-located pixel in the occupancy map.
[0016] According to an embodiment, one image structured data is a monochromatic geometry image representing the geometry of the point cloud, and the other image structured data is an occupancy map, where the pixel value in the occupancy map indicates whether a block of the texture and the monochromatic geometry image includes at least one orthographic projection point of the point cloud, and where the pixel value of the fourth channel is the product of the pixel value of the monochromatic geometry image and the value of the co-located pixel in the occupancy map.
[0017] According to an embodiment, when multiple monochromatic geometric shape images and multiple three-channel texture images need to be sent, the first channel, the second channel, and the third channel of at least two three-channel texture images are respectively packed into the first channel, the second channel, and the third channel of the four-channel image, and the combined information, which is obtained by combining the structured data of the at least two images together, is stored in the fourth channel of the four-channel image.
[0018] According to an embodiment, packing a three-channel texture image into a four-channel image includes replicating the texture image side by side.
[0019] According to an embodiment, packing a three-channel texture image into a four-channel image includes alternately interleaving the information represented by the three-channel texture image according to a pattern.
[0020] According to an embodiment, the fourth channel is the alpha channel in the RGBA image format.
[0021] One or more of at least one embodiment also provide a device, a computer program product, and a non-transitory computer-readable medium.
[0022] The specific nature of at least one embodiment in this embodiment and other objectives, advantages, features, and uses of the at least one embodiment will become apparent from the following description of examples in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In the drawings, examples of several embodiments are shown. The drawings show:
[0024] Figure 1 A schematic block diagram showing an example of a two-layer based point cloud encoding structure according to at least one embodiment in this embodiment;
[0025] Figure 2 A schematic block diagram showing an example of a two-layer based point cloud decoding structure according to at least one embodiment in this embodiment;
[0026] Figure 3 A schematic block diagram showing an example of an image-based point cloud encoder according to at least one embodiment in this embodiment;
[0027] Figure 3a An example of a canvas showing 2 patches and their 2D bounding boxes;
[0028] Figure 3b An example of two in-between 3D samples located between two 3D samples along a projection line;
[0029] Figure 4Schematically shows a block diagram of an example of an image-based point cloud decoder according to at least one embodiment of the present embodiment;
[0030] Figure 5 Schematically shows a syntax example of a bitstream representing a base layer BL according to at least one embodiment of the present embodiment;
[0031] Figure 6 Schematically shows a block diagram of an example of a system in which various aspects and embodiments are implemented;
[0032] Figure 7 Shows a flowchart of a method for transferring image structured data between memories according to at least one embodiment of the present embodiment; and
[0033] Figure 8 Shows an example of combining image structured data together according to an embodiment. Detailed Description
[0034] The following will describe at least one of the present embodiments more fully with reference to the accompanying drawings, in which examples of at least one embodiment of the present embodiment are shown. However, one embodiment may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Accordingly, it should be understood that the embodiments are not intended to be limited to the particular forms disclosed. Instead, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present application.
[0035] When the figure is presented in the form of a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when the figure is presented in the form of a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0036] Similar or identical elements in the figures are denoted by the same reference numerals.
[0037] The aspects described and contemplated below can be implemented in many different forms. The following Figures 1-7 provides some embodiments, but other embodiments are also contemplated, and the Figures 1-7 discussion of which does not limit the breadth of the implementation.
[0038] At least one of these aspects generally relates to point cloud encoding and decoding, and at least another aspect generally relates to transmitting a generated or encoded bitstream.
[0039] More precisely, the various methods and other aspects described herein can be used to implement a module. For example, this embodiment can be implemented by a mixer / combiner of a geometric shape RG (output of the geometric shape generation module GGM, module 4500) and the reconstructed point cloud RPCF (IRPCF) (output of the texture generation module TGM, module 4600).
[0040] Furthermore, this aspect is not limited to MPEG standards such as MPEG-I Part 5 related to point cloud compression, and can be applied to, for example, other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including MPEG-I Part 5). Unless otherwise specified or technically excluded, the aspects described in this application can be used alone or in combination.
[0041] Hereinafter, image data refers to data, for example, one or several arrays of 2D samples of a specific image / video format. The specific image / video format can specify information related to the pixel values of the image (or video). For example, the specific image / video format can also specify information that can be used by a display and / or any other device for visualizing and / or decoding the image (or video). An image typically includes a first component (also known as a channel), which is a 2D array of samples and usually represents the luminance (or luma) of the image. The image can also include a second component and a third component, which are 2D arrays of other samples and usually represent the chrominance (or chroma) of the image. Such images are typically represented as three-channel images, such as traditional three-color RGB images or YCbCr / YUV images.
[0042] In one or more embodiments, the pixel value is represented by a vector of C values, where C is the number of components (channels). Each value of the vector is typically represented by a plurality of bits, where the bits can define the dynamic range of the pixel value.
[0043] An image block refers to a group of pixels belonging to an image. The pixel values of an image block (or image block data) refer to the values of the pixels belonging to that image block. Although rectangles are common, an image block can have any shape.
[0044] A point cloud can be represented by a 3D sample data set in a 3D volume space, where these 3D samples have unique coordinates and may also have one or more attributes.
[0045] The 3D samples of the data set can be defined by their spatial location (X, Y, and Z coordinates in 3D space) and possibly by one or more associated attributes, such as color (e.g., represented in the RGB or YUV color space), transparency, reflectivity, a two-component normal vector, or any feature representing a feature of the sample. For example, a 3D sample can be defined by 6 components (X, Y, Z, R, G, B) or equivalently (X, Y, Z, y, U, V), where (X, Y, Z) define the coordinates of a point in 3D space, and (R, G, B) or (y, U, V) define the color of this 3D sample. Attributes of the same type can be presented multiple times. For example, multiple color attributes can provide color information from different viewpoints.
[0046] A point cloud can be static or dynamic, depending on whether the cloud changes over time. Instances of a static point cloud or a dynamic point cloud are typically represented as point cloud frames. It should be noted that in the case of a dynamic point cloud, the number of points is usually not constant; instead, it typically changes over time. More generally, if anything changes over time, such as, for example, the number of points, the position of one or more points, or any attribute of any point, then the point cloud can be considered dynamic.
[0047] For example, a 2D sample can be defined by 6 components (u, v, Z, R, G, B) or equivalently (u, v, Z, y, U, V). (u, v) define the coordinates of the 2D sample in the 2D space of the projection plane. Z is the depth value of the 3D sample projected onto this projection plane. (R, G, B) or (y, U, V) define the color of this 3D sample.
[0048] Figure 1 A schematic block diagram showing an example of a two-layer based point cloud coding structure 1000 according to at least one embodiment of the present embodiment is shown.
[0049] The two-layer based point cloud coding structure 1000 can provide a bitstream B representing an input point cloud frame IPCF. It is possible that the input point cloud frame IPCF represents a frame of a dynamic point cloud. Then, the frame of the dynamic point cloud can be encoded by the two-layer based point cloud coding structure 1000 independently of another frame.
[0050] Basically, the two-layer based point cloud coding structure 1000 can provide the ability to structure the bitstream B into a base layer BL and an enhancement layer EL. The base layer BL can provide a lossy representation of the input point cloud frame IPCF, and the enhancement layer EL can provide a higher quality (possibly lossless) representation by encoding the isolated points not represented by the base layer BL.
[0051] As Figure 3As shown, the base layer BL can be provided by the image-based encoder 3000. The image-based encoder 3000 can provide a geometry / texture image representing the geometry / attributes of the 3D samples of the input point cloud frame IPCF. It may allow discarding isolated 3D samples. The base layer BL can be as Figure 4 decoded by the image-based decoder 4000 shown, which can provide an intermediate reconstructed point cloud frame IRPCF.
[0052] Then, returning to Figure 1 in the two-layer point cloud coding 1000, the comparator COMP can compare the 3D samples of the input point cloud frame IPCF with the 3D samples of the intermediate reconstructed point cloud frame IRPCF to detect / localize lost / isolated 3D samples. Next, the encoder ENC can encode the lost 3D samples and can provide the enhancement layer EL. Finally, the base layer BL and the enhancement layer EL can be multiplexed together by the multiplexer MUX to generate the bitstream B.
[0053] According to one embodiment, the encoder ENC can include a detector that can detect the 3D reference sample R of the intermediate reconstructed point cloud frame IRPCF and associate it with the lost 3D sample M.
[0054] For example, according to a given metric, the 3D reference sample R associated with the lost 3D sample M can be the nearest neighbor of M.
[0055] According to an embodiment, the encoder ENC can then encode the spatial localization of the lost 3D sample M and its attributes as a difference determined according to the spatial localization and attributes of the 3D reference sample R.
[0056] In a variant, those differences can be encoded separately.
[0057] For example, for the lost 3D sample M with spatial coordinates x(M), y(M), and z(M), the x coordinate position difference Dx(M), y coordinate position difference Dy(M), z coordinate position difference Dz(M), R attribute component difference Dr(M), G attribute component difference Dg(M), and B attribute component difference Db(M) can be calculated as follows:
[0058] Dx(M) = x(M) - x(R),
[0059] where x(M) is the x coordinate of the 3D sample M, and respectively Figure 3 R in the provided geometry image,
[0060] Dy(M) = y(M) - y(R)
[0061] where y(M) is the y coordinate of the 3D sample M, and respectivelyFigure 3 R in the provided geometric shape image,
[0062] Dz(M) = z(M) - z(R)
[0063] where z(M) is the z - coordinate of the 3D sample M, respectively Figure 3 R in the provided geometric shape image,
[0064] Dr(M) = R(M) - R(R).
[0065] where R(M) and R(R) are the r - color components of the color attributes of the 3D samples M and R respectively,
[0066] Dg(M) = G(M) - G(R).
[0067] where G(M) and G(R) are the g - color components of the color attributes of the 3D samples M and R respectively,
[0068] Db(M) = B(M) - B(R).
[0069] where B(M) and B(R) are the b - color components of the color attributes of the 3D samples M and R respectively.
[0070] Figure 2 Fig. shows a schematic block diagram of an example of a two - layer based point cloud decoding structure 2000 according to at least one embodiment of the present embodiment.
[0071] The behavior of the two - layer based point cloud decoding structure 2000 depends on its capabilities.
[0072] The two - layer based point cloud decoding structure 2000 with limited capabilities can access only the base layer BL from the bitstream B by using a demultiplexer DMUX, and then can provide an accurate (but lossy) version IRCF of the input point cloud frame IRPCF by decoding the base layer BL by the Figure 4 shown point cloud decoder 4000.
[0073] The two - layer based point cloud decoding structure 2000 with full capabilities can access both the base layer BL and the enhancement layer EL from the bitstream B by using a demultiplexer DMUX. As Figure 4 shown, the point cloud decoder 4000 can determine the intermediate reconstructed point cloud frame IRPCF from the base layer BL. The decoder DEC can determine the complementary point cloud frame CPCF from the enhancement layer EL. Then, the combiner COMB can combine the intermediate reconstructed point cloud frame IRPCF and the complementary point cloud frame CPCF together, thereby providing a higher - quality (possibly lossless) representation (reconstruction) CRPCF of the input point cloud frame IPCF.
[0074] Figure 3 A schematic block diagram showing an example of an image-based point cloud encoder 3000 according to at least one embodiment of the present embodiment is shown.
[0075] The image-based point cloud encoder 3000 utilizes an existing video codec to compress the geometric shape and texture (attribute) information of a dynamic point cloud. This is achieved by essentially converting the point cloud data into a set of different video sequences.
[0076] In a particular embodiment, two videos can be generated and compressed using an existing video codec, one for capturing the geometric shape information of the point cloud data and the other for capturing the texture information. An example of an existing video codec is the HEVC Main Profile encoder / decoder (ITU (02 / 2018) ITU-T H.265 Telecommunication Standardization Sector, Series H: Audiovisual and Multimedia Systems, Audiovisual Service Infrastructure - Mobile Video Coding, High Efficiency Video Coding, ITU-T H.265 Recommendation).
[0077] Additional metadata for interpreting the two videos is typically also generated and compressed separately. Such additional metadata includes, for example, the occupancy map OM and / or the auxiliary patch information PI.
[0078] The generated video bitstreams and the metadata can then be multiplexed together to generate a combined bitstream.
[0079] It should be noted that the metadata typically represents a small portion of the overall information. Most of the information is in the video bitstream.
[0080] Performing the test model class 2 algorithm (also denoted as V-PCC) of the MPEG draft standard defined in ISO / IEC JTC1 / SC29 / WG11 MPEG2019 / w18180 (January 2019, Marrakech) gives an example of such a point cloud encoding / decoding process).
[0081] In step 3100, the module PGM can generate at least one patch by decomposing the 3D samples of the dataset representing the input point cloud frame IPCF into 2D samples on a projection plane using a strategy that provides optimal compression.
[0082] A patch can be defined as a group of 2D samples.
[0083] For example, in V-PCC, as described by Hoppe et al. (Hugues Hoppe, Tony DeRose, Tom Duchamp, John McDonald, Werner Stuetzle. Surface reconstruction from unorganized points. ACM SIGGRAPH 1992 Proceedings, 71-78), first, the normal of each 3D sample is estimated. Next, an initial cluster of the input point cloud frame IPCF is obtained by associating each 3D sample with one of the six oriented planes of a 3D bounding box, where the 3D bounding box encompasses the 3D samples of the input point cloud frame IPCF. More precisely, each 3D sample is clustered and associated with the oriented plane having the closest normal (i.e., maximizing the scalar product of the point normal and the plane normal). Then, the 3D samples are projected onto their associated planes. A group of 3D samples that form a connected region on their plane is called a connected component. A connected component is a set of at least one 3D sample having similar normals and the same associated oriented plane. Then, the initial cluster is refined by iteratively updating the cluster associated with each 3D sample based on the normal of each 3D sample and the cluster of its nearest neighbors. The last step includes generating a patch from each connected component, which is done by projecting the 3D samples of each connected component onto the oriented plane associated with the connected component. The patch is associated with auxiliary patch information PI, which represents the auxiliary patch information defined for each patch to explain the projected 2D samples corresponding to the geometry and / or attribute information.
[0084] For example, in V-PCC, the auxiliary patch information PI includes 1) information indicating one of the six oriented planes of the 3D bounding box that encompasses the 3D samples of the connected component; 2) information related to the plane normal; 3) information determining the 3D positioning of the connected component relative to the patch represented by depth, tangential offset, and bi-tangential offset; and 4) information such as coordinates (u0, v0, u1, v1) in the projection plane that defines the 2D bounding box encompassing the patch.
[0085] In step 3200, the patch packing module PPM can map (place) at least one generated patch onto a 2D grid (also called a canvas) without any overlap in a way that typically minimizes unused space, and can ensure that each T×T (e.g., 16×16) block of the 2D grid is associated with a unique patch. The given minimum block size T×T of the 2D grid can specify the minimum distance between different patches placed on the 2D grid. The resolution of the 2D grid can depend on the size of the input point cloud, and its width W and height H, as well as the block size T, can be sent to the decoder as metadata.
[0086] The auxiliary patch information PI may also include information related to the association between the blocks and patches of the 2D grid.
[0087] In V-PCC, the auxiliary information PI may include block-to-patch index information (BlockToPatch), where the block-to-patch index information (BlockToPatch) determines the association between the blocks and patch indices of the 2D grid.
[0088] Figure 3a An example of the canvas C is shown, where the canvas C includes two patches P1 and P2 and their associated 2D bounding boxes B1 and B2. Note that the two bounding boxes may overlap in the canvas C as Figure 3a shown. The 2D grid (the partitioning of the canvas) is represented only within those bounding boxes, but the partitioning of the canvas also exists outside those bounding boxes. The bounding boxes associated with the patches can be divided into T×T blocks, typically T = 16.
[0089] The T×T blocks containing the 2D samples belonging to the patch can be considered occupied blocks. Each occupied block of the canvas is represented by a specific pixel value in the occupancy map OM (a three-channel image) (e.g., 1), and each unoccupied block of the canvas is represented by another specific value (e.g., 0). Then, the pixel values of the occupancy map OM can indicate whether the T×T blocks of the canvas are occupied, i.e., contain 2D samples belonging to the patch.
[0090] In Figure 3a , the occupied blocks are represented by white blocks, and the light gray blocks represent unoccupied blocks. The image generation process (steps 3300 and 3400) utilizes the mapping of the at least one generated patch to the 2D grid computed during step 3200 to store the geometry and texture of the input point cloud frame IPCF as an image.
[0091] In step 3300, the geometry image generator GIG may generate at least one geometry image GI from the input point cloud frame IPCF, the occupancy map OM, and the auxiliary patch information PI. The geometry image generator GIG may utilize the occupancy map information to detect (locate) the occupied blocks, and thus detect (locate) the non-empty pixels in the geometry image GI.
[0092] The geometry image GI may represent the geometry of the input point cloud frame IPCF and may be W×H pixels, for example, a monochrome image represented in the YUV420 - 8bit (bit) format.
[0093] To better handle the case where multiple 3D samples are projected (mapped) onto the same 2D sample in the projection plane along the same projection direction (line), multiple images, called layers, can be generated. Thus, different depth values D1, …, Dn can be associated with the 2D samples of the patch, and then multiple geometry images can be generated.
[0094] In V-PCC, the 2D samples of the patch are projected onto two layers. The first layer, also called the near layer, can store, for example, the depth value D0 associated with the 2D sample having a smaller depth. The second layer, also called the far layer, can store, for example, the depth value D1 associated with the 2D sample having a larger depth. Alternatively, the second layer can store the difference between the depth values D1 and D0. For example, the information stored in the second depth image can be within the interval [0, Δ] corresponding to the depth values in the range [D0, D0 + Δ], where Δ is a user-defined parameter describing the surface thickness.
[0095] In this way, the second layer can contain significant contour-like high-frequency features. Thus, it is obvious that the second depth image may be difficult to encode and decode using a traditional video codec, and thus, the depth value is unlikely to be reconstructed from the decoded second depth image, which results in poor geometric quality of the reconstructed point cloud frame.
[0096] According to an embodiment, the geometry image generation module GIG can encode and decode (derive) the depth values associated with the 2D samples of the first and second layers by using the auxiliary patch information PI.
[0097] In V-PCC, the positioning of the 3D sample in the patch with the corresponding connected component can be represented by the depth δ(u, v), the tangential offset s(u, v), and the bi-tangential offset r(u, v) as follows:
[0098] δ(u, v) = δ0 + g(u, v)
[0099] s(u, v) = s0 - u0 + u
[0100] r(u, v) = r0 - v0 + v
[0101] where g(u, v) is the luminance component of the geometry image, (u, v) is the pixel associated with the 3D sample on the projection plane, (δ0, s0, r0) is the 3D positioning of the corresponding patch of the connected component to which the 3D sample belongs, and (u0, v0, u1, v1) are the coordinates in the projection plane that define the 2D bounding box covering the projection of the patch associated with the connected component.
[0102] Thus, the geometric shape image generation module GIG can encode (derive) the depth values associated with the 2D samples of the layer (the first layer or the second layer or both) into the luminance component g(u, v) given by: g(u, v) = δ(u, v) - δ0. Note that this relationship can be used to reconstruct the 3D sample location (δ0, s0, r0) from the reconstructed geometric shape image g(u, v) with the accompanying auxiliary patch information PI.
[0103] According to an embodiment, a projection mode can be used to indicate whether the first geometric shape image GI0 can store the depth values of the 2D samples of the first layer or the second layer, and whether the second geometric shape image GI1 can store the depth values associated with the 2D samples of the second layer or the first layer.
[0104] For example, when the projection mode is equal to 0, the first geometric shape image GI0 can store the depth values of the 2D samples of the first layer, while the second geometric shape image GI1 can store the depth values associated with the 2D samples of the second layer. Conversely, when the projection mode is equal to 1, the first geometric shape image GI0 can store the depth values of the 2D samples of the second layer, while the second geometric shape image GI1 can store the depth values associated with the 2D samples of the first layer.
[0105] According to an embodiment, a frame projection mode can be used to indicate whether a fixed projection mode is used for all patches or whether a variable projection mode is used, in which each patch can use a different projection mode.
[0106] The projection mode and / or the frame projection mode can be sent as metadata.
[0107] For example, a frame projection mode decision algorithm can be provided in Section 2.2.1.3.1 of V-PCC.
[0108] According to an embodiment, when the frame projection indicates that a variable projection mode can be used, a patch projection mode can be used to indicate the appropriate mode for (de)projecting the patch.
[0109] The patch projection mode can be sent as metadata and may be information included in the auxiliary patch information PI.
[0110] For example, a patch projection mode decision algorithm is provided in Section 2.2.1.3.2 of V-PCC.
[0111] According to an embodiment of step 3300, the pixel value in the first geometric shape image, e.g., GI0, corresponding to the 2D sample (u, v) of the patch, may represent the depth value of at least one in - between 3D sample defined along the projection line corresponding to the 2D sample (u, v). More precisely, the in - between 3D sample is on the projection line and shares the same coordinates of the 2D sample (u, v), and the depth value D1 of the 2D sample (u, v) is encoded and decoded in the second geometric shape image, e.g., GI1. Further, the in - between 3D sample may have a depth value between the depth value D0 and the depth value D1. A flag bit may be associated with each of the in - between 3D samples, and the flag bit is set to 1 if the in - between 3D sample exists, otherwise the flag bit is set to 0.
[0112] Figure 3b An example showing two in - between 3D samples P i1 and P i2 lying between two 3D samples P0 and P1 along the projection line PL is shown. The 3D samples P0 and P1 have depth values equal to D0 and D1 respectively. The depth values D i1 and P i2 of the two in - between 3D samples P i1 and D i2 are respectively greater than D0 and less than D1.
[0113] Then, all the flag bits along the projection line may be concatenated to form a codeword, hereinafter referred to as an Enhanced Occupancy Map (EOM) codeword. As Figure 3b shown, assuming the EOM codeword length is 8 bits, where 2 bits are equal to 1 to indicate the positions of the two 3D samples P i1 and P i2 . Finally, all the EOM codewords may be packed in an image, e.g., the occupancy map OM. In that case, at least one patch of the canvas may contain at least one EOM codeword. Such a patch is denoted as a reference patch and the block of the reference patch is denoted as an EOM reference block. Thus, the pixel value of the occupancy map OM may be equal to a first value, e.g., 0, to indicate an unoccupied block of the canvas, or another value, e.g., greater than 0, to indicate an occupied block of the canvas, e.g., when D1 - D0 <= 1, or e.g., when D1 - D0 > 1, to indicate an EOM reference block of the canvas.
[0114] The positions of the pixels in the occupancy map OM indicating the EOM reference block and the bit values of the EOM codewords obtained from the values of those pixels indicate the 3D coordinates of the in - between 3D samples.
[0115] In step 3400, the texture image generator TIG may generate at least one texture image TI from the input point cloud frame IPCF, the occupancy map OM, the auxiliary patch information PI, and the geometry of the reconstructed point cloud frame derived from at least one decoded geometry image DGI, and the output of the video decoder VDEC ( Figure 4 step 4200 in
[0116] The texture image TI is a three-channel image that may represent the texture of the input point cloud frame IPCF and may be an image of WxH pixels, for example, represented in YUV420-8-bit format or RGB444-8-bit format.
[0117] The texture image generator TG may utilize the occupancy map information to detect (locate) the occupied blocks, thereby detecting (locating) the non-empty pixels in the texture image.
[0118] The texture image generator TIG may be adapted to generate the texture image TI and associate it with each geometry image / layer DGI.
[0119] According to an embodiment, the texture image generator TIG may encode / decode (store) the texture (attribute) value T0 associated with the 2D samples of the first layer as the pixel values of the first texture image TI0, and encode / decode (store) the texture value T1 associated with the 2D samples of the second layer as the pixel values of the second texture image TI1.
[0120] Alternatively, the texture image generation module TIG may encode / decode (store) the texture value T1 associated with the 2D samples of the second layer as the pixel values of the first texture image TI0, and encode / decode (store) the texture value D0 associated with the 2D samples of the first layer as the pixel values of the second geometry image GI1.
[0121] For example, the color of the 3D samples may be obtained as described in Sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of V-PCC.
[0122] The texture values of two 3D samples are stored in the first or second texture image. However, the texture values of the intermediate 3D samples cannot be stored in either the first texture image TI0 or the second texture image TI1 because the location of the projected intermediate 3D samples corresponds to those already used for storage as Figure 3bOccupancy blocks of the texture values of another 3D sample (P0 or P1) as shown. Thus, the texture value of the intermediate 3D sample is stored in an EOM texture block at other positions in the first or second texture image located at a program-defined location (section 9.4.5 of V-PCC). In short, this process determines the location of the unoccupied blocks in the texture image and stores the texture value associated with the intermediate 3D sample as the pixel value of the unoccupied block of the texture image, represented as an EOM texture block.
[0123] According to one embodiment, a filling process may be applied to the geometry and / or texture image. This filling process can be used to fill the blank spaces between the patches to generate a piecewise smooth image suitable for video compression.
[0124] Sections 2.2.6 and 2.2.7 of V-PCC provide image filling examples.
[0125] In step 3500, the video encoder VENC may encode the generated image / layer TI and GI.
[0126] In step 3600, the encoder OMENC may encode the occupancy map as an image, as detailed, for example, in section 2.2.2 of V-PCC. Lossy or lossless encoding may be used.
[0127] According to an embodiment, the video encoder ENC and / or OMENC may be an HEVC-based encoder.
[0128] In step 3700, the encoder PIENC may encode the auxiliary patch information PI and possible additional metadata (such as the block size T, width W, and height H of the geometry / texture image).
[0129] According to an embodiment, the auxiliary patch information may be differentially encoded (as defined, for example, in section 2.4.1 of V-PCC).
[0130] In step 3800, a multiplexer may be applied to the generated outputs of steps 3500, 3600, and 3700, and as a result, these outputs may be multiplexed together to generate a bitstream representing the base layer BL. It should be noted that this metadata information represents a small part of the overall bitstream. Most of the information is compressed using the video codec.
[0131] Figure 4 A schematic block diagram showing an example of an image-based point cloud decoder 4000 according to at least one embodiment of the present embodiment is shown.
[0132] In step 4100, a demultiplexer DMUX may be applied to demultiplex the encoded information of the bitstream representing the base layer BL.
[0133] In step 4200, the video decoder VDEC may decode the encoded information to derive at least one decoded geometry image DGI and at least one decoded texture image DTI.
[0134] In step 4300, the decoder OMDEC may decode the encoded information to derive a decoded occupancy map DOM.
[0135] According to an embodiment, the video decoder VDEC and / or OMDEC may be an HEVC-based decoder.
[0136] In step 4400, the decoder PIDEC may decode the encoded information to derive auxiliary patch information DPI.
[0137] Possibly, metadata may also be derived from the bitstream BL.
[0138] In step 4500, the geometry generation module GGM may derive the geometry RG of the reconstructed point cloud frame IRPCF from the at least one decoded geometry image DGI, the decoded occupancy map DOM, the decoded auxiliary patch information DPI, and possibly additional metadata.
[0139] The geometry generation module GGM may utilize the decoded occupancy map information DOM to locate non-empty pixels in at least one decoded geometry image DGI.
[0140] As described above, based on the pixel value of the decoded occupancy information DOM and the value of D1 - D0, the non-empty pixels belong to an occupied block or an EOM reference block.
[0141] According to an embodiment of step 4500, the geometry generation module GGM may derive two of the 3D coordinates of an intermediate 3D sample from the coordinates of the non-empty pixels.
[0142] According to an embodiment of step 4500, when the non-empty pixel belongs to the EOM reference block, the geometry generation module GGM may derive the third of the 3D coordinates of the intermediate 3D sample from the bit value of the EOM codeword.
[0143] For example, according to Figure 3b the example of, the EOM codeword EOMC is used to determine the intermediate 3D sample P i1 and P i2 between the 3D coordinates. The third coordinate of the intermediate 3D sample P i1 may be derived, for example, from D0 through D i1 = D0 + 3, and the third coordinate of the reconstructed 3D sample P i2 may be derived, for example, from D0 through D i2= D0 + 5 is derived. This offset value (3 or 5) is the number of intervals along the projection line between D0 and D1.
[0144] According to an embodiment, when the non-empty pixel belongs to an occupied block, the geometry generation module GGM may derive the 3D coordinates of the reconstructed 3D sample from the coordinates of the non-empty pixel, the value of the non-empty pixel in one of the at least one decoded geometry image DGI, the decoded auxiliary patch information, and possibly from additional metadata.
[0145] The use of non-empty pixels is based on the relationship between 2D pixels and 3D samples. For example, using the projection in V-PCC, the 3D coordinates of the reconstructed 3D sample can be expressed as follows according to the depth δ(u, v), the tangential offset s(u, v), and the bi-tangential offset r(u, v):
[0146] δ(u, v) = δ0 + g(u, v)
[0147] s(u, v) = s0 u0 + u
[0148] r(u, v) = r0 v0 + v
[0149] where g(u, v) is the luminance component of the decoded geometry image DGI, (u, v) is the pixel associated with the reconstructed 3D sample, (δ0, s0, r0) is the 3D positioning of the connected component to which the reconstructed 3D sample belongs, and (u0, v0, u1, v1) are the coordinates in the projection plane defining the 2D bounding box that encompasses the projection of the patch associated with the connected component.
[0150] In step 4600, the texture generation module TGM may derive the texture of the reconstructed point cloud frame IRPCF from the geometry RG and the at least one decoded texture image DTI.
[0151] According to an embodiment of step 4600, the texture generation module TGM may derive the texture of the non-empty pixels belonging to the EOM reference block from the corresponding EOM texture block. The positioning of the EOM texture block in the texture image is defined by the program (section 9.4.5 of V-PCC)
[0152] According to an embodiment of step 4600, the texture generation module TGM may directly export the texture of the non-empty pixels belonging to the occupied block as the pixel values of the first texture image or the second texture image.
[0153] Figure 5 Schematically shows an example syntax of the bitstream representing the base layer BL according to at least one embodiment in this embodiment.
[0154] The bitstream includes a bitstream header SH and at least one group of frame streams GOFS (Group Of Frame Stream).
[0155] The group of frame streams GOFS includes a header HS, at least one syntax element OMS representing an occupancy map OM, at least one syntax element GVS representing at least one geometric shape image (or video), at least one syntax element TVS representing at least one texture image (or video), and at least one syntax element PIS representing auxiliary patch information and other additional metadata.
[0156] In one variant, the group of frame streams GOFS includes at least one frame stream.
[0157] Figure 6 A schematic block diagram is shown that illustrates an example of a system in which various aspects and embodiments are implemented.
[0158] The system 6000 can be embodied as one or more devices including the various components described below and is configured to perform one or more of the aspects described in this document. Examples of devices that can form all or part of the system 6000 include personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output of the video decoder, pre-processors that provide input to the video encoder, network servers, set-top boxes, and any other devices for processing point clouds, videos or images or other communication devices. The elements of the system 6000 can be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of the system 6000 can be distributed across multiple ICs and / or discrete components. In various embodiments, the system 6000 can be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, the system 6000 can be configured to implement one or more of the aspects described in this document.
[0159] The system 6000 may include at least one processor 6010 configured to execute instructions loaded therein to implement various aspects described, for example, in this document. The processor 6010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 6000 may include at least one memory 6020 (e.g., volatile storage devices and / or non-volatile storage devices). The system 6000 may include a storage device 6040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 6040 may include internal storage devices, additional storage devices, and / or network-accessible storage devices.
[0160] The system 6000 may include an encoder / decoder module 6030 configured to, for example, process data to provide encoded data or decoded data, and the encoder / decoder module 6030 may include its own processor and memory. The encoder / decoder module 6030 may represent a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 6030 may be implemented as a separate element of the system 6000 or may be incorporated within the processor 6010 as a combination of hardware and software known to those skilled in the art.
[0161] The program code to be loaded onto the processor 6010 or the encoder / decoder 6030 to execute various aspects described in this document may be stored in the storage device 6040 and subsequently loaded onto the memory 6020 for execution by the processor 6010. According to various embodiments, one or more of the processor 6010, the memory 6020, the storage device 6040, and the encoder / decoder module 6030 may store one or more of various items during the execution of the processes described in this document. Such stored items may include but are not limited to point cloud frames, encoded / decoded geometry / texture video / images or portions thereof, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.
[0162] In some embodiments, the memory internal to the processor 6010 and / or the encoder / decoder module 6030 may be used to store instructions and provide a working memory for the processing to be performed during encoding or decoding.
[0163] However, in other embodiments, a memory external to the processing device (e.g., the processing device can be the processor 6010 or the encoder / decoder module 6030) can be used for one or more of these functions. The external memory can be the memory 6020 and / or the storage device 6040, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory can be used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM can be used as a working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), HEVC (High Efficiency Video Coding), or VVC (Versatile Video Coding).
[0164] As shown in block 6130, input to the elements of the system 6000 can be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that can receive RF signals, e.g., transmitted over the air by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0165] In various embodiments, the input devices of block 6130 can have associated therewith corresponding input processing elements known in the art. For example, the RF section can be associated with elements necessary for the operations of (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band to select a signal frequency band that can be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF section of various embodiments can include one or more elements to perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section can include a tuner that performs various functions of these, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband.
[0166] In one set-top box embodiment, the RF section and its associated input processing elements can receive an RF signal transmitted through a wired (e.g., cable) medium. Then, the RF section can perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.
[0167] Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.
[0168] Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF portion may include an antenna.
[0169] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 6000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented within, for example, a separate input processing IC or within the processor 6010 as needed. Similarly, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 6010 as needed. The streams of demodulation, error correction, and demultiplexing may be provided to various processing elements, including, for example, the processor 6010 and the encoder / decoder 6030, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0170] The various elements of the system 6000 may be disposed within an integrated housing. Within the integrated housing, the various elements may be interconnected using a suitable connection arrangement 6140 and data may be transmitted therebetween, such as internal buses known in the art, including I2C buses, wiring, and printed circuit boards.
[0171] The system 6000 may include a communication interface 6050 capable of communicating with other devices via a communication channel 6060. The communication interface 6050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 6060. The communication interface 6050 may include, but is not limited to, a modem or a network card, and the communication channel 6060 may be implemented, for example, within a wired and / or wireless medium.
[0172] In various embodiments, data may be streamed to the system 6000 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments may be received via the communication channel 6060 and the communication interface 6050 suitable for Wi-Fi communication. The communication channel 6060 of these embodiments may generally be connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other cloud communications.
[0173] Other embodiments may use a set-top box that emits data via an HDMI connection to the input block 6130 to provide streaming data to the system 6000.
[0174] Other embodiments may provide streaming data to the system 6000 using the RF connection of the input block 6130.
[0175] It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. can be used to signal information to the corresponding decoder.
[0176] The system 6000 may provide output signals to various output devices, including the display 6100, the speaker 6110, and other peripheral devices 6120. In various examples of the embodiments, the other peripheral devices 6120 may include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices based on the output providing functions of the system 3000.
[0177] In various embodiments, control signals may be communicated between the system 6000 and the display 6100, the speaker 6110, or the other peripheral devices 6120 using signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that allow device-to-device control with or without user intervention.
[0178] The output devices may be communicatively coupled to the system 6000 via dedicated connections through their respective interfaces 6070, 6080, and 6090.
[0179] Alternatively, the output devices may be connected to the system 6000 via the communication interface 6050 using the communication channel 6060. The display 6100 and the speaker 6110 may be integrated into a single unit in an electronic device (such as, for example, a television) together with other components of the system 6000.
[0180] In various embodiments, the display interface 6070 may include a display driver, such as a timing controller (T Con) chip.
[0181] For example, if the RF portion of the input 6130 is part of a separate set-top box, the display 6100 and the speaker 6110 may alternatively be combined with or separate from one of the multiple other components. In various embodiments where the display 6100 and the speaker 6110 may be external components, output signals may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0182] When implementing an image-based point cloud decoder, such as Figure 4For the V-PCC decoder, before deriving the geometry (RG) and texture of the reconstructed point cloud, the reconstructed point cloud needs to first decode the geometry (DGI) and texture (DTI) images and additional image structured data (such as occupancy map (DOM) or block-to-patch index information (BlockToPatch), a part of auxiliary patch information (DPI)). The decoded texture and geometry images, the decoded occupancy map, and the block-to-patch index information are image structured data, that is, they refer to data organized in a 2D array. The values of the elements of these 2D arrays and the 2D positioning of these elements in these 2D arrays are all relevant information for reconstructing the point cloud. For example, the pixel value of the decoded occupancy map DOM indicates whether the TxT block is occupied or unoccupied, as described above ( Figure 3a ). The 2D positioning of the pixels in the 2D array indicates the 2D positioning of the TxT block in the decoded geometry and texture images. As another example, the pixel value of the block-to-patch index information (monochrome image) indicates the patch index of the block, and the 2D positioning of the pixel in the 2D array indicates the 2D positioning of the block in the decoded geometry and texture images.
[0183] Generally, decoding image structured data can be implemented using a dedicated processing unit / memory, and reconstructing the point cloud from the decoded image can be implemented by using other dedicated processing units / memories. Therefore, it is necessary to copy the image structured data between the interfaces of these dedicated processing units / memories. Generally, such interfaces use three image channels, usually in the YUV420-8-bit format (or RGB444-8-bit). On recent processing units, the same interfaces are usually used for parallel implementation, and these interfaces are very suitable for transmitting multiple (a multiple of 3) image channels, such as a three-channel image or two three-channel images, or 3 monochrome images, etc. However, when the number of image channels to be copied is not a multiple of 3, for example, when sending a three-channel image and 2 monochrome images (then 5 image channels must be copied), 2 interfaces are copied, each using a three-channel image, which causes a bandwidth problem because 2 of these 5 image channels are not used. For example, there will be problems in transmitting the decoded three-channel texture image DTI, the decoded occupancy map DOM (where each pixel value represents binary information), and the monochrome geometry image DGI (or the block-to-patch index information represented by a monochrome image), because one interface is used to copy the three channels of the three-channel texture image, and the second interface is used to copy the other two channels to be transmitted. For example, one of these two channels is used to copy the binary information of the decoded occupancy map DOM, and the other channel is used to copy the pixel values of the monochrome geometry image DGI, or one of these two channels is used to copy the binary information of the decoded occupancy map DOM and the other channel is used to copy the block-to-patch index information. However, the channels of the second interface are not used, causing a bandwidth problem.
[0184] Therefore, from the perspective of memory occupancy, copying such three-channel images representing such image structured data is inefficient and requires a large data path to transmit the image channels required for reconstructing the point cloud.
[0185] A straightforward solution to optimize memory transfer might be to unfold and copy the decoded image structured data into a 1D array. However, this solution destroys the initial shaping / formatting of the image structured data, thus losing the relevant 2D positioning information. This can be a problem because the V-PCC point cloud reconstruction process requires such 2D positioning information, e.g., the 2D positioning information of the pixels of the occupancy map DOM or the 2D positioning information of the decoded geometry image DGI or the 2D positioning information of the monochromatic image representing block-to-patch index information.
[0186] According to at least one of the present principles, transmitting a three-channel texture image and at least two image structured data for reconstructing the geometry of the point cloud between two processing units or memories includes copying the three channels of the three-channel texture image to three channels of a four-channel image and storing the combined information obtained by combining the at least two image structured data into the fourth channel of the four-channel image. Then the four-channel image is transmitted.
[0187] The transfer bandwidth between memories of the image structured data is thus optimized because, compared to typically transmitting a three-channel image for each image structured data, the image structured data only transmits (copies) a four-channel image while preserving the 2D structure of the transmitted data.
[0188] Figure 7 A flowchart showing a method of transmitting image structured data between processing units / memories according to at least one of the present embodiments is shown.
[0189] In step 71, the three channels of the texture image DGI are copied to three channels of a four-channel image.
[0190] In step 72, the combined information CI is stored in the fourth channel of the four-channel image. The combined information CI is obtained by combining the at least two image structured data.
[0191] In step 73, the four-channel image is transmitted.
[0192] According to an embodiment, the fourth channel is an alpha channel. Generally, the alpha channel is in the RGBA image format.
[0193] According to an embodiment, one image structured data is a monochromatic geometric shape image DGI, and the other image structured data is an occupancy map DOM. Then, the pixel value A(p) of the fourth channel is the product of the deviation value DGI(p) of the pixels of the monochromatic geometric shape image DGI and the value of the co-located pixel DOM(p) in the occupancy map DOM.
[0194] A(p) = DOM(p) x (DGI(p) + 1)
[0195] Therefore, the block occupancy information is saved and encoded / decoded as non-zero values (even if DGI(p) is a null value).
[0196] Conversely, the pixel values of the occupancy map DOM and the geometric shape image DGI can be retrieved as follows:
[0197] If A(p) = 0, then DOM(p) = 0 (unoccupied TxT block)
[0198] Otherwise, DOM(p) = 1 and DGI(p) = A(p) - 1
[0199] It may be noted that there is a risk of overflow in the value of the occupancy map DOM.
[0200] In a variant, a clip may be added to solve this problem.
[0201] According to an embodiment, one image structured data is a monochromatic geometric shape image DGI, and the other image structured data is an occupancy map DOM. The pixel value A(p) of the fourth channel A is the product of the value DGI(p) of the pixels of the geometric shape image DGI and the value of the co-located pixel DOM(p) in the occupancy map DOM:
[0202] A(p) = DOM(p) x DGI(p)
[0203] Overflow is not considered here, and when A(p) is not a null value and DOM(p) is a binary value, the pixel value of the geometric shape image DGI can be directly obtained from A(p).
[0204] It may be noted that the pixel value of DGI(p) must be strictly positive.
[0205] According to an embodiment, one image structured data is a monochromatic geometric shape image DGI, another image structured data is the occupancy map DOM and another image structured data is the block-to-patch index information (monochromatic image), and wherein the range of the fourth channel is divided into sub-intervals, each sub-interval i is associated with a patch index i when the pixel belongs to a block associated with the patch index i, and the pixel value of the monochromatic geometric shape image DGI is stored in the sub-interval i.
[0206] Figure 8 Shows a non - limiting example for combining image structured data together according to an embodiment.
[0207] In this example, the number of patch indices P is equal to 4, from 0 to 3. The range R of the fourth channel, typically 256, is divided into 4 sub - intervals. When the pixel belongs to the block associated with patch 0, the pixel value of the monochromatic geometric image DGI is represented by a value belonging to (0; M - 1) (usually M = 256 / 4 = 64), when the pixel belongs to the block associated with patch 1, the pixel value of the monochromatic geometric image DGI is represented by a value belonging to (M - 1; 2M - 1), when the pixel belongs to the block associated with patch 2, the pixel value of the monochromatic geometric image DGI is represented by a value belonging to (2M - 1; 3M - 1), and when the pixel belongs to the block associated with patch 3, the pixel value of the monochromatic geometric image DGI is represented by a value belonging to (3M - 1; R - 1).
[0208] Thus, when the pixel value of the occupancy map DOM is equal to 0, the block is unoccupied, otherwise the block is occupied. When a block is occupied, if the value A(p) of the co - located pixel of the fourth channel belongs to (M - 1; 2M - 1), then the pixel value of the monochromatic image DTI is equal to A(p)-(M - 1). This is the depth value of pixel p of the block of patch 1. The 2D localization of the pixel is given by the 2D localization of pixel p in the four - channel image.
[0209] According to an embodiment of the method, when it is necessary to transmit multiple monochromatic geometric images DGI and multiple three - channel texture images DTI, typically 2. Then, in step 70, the first, second, and third channels of the two texture images DTI are respectively packed into the first, second, and third channels of the four - channel image, and the combined information CI is stored in the fourth channel of the four - channel image.
[0210] The size of the four - channel image is equal to the product of the size of the three - channel texture image DGI and a factor (usually 2) depending on the number of three - channel texture images DGI.
[0211] The combined information CI can be obtained by packing the monochromatic geometric image DGI into a monochromatic image and combining the information represented by the monochromatic image with other image structured data (monochromatic images) as described above. Note that the size of the image structured data, such as the size of the monochromatic image representing the occupancy map DOM, may be different from the size of the four - channel image. Then, the image structured data of the three - channel texture image DGI is reused for another one. For example, the information of the occupancy map DOM associated with one three - channel texture image DGI can be reused for another (e.g., the second) three - channel texture image DGI.
[0212] According to an embodiment, packing a texture image into a four-channel image includes duplicating the texture image side by side.
[0213] According to an embodiment, packing a texture image into a four-channel image includes alternately interleaving / staggering the information represented by the texture image in a pattern.
[0214] For example, when the information represented by two texture images DTI0 and DTI1 is interleaved, the following patterns can be used:
[0215] The information represented by the second texture image.
[0216] DTI0(p), DTI1(p), DTI0(p + 1), DTI1(p + 1), etc. or a checkerboard pattern DTI0(p), DTI0(p + 1), DTI1(p), DTI1(p + 1),... where p is a pixel of the texture image.
[0217] This embodiment reduces the memory cache access requirements.
[0218] In Figures 1-8 , various methods are described herein, and each of the methods includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined.
[0219] Some examples are described with respect to block diagrams and operational flowcharts. Each block represents a circuit element, module, or section of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that in other embodiments, the functions noted in the blocks may not occur in the order indicated. For example, based on the functions involved, two consecutive blocks shown may actually be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order.
[0220] The embodiments and aspects described herein can be implemented in, for example, a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single form of embodiment (e.g., only as a method), the embodiments of the features discussed can be implemented in other forms (e.g., an apparatus or a computer program).
[0221] The method can be implemented, for example, in a processor, which generally refers to a processing device and includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.
[0222] In addition, these methods can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the embodiments) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product, which is embodied in one or more computer-readable media and has computer-readable program code executable by a computer embodied thereon. Considering the inherent ability to store information therein and the inherent ability to provide information retrieval therefrom, the computer-readable storage medium used herein can be considered a non-transitory storage medium. The computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be understood that the following, although providing more specific examples of computer-readable storage media to which the present embodiment can be applied, is merely illustrative and not an exhaustive list as would be readily understood by a person of ordinary skill in the art: portable computer floppy disks; hard disks; read-only memory (ROM); erasable programmable read-only memory (EPROM or flash memory); portable compact disc read-only memory (CD-ROM); optical storage devices; magnetic storage devices; or any suitable combination of the foregoing.
[0223] The instructions can form an application program tangibly embodied on a processor-readable medium.
[0224] For example, the instructions can be in hardware, firmware, software, or a combination. For example, the instructions can be found in an operating system, a separate application program, or a combination of both. Thus, a processor can be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (such as a storage device) having instructions for executing the process. In addition, in addition to or instead of the instructions, the processor-readable medium can store data values generated by the embodiments.
[0225] The apparatus can be implemented in, for example, appropriate hardware, software, and firmware. Examples of such apparatus include personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, head-mounted display devices (HMDs, see-through glasses), projectors, "visions" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output of the video decoder, pre-processors that provide input to the video encoder, web servers, set-top boxes, and any other devices for processing point clouds, videos, or images or other communication devices. It should be clear that the device can be mobile and can even be installed in a moving vehicle.
[0226] The computer software can be implemented by the processor 6010, or by hardware, or by a combination of hardware and software. As a non-limiting example, this embodiment can also be implemented by one or more integrated circuits. As a non-limiting example, the memory 6020 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, the processor 6010 can be of any type suitable for the technical environment and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0227] It will be apparent to those of ordinary skill in the art that the implementations can generate various signals that are formatted to carry information such as can be stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal can be formatted to carry a bitstream of the described embodiment. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.
[0228] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" may also be intended to include the plural forms. It will be further understood that when used in this specification, the terms "includes / comprises" and / or "including / comprising" may specify the stated features, integers, steps, operations, elements, and / or components, but do not preclude the existence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Additionally, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to other elements, no intervening elements are present.
[0229] It should be understood that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the symbols / terms " / " and "and / or" and "at least one of them" may be intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.
[0230] It should be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the teachings of the present application. There is no implied ordering between the first element and the second element.
[0231] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variants are often used to convey that specific features, structures, characteristics, etc. described in connection with the embodiment / implementation are included in at least one embodiment / implementation. Thus, the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation", and the occurrence of any other variants, that appear in various places in the present application do not necessarily all refer to the same embodiment.
[0232] Similarly, references in this document to "according to one embodiment / example / implementation" or "in one embodiment / example / implementation" and other variations thereof are often used to convey that a particular feature, structure, characteristic, etc. described in connection with the embodiment / example / implementation is included in at least one embodiment / example / implementation. Thus, the phrases "according to one embodiment / example / implementation" or "in one embodiment / example / implementation" that appear in various places in the specification do not necessarily all refer to the same embodiment / example / implementation, nor do separate or alternative embodiments / examples / implementations have to be mutually exclusive of other embodiments / examples / implementations.
[0233] Reference numerals that appear in the claims are for illustrative purposes only and should not limit the scope of the claims. Although not explicitly described, the embodiments / examples and variations thereof can be employed in any combination or sub-combination.
[0234] When a figure is presented in the form of a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented in block diagram form, it should be understood that it also provides a flowchart of the corresponding method / process.
[0235] Although some figures include arrows on communication paths to show the main direction of communication, it should be understood that communication can occur in the direction opposite to the depicted arrows.
[0236] Various embodiments relate to decoding. As used in this application, "decoding" can cover all or part of the execution process, e.g., when a received point cloud frame (which may include a bitstream encoding one or more point cloud frames) is processed to produce a final output suitable for display or further processing in the reconstructed point cloud domain. In various embodiments, such processes include one or more processes typically performed by an image-based decoder. In various embodiments, such processes also or alternatively include processes performed by the decoders of the various embodiments described in this application.
[0237] As a further example, in one embodiment, "decoding" can refer only to entropy decoding, in another embodiment, "decoding" can refer only to differential decoding, and in another embodiment, "decoding" can refer to a combination of entropy decoding and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, and it is believed to be well understood by those skilled in the art.
[0238] Various embodiments relate to encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application may cover, for example, all or part of the process performed on an input point cloud frame to generate an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an image-based decoder. In various embodiments, such processes also or alternatively include, for example, processes performed by the encoders of the various embodiments described in this application.
[0239] As a further example, in one embodiment, "encoding" may refer only to entropy encoding, in another embodiment, "encoding" may refer only to differential encoding, and in another embodiment, "encoding" may refer to a combination of differential encoding and entropy encoding. Based on the context of the specific description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, and it is believed to be well understood by those skilled in the art.
[0240] Various embodiments relate to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is typically considered, and limitations in computational complexity are also typically considered. Rate-distortion optimization can generally be formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different ways to solve the rate-distortion optimization problem. For example, these methods can be based on extensive testing of all encoding options, including all considered modes or codec parameter values, and a complete evaluation of their encoding and decoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save encoding complexity, especially in cases where an approximate distortion is calculated based on the predicted or prediction residual signal rather than the reconstructed signal. A hybrid of these two methods can also be used, for example, using approximate distortion only for some possible encoding options and complete distortion for other encoding options. Other methods only evaluate a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform the optimization, but the optimization does not necessarily involve a complete evaluation of the encoding and decoding costs and the associated distortion.
[0241] Additionally, this application may relate to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0242] Furthermore, this application may relate to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0243] Additionally, the present application may relate to "receiving" various information. Receiving, like "accessing", is a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Further, "receiving" typically involves, in one way or another, during an operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information.
[0244] In addition, as used herein, the word "signal" particularly refers to indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals a particular XXX. Thus, in one embodiment, the same parameters can be used on both the encoder side and the decoder side. So, for example, the encoder can send (explicitly signal) a particular parameter to the decoder such that the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter along with other parameters, signaling can be used without sending (implicitly signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmitting any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing related to the verb form of the word "signal", the word "signal" can also be used as a noun herein.
[0245] Numerous embodiments have been described. However, it should be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to yield other embodiments. Additionally, one of ordinary skill in the art will understand that other structures and processes can replace those disclosed, and the resulting embodiments will perform at least substantially the same functions in at least substantially the same way to achieve at least substantially the same results as the disclosed embodiments. Accordingly, the present application contemplates these and other embodiments.
Claims
1. A method for transmitting information, the method comprising: Copy information from three channels of a first three-channel texture image into a first channel, a second channel, and a third channel of a four-channel image, where the first three-channel texture image represents the texture of a point cloud; Store combined information in a fourth channel of the four-channel image, where the combined information includes a combination of at least two image structured data for reconstructing the geometry of the point cloud; And Transmit the four-channel image.
2. The method according to claim 1, wherein: One of the two image structured data includes a monochromatic geometry image representing the geometry of the point cloud; And The other of the two image structured data includes an occupancy map, where the pixel values of the occupancy map indicate whether blocks of the first three-channel texture image and the monochromatic geometry image include at least one orthographic projection point of the point cloud; Where the pixel values of the fourth channel are the product of the deviation values of the pixels of the monochromatic geometry image and the values of the co-located pixels in the occupancy map.
3. The method according to claim 1, wherein One of the two image structured data includes a monochromatic geometry image representing the geometry of the point cloud; And The other of the two image structured data includes an occupancy map, where the pixel values of the occupancy map indicate whether blocks of the first three-channel texture image and the monochromatic geometry image include at least one orthographic projection point of the point cloud; And Where the pixel values of the fourth channel are the product of the pixel values of the monochromatic geometry image and the values of the co-located pixels in the occupancy map.
4. The method according to claim 1, further comprising packing the first channel, the second channel, and the third channel of the second three-channel texture image into the first channel, the second channel, and the third channel of the four-channel image respectively.
5. The method according to claim 4, wherein Packing the first channel, the second channel, and the third channel of the second three-channel texture image into the four-channel image includes: copying the first channel, the second channel, and the third channel of the second three-channel texture image side by side with the first channel, the second channel, and the third channel of the first three-channel texture image.
6. The method according to claim 4, wherein Packing the first channel, the second channel, and the third channel of the second three-channel texture image into the four-channel image includes: alternately interleaving in a pattern the information represented by the first channel, the second channel, and the third channel of the first three-channel texture image and the first channel, the second channel, and the third channel of the second three-channel texture image.
7. The method according to claim 1, wherein The fourth channel of the four-channel image includes an alpha channel in RGBA image format.
8. The method according to claim 1, wherein the point cloud is reconstructed using the four-channel image.
9. An apparatus for transmitting information, the apparatus comprising one or more processors configured to: Copy information from three channels of a first three-channel texture image into the first channel, the second channel, and the third channel of a four-channel image, wherein the first three-channel texture image represents the texture of a point cloud; Store combined information in the fourth channel of the four-channel image, wherein the combined information includes a combination of at least two image structured data for reconstructing the geometry of the point cloud; and Transmit the four-channel image.
10. The apparatus according to claim 9, wherein: One of the two image structured data includes a monochromatic geometry image representing the geometry of the point cloud; And The other of the two image structured data includes an occupancy map, where the pixel values of the occupancy map indicate whether blocks of the first three-channel texture image and the monochromatic geometry image include at least one orthographic projection point of the point cloud; And Where the pixel values of the fourth channel are the product of the deviation values of the pixels of the monochromatic geometry image and the values of the co-located pixels in the occupancy map.
11. The apparatus according to claim 9, wherein: One of the two image structured data includes a monochromatic geometry image representing the geometry of the point cloud; And The other of the two image structured data includes an occupancy map, where the pixel value of the occupancy map indicates whether the block of the first three-channel texture image and the monochromatic geometry image includes at least one orthographic projection point of the point cloud; Wherein the pixel value of the fourth channel is the product of the pixel value of the monochromatic geometry image and the value of the co-located pixel in the occupancy map.
12. The apparatus according to claim 9, wherein the one or more processors are further configured to pack a first channel, a second channel, and a third channel of the second three-channel texture image into a first channel, a second channel, and a third channel of the four-channel image, respectively.
13. The apparatus according to claim 12, wherein, Packing the first channel, second channel, and third channel of the second three-channel texture image into the four-channel image includes: copying the first channel, second channel, and third channel of the second three-channel texture image side by side with the first channel, second channel, and third channel of the first three-channel texture image.
14. The apparatus according to claim 12, wherein, Packing the first channel, second channel, and third channel of the second three-channel texture image into the four-channel image includes: alternately interleaving the information represented by the first channel, second channel, and third channel of the second three-channel texture image according to a pattern.
15. The apparatus according to claim 9, wherein, The fourth channel of the four-channel image includes an alpha channel in RGBA image format.
16. The apparatus according to claim 9, wherein the point cloud is reconstructed using the four-channel image.
17. A non-transitory computer-readable medium comprising instructions for causing one or more processors to perform the following operations: Copy information from three channels of a first three-channel texture image into a first channel, a second channel, and a third channel of a four-channel image, wherein the first three-channel texture image represents the texture of a point cloud; Store combined information in a fourth channel of the four-channel image, wherein the combined information includes a combination of at least two image structured data for reconstructing the geometry of the point cloud; and Transmit the four-channel image.
Citation Information
Patent Citations
Mapping space using multi-directional camera
CN108369743A
A method for semantic segmentation of scene point cloud
CN109410307A