Intermediate points for processing a point cloud

By decomposing point cloud data into 2D texture patches and encoding them using a video codec, the problem of efficient compression and transmission of dynamic point cloud data at a limited bit rate is solved, improving the reconstruction quality and immersive experience of point cloud data.

CN113632486BActive Publication Date: 2025-12-16INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080022209.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-20
Filing Date
2020-01-27
Publication Date
2025-12-16
Estimated Expiration
2040-04-26

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently compress and transmit dynamic point cloud data while maintaining acceptable quality of experience, especially under limited bitrate conditions.

Method used

A two-layer point cloud coding structure is adopted to decompose point cloud data into 2D sample texture patches. Syntax elements representing texture patches, including 2D position, size, quantity and index, are sent by signal transmission. This is combined with existing video codecs for encoding and decoding.

Benefits of technology

It achieves efficient compression and transmission of dynamic point cloud data at a limited bit rate, improving the reconstruction quality and immersive experience of point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113632486B_ABST
    Figure CN113632486B_ABST
Patent Text Reader

Abstract

At least one embodiment is directed to signaling at least one texture patch representing a texture value of at least one intermediate 3D sample, the texture patch being a set of 2D samples representing texture values of 3D samples of a point cloud orthogonally projected onto a projection plane along a projection line, and the at least one intermediate 3D sample being a 3D sample of the point cloud having a depth value greater than a nearer 3D sample of the point cloud and less than a farther 3D sample of the point cloud, the at least one intermediate 3D sample and the nearer and farther 3D samples being projected along the same projection line.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment of the present invention relates primarily to the processing of point clouds. Background Technology

[0002] This section aims to introduce the reader to various aspects of the technology that may be related to at least one aspect of the embodiments of the invention described below and / or claimed. It is believed that this discussion will help provide the reader with background information to better understand the various aspects of at least one embodiment.

[0003] Point clouds can be used for various purposes, such as cultural heritage / architecture, where objects like statues or buildings are 3D scanned to share their spatial configuration without sending or accessing them. Furthermore, this is a way to preserve knowledge of an object in the event of potential damage (e.g., temple collapse due to an earthquake). Such point clouds are typically static, colored, and massive.

[0004] Another use case is in topographic mapping and cartography, where 3D representation allows maps to be not limited to flat surfaces and can include undulations. Google Maps is now a good example of 3D maps, but it uses a grid instead of point clouds. However, point clouds can be a suitable data format for 3D maps, and such point clouds are typically static, colored, and large.

[0005] The automotive industry and autonomous vehicles are also areas where point clouds can be used. Autonomous vehicles should be able to "detect" their environment to make sound driving decisions based on the reality of their immediate neighbors. Typical sensors like LiDAR (Light Detection and Ranging) generate dynamic point clouds used by decision engines. These point clouds are not intended for human viewing, and they are typically small, not necessarily colored, and dynamic, captured at a high frequency. These point clouds can possess other properties, such as reflectivity provided by LiDAR, as this property provides good information about the material of the sensed object and can aid in decision-making.

[0006] Virtual reality and immersive worlds have recently become hot topics and are widely predicted to be the future of 2D flat video. The basic idea is to immerse the viewer in an environment surrounding them, contrasting with standard TV where the viewer can only see a virtual world in front of them. Several levels of immersion exist based on the viewer's degree of freedom within the environment. Point clouds are a promising format candidate for distributing virtual reality (VR) worlds.

[0007] In many applications, it is important to be able to distribute dynamic point clouds to end users (or store them on servers) while maintaining an acceptable (or preferably very good) quality of experience, consuming only a reasonable amount of bitrate (or storage space used for storing the application). Efficient compression of these dynamic point clouds is crucial for making multi-immersive world distribution chains practical.

[0008] In view of the foregoing, at least one embodiment has been designed. Summary of the Invention

[0009] The following is a simplified overview of at least one embodiment of the invention to provide a basic understanding of some aspects of the invention. This overview is not a comprehensive summary of the embodiments. It is not intended to identify key or essential elements of the embodiments. The following description presents only some aspects of at least one embodiment of the invention in a simplified form, as a prelude to the more detailed descriptions provided elsewhere in this document.

[0010] According to a general aspect of at least one embodiment, a method is provided comprising: signaling at least one texture patch representing texture values ​​of at least one intermediate 3D sample, the texture patch being a set of 2D samples representing texture values ​​of 3D samples of a point cloud orthogonally projected onto a projection plane along a projection line, and the at least one intermediate 3D sample being a 3D sample in the point cloud having a depth value greater than that of a closer 3D sample of the point cloud and less than that of a farther 3D sample of the point cloud, the at least one intermediate 3D sample and the closer and farther 3D samples being projected along the same projection line.

[0011] According to an embodiment, signaling a texture patch representing the texture value of at least one intermediate 3D sample includes:

[0012] - Add at least one syntax element to the bitstream, the at least one syntax element representing the 2D position of the texture patch defined in the 2D grid and the size of the texture patch;

[0013] - Send the bit stream; and

[0014] - Retrieve the at least one syntax element from the bitstream, and extract the 2D position and size of the texture patch defined in the 2D grid from the at least one retrieved syntax element.

[0015] According to an embodiment, the at least one syntax element can also signal the number of 2D samples of the patch of the 2D mesh.

[0016] According to an embodiment, the at least one syntax element can also signal the number of texture patches in the 2D mesh and the offset for determining the starting position of the texture patches.

[0017] According to an embodiment, the at least one syntax element can also signal an index of a texture patch.

[0018] According to an embodiment, the syntax elements can also be signaled. The at least one syntax element added to the bitstream is signaled at different levels throughout the syntax representing the point cloud frame.

[0019] According to another general aspect of at least one embodiment, a method is provided, comprising: analyzing an orthogonal projection of a point cloud frame onto a projection plane to derive texture values ​​of at least one intermediate 3D sample; mapping the at least one texture value to at least one texture patch; packing the at least one texture patch into a texture image; and transmitting the at least one texture patch as a signal in a bitstream according to the method described above.

[0020] According to another general aspect of at least one embodiment, a method is provided comprising: deriving at least one texture patch from at least one syntax element signaled in a bitstream according to the above-described signaling method; deriving at least one texture value from the at least one texture patch; and assigning the at least one texture value to at least one intermediate 3D sample.

[0021] One or more of the at least one embodiment also provide a device, a computer program product, a non-transient computer-readable medium, and a signal.

[0022] The specific properties of at least one of the embodiments of the present invention, as well as other objects, advantages, features, and uses of the at least one of the embodiments of the present invention, will become apparent from the following illustrative description taken in conjunction with the accompanying drawings. Attached Figure Description

[0023] The accompanying drawings illustrate examples of several embodiments.

[0024] Figure 1 A schematic block diagram illustrating an example of a two-layer point cloud coding structure according to at least one embodiment of the present invention is shown.

[0025] Figure 2 A schematic block diagram illustrating an example of a two-layer point cloud decoding structure according to at least one embodiment of the present invention is shown;

[0026] Figure 3 A schematic block diagram of an example of an image-based point cloud encoder according to at least one embodiment of the present invention is shown;

[0027] Figure 3a An example of a canvas including two patches and their 2D bounding boxes is shown;

[0028] Figure 3b An example of two intermediate 3D samples located between two 3D samples along the projection line is shown;

[0029] Figure 4 A schematic block diagram of an example of an image-based point cloud decoder according to at least one embodiment of the present invention is shown;

[0030] Figure 5 An example of a syntax for representing a bitstream of a base layer BL according to at least one current embodiment is illustrated schematically;

[0031] Figure 6 A schematic block diagram illustrating an example of a system in which various aspects and embodiments are implemented is shown;

[0032] Figure 7 An example of a method for transmitting at least one EOM texture patch by signaling, according to at least one embodiment of the present invention, is shown;

[0033] Figure 8a -d illustrates an example of a syntax element according to at least one embodiment of step 710;

[0034] Figure 9a -c illustrates an example canvas according to at least one embodiment of step 710;

[0035] Figure 10 An example of a syntax element according to the embodiment of step 710 is shown;

[0036] Figure 11 An example of the syntax element of the variant according to the embodiment described in step 710 is shown;

[0037] Figure 12 Examples of syntactic elements of a variant of the embodiment according to step 710 are shown;

[0038] Figure 13 An example of a table defining patch patterns according to at least one current embodiment is shown;

[0039] Figure 14 Examples of syntactic elements of a variant of the embodiment according to step 710 are shown;

[0040] Figure 15 An example of a syntax element according to the embodiment of step 710 is shown;

[0041] Figure 16A block diagram of a method for encoding texture values ​​of intermediate 3D samples according to at least one embodiment of this invention is shown;

[0042] Figure 17 A block diagram of a method for decoding texture values ​​of intermediate 3D samples according to at least one embodiment of this invention is shown; and

[0043] Figure 18 An example is shown following the raster scan order of the blocks. Detailed Implementation

[0044] At least one embodiment of the invention is described more fully below with reference to the accompanying drawings, which illustrate examples of at least one embodiment. However, embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Therefore, it should be understood that the embodiments are not intended to be limited to the specific forms disclosed. Rather, this disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this application.

[0045] When the accompanying drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding devices. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding methods / processes.

[0046] Similar or identical elements in the figure are represented by the same reference numerals.

[0047] Some diagrams illustrate syntax tables widely used in V-PCC to define the structure of bitstreams conforming to V-PCC. In these syntax tables, the term '…' indicates an unchanged portion of the syntax relative to the original definition given in V-PCC and is removed from the diagram for ease of reading. Bold items in the diagram indicate that the value of that item was obtained by parsing the bitstream. The right column of the syntax table indicates the number of bits used to encode the data of the syntax elements. For example, u(4) indicates 4 bits are used to encode the data, u(8) indicates 8 bits, and ae(v) indicates that the integer value v is arithmetically encoded using, for example, CABAC (Context-Adaptive-Binary-Arithmetic Coding).

[0048] The aspects described and anticipated below can be realized in many different forms. Figure 1-18 Some embodiments have been provided, but other embodiments are conceivable, and Figure 1-18 The discussion does not limit the breadth of implementation.

[0049] At least one aspect generally involves point cloud encoding and decoding, and at least one other aspect generally involves transmitting the generated or encoded bit stream.

[0050] More precisely, the various methods and other aspects described in this paper can be used to modify modules, for example, Figure 3 Modules 3100, 3200, 3400, and 3700 can be modified for implementation. Figure 16 The method. Modules 4400 and 4600 can also be modified to implement this. Figure 17 The method.

[0051] Furthermore, aspects of the present invention are not limited to MPEG standards (e.g., MPEG-I Part 5, which relates to point cloud compression), but can be applied to, for example, other standards and recommendations (whether pre-existing or developed in the future) and any extensions to such standards and recommendations (including MPEG-I Part 5). Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.

[0052] In the following text, image data refers to data, such as one or more 2D sample arrays of a specific image / video format. A specific image / video format may specify information relating to the pixel values ​​of an image (or video). A specific image / video format may also specify information that can be used by a display and / or any other device, for example, to visualize and / or decode the image (or video). An image typically includes a first component, which is shaped like a first 2D sample array and typically represents the image's luminance (or luma). An image may also include second and third components, which are shaped like other 2D sample arrays and typically represent the image's chroma (or chroma). Some embodiments use a set of 2D color sample arrays to represent the same information, such as the traditional three-color RGB representation.

[0053] In one or more embodiments, pixel values ​​are represented by a vector with C values, where C is the number of components. Each value of the vector is typically represented by a number of bits that define the dynamic range of the pixel value.

[0054] An image patch is a group of pixels that belong to an image. The pixel value of an image patch (or image patch data) refers to the pixel value belonging to that image patch. Image patches can have any shape, although rectangles are common.

[0055] Point clouds can be represented by a dataset of 3D samples in a 3D volumetric space, each 3D sample having unique coordinates and one or more attributes.

[0056] The 3D samples in this dataset can be defined by their spatial location (X, Y, and Z coordinates in 3D space) and may be defined by one or more associated attributes, such as color (represented in, for example, RGB or YUV color spaces), transparency, reflectivity, a two-component normal vector, or any feature representing the characteristics of the sample. For example, a 3D sample can be defined by six components (X, Y, Z, R, G, B) or equivalently (X, Y, Z, y, U, V), where (X, Y, Z) defines the coordinates of a point in 3D space, and (R, G, B) or (y, U, V) defines the color of the 3D sample. Attributes of the same type can appear multiple times. For example, multiple color attributes can provide color information from different viewpoints.

[0057] Point clouds can be static or dynamic, depending on whether the cloud changes over time. Instances of static or dynamic point clouds are typically represented as point cloud frames. It should be noted that in the case of dynamic point clouds, the number of points is usually not constant, but rather changes over time. More generally, a point cloud can be considered dynamic if anything (e.g., the number of points, the location of one or more points, or any property of any point) changes over time.

[0058] As an example, a 2D sample can be defined by six components (u, v, Z, R, G, B) or equivalently (u, v, Z, y, U, V). (u, v) defines the coordinates of the 2D sample in the 2D space of the projection plane. Z is the depth value of the 3D sample projected onto the projection plane. (R, G, B) or (y, U, V) defines the color of the 3D sample.

[0059] Figure 1 A schematic block diagram of an example of a two-layer point cloud coding structure 1000 according to at least one embodiment of the present invention is shown.

[0060] The two-layer point cloud coding structure 1000 can provide a bitstream B representing an input point cloud frame IPCF. Possibly, the input point cloud frame IPCF represents a frame of a dynamic point cloud. This frame of the dynamic point cloud can then be encoded independently of another frame by the two-layer point cloud coding structure 1000.

[0061] Essentially, the two-layer point cloud coding structure 1000 provides the ability to structure a bitstream B into a base layer BL and an enhancement layer EL. The base layer BL can provide a lossy representation of the input point cloud frame IPCF, and the enhancement layer EL can provide a higher quality (potentially lossless) representation by encoding isolated points not represented by the base layer BL.

[0062] The basic layer BL can be composed of, for example... Figure 3The image-based encoder 3000 shown provides this. The image-based encoder 3000 can provide a geometric structure / texture image representing the geometry / attributes of 3D samples of the input point cloud frame IPCF. It allows discarding isolated 3D samples. The base layer BL can be provided by, for example... Figure 4 The image-based decoder 4000 shown can decode and provide intermediate reconstructed point cloud frames (IRPCF).

[0063] Then, return to Figure 1 The two-layer point cloud encoding 1000 uses a comparator COMP to compare 3D samples from the input point cloud frame IPCF with 3D samples from the intermediate reconstructed point cloud frame IRPCF to detect / locate missing / isolated 3D samples. Next, the encoder ENC encodes the missing 3D samples and provides an enhancement layer EL. Finally, the base layer BL and the enhancement layer EL are multiplexed together by a multiplexer MUX to produce the bitstream B.

[0064] According to an embodiment, the encoder ENC may include a detector that can detect a 3D reference sample R of an intermediate reconstructed point cloud frame IRPCF and associate it with a lost 3D sample M.

[0065] For example, a 3D reference sample R associated with a missing 3D sample M can be the nearest neighbor of M based on a given metric.

[0066] According to an embodiment, the encoder ENC can then encode the spatial location and attributes of the lost 3D sample M into a difference determined based on the spatial location and attributes of the 3D reference sample R.

[0067] In one variant, these differences can be encoded individually.

[0068] For example, for a missing 3D sample M, which has spatial coordinates x(M), y(M), and z(M), the position difference of x coordinate Dx(M), the position difference of y coordinate Dy(M), the position difference of z coordinate Dz(M), the difference of R attribute components Dr(M), the difference of G attribute components Dg(M), and the difference of B attribute components Db(M) can be calculated as follows:

[0069] Dx(M)=x(M)-x(R),

[0070] Where x(M) is Figure 3 The provided geometric image contains 3D samples M and their corresponding x-coordinates R.

[0071] Dy(M)=y(M)-y(R)

[0072] Where y(M) is Figure 3The y-coordinates of the 3D sample M and the corresponding R in the provided geometric image.

[0073] Dz(M)=z(M)-z(R)

[0074] Where z(M) is Figure 3 The provided geometric image contains 3D samples M and their corresponding z-coordinates R.

[0075] Dr(M) = R(M) - R(R).

[0076] Where R(M) and R(R) are the r-color components of the color attributes of the 3D sample M and the corresponding R, respectively.

[0077] Dg(M)=G(M)-G(R)

[0078] Where G(M) and G(R) are the g color components of the color attributes of the 3D sample M and the corresponding R, respectively.

[0079] Db(M) = B(M) - B(R).

[0080] Where B(M) and B(R) are the b color components of the color attributes of the 3D sample M and the corresponding R, respectively.

[0081] Figure 2 A schematic block diagram of an example of a two-layer point cloud decoding structure 2000 according to at least one embodiment of the present invention is shown.

[0082] The behavior of the two-layer point cloud decoding structure 2000 depends on its capabilities.

[0083] The two-layer point cloud decoding architecture 2000, with limited capabilities, can access only the basic layer BL from the bitstream B using a demultiplexer DMUX, and then can be further decoded by means of... Figure 4 The point cloud decoder 4000 shown decodes the base layer BL to provide a faithful (but lossy) version of the input point cloud frame IPCF, IRPCF.

[0084] The fully capable two-layer point cloud decoding architecture 2000 allows access to the basic layer BL and enhancement layer EL from the bitstream B using the demultiplexer DMUX. Figure 4 As shown, the point cloud decoder 4000 can determine the intermediate reconstructed point cloud frame IRPCF based on the base layer BL. The decoder DEC can determine the complementary point cloud frame CPCF from the enhancement layer EL. Then, the combiner COM can combine the intermediate reconstructed point cloud frame IRPCF and the complementary point cloud frame CPCF together to thus provide a higher quality (potentially lossless) representation (reconstruction) of the input point cloud frame IPCF (CRPCF).

[0085] Figure 3 A schematic block diagram of an example of an image-based point cloud encoder 3000 according to at least one embodiment of this invention is shown.

[0086] The image-based point cloud encoder 3000 utilizes existing video codecs to compress the geometric structure and texture (attribute) information of dynamic point clouds. This is achieved by essentially converting the point cloud data into a set of different video sequences.

[0087] In a particular embodiment, existing video codecs can be used to generate and compress two videos, one for capturing geometric information of point cloud data and the other for capturing texture information. An example of an existing video codec is the HEVC master profile encoder / decoder (ITU-T H.265 Telecommunications Standardization Sector, ITU (02 / 2018), H Series: Audiovisual and Multimedia Systems, Audiovisual Services Infrastructure – Coding of Motion Video, Efficient Video Coding, Recommended ITU-T H.265).

[0088] Additional metadata used to interpret the two videos is typically generated and compressed separately. Such additional metadata includes, for example, occupancy graph (OM) and / or supplementary patch information (PI).

[0089] The generated video bitstream and metadata can then be multiplexed together to generate a combined bitstream.

[0090] It should be noted that the metadata typically represents a small amount of overall information. A large amount of information is contained in the video bitstream.

[0091] An example of this point cloud encoding / decoding process is given by the Test Model Category 2 algorithm (also known as V-PCC), which implements the MPEG draft standard defined in ISO / IEC JTC1 / SC29 / WG11MPEG2019 / w18180 (January 2019, Mareshkak).

[0092] In step 3100, module PGM can generate at least one patch by decomposing 3D samples of the dataset representing the input point cloud frame IPCF into 2D samples on the projection plane using a strategy that provides optimal compression.

[0093] A patch can be defined as a set of 2D samples.

[0094] For example, in V-PCC, the normal at each 3D sample is first estimated as described by Hoppe et al. (Hugues Hoppe, Tony DeRose, Tom Duchamp, John McDonald, Werner Stuetzle. Surface reconstruction from unorganized points. ACM SIGGRAPH 1992 Proceedings, 71-78). Next, an initial cluster of the input point cloud frame IPCF is obtained by associating each 3D sample with one of six orientation planes of the 3D bounding box containing the 3D samples of the input point cloud frame IPCF. More precisely, each 3D sample is clustered and associated with the orientation plane that has the closest normal (i.e., maximizes the dot product of the point normal and the face normal). The 3D samples are then projected onto their associated planes. A set of 3D samples that form connected regions in their planes is called a connected component. A connected component is a set of at least one 3D sample with similar normals and the same associated orientation plane. The initial clusters are then refined by iteratively updating the clusters associated with each 3D sample based on the normals of each 3D sample and its nearest neighbors. The final step involves generating a patch from each connected component, which is accomplished by projecting the 3D sample of each connected component onto an orientation plane associated with that connected component. The patch is associated with auxiliary patch information PI, which represents auxiliary patch information defined for each patch to interpret the 2D sample corresponding to the projection of geometric and / or attribute information.

[0095] For example, in V-PCC, the auxiliary patch information PI includes 1) information indicating one of the six orientation planes of the 3D bounding box containing the connected components; 2) information relative to the plane normal; 3) information determining the 3D position of the connected components relative to the patch, which is represented by depth, tangential shift, and double tangential shift; and 4) information such as defining the coordinates (u0, v0, u1, v1) of the 2D bounding box containing the patch in the projection plane.

[0096] In step 3200, the patch packing module (PPM) can map (place) at least one generated patch onto a 2D grid (also called a canvas) in a manner that typically minimizes unused space without any overlap, and can guarantee that each TxT (e.g., 16×16) block of the 2D grid is associated with a unique patch. A given minimum block size TxT of the 2D grid can specify the minimum distance between different patches placed on that 2D grid. The 2D grid resolution can depend on the input point cloud size and its width W and height H, and the block size T can be sent as metadata to the decoder.

[0097] The auxiliary patch information PI may further include information about the association between blocks and patches of the 2D mesh.

[0098] In V-PCC, the auxiliary information PI may include block-to-patch index information (BlockToPatch), which determines the association between block and patch indices of the 2D mesh.

[0099] Figure 3a An example of canvas C is shown, which includes two patches P1 and P2 and their associated 2D bounding boxes B1 and B2. Note that, as Figure 3a As shown, two bounding boxes in canvas C can overlap. The 2D mesh (canvas segmentation) is represented only inside the bounding boxes, but canvas segmentation also occurs outside those bounding boxes. The bounding box associated with the patch can be segmented into TxT blocks, typically T = 16.

[0100] A TxT block containing a 2D sample belonging to a patch can be considered an occupied block. Each occupied block of the canvas is represented by a specific pixel value (e.g., 1) in the Occupation Map OM, while each unoccupied block of the canvas is represented by another specific value (e.g., 0). The pixel value of the Occupation Map OM can then indicate whether a TxT block of the canvas is occupied, i.e., whether it contains a 2D sample belonging to a patch.

[0101] exist Figure 3a In the image, occupied blocks are represented by white blocks, while light gray blocks represent unoccupied blocks. The image generation process (steps 3300 and 3400) utilizes at least one generated patch mapping to a 2D mesh calculated during step 3200 to store the geometry and texture of the input point cloud frame IPCF as an image.

[0102] In step 3300, the geometric structure image generator GIG can generate at least one geometric structure image GI based on the input point cloud box IPCF, occupancy map OM, and auxiliary patch information PI. The geometric structure image generator GIG can use the occupancy map information to detect (locate) occupancy blocks and non-empty pixels in the geometric structure image GI.

[0103] The geometric structure image (GI) can represent the geometry of the input point cloud frame (IPCF) and can be, for example, a monochrome image of W x H pixels represented in YUV420-8 bit format.

[0104] To better handle the situation where multiple 3D samples are projected (mapped) onto the same 2D sample (along the same projection direction (line)) onto the projection plane, multiple images called layers can be generated. Therefore, different depth values ​​D1, ..., Dn can be associated with the 2D samples of the patch, and multiple geometric images can then be generated.

[0105] In V-PCC, 2D samples of the patch are projected onto two layers. The first layer, also called the near layer, can store, for example, a depth value D0 associated with a 2D sample having a smaller depth. The second layer, called the far layer, can store, for example, a depth value D1 associated with a 2D sample having a larger depth. Alternatively, the second layer can store the difference between depth values ​​D1 and D0. For example, the information stored by the second depth image can be within the interval [0, Δ] corresponding to depth values ​​in the range [D0, D0+Δ], where Δ is a user-defined parameter describing the surface thickness.

[0106] In this way, the second layer can contain significant contour-like high-frequency features. Therefore, it is clear that the second depth image may be difficult to encode using a conventional video encoder, and consequently, the depth values ​​may be poorly reconstructed from the decoded second depth image, resulting in poor quality geometry in the reconstructed point cloud frame.

[0107] According to an embodiment, the geometry image generation module GIG can encode (derive) depth values ​​associated with 2D samples of the first and second layers by using auxiliary patch information PI.

[0108] In V-PCC, the position of a 3D sample with a corresponding connected component in the patch can be represented by depth δ(u,v), tangential shift s(u,v), and double tangential shift r(u,v) as follows:

[0109] δ(u,v)=δ0+g(u,v)

[0110] s(u,v)=s0–u0+u

[0111] r(u,v)=r0–v0+v

[0112] Where g(u,v) is the luminance component of the geometric structure image, (u,v) is the pixel associated with the 3D sample on the projection plane, (δ0,s0,r0) is the 3D position of the corresponding patch of the connected component to which the 3D sample belongs, and (u0,v0,u1,v1) are the coordinates in the projection plane, which define a 2D bounding box containing the projection of the patch associated with the connected component.

[0113] Therefore, the geometry image generation module GIG can encode (derive) the depth values ​​associated with 2D samples of a layer (the first layer, the second layer, or both) as a brightness component g(u,v), given by: g(u,v) = δ(u,v) - δ0. It should be noted that this relationship can be used to reconstruct the 3D sample positions (δ0, s0, r0) from the reconstructed geometry image g(u,v) with accompanying auxiliary patch information PI.

[0114] According to an embodiment, the projection mode can be used to indicate whether the first geometric structure image GI0 can store the depth value of the 2D sample of the first layer or the second layer, and whether the second geometric structure image GI1 can store the depth value associated with the 2D sample of the second layer or the first layer.

[0115] For example, when the projection mode is equal to 0, the first geometric structure image GI0 can store the depth values ​​of the 2D samples of the first layer, and the second geometric structure image GI1 can store the depth values ​​associated with the 2D samples of the second layer. Conversely, when the projection mode is equal to 1, the first geometric structure image GI0 can store the depth values ​​of the 2D samples of the second layer, and the second geometric structure image GI1 can store the depth values ​​associated with the 2D samples of the first layer.

[0116] According to an embodiment, the frame projection mode can be used to indicate whether a fixed projection mode is used for all patches, or whether a variable projection mode is used, in which each patch can use a different projection mode.

[0117] The projection mode and / or frame projection mode can be sent as metadata.

[0118] For example, a frame projection mode determination algorithm can be provided in Section 2.2.1.3.1 of V-PCC.

[0119] According to an embodiment, when the frame projection indication can use a variable projection mode, a patch projection mode can be used to indicate the appropriate mode for (de)projecting the patch.

[0120] The patch projection pattern can be sent as metadata and may be information included in the auxiliary patch information (PI).

[0121] For example, the patch projection mode decision algorithm is provided in section 2.2.1.3.2 of V-PCC.

[0122] According to an embodiment of step 3300, the pixel value of a 2D sample (u, v) corresponding to a patch in the first geometric structure image (e.g., GI0) can represent the depth value of at least one intermediate 3D sample defined along a projection line corresponding to the 2D sample (u, v). More precisely, the intermediate 3D sample is located along the projection line and shares the same coordinates as the 2D sample (u, v), and its depth value D1 is encoded in the second geometric structure image, e.g., GI1. Furthermore, the intermediate 3D sample may have a depth value between the depth value D0 and the depth value D1. A specified bit can be associated with each of the intermediate 3D samples; if an intermediate 3D sample exists, the bit is set to 1, otherwise it is set to 0.

[0123] Figure 3b This shows two intermediate 3D samples P1 located between two 3D samples P0 and P1 along the projection line PL. i1 and P i2 Example. 3D samples P0 and P1 have depth values ​​equal to D0 and D1, respectively. The two intermediate 3D samples P i1 With P i2 Depth value D i1 With D i2 They are both greater than D0 and less than D1.

[0124] Then, all the designated bits along the projection line can be concatenated to form a codeword, which is subsequently represented as an Enhanced Occupancy Map (EOM) codeword. Figure 3b As shown, assume the EOM codeword is 8 bits long, with 2 bits equal to 1, to represent two 3D samples P. i1 and P i2 The position. Finally, all EOM codewords can be packed into the image, for example, the occupancy map OM. In this case, at least one patch of the canvas can contain at least one EOM codeword. Such a patch is represented as a reference patch, and the block of the reference patch is represented as an EOM reference block. Thus, for example, when D1-D0<=1, the pixel value of the occupancy map OM can be equal to a first value (e.g., 0) to indicate an unoccupied block of the canvas, or equal to another value (e.g., greater than 0) to indicate an occupied block of the canvas, or for example, when D1-D0>1, to indicate an EOM reference block of the canvas.

[0125] The position of the pixel in the occupancy map OM of the EOM reference block and the bit value of the EOM codeword obtained from the value of those pixels indicate the 3D coordinates of the intermediate 3D sample.

[0126] In step 3400, the texture image generator TIG can obtain the input point cloud frame IPCF, occupancy map OM, auxiliary patch information PI, and ( Figure 4In step 4200), the geometry of the reconstructed point cloud frame is derived from at least one decoded geometry image DGI output from the video decoder VDEC, generating at least one texture image TI.

[0127] The texture image TI can represent the texture of the input point cloud frame IPCF, and can be, for example, an image of WxH pixels represented in YUV420-8-bit format.

[0128] The texture image generator TG can use the occupancy map information to detect (locate) occupancy blocks, thereby detecting (locating) non-empty pixels in the texture image.

[0129] The texture image generator TIG can be adapted to generate texture images TI and associate them with each geometry image / layer DGI.

[0130] According to an embodiment, the texture image generator TIG can encode (store) the texture (attribute) value T0 associated with the 2D sample of the first layer as the pixel value of the first texture image TI0, and encode (store) the texture value T1 associated with the 2D sample of the second layer as the pixel value of the second texture image TI1.

[0131] Alternatively, the texture image generation module TIG can encode (store) the texture value T1 associated with the 2D sample of the second layer as the pixel value of the first texture image TI0, and encode (store) the texture value D0 associated with the 2D sample of the first layer as the pixel value of the second geometric structure image GI1.

[0132] For example, the colors of a 3D sample can be obtained as described in Sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of the V-PCC.

[0133] The texture values ​​of two 3D samples are stored in either the first or second texture image. However, the texture value of the intermediate 3D sample cannot be stored in either the first texture image TI0 or the second texture image TI1 because the position of the projected intermediate 3D sample corresponds to an already occupied block used to store the texture value of another 3D sample (P0 or P1). Figure 3b As shown. Therefore, the texture values ​​of the intermediate 3D sample are stored in a procedurally defined location within an EOM texture block located elsewhere in the first or second texture image (section 9.4.5 of V-PCC). In short, the process determines the location of an unoccupied block in the texture image and stores the texture value associated with the intermediate 3D sample as the pixel value of that unoccupied block in the texture image, which is represented as an EOM texture block.

[0134] According to an embodiment, a padding process can be applied to geometric structures and / or texture images. This padding process can be used to fill the blank spaces between patches to generate a segmented, smooth image suitable for video compression.

[0135] Image filling examples are provided in sections 2.2.6 and 2.2.7 of V-PCC.

[0136] In step 3500, the video encoder VENC can encode the generated image / layer TI and GI.

[0137] In step 3600, the encoder OMENC can encode the occupancy map into an image, for example, as described in section 2.2.2 of V-PCC. Lossy or lossless encoding can be used.

[0138] According to an embodiment, the video encoder ENC and / or OMENC can be HEVC-based encoders.

[0139] In step 3700, the encoder PIENC can encode the auxiliary patch information PI and possible additional metadata, such as the block size T, the width W and height H of the geometry / texture image.

[0140] According to one embodiment, the auxiliary patch information can be differentially encoded (e.g., as defined in section 2.4.1 of V-PCC).

[0141] In step 3800, a multiplexer can be applied to the outputs generated in steps 3500, 3600, and 3700, and as a result, these outputs can be multiplexed together to generate a bitstream representing the base layer (BL). It should be noted that the metadata information represents a small portion of the entire bitstream. A significant amount of information is compressed using a video codec.

[0142] Figure 4 A schematic block diagram of an example of an image-based point cloud decoder 4000 according to at least one embodiment of the present invention is shown.

[0143] In step 4100, the demultiplexer DMUX can be used to demultiplex the encoded information of the bit stream representing the basic layer BL.

[0144] In step 4200, the video decoder VDEC can decode the encoded information to obtain at least one decoded geometry image DGI and at least one decoded texture image DTI.

[0145] In step 4300, the decoder OMDEC can decode the encoded information to export the decoded occupancy graph DOM.

[0146] According to one embodiment, the video decoder VDEC and / or OMDEC may be an HEVC-based decoder.

[0147] In step 4400, the decoder PIDEC can decode the encoded information to derive the auxiliary patch information DPI.

[0148] Metadata may also be derived from the bitstream BL.

[0149] In step 4500, the geometry generation module GGM can derive the geometry RG of the reconstructed point cloud frame IRPCF from at least one decoded geometry image DGI, decoded occupancy map DOM, decoded auxiliary patch information DPI, and possible additional metadata.

[0150] The Geometry Generation Module (GGM) can use the decoded Occupancy Map Information (DOM) to locate at least one non-empty pixel in a decoded geometry image (DGI).

[0151] The non-empty pixels belong to either the occupied block or the EOM reference block, depending on the pixel values ​​of the decoded occupancy information DOM and the values ​​of D1-D0 as described above.

[0152] According to the embodiment of step 4500, the geometry generation module GGM can derive two of the 3D coordinates of the intermediate 3D sample from the coordinates of the non-empty pixels.

[0153] According to the embodiment of step 4500, when the non-empty pixel belongs to the EOM reference block, the geometry generation module GGM can derive the third of the 3D coordinates of the intermediate 3D sample from the bit value of the EOM codeword.

[0154] For example, according to Figure 3b For example, the EOM codeword EOMC is used to determine the intermediate 3D sample P. i1 and P i2 3D coordinates. For example, it can be derived from D. i1 =D0+3 The intermediate 3D sample P is derived from D0. i1 The third coordinate, and for example, can be obtained from D. i2 =D0+5 The reconstructed 3D sample P is derived from D0. i2 The third coordinate. The offset value (3 or 5) is the number of intervals between D0 and D1 along the projection line.

[0155] According to an embodiment, when the non-empty pixel belongs to an occupied block, the geometry generation module GGM can derive the 3D coordinates of the reconstructed 3D sample from the coordinates of the non-empty pixel, the value of the non-empty pixel in at least one decoded geometry image DGI, decoded auxiliary patch information, and possibly from additional metadata.

[0156] The use of non-empty pixels is based on their relationship to the 2D pixels of the 3D sample. For example, for the projection in V-PCC, the 3D coordinates of the reconstructed 3D sample can be represented by depth δ(u,v), tangential shift s(u,v), and double tangential shift r(u,v) as follows:

[0157] δ(u,v)=δ0+g(u,v)

[0158] s(u,v)=s0–u0+u

[0159] r(u,v)=r0–v0+v

[0160] Where g(u,v) is the luminance component of the decoded geometric structure image DGI, (u,v) is the pixel associated with the reconstructed 3D sample, (δ0,s0,r0) is the 3D position of the connected component to which the reconstructed 3D sample belongs, and (u0,v0,u1,v1) are the coordinates in the projection plane that defines the 2D bounding box, which contains the projection of the patch associated with the connected component.

[0161] In step 4600, the texture generation module TGM can derive the texture of the reconstructed point cloud frame IRPCF from the geometry RG and at least one decoded texture image DTI.

[0162] According to the embodiment of step 4600, the texture generation module TGM can derive the texture of non-empty pixels belonging to the EOM reference block from the corresponding EOM texture block. The position of the EOM texture block in the texture image is programmatically defined (Section 9.4.5 of V-PCC).

[0163] According to the embodiment of step 4600, the texture generation module TGM can directly export the texture of non-empty pixels belonging to the occupied block as the pixel value of the first or second texture image.

[0164] Figure 5 An example syntax for representing a bitstream of a basic layer BL according to at least one of this embodiments is illustrated schematically.

[0165] The bitstream includes a bitstream header SH and at least one frame stream group GOFS.

[0166] The frame stream group GOFS includes a header HS, at least one syntax element OMS representing the occupancy map OM, at least one syntax element GVS representing at least one geometric structure image (or video), at least one syntax element TVS representing at least one texture image (or video), and at least one syntax element PIS representing auxiliary patch information and other additional metadata.

[0167] In one variant, the Frame Stream Group GOFS comprises at least one frame stream.

[0168] Figure 6 A schematic block diagram illustrating an example of a system in which various aspects and embodiments are implemented is shown.

[0169] System 6000 can be implemented as one or more devices comprising the various components described below and configured to perform one or more aspects described herein. Examples of devices that can form all or part of System 6000 include personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from the video decoder, pre-processors that provide input to the video encoder, web servers, set-top boxes, and any other devices or other communication devices for processing point clouds, video, or images. Elements of System 6000 can be implemented individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 6000 can be distributed across multiple ICs and / or discrete components. In various embodiments, system 6000 may be communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 6000 may be configured to implement one or more aspects described in this document.

[0170] System 6000 may include at least one processor 6010 configured to execute instructions loaded thereon to implement various aspects described herein, such as those described herein. Processor 6010 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 6000 may include at least one memory 6020 (e.g., a volatile memory device and / or a non-volatile memory device). System 6000 may include a storage device 6040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 6040 may include internal storage devices, additional storage devices, and / or network-accessible storage devices.

[0171] System 6000 may include an encoder / decoder module 6030 configured to, for example, process data to provide encoded or decoded data, and the encoder / decoder module 6030 may include its own processor and memory. The encoder / decoder module 6030 may represent one or more modules that may be included in the device to perform encoding and / or decoding functions. As is known, the device may include one or both encoding and decoding modules. Alternatively, the encoder / decoder module 6030 may be implemented as a separate element of system 6000, or it may be incorporated within processor 6010 as a combination of hardware and software known to those skilled in the art.

[0172] Program code to be loaded onto processor 6010 or encoder / decoder 6030 to execute the various aspects described herein may be stored in storage device 6040 and subsequently loaded onto memory 6020 for execution by processor 6010. According to various embodiments, one or more of processor 6010, memory 6020, storage device 6040, and encoder / decoder module 6030 may store one or more of various items during the execution of the processes described herein. Such stored items may include, but are not limited to, point cloud frames, encoded / decoded geometry / texture video / images or portions thereof, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.

[0173] In several embodiments, the memory within the processor 6010 and / or encoder / decoder module 6030 may be used to store instructions and provide working memory for processing that can be performed during encoding or decoding.

[0174] However, in other embodiments, external memory (e.g., the processing device may be processor 6010 or encoder / decoder module 6030) may be used for one or more of these functions. The external memory may be memory 6020 and / or storage device 6040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory may be used to store the television's operating system. In at least one embodiment, fast external dynamic volatile memory, such as RAM, may be used as working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 Video), HEVC (High-Efficiency Video Coding), or VVC (Various Video Coding).

[0175] As shown in box 6130, input to the components of system 6000 can be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that can receive, for example, RF signals transmitted over the air by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0176] In various embodiments, the input device of block 6130 may have associated corresponding input processing elements known in the art. For example, the RF section may be associated with the following necessary elements: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting it to a narrower band to select, for example, a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments may include one or more elements to perform these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various of these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband.

[0177] In one set-top box embodiment, the RF section and its associated input processing elements can receive RF signals transmitted via a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band.

[0178] Various embodiments may rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0179] Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF portion may include an antenna.

[0180] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 6000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented as needed, for example, within a separate input processing IC or within the processor 6010. Similarly, aspects of USB or HDMI interface processing can be implemented as needed within a separate interface IC or within the processor 6010. Demodulation, error correction, and demultiplexing streams can be provided to various processing elements, including, for example, the processor 6010 and the encoder / decoder 6030, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.

[0181] Various components of the system 6000 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected using a suitable connection arrangement 6140 and data can be transmitted therebetween, such as an internal bus known in the art, including an I2C bus, wiring, and printed circuit boards.

[0182] The system 6000 may include a communication interface 6050 that enables communication with other devices via a communication channel 6060. The communication interface 6050 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 6060. The communication interface 6050 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 6060 may be implemented, for example, in a wired and / or wireless medium.

[0183] In various embodiments, data can be streamed to system 6000 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals in these embodiments can be received via a communication channel 6060 and a communication interface 6050 suitable for Wi-Fi communication. The communication channel 6060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications.

[0184] Other embodiments may use a set-top box that delivers data via an HDMI connection through input box 6130 to provide streaming data to system 6000.

[0185] Other embodiments may use the RF connection of input box 6130 to provide streaming data to system 6000.

[0186] It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc., can be used to signal information to the corresponding decoder.

[0187] System 6000 can provide output signals to various output devices, including display 6100, speaker 6110, and other peripheral devices 6120. In various examples of embodiments, the other peripheral devices 6120 may include one or more of the following: a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 6000.

[0188] In various embodiments, signaling can be used to transmit control signals between system 6000 and display 6100, speaker 6110, or other peripheral devices 6120. This signaling may be AV.Link (audio / video link), CEC (consumer electronics control), or other communication protocols that enable device-to-device control with or without user intervention.

[0189] Output devices can be communicatively coupled to system 6000 via dedicated connections through their respective interfaces 6070, 6080 and 6090.

[0190] Alternatively, the output device can be connected to the system 6000 via communication interface 6050 using communication channel 6060. The display 6100 and speaker 6110 can be integrated with other components of the system 6000 into a single unit in an electronic device such as a television.

[0191] In various embodiments, the display interface 6070 may include a display driver, such as a timing controller (T Con) chip.

[0192] For example, if the RF section of input 6130 is part of a separate set-top box, then display 6100 and speaker 6110 may alternatively be separated from one or more other components. In various embodiments where display 6100 and speaker 6110 can be external components, the output signal may be provided via a dedicated output connection, which includes, for example, an HDMI port, a USB port, or a COMP output.

[0193] Retrieving the texture values ​​of intermediate 3D samples from EOM texture blocks as defined in V-PCC does not require additional syntax to indicate the location of the intermediate 3D samples, because these locations are derived from the EOM codewords of the EOM reference blocks occupying the graph OM and indicate the locations of the EOM texture blocks defined by sequential processing. However, random access to the texture values ​​of a particular intermediate 3D sample is not possible, because all EOM texture blocks will then be systematically processed to determine the location of a particular EOM texture block.

[0194] Additionally, embedding EOM texture blocks between patches of a texture image may reduce the compression efficiency of the texture image because the content of the texture image may therefore have low spatial correlation (redundancy).

[0195] According to a general aspect of at least one embodiment, a method is provided comprising: signaling at least one EOM (Enhanced Occupancy Map) texture patch representing a texture value of at least one intermediate 3D sample.

[0196] Signaling EOM texture patches allows for random (direct, fast) access to the texture values ​​of specific intermediate 3D samples. In other words, the texture corresponding to the intermediate 3D sample reconstructed from the EOM codeword can be decoded independently. This allows for additional features such as spatial scalability or parallel decoding.

[0197] Figure 7 An example of a method for transmitting at least one EOM texture patch by signaling is shown according to at least one embodiment of the present invention.

[0198] In step 710, the module may add at least one syntax element SE1 to the bitstream, which represents the 2D position of the EOM texture patch (indexed by patchIndex) in the canvas of the point cloud frame (indexed by frmIdx), and the size (height, width) of the EOM texture patch.

[0199] Canvas examples in Figure 3a As shown in the image.

[0200] In step 720, the bit stream can be sent.

[0201] In step 730, the module may retrieve (read) at least one syntax element SE1 from the bit stream (received bit stream), and may extract from the at least one syntax element SE1 the 2D position of the EOM texture patch in the canvas of the point cloud frame and the size (height, width) of the EOM texture patch.

[0202] According to an embodiment, the 2D position of the EOM texture patch in the canvas can be signaled (indexed by patchIndex) using the horizontal coordinate etpdu_2d_shift_u and the vertical coordinate etpdu_2d_shift_v defined in the canvas's 2D coordinate system.

[0203] This embodiment provides the flexibility to place EOM texture patches on the canvas.

[0204] According to an embodiment, the size (height, width) of the EOM texture patch (indexed by patchIndex) can be transmitted by signals representing the height and width of the EOM texture patch, respectively.

[0205] This embodiment allows for adjustment of the size of the EOM texture patch.

[0206] According to the embodiment of step 710, such as Figure 8a As shown, the syntax element SE1 can also be an element etpdu_points representing the number of 2D samples of the patch (indexed by patchIndex) of the canvas sent by signals.

[0207] This embodiment is advantageous because it requires sending a small amount of data.

[0208] However, even if the patch is not an EOM texture patch, that is, even if the patch does not carry any texture values ​​of intermediate 3D samples (in which case etpdu_points is set to 0), the number of 2D samples of the patch is also sent by signal.

[0209] Furthermore, this embodiment requires that the patch information in the EOM texture patch be in the same order as the patch in the current point cloud frame.

[0210] Figure 9a An example of a textured image of a canvas having four texture patches and a single EOM texture patch EOMP is shown according to the embodiment of step 710.

[0211] Each of the texture patches #1, #2, #3, and #4 stores the texture value of a 3D sample of the point cloud, and the EOM texture patch EOMP stores the texture value of an intermediate sample relative to texture patches #1, #2, and #4. First, a texture value relative to at least one intermediate 3D sample of texture patch #1 is added from the top left corner of the EOM texture patch EOMP, then a texture value (one or more) relative to at least one intermediate 3D sample of texture patch #2 is added, followed by a texture value (one or more) relative to at least one intermediate 3D sample of texture patch #4.

[0212] This embodiment does not support multiple EOM texture patches.

[0213] This embodiment does not allow, for example Figure 9a There is a “gap” between the texture values ​​of the two consecutive patches shown.

[0214] According to the embodiment of step 710, such as Figure 8bAs shown, the syntax element SE1 can also be the element etpdu_patch_count, which represents the number of reference patches in the EOM texture patch. Another syntax element SE1 can also be the element etpdu_ref_index, which represents the index of the "p"th reference patch. Yet another syntax element SE1 can also be the element etpdu_offset, which represents the offset (in pixels) used to determine the starting position of the "p"th (current) reference patch (after the previous reference patch "p-1").

[0215] Figure 9b An example of a textured image of a canvas having four texture patches and a single EOM texture patch EOMP is shown according to the embodiment of step 710.

[0216] Each of the texture patches #1, #2, #3, and #4 stores the texture value of a 3D sample of the point cloud, and the EOM texture patch EOMP stores the texture value of intermediate samples relative to texture patches #1, #2, and #4. First, the texture value (one or more) of at least one intermediate 3D sample relative to texture patch #1 is added from the top left corner of the EOM texture patch EOMP. Next, the starting position S1 of texture patch #2 is determined from the element etpdu_offset relative to texture patch #2. Then, the texture value (one or more) of at least one intermediate 3D sample associated with texture patch #2 is stored. Next, the starting position S2 of texture patch #4 is determined from the element etpdu_offset relative to texture patch #4. Then, the texture value (one or more) of at least one intermediate 3D sample associated with texture patch #4 is stored.

[0217] This embodiment provides a highly flexible solution and can support, for example... Figure 9c The example shows multiple EOM texture patches, but compared to previous embodiments, more data needs to be transmitted.

[0218] According to the embodiment of step 710, such as Figure 8c As shown, the syntax element SE1 can also be an element etpdu_patch_count representing the number of reference patches in the EOM texture patch, and an element etpdu_offset representing the offset (in pixels) used to determine the starting position of the "p" (current) reference patch (after the previous reference patch "p-1").

[0219] and Figure 8a Compared to the embodiment shown in -b, Figure 8c The example shown adds a little complexity to the parsing of the syntax ("if" statement), but provides a very close approximation. Figure 8bThe embodiments shown offer a level of flexibility (only the order of the patch blocks cannot be changed).

[0220] Figure 8c The embodiment shown supports multiple EOM texture reference patches, but it requires that the indexes of the EOM texture reference patches follow the same order as the indexes of the patches of the point cloud on all EOM texture patches; that is, if there are N regular patches, the textures of the first m1 patches should be in the first EOM texture patch, the textures of the next m2 patches should be in the next EOM texture patch, and so on, where the sum of m1...mX should be equal to or less than N.

[0221] According to the embodiment of step 710, such as Figure 8d As shown, the syntax element SE1 can also be the element etpdu_mode, which indicates the following: Figure 8a The specific syntax for EOM texture patches defined in one of the embodiments of step 710 shown in -c.

[0222] This embodiment allows multiple variants to be combined into a single syntax SE1.

[0223] According to the embodiment of step 710, the module may also add at least one other syntax element SE2 to the bitstream to signal the at least one syntax element SE1 at different levels throughout the syntax representing the point cloud frame.

[0224] According to an embodiment of step 710, at least one syntax element SE1 can be signaled at the sequence level.

[0225] For example, the at least one syntax element SE1 can be sent by a signal in a syntax element such as sequence_parameter_set() defined in V-PCC.

[0226] According to a variation of the embodiment described in step 710, such as Figure 10 As shown, the second syntax element SE2 can be the syntax element sps_enhanced_occupancy_map_texture_patch_present_flag of the sequence parameter set syntax defined in V-PCC.

[0227] The syntax element `sps_enhanced_occupancy_map_texture_patch_present_flag` indicates whether an EOM texture patch exists for a sequence of point cloud frames.

[0228] When the syntax element `sps_enhanced_occupancy_map_texture_patch_present_flag` equals 0, the EOM texture patch does not exist. When the syntax element `sps_enhanced_occupancy_map_texture_patch_present_flag` equals 1, the EOM texture patch exists.

[0229] The syntax element sps_enhanced_occupancy_map_texture_patch_present_flag can also be combined with the syntax element sps_enhanced_occupancy_map_depth_for_enabled_flag, as defined in V-PCC, to indicate whether an existing EOM texture patch exists in texture image TI0 or TI1 or whether the existing EOM texture patch exists in another bitstream.

[0230] According to a variation of the embodiment described in step 710, the syntax element sequence_parameter_set() may optionally use different videos for the EOM texture patch; that is, the PCM patch (parts 7.3.33 and 7.4.33 in V-PCC) and the EOM texture patch will be in different texture images (video bitstreams).

[0231] While this increases the number of sub-bitstreams in the global V-PCC bitstream, it allows for better tuning of the encoding parameters for each bitstream and primarily provides better scalability features—for example, if only the texture of the PCM patch is needed, there is no need to decode the EOM texture patch, and vice versa.

[0232] Figure 11 An example of the syntax table for the syntax element sequence_parameter_set() of the variant according to the embodiment of step 710 is shown.

[0233] The syntax element `sps_eom_texture_patch_separate_video_present_flag` explicitly indicates whether a separate video is used for an EOM texture patch.

[0234] According to the embodiment of step 710, such as Figure 12 and 14 As explained in the document, at least one syntax element SE1 is emitted at the frame level using a signal.

[0235] For example, the at least one syntax element SE1 is sent by signal in the syntax element patch_frame_data_unit() defined in V-PCC.

[0236] According to a variation of the embodiment described in step 710, such as Figure 12 As described in the document, the second syntax element SE2 can be the syntax element sps_enhanced_occupancy_map_texture_patch_present_flag, which uses signals to send specific syntax for representing EOM texture patches.

[0237] The syntax element `sps_enhanced_occupancy_map_texture_patch_present_flag` indicates whether the EOM texture patch exists in the bitstream.

[0238] When the syntax element `sps_enhanced_occupancy_map_texture_patch_present_flag` equals 0, the EOM texture patch does not exist. When the syntax element `sps_enhanced_occupancy_map_texture_patch_present_flag` equals 1, the EOM texture patch exists, and data related to the EOM texture patch is retrieved from the bitstream due to the function `patch_information_data(.)` as defined in V-PCC.

[0239] The function depends on, for example Figure 13 The patch pattern is defined in the table.

[0240] The patch mode I_EOMT (for intra-frame) identifies the patch as an enhanced occupancy map texture patch in an intra-frame as defined in V-PCC, while the patch mode P_EOMT (for inter-frame or predictive frames) identifies the patch as an enhanced occupancy map texture patch in an inter-frame as defined in V-PCC.

[0241] according to Figure 14 In a variation of the embodiment shown in step 710, the second element SE2 can also be a syntax element pfdu_eom_texture_patch_count indicating how many EOM texture patches exist. The number of EOM texture patches must be equal to or less than pfdu_patch_count_minus1+1.

[0242] The variant is in Figure 14 As shown in the diagram. For each of the obtained data associated with an EOM texture patch, the process is repeated over the number of EOM texture patches.

[0243] Using multiple EOM texture patches requires slightly more bit rate than using a single patch, but allows for more compact patch packing (several small EOM texture patches are easier to fit into the texture canvas than a single large EOM texture patch).

[0244] according to Figure 15 In the embodiment of step 710 shown, the second element SE2 can also be a syntax element patch_mode that indicates the patch type; for example, a regular intra-frame patch, I_INTRA, or P_INTRA.

[0245] Figure 16 A block diagram is shown of a method for encoding texture values ​​of intermediate 3D samples according to at least one current embodiment.

[0246] In step 1610, the module can analyze the orthogonal projection of the point cloud frame PCF onto the projection plane to derive the texture value TV of at least one intermediate 3D sample.

[0247] In step 1620, the module can map the at least one texture value TV to at least one EOM texture patch EOMP.

[0248] In step 1630, the module can package the at least one EOM texture patch EOMTP into a texture image.

[0249] For example, steps 1610 and 1620 may be part of step 3100, where the module PGM can generate an additional patch for each EOM texture patch EOMTP. Step 1630 may then be part of steps 3200 and 3400. In step 3200, the at least one additional patch is packaged together with other generated patches in the canvas, and in step 3400, the texture image generator TG can encode (map) a texture value TV in at least one EOM texture patch EOMTP, which is located in the same position in the texture image as the at least one additional patch.

[0250] Such as about Figure 7 As described, in step 1640, the module can send the at least one packed EOM texture patch EOMTP in the bitstream by signaling.

[0251] For example, step 1640 can be part of step 3700, where the encoder PIENC can follow the rules regarding... Figure 10-15 The described syntax encodes at least one syntax element SE1 and possibly SE2 representing the at least one EOM texture patch EOMTP.

[0252] According to the embodiment of step 1620, mapping the texture value TV to at least one EOM texture patch EOMTP includes: as follows Figure 16 The two sub-steps 1621 and 1622 are shown.

[0253] In substep 1621, the module can check whether the patch of the occupancy map OM of the point cloud frame contains the EOM reference block EOMB.

[0254] In step 1622, the module may embed the texture value TV of at least one intermediate 3D sample of each EOM reference block into at least one EOM texture patch EOPM. Therefore, each EOM texture patch is associated with at least one EOM reference block.

[0255] Basically, for each path that includes at least one EOM reference block, a sorted list of texture values ​​TV is formed, and then the sorted list is rasterized into at least one EOM texture patch.

[0256] According to the embodiment of step 1622, a sorted list of texture values ​​TV can be formed by sequentially scanning the pixels of the patch of the canvas occupying the OM. If a pixel value corresponds to an EOM codeword, the texture value TV of the corresponding intermediate 3D sample is concatenated to the end of the sorted list.

[0257] Examples of scanning methods include raster scanning, Z-order scanning, and 2D-Hilbert curve scanning. Raster scanning can scan the EOM texture patch from left to right and from top to bottom. Block-based scanning or block-by-block scanning of the EOM texture patch can also be used, where the blocks are raster scanned.

[0258] This embodiment of step 1622 is the most straightforward method because it does not require scanning 3D space.

[0259] According to the embodiment of step 1622, the sorted list of texture values ​​TV can be formed as follows:

[0260] First, all intermediate 3D samples of the patch are reconstructed. Then, using a 3D curve to scan the 3D space, the texture values ​​(DV) of the intermediate 3D samples of the patch are concatenated into the list in the order found in the 3D curve.

[0261] Examples of the 3D curves mentioned are Hilbert curves, Z-order curves, or any position-preserving curves.

[0262] This embodiment is more complex due to 3D scanning, but it increases the correlation between adjacent samples of the EOM texture patch, thereby increasing coding efficiency.

[0263] According to the embodiment of step 1622, the sorted list of texture values ​​TV can be formed as follows:

[0264] First, all intermediate 3D samples of the patch are reconstructed and represented by a tree. This tree is created by first mapping all intermediate samples to 3D space, and then recursively partitioning the 3D space in, for example, octets (to form an octree) or binary octets (to form a KD tree). Then, a sorted list of texture values ​​TV is formed by traversing this tree.

[0265] Examples of trees are octrees and KD-trees. Traversal can be performed in either depth-first or breadth-first order.

[0266] This embodiment represents a compromise between 2D and 3D scanning.

[0267] According to the embodiment of step 710, the syntax element SE1 can also be an element indicating how the list of texture values ​​is sorted.

[0268] According to the embodiment of step 710, the syntax element SE1 can also be an element indicating how to use the rasterization type.

[0269] According to one embodiment, the classification and rasterization types are fixed and known to both the encoder and the decoder.

[0270] Figure 17 A block diagram is shown of a method for decoding texture values ​​of intermediate 3D samples according to at least one current embodiment.

[0271] In step 1710, when based on the information regarding Figure 7 When the described method sends at least one EOM texture patch EOMTP with a signal, the module can derive the at least one EOM texture patch EOMTP from the bitstream.

[0272] For example, step 1710 is part of step 4400, where the decoded PIDEC can be determined according to the following... Figure 10-15 The described syntax includes at least one syntax element SE1 and possibly SE2, and decodes at least one EOM texture patch EOMTP information.

[0273] In step 1720, the module can derive the texture value TV from the at least one EOM texture patch EOMTP.

[0274] In step 1730, the module may assign at least one texture value to at least one intermediate 3D sample.

[0275] For example, step 1720 can be part of step 4600, wherein the texture generation module TGM can assign texture value TV to at least one intermediate 3D sample.

[0276] According to the embodiment of step 1720, deriving the texture value TV from the EOM texture patch EOMTP is a modification of step 6 of the reconstruction process as described in section 9.4.5 of V-PCC.

[0277] More precisely, in sub-step 1721, the (u, v) coordinates of the first pixel of the reference patch in the EOM texture patch are determined. This pixel value provides a texture value TV to the first intermediate 3D sample. Then, in sub-step 1722, the texture value TV of at least one subsequent intermediate 3D sample is retrieved from the (u, v) coordinates of the first pixel.

[0278] Assuming Figure 8b The richer syntax of the EOM texture patch shown provides a description of these sub-steps. However, other alternative embodiments of these sub-steps relative to other syntaxes of the EOM texture patch as described above can also be derived.

[0279] In the following text, it is assumed that the bitstream has been parsed to decode the syntax of EOM texture patches in a point cloud frame (with an indexed frame frameIdx). `pfdu_eom_texture_patch_count` refers to the number of EOM texture patches in the point cloud frame frameIdx. A point in frameIdx that belongs to patchIdx is the first intermediate point in the decoding order of patchIdx.

[0280] According to the embodiment of sub-step 1721, the (u,v) coordinates of the first pixel of the reference patch in the EOM texture patch are calculated as follows:

[0281] - Scan the EOM texture patch until the current index p equals the target patch index patchIdx;

[0282] as well as

[0283] - Initialize the coordinates (u, v) to the starting position of the current patch (current P) by offsetting (etpdu_2d_shift_u and etpdu_2d_shift_v).

[0284] Below is the pseudocode for implementing the algorithm in this embodiment.

[0285]

[0286] According to the embodiment used for EOM texture patch syntax, some of the above parameters can be set to their default values.

[0287] For example:

[0288] If no syntax element is available for the offset, then in the above algorithm, set etpdu_offset[frmIdx][p][r] = 0.

[0289] If there is no syntax element pfdu_eom_texture_patch_count for reference patch counting, then set pfdu_eom_texture_patch_count[frmIdx][p] to pfdu_patch_count_minus1+1.

[0290] According to an embodiment of sub-step 1722, based on a scan of a reference patch having a first point with (u, v) coordinates of the first pixel, the texture value TV for at least one subsequent intermediate 3D sample is retrieved from the (u, v) coordinates of the first pixel (step 1721).

[0291] For example, such a scan can be accomplished by the function coordinate_advance_raster(u,v,n), which advances the coordinates (u,v) by n positions according to the given raster scan order (e.g., one of the rasterization modes listed in Section 3.1.2 of V-PCC) and the size on the EOM texture patch.

[0292] Figure 18 The text describes a typical raster scan.

[0293] exist Figure 1-18 This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined.

[0294] Some examples are described in terms of block diagrams and operation flowcharts. Each block represents a circuit element, module, or code section, which includes one or more executable instructions for implementing a specified logical function(s). It should also be noted that in other implementations, the functions(s) marked in the blocks may not occur in the specified order. For example, depending on the functions involved, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order.

[0295] The implementations and aspects described herein can be implemented in, for example, methods or procedures, apparatus, computer programs, data streams, bit streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the features in question can also be implemented in other forms (e.g., apparatus or computer programs).

[0296] The method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. The processor also includes communication devices.

[0297] Additionally, the method can be implemented by instructions executable by a processor, and such instructions (and / or data values ​​generated by the implementation) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product contained in one or more computer-readable media and having computer-readable program code contained thereon, which can be executed by a computer. As used herein, a computer-readable storage medium can be considered a non-transitory storage medium having the inherent ability to store information therein and to provide the inherent ability to retrieve information therefrom. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. It should be understood that although more specific examples of computer-readable storage media to which this embodiment can be applied are provided, the following are merely illustrative and not an exhaustive list as readily understood by those skilled in the art: portable computer disks; hard disks; read-only memory (ROM); erasable programmable read-only memory (EPROM or flash memory); portable optical disc read-only memory (CD-ROM); optical storage devices; magnetic storage devices; or any suitable combination thereof.

[0298] The instructions can form an application program that is tangibly embodied on a processor-readable medium.

[0299] Instructions can be, for example, hardware, firmware, software, or a combination thereof. Instructions can be found, for example, in an operating system, a standalone application, or a combination of both. Therefore, a processor can be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (such as a storage device) having instructions for executing the process. Furthermore, the processor-readable medium can store data values ​​generated by the implementation, as a supplement to or replacement of the instructions.

[0300] For example, the device can be implemented with appropriate hardware, software, and firmware. Examples of such devices include personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from the video decoder, pre-processors that provide input to the video encoder, web servers, set-top boxes, and any other devices or other communication devices used to process point clouds, video, or images. It should be understood that the device can be mobile and can even be installed in a mobile vehicle.

[0301] The computer software can be implemented by the processor 6010 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiment can also be implemented by one or more integrated circuits. The memory 6020 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 6010 can be of any type suitable for the technical environment and can include one or more of microprocessors, general-purpose computers, special-purpose computers, and multi-core architecture-based processors, as non-limiting examples.

[0302] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry information, such as information that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. The formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.

[0303] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “comprising / including” as used in this specification may specify the presence of, for example, features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. Furthermore, when an element is referred to as “responding” or “connected” to another element, it may directly respond to or be connected to the other element, or there may be intermediate elements. Conversely, when an element is referred to as “directly responding” or “directly connected” to another element, there are no intermediate elements.

[0304] It should be understood that the use of any of the symbols / terms “ / ”, “and / or”, and “at least one”, such as in “A / B”, “A and / or B”, and “at least one of A and B”, may be intended to include only selecting the first listed option (A), or only selecting the second listed option (B), or selecting both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such wording is intended to include only selecting the first listed option (A), or only selecting the second listed option (B), or only selecting the third listed option (C), or only selecting the first and second listed options (A and B), or only selecting the first and third listed options (A and C), or only selecting the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to multiple listed items.

[0305] Various numerical values ​​may be used in this application. Specific values ​​may be, for example, for purposes, and the aspects described are not limited to these specific values.

[0306] It should be understood that although the terms first, second, etc., may be used herein to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the teachings of this application, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. No order is implied between the first element and the second element.

[0307] References to "one embodiment," "one implementation," "one implementation," or "an implementation" and other variations are frequently used to indicate that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / implementation) is included in at least one embodiment / implementation. Therefore, the phrases "in one embodiment," "in one embodiment," "in one implementation," or "in one implementation" appearing in various places throughout this application, as well as any other variations, do not necessarily refer to the same embodiment.

[0308] Similarly, references to "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" and other variations herein are frequently used to convey that a particular feature, structure, or characteristic (described in conjunction with the embodiment / example / implementation description) may include at least one embodiment / example / implementation. Therefore, the expressions "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" appearing in various places in the specification do not necessarily all refer to the same embodiment / example / implementation, nor are they necessarily separate or alternative embodiments / examples / implementations that are mutually exclusive with other embodiments / examples / implementations.

[0309] Reference numerals appearing in the claims are for illustrative purposes only and should not be construed as limiting the scope of the claims. Although not explicitly described, embodiments / examples and variations of the invention may be employed in any combination or sub-combination.

[0310] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding apparatus are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.

[0311] Although some diagrams include arrows along the communication path to indicate the main direction of communication, it should be understood that communication can occur in the opposite direction to the arrows depicted.

[0312] Various implementations involve decoding. As used herein, "decoding" can include, for example, all or part of a process performed on received point cloud frames (which may include received bitstreams that encode one or more point cloud frames) to produce a final output suitable for display or suitable for further processing in the reconstructed point cloud domain. In various embodiments, such processes include one or more of the processes typically performed by an image-based decoder.

[0313] As a further example, in one embodiment, "decoding" may refer only to entropy decoding; in another embodiment, "decoding" may refer only to differential decoding; and in yet another embodiment, "decoding" may refer to a combination of entropy decoding and differential decoding. It will be clear, and is believed to be fully understood by those skilled in the art, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, based on the context of the specific description.

[0314] Various implementations involve encoding. In a manner similar to the above discussion of “decoding,” “encoding,” as used herein, can include all or part of a process performed, for example, on an input point cloud frame to produce an encoded bitstream. In various embodiments, such a process includes one or more processes typically performed by an image-based decoder.

[0315] As a further example, in one embodiment, "encoding" may refer only to entropy encoding; in another embodiment, "encoding" may refer only to differential encoding; and in yet another embodiment, "encoding" may refer to a combination of differential and entropy encoding. It will be clear, and is believed to be fully understood by those skilled in the art, whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, based on the context of the specific description.

[0316] Note that the syntax elements used in this article, such as etpdu_2d_shift_u, etpdu_2d_shift_v, etpdu_2d_delta_size_u, etpdu_2d_delta_size_v, etpdu_points, etpdu_patch_count, etpdu_ref_index, etpdu_offset, etpdu_mode, sps_enhanced_occupancy_map_texture_patch_present_flag, sps_enhanced_occupancy_map_depth_for_enabled_flag, sps_eom_texture_patch_separate_video_present_flag, pfdu_eom_texture_patch_count, pfdu_patch_count_minus1, and patch_mode, are descriptive terms. Therefore, they do not preclude the use of other syntax element names.

[0317] Various embodiments involve rate distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often with constraints on computational complexity. Rate distortion optimization can generally be formulated as minimizing a rate distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate distortion optimization problem. For example, these approaches can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, to fully evaluate their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save encoding complexity, particularly by calculating approximate distortion based on predicting or predicting the residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, for example, by using approximate distortion only for some possible encoding options and full distortion for others. Other methods evaluate only a subset of possible encoding options. More generally, many methods employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a full evaluation of both encoding costs and associated distortion.

[0318] Additionally, this application may involve "determining" various types of information. Determining such information may include, for example, one or more of estimation information, calculation information, prediction information, or information retrieved from memory.

[0319] Furthermore, this application may relate to "accessing" various types of information. Accessing such information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0320] Additionally, this application can refer to "receiving" various types of information. Like "access," receiving is intended to be a broad term. Receiving such information can include, for example, accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" is generally referred to in one or more ways during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0321] Furthermore, as used herein, the term "signal" specifically refers to indicating something to the corresponding decoder. For example, in some embodiments, the encoder signals a specific syntax element SE1 and possibly signals a syntax element SE2. Thus, in one embodiment, the same parameters can be used on both the encoder and decoder sides. Therefore, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the foregoing refers to the verb form of the term "signal," the term "signal" can also be used as a noun herein.

[0322] Many implementations have been described. However, it should be understood that various modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to produce other implementations. Furthermore, those skilled in the art will understand that other structures and processes can replace those disclosed, and the resulting implementations will perform at least substantially the same functions (one or more) in at least substantially the same manner (one or more) to achieve at least substantially the same results (one or more) as the disclosed implementations. Therefore, this application contemplates these and other implementations.

Claims

1. A method comprising transmitting at least one texture patch, the at least one texture patch representing texture values of at least one intermediate 3D sample of a point cloud frame, wherein transmitting the texture patch comprises: The transmission represents at least one first syntax element representing the 2D position and size of the texture patch defined in a 2D mesh, wherein the at least one intermediate 3D sample is a 3D sample of the point cloud frame located between a first 3D sample and a second 3D sample of the point cloud frame along the same projection line, and the position of the at least one intermediate 3D sample along the projection line is indicated by a bit of a codeword stored in a reference patch of the occupancy map.

2. A method comprising: Decoding point cloud frames, which includes: - Decode at least one first syntax element representing the 2D position and size of at least one texture patch defined in a 2D mesh, the at least one texture patch representing the texture value of at least one intermediate 3D sample of the point cloud frame, wherein the at least one intermediate 3D sample is a 3D sample of the point cloud frame located along the same projection line between a first 3D sample and a second 3D sample of the point cloud frame, and the position of the at least one intermediate 3D sample along the projection line is indicated by a bit of a codeword stored in a reference patch of the occupancy map. - Retrieve the at least one texture patch based on the 2D position and size of the texture patch.

3. The method according to claim 1 or 2, wherein the at least one intermediate 3D sample is a 3D sample of the point cloud frame, the 3D sample having a depth value greater than the first 3D sample of the point cloud frame and lower than the second 3D sample of the point cloud frame, the at least one intermediate 3D sample and the first and second 3D samples being projected along the same projection line.

4. The method of claim 2, wherein decoding a point cloud frame further comprises: Decode at least one second syntax element representing the number of 2D samples in the texture patch of the 2D mesh.

5. The method of claim 2, wherein decoding a point cloud frame further comprises: Decode at least one third syntax element, which includes: The number of reference patches in the 2D mesh, each reference patch storing bits of its texture value stored in the intermediate 3D sample of the at least one texture patch, and An offset, used to determine the starting position of the texture value of an intermediate 3D sample stored in a given reference patch in the at least one texture patch.

6. The method of claim 5, wherein the at least one third syntax element comprises an index of the given reference patch.

7. The method of claim 2, wherein the method further comprises: Receive another syntax element, which instructs the at least one first syntax element to be decoded at a different level representing the overall syntax of the point cloud frame.

8. An apparatus comprising one or more processors configured for transmitting at least one texture patch representing texture values of at least one intermediate 3D sample of a point cloud frame, wherein transmitting the texture patch comprises: The transmission represents at least one first syntax element representing the 2D position and size of the texture patch defined in a 2D mesh, wherein the at least one intermediate 3D sample is a 3D sample of the point cloud frame located between a first 3D sample and a second 3D sample of the point cloud frame along the same projection line, and the position of the at least one intermediate 3D sample along the projection line is indicated by a bit of a codeword stored in a reference patch of the occupancy map.

9. The apparatus of claim 8, wherein transmitting a table of the texture patches further comprises: Transmit at least one second syntax element representing the number of 2D samples in the texture patch of the 2D mesh.

10. The apparatus of claim 8, wherein transmitting a table of the texture patches further comprises: Transmit at least one third syntax element, which includes: The number of reference patches in the 2D mesh, each reference patch storing bits of its texture value stored in the intermediate 3D sample of the at least one texture patch, and An offset, used to determine the starting position of the texture value of an intermediate 3D sample stored in a given reference patch in the at least one texture patch.

11. The apparatus of claim 10, wherein the at least one third syntax element comprises an index of the given reference patch.

12. The apparatus of claim 8, wherein the one or more processors are further configured to: receive another syntax element, the other syntax element instructing the at least one first syntax element to be decoded at a different level representing the overall syntax of the point cloud frame.

13. An apparatus comprising one or more processors configured to decode point cloud frames, comprising: - Decode at least one first syntax element representing the 2D position and size of at least one texture patch defined in a 2D mesh, the at least one texture patch representing the texture value of at least one intermediate 3D sample of the point cloud frame, wherein the at least one intermediate 3D sample is a 3D sample of the point cloud frame located along the same projection line between a first 3D sample and a second 3D sample of the point cloud frame, and the position of the at least one intermediate 3D sample along the projection line is indicated by a bit of a codeword stored in a reference patch of the occupancy map. - Retrieve the at least one texture patch based on the 2D position and size of the texture patch.

14. The apparatus of claim 8 or 13, wherein the at least one intermediate 3D sample is a 3D sample of the point cloud frame having a depth value greater than that of the first 3D sample of the point cloud frame and lower than that of the second 3D sample of the point cloud frame, and the at least one intermediate 3D sample and the first and second 3D samples are projected along the same projection line.

15. The apparatus of claim 13, wherein decoding a point cloud frame further comprises: Decode at least one second syntax element representing the number of 2D samples in the texture patch of the 2D mesh.

16. The apparatus of claim 13, wherein decoding a point cloud frame further comprises: Decode at least one third syntax element, which includes: The number of reference patches in the 2D mesh, each reference patch storing bits of its texture value stored in the intermediate 3D sample of the at least one texture patch, and An offset, used to determine the starting position of the texture value of an intermediate 3D sample stored in a given reference patch in the at least one texture patch.

17. The apparatus of claim 16, wherein the at least one third syntax element comprises an index of the given reference patch.

18. The apparatus of claim 16, wherein the one or more processors are further configured to: receive another syntax element, the other syntax element instructing the at least one first syntax element to be decoded at a different level representing the overall syntax of the point cloud frame.

19. A non-transitory computer-readable medium containing instructions for causing one or more processors to perform transmitting at least one texture patch, the at least one texture patch representing a texture value of at least one in-between 3D sample of a point cloud frame, wherein transmitting the texture patch comprises: The transmission represents at least one first syntax element representing the 2D position and size of the texture patch defined in a 2D mesh, wherein the at least one intermediate 3D sample is a 3D sample of the point cloud frame located between a first 3D sample and a second 3D sample of the point cloud frame along the same projection line, and the position of the at least one intermediate 3D sample along the projection line is indicated by a bit of a codeword stored in a reference patch of the occupancy map.

20. A non-transitory computer-readable medium comprising instructions for causing one or more processors to perform a decoding of a point cloud frame, the decoding point cloud frame comprising: - Decode at least one first syntax element representing the 2D position and size of at least one texture patch defined in a 2D mesh, the at least one texture patch representing the texture value of at least one intermediate 3D sample of the point cloud frame, wherein the at least one intermediate 3D sample is a 3D sample of the point cloud frame located along the same projection line between a first 3D sample and a second 3D sample of the point cloud frame, and the position of the at least one intermediate 3D sample along the projection line is indicated by a bit of a codeword stored in a reference patch of the occupancy map. - Retrieve the at least one texture patch based on the 2D position and size of the texture patch.