Point cloud processing
By projecting and reconstructing point clouds onto 2D planes and using occupancy maps, the method addresses the challenge of reducing dynamic point cloud size, enhancing efficiency in data storage and distribution for applications like autonomous vehicles and virtual reality.
Patent Information
- Application Number
- JP2020572374
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-01-16
- Filing Date
- 2019-07-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2039-07-10
AI Technical Summary
Existing point cloud processing technologies face challenges in efficiently reducing the size of dynamic point clouds while maintaining quality, particularly in applications like autonomous vehicles and virtual reality, where high-frequency capture generates large datasets that require efficient compression to be practical for distribution.
A method and apparatus for processing point clouds by projecting 3D points onto a 2D plane, dividing the patch into small blocks, determining pixel counts, and using an updated occupancy map to reconstruct a low-density representation, which includes encoding and decoding processes to reduce unnecessary points.
This approach effectively reduces the density of point clouds, minimizing data size without significantly compromising quality, making it suitable for efficient distribution and storage in applications like autonomous vehicles and virtual reality.
Smart Images

Figure 0007704531000003 
Figure 0007704531000004 
Figure 0007704531000005
Abstract
Description
Technical Field
[0001] Technical Field At least one of the present embodiments generally relates to the processing of point clouds, and more particularly to a method and apparatus for efficiently processing point clouds by removing unimportant points from the point clouds.
Background Art
[0002] Background This chapter is intended to introduce the reader to various aspects of technologies that may be related to various aspects of at least one embodiment of the present embodiment described and / or claimed below. This discussion is thought to be useful in providing the reader with background information to facilitate a better understanding of the various aspects of at least one embodiment.
[0003] Point clouds of cultural heritage / buildings, etc. can be used for various purposes. For example, an object such as a statue or a building is scanned in 3D to share the spatial shape of the object without sending or visiting the object. Also, this is a way to ensure the preservation of knowledge about the object in case the object is destroyed (e.g., a temple is destroyed by an earthquake). Such point clouds are usually static, colored, and huge.
[0004] Another use case is in topography and cartography where the use of 3D representation enables a map that can include relief and not be limited to a flat plane. Google Maps is now a good example of a 3D map, but it uses a mesh instead of a point cloud. Nevertheless, a point cloud can be a suitable data format for a 3D map, and such point clouds are usually static, colored, and huge.
[0005] The automotive industry and autonomous cars are also areas where point clouds can be used. An autonomous car should be able to "explore" its environment in order to make good driving decisions based on the reality right next to it. Typical sensors such as LIDAR (Light Detection And Ranging) generate dynamic point clouds that are used by the decision-making engine. These point clouds are not intended to be viewed by humans, are usually small, not necessarily colored, and are dynamic due to high-frequency capture. These point clouds can have other attributes such as the reflectivity provided by LIDAR. This attribute provides good information about the material of the perceived object and can be useful in making judgments.
[0006] Virtual reality and immersive worlds have recently become a hot topic and are predicted by many to be the future of 2D flat video. The basic idea is to immerse the viewer within an environment that surrounds the viewer, as opposed to a standard TV where the viewer can only see the virtual world in front of the viewer. There are several levels of immersion depending on the freedom of the viewer within the environment. Point clouds are a good candidate format for distributing virtual reality (VR) worlds.
[0007] In many applications, it is important that dynamic point clouds can be distributed to end users (or stored on a server) while consuming a reasonable amount of bitrate (or storage space for storage applications) while maintaining an acceptable (or preferably very good) quality experience. Efficient compression of these dynamic point clouds is a key point for making many immersive world distribution chains practical.
[0008] At least one embodiment has been devised in view of the above. SUMMARY OF THE INVENTION
[0009] Summary The following presents a simplified overview of at least one embodiment of the present embodiments to provide a basic understanding of some aspects of the present disclosure. This overview is not an extensive overview of the embodiments. This overview is not intended to identify key elements or critical elements of the embodiments. The following overview merely presents some aspects of at least one embodiment of the present embodiments in a simplified form as a prelude to a more detailed description provided elsewhere in this specification.
[0010] According to a general aspect of at least one embodiment, a method for reducing a point cloud representing an image, comprising: obtaining a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the 2D patch having a plurality of pixels; dividing the patch into a plurality of small blocks; determining the number of pixels within each block of the plurality of small blocks; obtaining an updated occupancy map based on the determined number of pixels within each block of the plurality of small blocks; and obtaining a reconstructed point cloud based on the updated occupancy map, the reconstructed point cloud being a low-density representation of the point cloud.
[0011] According to another general aspect of at least one embodiment, an apparatus for reducing a point cloud representing an image, comprising: means for obtaining a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the 2D patch having a plurality of pixels; means for dividing the patch into a plurality of small blocks; means for determining the number of pixels within each block of the plurality of small blocks; means for obtaining an updated occupancy map based on the determined number of pixels within each block of the plurality of small blocks; and means for obtaining a reconstructed point cloud based on the updated occupancy map, the reconstructed point cloud being a low-density representation of the point cloud.
[0012] According to another general aspect of at least one embodiment, an apparatus for reducing a point cloud representing an image is provided, the apparatus including one or more processors. The one or more processors are configured to obtain a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the 2D patch having a plurality of pixels; divide the patch into a plurality of small blocks; determine the number of pixels within each block of the plurality of small blocks; obtain an updated occupancy map based on the determined number of pixels within each block of the plurality of small blocks; and obtain a reconstructed point cloud based on the updated occupancy map, the reconstructed point cloud being a low-density representation of the point cloud.
[0013] According to another embodiment, a bitstream including a reconstructed point cloud is provided. The bitstream is formed by performing: obtaining a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the 2D patch having a plurality of pixels; dividing the patch into a plurality of small blocks; determining the number of pixels within each block of the plurality of small blocks; obtaining an updated occupancy map based on the determined number of pixels within each block of the plurality of small blocks; and obtaining a reconstructed point cloud based on the updated occupancy map, the reconstructed point cloud being a low-density representation of the point cloud.
[0014] At least one embodiment also provides a device, a computer program product, a non-transitory computer-readable medium, and a signal.
[0015] The particularity of at least one embodiment of this embodiment, as well as other objects, advantages, features, and uses of at least one embodiment of this embodiment, will become apparent from the following description of the examples taken in conjunction with the accompanying drawings.
[0016] Brief Description of the Drawings Examples of several embodiments are shown in the accompanying drawings.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 8a
Figure 9
Modes for Carrying Out the Invention
[0018] Detailed Description At least one of the embodiments is described in more detail below with reference to the accompanying drawings, which illustrate examples of at least one of the embodiments. However, the embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Accordingly, it should be understood that there is no intention to limit the embodiments to the particular forms disclosed. On the contrary, this disclosure is intended to cover all modifications, equivalents, and alternative embodiments falling within the spirit and scope of this application.
[0019] It should be understood that when a figure is presented as a flowchart, the figure also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, the figure also provides a flowchart of the corresponding method / process.
[0020] Like or identical elements in the drawings are referenced by the same reference numerals.
[0021] The aspects described and contemplated below may be implemented in many different forms. The following FIGS. 1-9 provide some embodiments, but other embodiments are contemplated, and thus the discussion of FIGS. 1-9 does not limit the scope of the embodiments.
[0022] At least one of the aspects generally relates to point cloud encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream that is generated or encoded.
[0023] More precisely, the various methods and other aspects described herein may be used to modify modules (e.g., the image-based encoder 3000 and decoder 4000 shown in FIGS. 3 and 4, respectively).
[0024] Furthermore, this aspect is not limited to MPEG standards such as MPEG-I Part 5 regarding point cloud compression, and can be applied, for example, to other standard specifications and recommendations, whether existing or to be developed in the future, and to extended versions (including MPEG-I Part 5) of any such standard specifications and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application may be used individually or in combination.
[0025] Hereinafter, image data refers to data (e.g., one or several arrays of 2D samples of a specific image / video format). The specific image / video format may define information related to the pixel values of the image (or video). The specific image / video format may also define information that can be used, for example, by a display and / or any other device to visualize and / or decode the image (or video). An image typically includes a first component (in the form of a first 2D array of samples) that usually represents the luminance (or luma) of the image. The image may also include a second component and a third component in the shape of other 2D arrays of samples (usually representing the chrominance (or chroma) of the image). Some embodiments use a set of 2D arrays of color samples, such as the traditional three-color RGB representation, to represent the same information.
[0026] In one or more embodiments, the pixel value is represented by a vector of C values, where C is the number of components. Each value of the vector is usually represented by a number of bits that can define the dynamic range of the pixel value.
[0027] An image block means a set of pixels belonging to an image. The pixel values (or image block data) of an image block refer to the values of the pixels belonging to this image block. The image block generally has a rectangular shape but can have an arbitrary shape.
[0028] A point cloud can be represented by a dataset of 3D samples in a 3D volume space that has unique coordinates and can have one or more attributes.
[0029] The 3D samples of this dataset can be defined by their spatial positions (X, Y, and Z coordinates in 3D space) and perhaps one or more associated attributes such as color, transparency, reflectivity, a two-component normal vector, or any feature representing a feature of this sample, expressed, for example, in the RGB or YUV color space. For example, a 3D sample can be defined by six components (X, Y, Z, R, G, B) or equivalently (X, Y, Z, y, U, V), where (X, Y, Z) define the coordinates of a point in 3D space, and (R, G, B) or (y, U, V) define the color of this 3D sample. Attributes of the same type can exist multiple times. For example, multiple color attributes can provide color information from different viewpoints.
[0030] A point cloud (point set) can be static or dynamic depending on whether it changes over time. A static point cloud, or an instance of a dynamic point cloud, is commonly referred to as a point cloud frame. It should be noted that in the case of a dynamic point cloud, the number of points is generally not constant, but rather generally changes over time. More generally, a point cloud can be considered dynamic if anything (such as the number of points, the position of one or more points, or any attribute of any point) changes over time.
[0031] As an example, a 2D sample can be defined by six components (u, v, Z, R, G, B) or equivalently (u, v, Z, y, U, V). (u, v) define the coordinates of the 2D sample in the 2D space of the projection plane. Z is the depth value of the projected 3D sample onto this projection plane. (R, G, B) or (y, U, V) define the color of this 3D sample.
[0032] FIG. 1 shows a schematic block diagram of an example of a two-layer base point cloud encoding structure 1000 according to at least one embodiment of this embodiment.
[0033] The two - layer base point cloud encoding structure 1000 may provide a bitstream B representing an input point cloud frame IPCF (input point cloud frame). Perhaps, the input point cloud frame IPCF represents a frame of a dynamic point cloud. Next, the frame of the dynamic point cloud may be encoded by the two - layer base point cloud encoding structure 1000 independently of another frame.
[0034] Basically, the two - layer base point cloud encoding structure 1000 may provide the ability to structure the bitstream B as a base layer BL (Base Layer) and an enhancement layer EL (Enhancement Layer). The base layer BL may provide a lossy representation of the input point cloud frame IPCF, and the enhancement layer EL may provide a high - quality (perhaps lossless) representation by encoding the isolated points not represented by the base layer BL.
[0035] The base layer BL may be provided by an image - based encoder 3000 as shown in FIG. 3. The image - based encoder 3000 may provide a geometry / texture image representing the geometry / attributes of the 3D samples of the input point cloud frame IPCF. The image - based encoder 3000 may make it possible for isolated 3D samples to be discarded. The base layer BL may be decoded by an image - based decoder 4000 (which may provide an intermediate reconstructed point cloud frame IRPCF) as shown in FIG. 4.
[0036] Next, returning to the two - layer base point cloud decoder 1000 of FIG. 1, the comparator COMP may compare the 3D samples of the input point cloud frame IPCF and the 3D samples of the intermediate reconstructed point cloud frame IRPCF to detect / find missed / isolated 3D samples. Next, the encoder ENC may encode the missed 3D samples and may provide the enhancement layer EL. Finally, the base layer BL and the enhancement layer EL may be multiplexed by a multiplexer MUX to generate the bitstream B.
[0037] According to one embodiment, the encoder ENC may include a detector that can detect the 3D reference sample R of the intermediate reconstruction point cloud frame IRPCF and associate it with the missing 3D sample M.
[0038] For example, the 3D reference sample R associated with the missing 3D sample M may be the nearest neighbor sample of the 3D sample M based on a given metric.
[0039] According to one embodiment, next, the encoder ENC may encode the spatial positions and their attributes of the missing 3D samples M as differences determined according to the spatial positions and attributes of the 3D reference samples R.
[0040] In a variation, those differences may be encoded separately.
[0041] For example, for the missing 3D sample M, the x - coordinate position difference Dx(M), y - coordinate position difference Dy(M), z - coordinate position difference Dz(M), R attribute component difference Dr(M), G attribute component difference Dg(M), and B attribute component difference Db(M) may be calculated as follows according to the spatial coordinates x(M), y(M), z(M): Dx(M)=x(M)-x(R), where x(M) and x(R) are the x - coordinates of the 3D samples M and R respectively in the geometry image provided by FIG. 3, Dy(M)=y(M)-y(R) where y(M) and y(R) are the y - coordinates of the 3D samples M and R respectively in the geometry image provided by FIG. 3, Dz(M)=z(M)-z(R) where z(M) and z(R) are the z - coordinates of the 3D samples M and R respectively in the geometry image provided by FIG. 3, Dr(M)=R(M)-R(R). where R(M) and R(R) are the r - color components of the color attributes of the 3D samples M and R respectively, Dg(M)=G(M)-G(R) where G(M) and G(R) are the g - color components of the color attributes of the 3D samples M and R respectively, Db(M) = B(M) - B(R) Here, B(M) and B(R) are the b color components of the color attributes of the 3D samples M and R, respectively.
[0042] Figure 2 shows a schematic block diagram of an example of a two-layer base point cloud decoding structure 2000 according to at least one embodiment of the present embodiment.
[0043] The behavior of the two-layer base point cloud decoding structure 2000 depends on its capabilities.
[0044] The two-layer base point cloud decoding structure 2000 with limited capabilities can access only the base layer BL from the bitstream B by using a de-multiplexer (DMUX). Next, it can provide a faithful (but lossy) version IRPCF of the input point cloud frame IPCF by decoding the base layer BL with a point cloud decoder 4000 as shown in FIG. 4.
[0045] The two-layer base point cloud decoding structure 2000 with sufficient capabilities can access both the base layer BL and the enhancement layer EL from the bitstream B by using a de-multiplexer (DMUX). A point cloud decoder 4000 as shown in FIG. 4 can determine an intermediate reconstructed point cloud frame IRPCF from the base layer BL. A decoder DEC can determine a complementary point cloud frame CPCF from the enhancement layer EL. Next, a synthesizer COM can synthesize the intermediate reconstructed point cloud frame IRPCF and the complementary point cloud frame CPCF to provide a high-quality (possibly lossless) representation (reconstruction) CRPCF of the input point cloud frame IPCF.
[0046] Figure 3 shows a schematic block diagram of an example of an image-based point cloud encoder 3000 according to at least one embodiment of the present embodiment.
[0047] The image-based point cloud coder 3000 utilizes an existing video codec to compress the geometry and texture (attribute) information of the dynamic point cloud. This is achieved by essentially converting the point cloud data into a set of different video sequences.
[0048] In certain embodiments, two videos (one for capturing the geometry information of the point cloud data and one for capturing the texture information) can be generated and compressed using an existing video codec. Examples of existing video codecs are the HEVC Main profile encoder / decoder (ITU-T H.265 Telecommunication standardization sector of ITU (02 / 2018), series H: audiovisual and multimedia systems, infrastructure of audiovisual services - coding of moving video, High efficiency video coding, Recommendation ITU-T H.265).
[0049] The additional metadata used to interpret the two videos is also typically generated and compressed separately. Such additional metadata includes, for example, an occupancy map (OM) and / or patch information (PI).
[0050] Next, the generated video bitstream and metadata can be multiplexed to generate a composite bitstream.
[0051] It should be noted that the metadata typically represents a small amount of the overall information. Most of this information is present within the video bitstream.
[0052] An example of such point cloud encoding / decoding processing is given by the Test model Category 2 algorithm (also referred to as Video-based Point Cloud Compression (abbreviated as V-PCC)) that implements the MPEG standard proposal defined in ISO / IEC JTC1 / SC29 / WG11 MPEG2019 / w18180 January 2019, Marrakesh.
[0053] In step 3100, the module PGM can generate at least one patch by decomposing the 3D samples of the dataset representing the input point cloud frame IPCF for the 2D samples on the projection plane by using a strategy that provides the best compression.
[0054] The patch can be defined as a set of 2D samples.
[0055] For example, in V-PCC, as described, for example, in Hoppe et al. (Hugues Hoppe, Tony DeRose, Tom Duchamp, John McDonald, Werner Stuetzle. Surface reconstruction from unorganized points. ACM SIGGRAPH 1992 Proceedings, 71-78), the normal in every 3D sample is first estimated. Next, an initial clustering of the input point cloud frame IPCF is obtained by associating each 3D sample with one of the six oriented planes of the 3D bounding box that encloses the 3D samples of the input point cloud frame IPCF. More precisely, each 3D sample is clustered and associated with the oriented plane that has the closest normal (i.e., maximizes the dot product of the point normal and the plane normal). Next, the 3D samples are projected onto their associated planes. A set of 3D samples that form a connected region within those planes are called connected components. A connected component is a set of at least one 3D sample that has similar normals and the same associated oriented plane. Next, the initial clustering is refined by iteratively updating the cluster associated with each 3D sample based on its normal and the cluster of its nearest neighbor samples. The final step consists of generating one patch from each connected component by projecting the 3D samples of each connected component onto the oriented plane associated with the connected component. The patch is associated with auxiliary patch information PI that represents auxiliary patch information defined for each patch to interpret the projected 2D samples corresponding to the geometry and / or attribute information.
[0056] In V-PCC, for example, the auxiliary patch information PI includes: 1) information indicating one of the six oriented planes of the 3D bounding box that encloses the 3D samples of the connected component; 2) information regarding the plane normal; 3) information for determining the 3D position of the connected component with respect to the patch represented in terms of depth, tangent direction shift, and bitangent direction shift; and 4) information such as the coordinates (u0, v0, u1, v1) within the projection plane that defines the 2D bounding box enclosing the patch.
[0057] In Project 3200, a patch packing module (PPM) can map at least one generated patch onto a 2D grid (also called a canvas) without overlap in a way that typically minimizes unused space, and can ensure that every T×T (e.g., 16×16) block of the 2D grid is associated with a unique patch. A given minimum block size T×T of the 2D grid can define the minimum distance between individual patches placed on this 2D grid. The 2D grid resolution can depend on the input point cloud size as well as its width W and height H, and the block size T can be sent to the decoder as metadata.
[0058] The auxiliary patch information PI can further include information regarding the association between the blocks and patches of the 2D grid.
[0059] In V-PCC, the auxiliary information PI can include block-to-patch index information (BlockToPatch) that determines the association between the blocks of the 2D grid and the patch indices.
[0060] Figure 5 shows an example of a canvas C that includes two patches P1, P2 and their associated 2D bounding boxes B1, B2. Note that the two bounding boxes can overlap within the canvas C as shown in Figure 5. The 2D grid (division of the canvas) is represented only inside the bounding boxes, but the division of the canvas also occurs outside those bounding boxes. The bounding boxes associated with the patches can be divided into T×T blocks (usually T = 16).
[0061] A T×T block containing 2D samples belonging to a patch can be considered an occupied block. Each occupied block on the canvas is represented by a specific pixel value (e.g., 1) in the occupancy map OM, and unoccupied blocks on the canvas are represented by another specific value (e.g., 0). Next, the pixel values of the occupancy map OM can indicate whether a T×T block on the canvas is occupied (i.e., whether it contains 2D samples belonging to a patch).
[0062] In FIG. 5, occupied blocks are represented by white blocks, and lightly gray blocks represent unoccupied blocks. The image generation process (steps 3300, 3400 in FIG. 3) utilizes the mapping of at least one generated patch onto a 2D grid calculated during step 3200 to store the geometry and texture of the input point cloud frame IPCF as an image.
[0063] In step 3300, a geometry image generator GIG can generate at least one geometry image GI from the input point cloud frame IPCF, the occupancy map OM, and the auxiliary patch information PI. The geometry image generator GIG can utilize the occupancy map information to detect (find) occupied blocks and thus non-empty pixels within the geometry image GI.
[0064] The geometry image GI can represent the geometry of the input point cloud frame IPCF and can be a monochromatic image of W×H pixels represented, for example, in the YUV420 - 8-bit format.
[0065] To handle well the case where multiple 3D samples are projected (mapped) onto the same 2D sample of a projection plane (along the same projection direction (line)), multiple images called layers can be generated. Thus, various depth values D1,..., Dn can be associated with the 2D samples of a patch, and then multiple geometry images can be generated.
[0066] In V-PCC, 2D samples of patches are projected onto two layers. The first layer, also called the nearby layer, may store a depth value D0 associated with 2D samples having, for example, a shallow depth. The second layer, also called the distant layer, may store a depth value D1 associated with 2D samples having, for example, a deep depth. Alternatively, the second layer may store the difference value between the depth values D1 and D0. For example, the information stored by the second depth image may be within the interval [0,Δ] corresponding to depth values within the range [D0,D0+Δ], where Δ is a user-defined parameter describing the surface thickness.
[0067] In this way, the second layer may contain significant contour-like high-frequency features. Thus, it would seem obvious that the second depth image may be difficult to encode by using a legacy video coder, and thus the depth values may be poorly reconstructed from the decoded second depth image, resulting in poor quality of the geometry of the reconstructed point cloud frame.
[0068] According to one embodiment, the geometry image generation module GIG may encode (derive) the depth values associated with the 2D samples of the first and second layers by using the auxiliary patch information PI.
[0069] In V-PCC, the positions of 3D samples within a patch having corresponding connected components may be expressed as follows in terms of the depth δ(u,v), the tangent direction shift s(u,v), and the bi-tangent direction shift r(u,v): δ(u,v)=δ0+g(u,v) s(u,v)=s0-u0+u r(u,v)=r0-v0+v Here, g(u,v) is the luma component of the geometry image, (u,v) is the pixel associated with the 3D sample on the projection plane, (δ0,s0,r0) is the 3D position of the corresponding patch of the connected component to which the 3D sample belongs, and (u0,v0,u1,v1) are the coordinates in the projection plane defining the 2D bounding box encompassing the projection of the patch associated with the connected component.
[0070] Therefore, the geometry image generation module GIG can encode (derive) the depth values associated with the 2D samples of the layer (the first or second layer, or both layers) as the luminance component g(u,v) given by the following equation: g(u,v) = δ(u,v) - δ0. It should be noted that this relational expression can be adopted to reconstruct the 3D sample positions (δ0, s0, r0) from the reconstructed geometry image g(u,v) with the accompanying auxiliary patch information PI.
[0071] According to one embodiment, the projection mode can be used to indicate whether the first geometry image GI0 can store the depth values of the 2D samples of either the first or second layer and whether the second geometry image GI1 can store the depth values associated with the 2D samples of either the first layer or the second layer.
[0072] For example, when the projection mode is equal to 0, the first geometry image GI0 can store the depth values of the 2D samples of the first layer, and the second geometry image GI1 can store the depth values associated with the 2D samples of the second layer. Similarly, when the projection mode is equal to 1, the first geometry image GI0 can store the depth values of the 2D samples of the second layer, and the second geometry image GI1 can store the depth values associated with the 2D samples of the first layer.
[0073] According to one embodiment, the frame projection mode can be used to indicate whether the fixed projection mode is used for all patches or whether the variable projection mode is used for all patches. Here, each patch can use a different projection mode.
[0074] The projection mode and / or the frame projection mode can be transmitted as metadata.
[0075] A frame projection mode determination algorithm (e.g., Section 2.2.1.3.1 of V-PCC) can be provided.
[0076] According to one embodiment, when the frame projection mode determination algorithm indicates that a variable projection mode can be used, the patch projection mode can be used to indicate the appropriate mode to be used for (inverse) projecting the patch.
[0077] The patch projection mode can be transmitted as metadata and may be information possibly included in the auxiliary patch information PI.
[0078] A patch projection mode determination algorithm (e.g., section 2.2.1.3.2 of V-PCC) is provided.
[0079] According to one embodiment of step 3300, the pixel value in the first geometry image (e.g., GI0) corresponding to the 2D sample (u, v) of the patch can represent the depth value of at least one intermediate 3D sample defined along the projection line corresponding to the 2D sample (u, v). More precisely, the intermediate 3D sample is along the projection line, and its depth value D1 shares the same coordinates of the 2D sample (u, v) encoded in the second geometry image (e.g., GI1). Further, the intermediate 3D sample can have a depth value between the depth value D0 and the depth value D1. A specified bit can be associated with each intermediate 3D sample and is set to 1 if the intermediate 3D sample exists and 0 otherwise.
[0080] FIG. 6 shows an example of two intermediate 3D samples P i1 , P i2 located between two 3D samples P0 and P1 along the projection line PL. The 3D samples P0, P1 have depth values equal to D0, D1 respectively. The depth values D i1 , D i2 of the two intermediate 3D samples P i1 , D i2 are greater than D0 and lower than D1.
[0081] Next, all designated bits along the projection lines can be concatenated to form a signature word (hereinafter referred to as the Enhanced - Occupancy map (EOM) signature word). As shown in FIG. 6, assuming an EOM signature word of 8 - bit length, 2 bits are equal to 1 to indicate the positions of two 3D samples P i1 , P i2 . Finally, all EOM signature words can be packed into the image (e.g., the occupancy map OM). In this case, at least one patch of the canvas can contain at least one EOM signature word. Such a patch is called a reference patch, and a block of reference patches is called an EOM reference block. Thus, the pixel value of the occupancy map OM may be equal to a first value (e.g., 0) to indicate an unoccupied block of the canvas, or may be equal to another value greater than 0 for any of, for example, indicating an occupied block of the canvas when D1 - D0 <= 1 or indicating an EOM reference block of the canvas when D1 - D0 > 1.
[0082] The bit values of the EOM signature word obtained from the positions of the pixels in the occupancy map OM that indicate the EOM reference block and the values of those pixels indicate the 3D coordinates of the intermediate 3D samples.
[0083] In step 3400, a texture image generator (TIG) can generate at least one texture image (TI) from the geometry of the reconstructed point cloud frame derived from the input point cloud frame (IPCF), the occupancy map (OM), the auxiliary patch information (PI), and at least one decoded geometry image (DGI, the output of the video decoder (VDEC) in step 4200 of FIG. 4).
[0084] The texture image TI can represent the texture of the input point cloud frame IPCF and can be, for example, an image of W×H pixels represented in the YUV420 - 8 - bit format.
[0085] The texture image generator TG can utilize occupancy map information to detect (find) occupied blocks and thus non-empty pixels within the texture image.
[0086] The texture image generator TIG can be configured to generate a texture image TI and associate it with each geometry image / layer DGI.
[0087] According to one embodiment, the texture image generator TIG can encode (store) the texture (attribute) value T0 associated with the 2D samples of the first layer as the pixel values of the first texture image TI0, and the texture value T1 associated with the 2D samples of the second layer as the pixel values of the second texture image TI1.
[0088] Alternatively, the texture image generation module TIG can encode (store) the texture value T1 associated with the 2D samples of the second layer as the pixel values of the first texture image TI0 and the texture value D0 associated with the 2D samples of the first layer as the pixel values of the second geometry image GI1.
[0089] For example, the color of the 3D samples can be obtained as described in Sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of V-PCC.
[0090] The texture values of the two 3D samples are stored in either the first or the second texture image. However, the texture value of the intermediate 3D sample cannot be stored in either the first texture image TI0 or the second texture image TI1 because the position of the projected intermediate 3D sample corresponds to an occupied block that has already been used to store the texture value of another 3D sample (P0 or P1) as shown in FIG. 6. Therefore, the texture value of the intermediate 3D sample is stored in an EOM texture block located elsewhere in either the first or the second texture image at a position defined according to the procedure (Chapter 9.4.5 of V-PCC). In summary, this process determines the position of the unoccupied blocks in the texture image and stores the texture value associated with the intermediate 3D sample as the pixel value of the unoccupied block (referred to as the EOM texture block) in the texture image.
[0091] According to one embodiment, a padding process may be applied to the geometry and / or the texture image. The padding process may be used to fill the empty space between patches to generate a segmented smooth image suitable for video compression.
[0092] Examples of image padding are provided in Sections 2.2.6 and 2.2.7 of V-PCC.
[0093] In step 3500, the video encoder VENC may encode the generated image / layer TI and GI.
[0094] In step 3600, the encoder OMENC may encode the occupancy map as an image, for example, as detailed in Section 2.2.2 of V-PCC. Lossy or lossless encoding may be used.
[0095] According to one embodiment, the video encoder ENC and / or OMENC may be an HEVC-based encoder.
[0096] In operation 3700, the coder PIENC may encode auxiliary patch information PI such as the block size T, width W, and height H of the geometry / texture image, and possibly additional metadata.
[0097] According to one embodiment, the auxiliary patch information may be differentially encoded (e.g., as defined in section 2.4.1 of V-PCC).
[0098] In operation 3800, a multiplexer may be applied to the generated outputs of operations 3500, 3600, 3700, and as a result, these outputs may be multiplexed to generate a bitstream representing the base layer BL. It should be noted that the metadata information represents only a very small part of the entire bitstream. Most of the information is compressed using a video codec.
[0099] FIG. 4 shows a schematic block diagram of an example of an image-based point cloud decoder 4000 according to at least one embodiment of this embodiment.
[0100] In operation 4100, a demultiplexer DMUX may be applied to demultiplex the encoded information of the bitstream representing the base layer BL.
[0101] In operation 4200, a video decoder VDEC may decode the encoded information to derive at least one decoded geometry image DGI and at least one decoded texture image DTI.
[0102] In operation 4300, a decoder OMDEC may decode the encoded information to derive a decoded occupancy map DOM.
[0103] According to one embodiment, the video decoder VDEC and / or OMDEC may be an HEVC-based decoder.
[0104] In operation 4400, a decoder PIDEC may decode the encoded information to derive auxiliary patch information DPI.
[0105] Perhaps, metadata can also be derived from the bitstream BL.
[0106] In step 4500, the geometry generation module GGM can derive the geometry RG of the reconstructed point cloud frame IRPCF from at least one decoded geometry image DGI, decoded occupancy map DOM, decoded auxiliary patch information DPI, and perhaps additional metadata.
[0107] The geometry generation module GGM can utilize the decoded occupancy map information DOM to find non-empty pixels in at least one decoded geometry image DGI.
[0108] As described above, non-empty pixels belong to either an occupied block or an EOM reference block depending on the pixel value of the decoded occupancy information DOM and the values of D1 to D0.
[0109] According to one embodiment of step 4500, the geometry generation module GGM can derive two of the 3D coordinates of the intermediate 3D samples from the coordinates of the non-empty pixels.
[0110] According to one embodiment of step 4500, if a non-empty pixel belongs to an EOM reference block, the geometry generation module GGM can derive the third coordinate of the 3D coordinates of the intermediate 3D sample from the bit values of the EOM codeword.
[0111] For example, according to the example of FIG. 6, the EOM codeword EOMC is used to determine the 3D coordinates of the intermediate 3D samples P i1 , P i2 . The third coordinate of the intermediate 3D sample P i1 can be derived from D0, for example, by D i1 = D0 + 3, and the third coordinate of the reconstructed 3D sample P i2 can be derived from D0, for example, by D i2 = D0 + 5. The offset value (3 or 5) is the number of intervals between D0 and D1 along the projection line.
[0112] According to one embodiment, when a non-empty pixel belongs to an occupied block, the geometry generation module GGM may derive the 3D coordinates of the reconstructed 3D sample from the coordinates of the non-empty pixel, the value of one non-empty pixel of at least one decoded geometry image DGI, the decoded auxiliary patch information, and possibly additional metadata.
[0113] The use of non-empty pixels is based on the relational expression between 2D pixels and 3D samples. For example, by the projection in V-PCC, the 3D coordinates of the reconstructed 3D sample can be expressed as follows in terms of depth δ(u,v), tangent direction shift s(u,v), and bi-tangent direction shift r(u,v): δ(u,v)=δ0+g(u,v) s(u,v)=s0 - u0 + u r(u,v)=r0 - v0 + v Here, g(u,v) is the luma component of the decoded geometry image DGI, (u,v) is the pixel associated with the reconstructed 3D sample, (δ0,s0,r0) is the 3D position of the connected component to which the reconstructed 3D sample belongs, and (u0,v0,u1,v1) are the coordinates in the projection plane that define the 2D bounding box encompassing the projection of the patch associated with the connected component.
[0114] In step 4600, the texture generation module TGM may derive the texture of the reconstructed point cloud frame IRPCF from the geometry RG and at least one decoded texture image DTI.
[0115] According to one embodiment of step 4600, the texture generation module TGM may derive the texture of non-empty pixels belonging to the EOM reference block from the corresponding EOM texture block. The position of the EOM texture image within the texture image is defined according to a procedure (section 9.4.5 of V-PCC).
[0116] According to one embodiment of Project 4600, the texture generation module TGM may derive the texture of non-empty pixels directly belonging to the occupied blocks as the pixel values of either the first or the second texture image.
[0117] FIG. 7 schematically shows an exemplary syntax of a bitstream 7000 representing the base layer BL according to at least one embodiment of the present embodiment.
[0118] The bitstream includes a bitstream header SH7100 and at least frame stream groups (Group Of Frame Stream) GOFS7110,... 7120,... 7130, etc.
[0119] The frame stream group GOFS includes a header HS7121, at least one syntax element OMS7122 representing an occupancy map OM, at least one syntax element GVS7123 representing at least one geometry image (or video), at least one syntax element TVS7125 representing at least one texture image (or video), at least one syntax element PIS7124 representing auxiliary patch information, and other additional metadata.
[0120] In a variation, the frame stream group GOFS includes at least one frame stream.
[0121] In V-PCC, the metadata can be divided into the following two categories: ● Per-patch metadata describes the coordinates of each patch within the 2D depth and texture (color) images (U0 and V0) and within the 3D space (U1, V1, D1), as well as the width and height of each patch (deltaSizeU0 and deltaSizeV0). ● Per-block metadata provides information indicating the following for each N×N block of the depth and texture images: ○ To which patch the current block (block-to-patch index information) belongs; ○ Which pixels within the current block correspond to the projection points (occupancy map).
[0122] The patch-by-patch metadata required to back-project each 2D patch into 3D space is relatively small. There are only a few (7 in the current version of V-PCC) parameters to be sent per patch, and the number of patches is usually small (in the hundreds).
[0123] However, the occupancy map and block-to-patch index metadata are required for all pixels of the patch image. For a point cloud using 10 bits per spatial coordinate, the size of the patch image is usually 1280×1280 pixels, generating a large amount of data to be encoded.
[0124] In V-PCC, the method attempts to reduce the size of the encoded occupancy map and block-to-patch metadata by (1) mixing the encoding of the occupancy map and block-to-patch index metadata and (2) reducing the accuracy of both. Here, instead of encoding the information for each pixel, it is encoded only once for an N×N pixel block. This slightly reduces the encoding efficiency of the depth and color patch images (since these images are larger otherwise), but it greatly reduces the size of the metadata. For the block-to-patch index, N is usually 16, which reduces the data volume by a factor of 256. For the occupancy map, N can be 1, 2, 4, 8, or 16. 4, which reduces the data volume by a factor of 16, is common.
[0125] Define as follows: 1. Occupancy map with sufficient accuracy: The occupancy map contains information for each pixel. 2. The occupancy map is only available at the encoder level. Occupancy map at block accuracy (usually 4): This contains information for the blocks at the occupancy map accuracy. The accuracy is called "small block". This is available on the encoder side and is sent to the decoder. 3. Block-to-patch index at block resolution (usually 16): This contains information for the blocks at the occupancy map resolution. This is available on the encoder side and is sent to the decoder.
[0126] This embodiment provides a method and / or apparatus for removing unimportant or less important points of a point cloud by changing some pixel values in a block-to-index map of per-block metadata provided for blocks of depth and texture images obtained by projecting points in an occupancy map and / or a point cloud onto a projection plane.
[0127] First Embodiment FIG. 8 shows a first embodiment of how an original point cloud (8100 in FIG. 8: 32×32 full resolution) can be processed and reconstructed into a reconstructed point cloud (8400 in FIG. 8 - reduced resolution) using an occupancy map (8200 in FIG. 8 - reduced by using 4×4 blocks) and a block-to-patch metric (8300 in FIG. 8 - further reduced by using 16×16 blocks).
[0128] This embodiment proposes a method for reducing occupancy map-to-patch metric data when there are few points in a block on the encoder side to avoid the generation of many unhelpful points (3D samples) into the reconstructed point cloud. This method reduces the data to be encoded as follows:
[0129] Update the occupancy map of small blocks.
[0130] For each small block (usually 4×4), do the following: a. Count the number of points NP small_block in the source point cloud using the full occupancy map accuracy. b. If NP small_block ≦Th small_block (where Th small_block is a given value), set the occupancy map to unoccupied; otherwise, set the block occupancy to 1. The standard value of Th small_block is 1.
[0131] OM reduce is an updated occupancy map of the source point cloud at full resolution.
[0132] According to one variation, the first embodiment of the method further includes the following: OM reduce Update the block-to-patch metric at block resolution. For each block, do the following: c. Count the number NB of occupied blocks block : d. If NB block = 0, mark the unoccupied blocks to the block-to-patch metric. Otherwise, mark the block as occupied.
[0133] However, as shown in FIG. 8, if a block is occupied, the result is that a large number of points will be reconstructed during the decoding process (if the occupancy accuracy is set to 4, the number of reconstructed points per block is 16). In particular, if one point occupies a small 4×4 block on the encoder side, the block-to-patch metric indicates that the block is occupied. During decoding, 16 points instead of 1 point will be generated.
[0134] Therefore, FIG. 8a shows another alternative of the first embodiment. From the source point cloud (8100a in FIG. 8a: full occupancy map accuracy), do the following: a. Count the number NP of points in the block at block resolution block : b. If NP block ≤ Th block (where Th block is a given value), remove the block from the block-to-patch metric. The standard value of Th block is 4.
[0135] Finally, obtain the updated occupancy map of the small block (8200 in FIG. 8a) and the updated patch-to-index of the block resolution (8300 in FIG. 8a). These indicate the data to be transmitted to the decoder side.
[0136] Therefore, a further improvement in efficiency compared to the embodiment of FIG. 8 is shown in FIG. 8a. As can be seen, the further improvement in FIG. 8a results in further suppression of small regions of the source point cloud that are not important or not very important. Even if some information is missing, the reconstructed point cloud is less dense than that shown in FIG. 8 and is more faithful to the original point cloud.
[0137] Second Embodiment The second embodiment is based on calculating the distance for each block to evaluate whether the reconstructed point cloud is close to or far from the source point cloud. This distance is the distance between two point clouds (point-to-point).
[0138] Assume that A and B are two sets of points in 3D space. The distance from A to B is defined as follows.
Equation
Equation
[0139] This embodiment avoids transmitting the reconstructed point cloud having a calculated distance greater than the threshold distance, and thus avoids overly dense point cloud reconstruction. Therefore, this embodiment also reduces the amount of data to be compressed.
[0140] Regarding gains: We observe a 1% gain in the metric (however, this was done as above for other algorithms, but needs to be done only for V-PCC).
[0141] Figure 9 shows a schematic block diagram illustrating an example of a system in which various aspects and embodiments are implemented.
[0142] System 9000 can be embodied as one or more devices including various components described below and is configured to perform one or more of the aspects described herein. Examples of devices that can form all or part of System 9000 include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital TV receivers, personal video recording systems, connected home devices, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "cables" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, set-top boxes, point clouds, any other device for processing video or images, or other communication devices. Elements of System 9000, alone or in combination, can be embodied as a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 9000 can be distributed across multiple ICs and / or discrete components. In various embodiments, System 9000 can be communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, System 9000 can be configured to implement one or more of the aspects described herein.
[0143] System 9000 may include at least one processor 9010 configured to execute instructions loaded therein to implement various aspects described herein, for example. The processor 9010 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 9000 may include at least one memory 9020 (e.g., volatile memory device and / or non-volatile memory device).
[0144] System 9000 may include a storage device 9040 that may include non-volatile memory and / or volatile memory. Non-volatile memory and / or volatile memory includes, but is not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk devices, and / or optical disk drives. The storage device 9040 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0145] System 9000 may include an encoder / decoder module 9030 configured to process data to provide, for example, encoded data or decoded data. The encoder / decoder module 9030 may include its own processor and memory. The encoder / decoder module 9030 may represent a module that may be included in a device that performs encoding and / or decoding functions. As is known, this device may include one or both of an encoding module and a decoding module. In addition, the encoder / decoder module 9030 may be implemented as another element of the system 9000 or may be incorporated into the processor 9010 as a combination of hardware and software, as is known to those skilled in the art.
[0146] The program code loaded into the processor 9010 or the coder / decoder 9030 to perform the various aspects described herein is stored in the storage device 9040 and may then be loaded into the memory 9020 for execution by the processor 9010. According to various embodiments, one or more of the processor 9010, the memory 9020, the storage device 9040, and the coder / decoder module 9030 may store one or more various items during the performance of the processes described herein. Such stored items may include, but are not limited to, point cloud frames, encoded / decoded geometry / texture video / images or portions thereof, bitstreams, matrices, variables, intermediate or final results from the processing of mathematical formulas, formulas, operations, and operation logics.
[0147] In some embodiments, the internal memory of the processor 9010 and / or the coder / decoder module 9030 may be used to store instructions and to provide a working memory for processes that may be performed during encoding or decoding.
[0148] However, in other embodiments, memory external to the processing device (e.g., the processing device may be either the processor 9010 or the coder / decoder module 9030) may be used for one or more of these functions. The external memory may be the memory 9020 and / or the storage device 9040, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory may be used to store the operating system of a television. In at least one embodiment, a high-speed external dynamic volatile memory such as RAM may be used as a working memory for video encoding and decoding operations such as MPEG2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, and also known as MPEG2 video), HEVC (High Efficiency Video Coding), or VVC (Versatile Video Coding).
[0149] Inputs to the elements of system 9000 can be provided via various input devices shown within block 9130. Such input devices can include, but are not limited to, (i) an RF section that can receive RF signals transmitted wirelessly, e.g., by a broadcaster, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.
[0150] In various embodiments, the input devices of block 9130 can each have associated input processing elements as are known in the art. For example, the RF section can be associated with elements necessary to (i) select a desired frequency (also called selecting a signal or band-limiting a signal to a frequency band), (ii) down-convert the selected signal, (iii) in some embodiments, band-limit again to a narrow-band frequency to select a signal frequency band, which can be called a channel in some cases, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired stream of data packets. The RF section of various embodiments can include one or more elements (e.g., a frequency selector, signal discriminator, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer) to perform these functions. The RF section can include a tuner that performs various of these functions including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near baseband frequency) or to baseband.
[0151] In one set-top box embodiment, the RF section and its associated input processing elements can receive RF signals transmitted over a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.
[0152] Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.
[0153] Adding an element can include inserting an element between existing elements (such as inserting an amplifier and an analog / digital converter). In various embodiments, the RF section can include an antenna.
[0154] In addition, the USB and / or HDMI terminals can each include an interface processor for connecting the system 9000 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing (such as Reed-Solomon error correction) can be implemented, for example, within a separate input processing IC or within the processor 9010 as needed. Similarly, aspects of USB or HDMI interface processing can be implemented within a separate interface IC or within the processor 9010 as needed. Demodulation, error correction, and de-multiplexing streams can be provided to various processing elements (including, for example, the processor 9010 and an encoder / decoder 9030 operating in combination with memory and storage elements) to process the data stream as needed for presentation on the output device.
[0155] The various elements of the system 9000 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and can transmit data therebetween using a suitable connection arrangement 9140 (such as an I2C bus, wiring, and an internal bus known in the art including a printed circuit board).
[0156] System 9000 may include a communication interface 9050 that enables communication with other devices via a communication channel 9060. The communication interface 9050 may include, without limitation, a transceiver configured to transmit and receive data over the communication channel 9060. The communication interface 9050 may include, without limitation, a modem or a network card, and the communication channel 9060 may be implemented, for example, in a wired and / or wireless medium.
[0157] In various embodiments, data may be streamed to the system 9000 by using a Wi-Fi network such as IEEE802.11. The Wi-Fi signals of these embodiments may be received on a communication channel 9060 and a communication interface 9050 adapted for Wi-Fi communication. The communication channel 9060 of these embodiments may typically be connected to an access point or a router that provides access to an external network including the Internet that enables streaming applications and other over-the-top communications.
[0158] Other embodiments may provide streamed data to the system 9000 by using a set-top box that delivers data on the HDMI connection of the input block 9130.
[0159] Still other embodiments may provide streamed data to the system 9000 by using the RF connection of the input block 9130.
[0160] It should be recognized that signaling can be accomplished in various ways. For example, one or more syntax elements, flags, etc. may be used to signal information to a corresponding decoder in various embodiments.
[0161] System 9000 can provide output signals to various output devices including a display 9100, a speaker 9110, and other peripheral devices 9120. The other peripheral devices 9120 can include, in various examples of the embodiments, one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of the system 9000.
[0162] In various embodiments, the control signal can be transmitted between the system 9000 and the display 9100, the speaker 9110, or the other peripheral devices 9120 by using signal transmission such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control regardless of the presence or absence of user intervention.
[0163] The output devices can be communicatively coupled to the system 9000 via dedicated connections through respective interfaces 9070, 9080, 9090.
[0164] Alternatively, the output devices can be connected to the system 9000 by using a communication channel 9060 via a communication interface 9050. The display 9100 and the speaker 9110 can be incorporated into a single unit together with other components of the system 9000 in an electronic device such as a television.
[0165] In various embodiments, the display interface 9070 can include a display driver such as a timing controller (T Con) chip.
[0166] Alternatively, the display 9100 and the speaker 9110 can be separated from one or more of the other components, for example, if the RF section of the input block 9130 is part of a separate set-top box. In various embodiments where the display 9100 and the speaker 9110 can be external components, the output signal can be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0167] According to another embodiment, a method for reducing a point cloud representing an image, the method comprising: obtaining a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the 2D patch having a plurality of pixels; dividing the patch into a plurality of small blocks; determining the number of pixels within each block of the plurality of small blocks; obtaining an updated occupancy map based on the determined number of pixels within each block of the plurality of small blocks; and obtaining a reconstructed point cloud based on the updated occupancy map, the reconstructed point cloud being a low-density representation of the point cloud.
[0168] According to another embodiment, an apparatus for reducing a point cloud representing an image, the apparatus comprising: means for obtaining a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the 2D patch having a plurality of pixels; means for dividing the patch into a plurality of small blocks; means for determining the number of pixels within each block of the plurality of small blocks; means for obtaining an updated occupancy map based on the determined number of pixels within each block of the plurality of small blocks; and means for obtaining a reconstructed point cloud based on the updated occupancy map, the reconstructed point cloud being a low-density representation of the point cloud.
[0169] According to another embodiment, there is provided an apparatus for reducing a point cloud representing an image, the apparatus including one or more processors, the one or more processors configured to: obtain a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the 2D patch having a plurality of pixels; divide the patch into a plurality of small blocks; determine the number of pixels within each block of the plurality of small blocks; obtain an updated occupancy map based on the determined number of pixels within each block of the plurality of small blocks; and obtain a reconstructed point cloud based on the updated occupancy map, the reconstructed point cloud being a low-density representation of the point cloud.
[0170] According to another embodiment, there is provided a bitstream including a reconstructed point cloud, the bitstream formed to: obtain a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the 2D patch having a plurality of pixels; divide the patch into a plurality of small blocks; determine the number of pixels within each block of the plurality of small blocks; obtain an updated occupancy map based on the determined number of pixels within each block of the plurality of small blocks; and obtain a reconstructed point cloud based on the updated occupancy map, the reconstructed point cloud being a low-density representation of the point cloud.
[0171] According to another embodiment, the embodiment further includes obtaining an updated block-to-patch index based on the updated occupancy map, the resolution of the updated occupancy map being higher than the resolution of the updated block-to-patch index.
[0172] According to another embodiment, the embodiment further includes comparing the number of pixels within each small block of the plurality of small blocks with a value.
[0173] According to another embodiment, the embodiment further includes setting each small block as unoccupied if the number of pixels in each small block is less than a threshold value.
[0174] According to another embodiment, the value is 1 or 4.
[0175] According to another embodiment, the plurality of small blocks are 4×4 blocks.
[0176] According to another embodiment, the resolution of the updated block-to-patch metric is 16×16.
[0177] In addition, one embodiment provides a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding method or a decoding method according to any of the above-described embodiments. One or more embodiments of the present embodiment also provide a computer-readable storage medium storing thereon instructions for decoding or encoding video data according to the above method. One or more embodiments also provide a computer-readable storage medium storing thereon the bitstream generated according to the above method. One or more embodiments also provide a method and an apparatus for transmitting or receiving the bitstream generated according to the above method.
[0178] Various methods are described herein, and each of the methods includes one or more steps or acts for implementing the above method. The order and / or use of the specific steps and / or acts can be modified or combined as long as the order of the specific steps and / or acts is not necessary for the proper operation of the method.
[0179] Several examples have been described with respect to block diagrams and flowcharts. Each block represents a circuit element, a module, or a portion of code that includes one or more executable instructions for implementing a specified logical function(s). It should also be noted that in other embodiments, the functions(s) shown within a block may occur out of the order shown. For example, two blocks shown in succession may actually be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order, depending on the functions involved.
[0180] Embodiments and aspects described herein may be implemented, for example, as a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if described in the context of a single form of embodiment (e.g., only as a method), the embodiments of the features described may also be implemented in other forms (e.g., in an apparatus or a program).
[0181] The method may be implemented, for example, within a processor, which refers to, for example, a processing device in general (including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device). The processor may also include a communication device.
[0182] Additionally, the method may be implemented by the instructions being performed by a processor, and such instructions (and / or data values generated by the embodiments) may be stored on a computer-readable storage medium. The computer-readable storage medium may take the form of a computer-readable program product embodied in one or more computer-readable media having computer-readable program code embodied thereon that is executable by a computer. The computer-readable storage medium used herein may be considered a non-transitory storage medium that provides the inherent ability to store information therein and the inherent ability to retrieve information therefrom. The computer-readable storage medium can be, by way of example and not limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any suitable combination of the foregoing. The following provide more specific examples of computer-readable storage media to which the present embodiment may be applied, but it should be recognized that these are merely illustrative and not an exhaustive list, as will be readily understood by those skilled in the art: portable computer diskettes; hard disks; read only memory (ROM); erasable programmable ROM (EPROM or flash memory); portable compact disk read only memory (CD-ROM); optical storage devices; magnetic storage devices; or any suitable combination of the foregoing.
[0183] The instructions may form an application program tangibly embodied on a processor-readable medium.
[0184] The instructions can be, for example, hardware, firmware, software, or a combination thereof. The instructions can be found, for example, in an operating system, a separate application, or a combination of the two. Thus, a processor can be characterized as both, for example, a device configured to perform processing and a device (such as a storage device) that includes a processor-readable medium having instructions for performing the processing. Further, the processor-readable medium can store data values generated by an embodiment, in addition to or instead of the instructions.
[0185] The apparatus can be implemented, for example, with suitable hardware, software, and firmware. Examples of such apparatus include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital TV receivers, personal video recording systems, connected home appliances, head-mounted display devices (HMDs: head mounted display devices, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, set-top boxes, point clouds, any other device for processing video or images, or other communication devices. As is apparent, these devices are mobile and can even be installed within a moving vehicle.
[0186] Computer software can be implemented by processor 9010, or by hardware, or by a combination of hardware and software. As a non-limiting example, embodiments can also be implemented by one or more integrated circuits. Memory 9020 can be of any type suitable for the technical environment, and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. Processor 9010 can be of any type suitable for the technical environment, and can include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0187] As will be apparent to those skilled in the art, embodiments can generate various signals that are formatted, for example, to carry information that can be stored or transmitted. The information can include, for example, instructions for executing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry a bitstream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the high-frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be analog or digital information, for example. The signal can be transmitted over a variety of wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0188] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a", "an" and "the" may be intended to include the plural as well, unless the context clearly dictates otherwise. The terms "includes" / "comprises" and / or "including" / "comprising", as used herein, may specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Further, when an element is referred to as being "responsive" or "connected" to another element, the element may be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to another element, no intervening elements are present.
[0189] For example, the use of any of the symbols / terms " / " ," and / or "and" at least one of "in cases such as "A / B", "A and / or B" and "at least one of A and B" may be intended to include the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As another example, in the case of "A, B and / or C" and "at least one of A, B and C", such expressions are the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or the selection of only the first and second listed options (A and B), or the selection of only the first and third listed options (A and C), or the selection of only the second and third listed options (B and C), or the selection of all three options (A and B and C). This can be extended to as many items as are enumerated, as will be apparent to those skilled in the art and related arts.
[0190] Various numerical values (for example, threshold values 1 or 4 for comparison with the number of pixels in each small block of a patch) can be used in this application. The specific values can be for illustrative purposes only, and thus the described aspects are not limited to these specific values.
[0191] Terms such as first, second, etc. can be used in this specification to describe various elements, but it is understood that these elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the teachings of this application, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. No ordering is implied between the first element and the second element.
[0192] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", and other variations, are frequently used to convey that certain functions, structures, features, etc. (described in relation to the embodiment / implementation) are included in at least one embodiment / implementation. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation", and "in other variations" throughout this application do not necessarily refer to the same embodiment.
[0193] Similarly, references in this specification to "according to an embodiment / example / implementation" or "in an embodiment / example / implementation", and other variations, are frequently used to convey that certain functions, structures, or features (described in relation to the embodiment / example / implementation) may be included in at least one embodiment / example / implementation. Thus, the appearances of the expressions "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" at various places in this specification do not necessarily refer to the same embodiment / example / implementation, nor do they necessarily refer to distinct or alternative embodiments / examples / implementations that are mutually exclusive of each other.
[0194] The reference signs appearing in the claims are for illustrative purposes only and thus do not have any limiting effect on the claims. Although not explicitly stated, the present embodiments / examples and variations can be adopted in any combination or sub-combination.
[0195] When the figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, when the figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.
[0196] It should be understood that some of the diagrams include arrows on the communication path to indicate the main direction of communication, but the communication can occur in the direction opposite to the depicted arrows.
[0197] Various embodiments are involved in decoding. As used in this application, "decoding" may include all or part of the processing performed on, for example, a received point cloud frame (possibly including a received bitstream encoding one or more point cloud frames) to generate a final output suitable for display or further processing in the reconstructed point cloud region. In various embodiments, such processing may include one or more of the processing typically performed by an image-based decoder. In various embodiments, such processing may also or alternatively include the processing performed by the decoders of the various embodiments described in this application (e.g., by decoder 2000 of FIG. 2 or decoder 4000 of FIG. 4).
[0198] As another example, in one embodiment, "decoding" may refer only to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will be apparent based on the context of the specific description and is thus considered to be fully understood by those skilled in the art.
[0199] Various embodiments are involved in encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application may include all or part of the processing performed on, for example, an input point cloud frame to generate an encoded bitstream. In various embodiments, such processing includes one or more of the processing normally performed by an image-based decoder. In various embodiments, such processing may also or alternatively include the processing performed by the decoders of the various embodiments described in this application (e.g., by decoder 1000 of FIG. 1 or decoder 3000 of FIG. 3).
[0200] In addition, this application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0201] Furthermore, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0202] In addition, this application may refer to "receiving" various information. Receiving is intended to be a broad term similar to "accessing". Receiving information may include, for example, accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" is usually involved in some way during operations such as storing information, processing information, transmitting information, moving information, copying information, deleting information, calculating information, determining information, predicting information, or estimating information, etc.
[0203] Numerous embodiments have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of various embodiments may be combined, modified, supplemented, or removed to produce other embodiments. Additionally, those skilled in the art will understand that "other structures and processes may replace those disclosed" and that "the resulting embodiments will perform at least substantially the same functions in at least substantially the same ways to achieve at least substantially the same results as the disclosed embodiments." Accordingly, these and other embodiments are contemplated by this application.
Claims
1. A method comprising encoding a point cloud, wherein encoding the point cloud comprises: obtaining an occupancy map including data indicating whether pixels of a depth image or a texture image are occupied; determining a first number of occupied pixels for at least one first block having a first size of the occupancy map; updating the occupancy map by setting the pixels of the first block in the occupancy map to unoccupied in response to a determination that the first number of occupied pixels in the at least one first block is less than or equal to a first given value; determining a second number of occupied pixels for at least one second block having a second size greater than the first size of the occupancy map; removing the second block from block-to-patch index data in response to a determination that the second number of occupied pixels in the second block is less than or equal to a second given value, wherein the block-to-patch index data determines a relevance between blocks of a 2D grid and patch indices and provides information regarding blocks of the second size; encoding the updated occupancy map.
2. An apparatus including one or more processors configured to encode a point cloud, wherein encoding the point cloud comprises: obtaining an occupancy map including data indicating whether pixels of a depth image or a texture image are occupied; determining a first number of occupied pixels for at least one first block having a first size of the occupancy map; updating the occupancy map by setting the pixels of the first block in the occupancy map to unoccupied in response to a determination that the first number of occupied pixels in the at least one first block is less than or equal to a first given value; determining a second number of occupied pixels for at least one second block having a second size greater than the first size of the occupancy map; Removing the second block from the block-to-patch index data in response to a determination that the second number of occupied pixels within the second block is less than or equal to a second given value, wherein the block-to-patch index data determines a relevance between blocks of a 2D grid and patch indices and provides information regarding blocks of the second size, the removing An apparatus comprising encoding the updated occupancy map **Claim 3** The method of claim 1, further comprising encoding the updated block-to-patch index data **Claim 4** The method according to claim 1 or 3, wherein obtaining the occupancy map comprises obtaining a 2D patch of the point cloud by projecting 3D points of the point cloud onto a projection plane, the patch having a plurality of pixels **Claim 5** The method of claim 4, wherein the patch has a higher resolution than the updated occupancy map **Claim 6** The method according to claim 1 or 3, wherein the resolution of the updated occupancy map is higher than the resolution of the updated block-to-patch index **Claim 7** The method according to claim 1 or any one of claims 3 to 6, wherein the first block having a first size is a 4×4 block **Claim 8** The method according to claim 1 or any one of claims 3 to 7, wherein the second block having a second size is 16×16 **Claim 9** The method according to claim 1 or 3, wherein the resolution of the updated block-to-patch index is 16×16 **Claim 10** A computer program comprising instructions for performing the method according to claim 1 or any one of claims 3 to 9 when executed by one or more processors **Claim 11** The apparatus of claim 2, wherein the resolution of the updated occupancy map is higher than the resolution of the updated block-to-patch index **Claim 12** The apparatus according to claim 2 or 11, wherein the first block having a first size is a 4×4 block **Claim 13** The apparatus according to claim 2 or any one of claims 11 to 12, wherein the second block having a second size is 16×16 **Claim 14** The apparatus of claim 2, wherein the resolution of the updated block-to-patch index is 16×16
Citation Information
Patent Citations
Content-based video compression
US6026183A