A method and apparatus for mesh decompression
Patent Information
- Application Number
- CN202280008556.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-14
- Filing Date
- 2022-09-16
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-09-16
Smart Images

Figure CN116711305B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Patent Application No. 17 / 945,013, filed September 14, 2022, entitled “METHOD AND APPARATUS OF ADAPTIVE SAMPLING FOR MESH COMPRESSION BY DECODERS,” and U.S. Provisional Application No. 63 / 252,084, filed October 4, 2021, entitled “Method and Apparatus of Adaptive Sampling for Mesh Compression by Decoders.” The entire disclosure of these two earlier applications is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure describes embodiments that generally involve mesh encoding and decoding. Background Technology
[0004] The background description provided herein is intended to present the overall context of this disclosure. The extent of the work of the currently attributed inventors described in the background section and in various aspects of this specification does not imply that it was prior art at the time of this disclosure's filing, nor is it expressly or implied that it was acknowledged as prior art to this disclosure.
[0005] Various techniques have been developed to capture and represent the world, such as objects and environments within a three-dimensional (3D) world. A 3D representation of the world enables more immersive interaction and communication. In some examples, point clouds and meshes can be used as 3D representations of the world. Summary of the Invention
[0006] This disclosure provides methods and apparatus for mesh encoding and decoding (e.g., compression and decompression). In some examples, the apparatus for mesh encoding and decoding includes processing circuitry. The processing circuitry decodes multiple maps in two-dimensional (2D) form from a bitstream carrying a mesh frame. The mesh frame uses polygons to represent the surfaces of objects. The multiple maps in 2D include at least a decoded geometry map and a decoded attribute map applied to an adaptive 2D atlas sampling. The processing circuitry determines at least a first sampling rate and a second sampling rate based on a syntax transmitted by signaling in the bitstream. The first sampling rate is applied to a first region of the mesh frame during the adaptive 2D atlas sampling, and the second sampling rate is applied to a second region of the mesh frame during the adaptive 2D atlas sampling. The first sampling rate differs from the second sampling rate. The processing circuitry reconstructs at least a first vertex of the mesh frame based on the multiple maps according to the first sampling rate, and reconstructs a second vertex of the mesh frame according to the second sampling rate.
[0007] In some embodiments, the plurality of maps includes a decoded occupancy map applied to the adaptive 2D atlas sampling, and the processing circuitry obtains the initial UV coordinates of an occupied point in a first sampled region of the decoded occupancy map corresponding to the first region of the grid frame, the occupied point corresponding to the first vertex; and determines the recovered UV coordinates of the first vertex based on the initial UV coordinates and the first sampling rate. In some examples, the processing circuitry decodes a first UV offset of the first sampled region from the bitstream; and determines the recovered UV coordinates of the first vertex based on the initial UV coordinates, the first sampling rate, and the first UV offset. In some examples, the processing circuitry determines the recovered 3D coordinates of the first vertex based on pixels at the initial UV coordinates in the decoded geometry; and determines the recovered attribute values of the first vertex based on pixels at the initial UV coordinates in the decoded attribute map.
[0008] In some examples, the multiple maps do not include an occupancy map, and the processing circuit decodes information indicating a first boundary vertex of the first region from the bitstream. The processing circuit infers a first occupied region corresponding to the first region in the occupancy map based on the first boundary vertex; and obtains the UV coordinates of an occupied point within the first occupied region. The occupied point corresponds to the first vertex. The processing circuit converts the UV coordinates to sampled UV coordinates at least according to the first sampling rate; and reconstructs the first vertex based on the sampled UV coordinates using the multiple maps. In one example, the processing circuit determines the recovered 3D coordinates of the first vertex based on pixels at the sampled UV coordinates in the decoded geometry map; and determines the recovered attribute values of the first vertex based on pixels at the sampled UV coordinates in the decoded attribute map.
[0009] In order to convert UV coordinates to sampled UV coordinates, in some examples, the processing circuitry decodes a first UV offset associated with the first region from the bitstream; and converts the UV coordinates to the sampled UV coordinates according to the first sampling rate and the first UV offset.
[0010] In one example, the processing circuit directly decodes the values of at least the first sampling rate and the second sampling rate from the bitstream. In another example, the processing circuit decodes at least a first index and a second index from the bitstream, the first index indicating the selection of the first sampling rate from the set of sampling rates, and the second index indicating the selection of the second sampling rate from the set of sampling rates. In another example, the processing circuit predicts the first sampling rate based on a pre-established set of rates. In another example, the processing circuit predicts the first sampling rate based on the previously used sampling rate of a decoded region of the grid frame. In yet another example, the processing circuit predicts the first sampling rate based on the previously used sampling rate of a decoded region in another grid frame decoded prior to this grid frame.
[0011] In some examples, the processing circuitry decodes a first syntax value indicating whether the first sampling rate is signaled or predicted. In response to the first syntax value indicating signaling of the first sampling rate, in one example, the processing circuitry decodes the value of the first sampling rate directly from the bitstream; or decodes an index from the bitstream that may indicate the selection of the first sampling rate from a set of sampling rates. In response to the first syntax value indicating prediction of the first sampling rate, in one example, the processing circuitry decodes a second syntax from the bitstream, indicating a predictor for predicting the first sampling rate. Furthermore, in one example, the processing circuitry determines a prediction residual based on the syntax value decoded from the bitstream; and determines the first sampling rate based on the predictor and the prediction residual.
[0012] In some examples, the processing circuitry decodes the basic sampling rate from the bitstream; and determines, based on the basic sampling rate, at least the first sampling rate and the second sampling rate.
[0013] In some examples, the processing circuitry decodes a control flag that indicates the enabling of adaptive 2D atlas sampling; determines multiple regions within the grid frame; and determines the sampling rate for each of the multiple regions.
[0014] In some embodiments, the processing circuitry determines a first UV offset associated with the first region from the bitstream; and reconstructs the first vertex of the grid frame based on the plurality of maps, according to the first sampling rate and the first UV offset. In one example, the processing circuitry directly decodes the value of the first UV offset from the bitstream. In another example, the processing circuitry predicts the first UV offset based on a pre-established set of UV offsets. In yet another example, the processing circuitry predicts the first UV offset based on previously used UV offsets of a decoded region of the grid frame. In yet another example, the processing circuitry predicts the first UV offset based on previously used UV offsets of a decoded region in another grid frame decoded prior to this grid frame.
[0015] In some examples, the processing circuitry decodes a first syntax value indicating whether the first UV offset is signaled or predicted. In response to the first syntax value indicating that the first UV offset is signaled, in one example, the processing circuitry directly decodes the value of the first UV offset from the bitstream; and infers the sign of the first UV offset based on a comparison of the first sampling rate with the basic sampling rate.
[0016] In another example, in response to the first syntax value indicating the prediction of the first UV offset, the processing circuitry decodes a second syntax from the bitstream, the second syntax indicating a predictor for predicting the first UV offset. Furthermore, in one example, the processing circuitry determines a prediction residual based on the syntax value decoded from the bitstream; and determines the first UV offset based on the predictor and the prediction residual.
[0017] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any one or a combination of methods for mesh encoding / decoding. Attached Figure Description
[0018] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, wherein:
[0019] Figure 1 Block diagrams of communication systems in some examples are shown.
[0020] Figure 2 Block diagrams of streaming media systems in some examples are shown.
[0021] Figure 3 A block diagram of an encoder used to encode point cloud frames is shown in some examples.
[0022] Figure 4A block diagram of a decoder used in some examples to decode a compressed bitstream corresponding to a point cloud frame is shown.
[0023] Figure 5 A block diagram of a video decoder is shown in some examples.
[0024] Figure 6 A block diagram of a video encoder is shown in some examples.
[0025] Figure 7 A block diagram of an encoder used to encode point cloud frames is shown in some examples.
[0026] Figure 8 A block diagram of a decoder used in some examples to decode a compressed bitstream carrying point cloud frames is shown.
[0027] Figure 9 The diagram illustrates a mapping from a grid to an atlas in some examples.
[0028] Figure 10 A schematic diagram illustrating downsampling is shown in some examples.
[0029] Figure 11 A schematic diagram of a framework for mesh compression according to some embodiments of the present disclosure is shown.
[0030] Figure 12 A schematic diagram of adaptive sampling is shown in some examples.
[0031] Figure 13 A schematic diagram of adaptive sampling is shown in some examples.
[0032] Figure 14 A flowchart outlining an example of a process is shown in some examples.
[0033] Figure 15 A flowchart outlining an example of a process is shown in some examples.
[0034] Figure 16 These are schematic diagrams of computer systems in some examples. Detailed Implementation
[0035] This disclosure provides techniques in the field of three-dimensional (3D) media processing.
[0036] Technological advancements in 3D media processing, such as progress in 3D capture, 3D modeling, and 3D rendering, have facilitated the widespread availability of 3D media content across multiple platforms and devices. In one example, a baby's first steps can be captured on one continent, and media technology could allow grandparents on another continent to watch (and perhaps interact with) and enjoy an immersive experience with the baby. According to one aspect of this disclosure, to enhance the immersive experience, 3D models are becoming increasingly complex, and the creation and consumption of 3D models consume significant data resources, such as data storage resources and data transmission resources.
[0037] According to some aspects of this disclosure, point clouds and meshes can be used as 3D models to represent immersive content.
[0038] A point cloud typically refers to a collection of points in 3D space, where each point has associated attributes such as color, material properties, texture information, intensity, reflectivity, motion-related attributes, modal attributes, and various other attributes. Point clouds can be used to reconstruct objects or scenes as combinations of these points.
[0039] An object's mesh (also known as a mesh model) can include polygons describing the object's surface. Each polygon can be defined by its vertices in 3D space and information about how those vertices connect to the polygon. Information about how the vertices connect is called connectivity information. In some examples, the mesh may also include properties associated with the vertices, such as color and normals.
[0040] According to some aspects of this disclosure, some codec tools used for point cloud compression (PCC) can be used for mesh compression. For example, a mesh can be remeshed to generate a new mesh, the connectivity information of which can be inferred. The vertices of the new mesh and the properties associated with the vertices of the new mesh can be considered as points in the point cloud and compressed using a PCC codec.
[0041] Point clouds can be used to reconstruct objects or scenes as combinations of points. These points can be captured using multiple cameras, depth sensors, or LiDAR in various settings and can consist of thousands to billions of points to realistically represent the reconstructed scene or object. A patch can typically refer to a connected subset of a surface described by a point cloud. In one example, a patch consists of points whose surface normal vectors are offset from each other by less than a threshold amount.
[0042] PCC can be implemented using various schemes, such as a geometry-based scheme known as G-PCC and a video codec-based scheme known as V-PCC. According to some aspects of this disclosure, G-PCC directly encodes 3D geometry and is a purely geometry-based method, sharing little in common with video codecs, while V-PCC is largely based on video codecs. For example, V-PCC can map points of a 3D cloud to pixels of a 2D mesh (image). The V-PCC scheme can utilize general-purpose video codecs for point cloud compression. The PCC codec (encoder / decoder) in this disclosure can be a G-PCC codec (encoder / decoder) or a V-PCC codec.
[0043] According to one aspect of this disclosure, the V-PCC scheme can use existing video codecs to compress the geometry, occupancy, and texture of a point cloud into three separate video sequences. The additional metadata required to interpret these three video sequences is compressed separately. A small portion of the entire bitstream is metadata, which, in one example, can be efficiently encoded / decoded using software. The majority of the information is processed by the video codec.
[0044] Figure 1 A block diagram of a communication system (100) is illustrated in some examples. The communication system (100) includes multiple terminal devices that can communicate with each other via, for example, a network (150). For example, the communication system (100) includes a pair of terminal devices (110) and (120) interconnected via the network (150). Figure 1 In the example, the first pair of terminal devices (110) and (120) can perform unidirectional transmission of point cloud data. For example, terminal device (110) can compress a point cloud (e.g., points representing structures) captured by a sensor (105) connected to terminal device (110). The compressed point cloud can be transmitted, for example, as a bitstream via a network (150) to another terminal device (120). Terminal device (120) can receive the compressed point cloud from the network (150), decompress the bitstream to reconstruct the point cloud, and display the reconstructed point cloud appropriately. Unidirectional data transmission can be common in media service applications, etc.
[0045] exist Figure 1In the examples, terminal device (110) and terminal device (120) can be exemplified as a server and a personal computer, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure can be applied to laptops, tablets, smartphones, gaming consoles, media players, and / or dedicated three-dimensional (3D) devices. Network (150) refers to any number of networks transmitting compressed point clouds between terminal device (110) and terminal device (120). Network (150) can include, for example, wired communication networks and / or wireless communication networks. Network (150) can exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), the Internet, etc.
[0046] Figure 2 A block diagram of a streaming media system (200) is illustrated in some examples. The streaming media system (200) is one application of point clouds. The disclosed subject matter can also be applied to other point cloud-enabled applications, such as 3D telepresence applications, virtual reality applications, etc.
[0047] The streaming media system (200) may include a capture subsystem (213). The capture subsystem (213) may include a point cloud source (201), such as a light detection and ranging (LiDAR) system, a 3D camera, a 3D scanner, a graphics generation component in software that generates an uncompressed point cloud, such as an uncompressed point cloud (202). In one example, the point cloud (202) comprises points captured by a 3D camera. The point cloud (202) is depicted with thick lines to emphasize the high data volume compared to a compressed point cloud (204) (a bitstream of the compressed point cloud). The compressed point cloud (204) may be generated by an electronics device (220) that includes an encoder (203) coupled to the point cloud source (201). The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter as described in more detail below. The compressed point cloud (204) (or the bitstream of the compressed point cloud (204)) is depicted with thin lines to emphasize its lower data volume compared to the point cloud stream (202) and can be stored on a streaming media server (205) for future use. One or more streaming media client subsystems, such as Figure 2 The client subsystems (206) and (208) can access the streaming media server (205) to retrieve copies (207) and (209) of the compressed point cloud (204). The client subsystem (206) may include, for example, a decoder (210) in an electronic device (230). The decoder (210) decodes the incoming copy (207) of the compressed point cloud and creates an outgoing stream of the reconstructed point cloud (211) that can be rendered on the rendering device (212).
[0048] It should be noted that electronic devices (220) and (230) may include other components (not shown). For example, electronic device (220) may include a decoder (not shown), and electronic device (230) may also include an encoder (not shown).
[0049] In some streaming media systems, compressed point clouds (204), (207), and (209) (e.g., the bitstream of the compressed point cloud) can be compressed according to certain standards. In some examples, video codec standards are used for point cloud compression. Examples of these standards include High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), etc.
[0050] Figure 3 A block diagram of a V-PCC encoder (300) for encoding point cloud frames is shown according to some embodiments. In some embodiments, the V-PCC encoder (300) can be used in a communication system (100) and a streaming media system (200). For example, an encoder (203) can be configured and operated in a similar manner to the V-PCC encoder (300).
[0051] The V-PCC encoder (300) receives point cloud frames as uncompressed input and generates a bitstream corresponding to compressed point cloud frames. In some embodiments, the V-PCC encoder (300) may receive point cloud frames from a point cloud source, such as a point cloud source (201).
[0052] exist Figure 3 In the example, the V-PCC encoder (300) includes a patch generation module (306), a patch packing module (308), a geometry image generation module (310), a texture image generation module (312), a patch information module (304), a occupancy map module (314), a smoothing module (336), an image filling module (316) and an image filling module (318), a group expansion module (320), a video compression module (322), a video compression module (323) and a video compression module (332), an auxiliary patch information compression module (338), an entropy compression module (334), and a multiplexer (324).
[0053] According to one aspect of this disclosure, a V-PCC encoder (300) converts a 3D point cloud frame along with some metadata (e.g., occupancy map and piece information) into an image-based representation used to convert the compressed point cloud back into a decompressed point cloud. In some examples, the V-PCC encoder (300) may convert the 3D point cloud frame into a geometry image, a texture image, and an occupancy map, and then encode the geometry image, texture image, and occupancy map into a bitstream using video encoding / decoding techniques. Typically, a geometry image is a 2D image in which pixels are filled with geometric structure values associated with points projected onto the pixels, and pixels filled with geometric structure values may be referred to as geometry samples. A texture image is a 2D image in which pixels are filled with texture values associated with points projected onto the pixels, and pixels filled with texture values may be referred to as texture samples. An occupancy map is a 2D image in which pixels are filled with values indicating whether a piece is occupied or not.
[0054] The patch generation module (306) segments the point cloud into a set of patches (e.g., a patch is defined as a connected subset of surfaces described by the point cloud), which may overlap or not overlap, such that each patch can be described by a depth field relative to a plane in 2D space. In some embodiments, the patch generation module (306) aims to decompose the point cloud into a minimum number of patches with smooth boundaries while also minimizing reconstruction errors.
[0055] In some examples, the slice information module (304) can collect slice information indicating the size and shape of the slice. In some examples, the slice information can be packed into an image frame and then encoded by the auxiliary slice information compression module (338) to generate compressed auxiliary slice information.
[0056] In some examples, the slice packing module (308) is configured to map the extracted slices onto a 2D grid while minimizing unused space and ensuring that each M×M (e.g., 16x16) block of the grid is associated with a unique slice. Effective slice packing can directly impact compression efficiency by minimizing unused space or ensuring temporal consistency.
[0057] The geometry image generation module (310) generates a 2D geometry image associated with the geometry of the point cloud at a given patch location. The texture image generation module (312) generates a 2D texture image associated with the texture of the point cloud at a given patch location. The geometry image generation module (310) and the texture image generation module (312) utilize a 3D-to-2D mapping computed during the packing process to store the geometry and texture of the point cloud as images. To better handle the case where multiple points are projected onto the same sample, each patch is projected onto two images, called layers. In one example, the geometry image is represented by a WxH monochrome frame in YUV420-8bit format. To generate the texture image, the texture generation process utilizes the reconstructed / smoothed geometry to compute the colors that will be associated with the resampled points.
[0058] The occupancy map module (314) can generate an occupancy map that describes the filling information at each cell. For example, the occupancy map includes a binary map that indicates for each cell of the grid whether the cell belongs to free space or to the point cloud. In one example, the occupancy map uses binary information to describe whether each pixel is filled. In another example, the occupancy map uses binary information to describe whether each pixel block is filled.
[0059] The occupancy map generated by the occupancy map module (314) can be compressed using either lossless or lossy encoding. When lossless encoding is used, the entropy compression module (334) is used to compress the occupancy map. When lossy encoding is used, the video compression module (332) is used to compress the occupancy map.
[0060] It is important to note that the slice packing module (308) may leave some free space between the 2D slices packed in the image frame. The image padding module (316) and image padding module (318) may fill this free space (referred to as padding) to generate an image frame suitable for 2D video and image codecs. Image padding, also known as background padding, fills unused space with redundant information. In some examples, good background padding increases the bit rate to a minimum without introducing significant coding distortion around slice boundaries.
[0061] The video compression modules (322), (323), and (332) can encode 2D images such as filled geometric images, filled texture images, and occupancy maps based on suitable video codec standards (such as HEVC, VVC, etc.). In one example, the video compression modules (322), (323), and (332) are separate components that operate independently. It should be noted that in another example, the video compression modules (322), (323), and (332) can be implemented as a single component.
[0062] In some examples, the smoothing module (336) is configured to generate a smoothed image of the reconstructed geometry. The smoothed image can be provided to the texture image generation module (312). The texture image generation module (312) can then adjust the generation of the texture image based on the reconstructed geometry. For example, when the sheet shape (e.g., geometry) is slightly distorted during encoding and decoding, this distortion can be taken into account when generating the texture image to correct for distortion in the sheet shape.
[0063] In some embodiments, the group expansion module (320) is configured to fill pixels around the object boundary with redundant low-frequency content in order to improve coding gain and visual quality of the reconstructed point cloud.
[0064] The multiplexer (324) can multiplex compressed geometric images, compressed texture images, compressed occupancy maps and compressed auxiliary slice information into a compressed bitstream.
[0065] Figure 4 A block diagram of a V-PCC decoder (400) for decoding compressed bitstreams corresponding to point cloud frames is shown in some examples. In some examples, the V-PCC decoder (400) can be used in communication systems (100) and streaming media systems (200). For example, a decoder (210) can be configured to operate in a similar manner to the V-PCC decoder (400). The V-PCC decoder (400) receives the compressed bitstream and generates a reconstructed point cloud based on the compressed bitstream.
[0066] exist Figure 4 In the example, the V-PCC decoder (400) includes a demultiplexer (432), a video decompression module (434) and a video decompression module (436), a occupancy map decompression module (438), an auxiliary slice information decompression module (442), a geometry reconstruction module (444), a smoothing module (446), a texture reconstruction module (448), and a color smoothing module (452).
[0067] The demultiplexer (432) can receive the compressed bit stream and separate it into a compressed texture image, a compressed geometric image, a compressed occupancy map, and compressed auxiliary slice information.
[0068] The video decompression module (434) and the video decompression module (436) can decode compressed images according to appropriate standards (e.g., HEVC, VVC, etc.) and output decompressed images. For example, the video decompression module (434) decodes compressed texture images and outputs decompressed texture images; the video decompression module (436) decodes compressed geometric images and outputs decompressed geometric images.
[0069] The occupancy map decompression module (438) can decode the compressed occupancy map according to a suitable standard (e.g., HEVC, VVC, etc.) and output the decompressed occupancy map.
[0070] The auxiliary chip information decompression module (442) can decode the compressed auxiliary chip information according to a suitable standard (e.g., HEVC, VVC, etc.) and output the decompressed auxiliary chip message.
[0071] The geometric reconstruction module (444) can receive the decompressed geometric image and generate the reconstructed point cloud geometry based on the decompressed occupancy map and decompressed auxiliary patch information.
[0072] The smoothing module (446) can smooth out inconsistencies at the edge of the sheet. The smoothing process is designed to mitigate potential discontinuities that may occur at the sheet boundaries due to compression artifacts. In some embodiments, a smoothing filter can be applied to pixels located on the sheet boundaries to mitigate distortion that may be caused by compression / decompression.
[0073] The texture reconstruction module (448) can determine the texture information of points in the point cloud based on the decompressed texture image and smoothed geometry.
[0074] The color smoothing module (452) can smooth out inconsistencies in coloring. In 2D video, non-adjacent pieces in 3D space are often packed side-by-side. In some examples, pixel values from non-adjacent pieces may be mixed by a block-based video codec. The goal of color smoothing is to reduce visible artifacts (spots) that appear at piece boundaries.
[0075] Figure 5 A block diagram of a video decoder (510) is shown in some examples. The video decoder (510) can be used in a V-PCC decoder (400). For example, video decompression modules (434) and (436), and occupancy graph decompression module (438) can be similarly configured as video decoders (510).
[0076] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from compressed images (such as encoded video sequences). The categories of these symbols include information used to manage the operation of the video decoder (510). The parser (520) may perform parsing / entropy decoding on the received encoded video sequence. The encoding and decoding of the encoded video sequence may be based on video encoding / decoding technologies or standards and may follow various principles, including variable-length coding with or without context sensitivity, Huffman coding, arithmetic coding, etc. The parser (520) may extract a set of subgroup parameters from the encoded video sequence for at least one subgroup of pixel subgroups in the video decoder, based on at least one parameter corresponding to a group. Subgroups may include Groups of Pictures (GOP), pictures, chunks, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The parser (520) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc., from the encoded video sequence.
[0077] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory to create symbols (521).
[0078] The reconstruction of the symbol (521) may involve multiple different units, depending on the type of the encoded video picture or its parts (such as inter-frame pictures and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (520). For clarity, the flow of such subgroup control information between the parser (520) and the multiple units described below is not depicted.
[0079] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical implementations operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated. However, for the sake of describing the disclosed subject matter, the following conceptual subdivision of the functional units is appropriate.
[0080] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantized transform coefficients as symbols (521) from the parser (520) and control information, including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block containing sample values that can be input into the aggregator (555).
[0081] In some cases, the output samples of the scaler / inverse transform (551) may belong to intra-coded blocks; that is, blocks that do not use prediction information from previously reconstructed images but can use prediction information from previously reconstructed portions of the current image. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses surrounding reconstructed information obtained from the current picture buffer (558) to generate blocks with the same size and shape as the blocks in the reconstruction. The current picture buffer (558) buffers, for example, partially reconstructed current images and / or fully reconstructed current images. In some cases, the aggregator (555) adds the prediction information already generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0082] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (553) can access the reference image memory (557) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (521) belonging to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (referred to as residual samples or residual signals in this case) to generate output sample information. The motion compensation prediction unit (553) can obtain the predicted samples from the address in the reference image memory (557) under the control of motion vectors, which are available to the motion compensation prediction unit (553) in the form of symbols (521) that may have, for example, X, Y and reference image components. Motion compensation may also include interpolation of sample values obtained from the reference image memory (557) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0083] The output samples of the aggregator (555) can be employed by various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and these parameters can be used as symbols (521) from the parser (520) in the loop filter unit (556). However, video compression techniques may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0084] The output of the loop filter unit (556) can be a sample stream, which can be output to a rendering device or stored in a reference image memory (557) for future inter-frame image prediction.
[0085] Once certain encoded images have been fully reconstructed, they can be used as reference images for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and that encoded image has been identified as a reference image (by, for example, the parser (520)), the current image buffer (558) can become part of the reference image memory (557), and a completely new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0086] The video decoder (510) can perform decoding operations according to a predetermined video compression technique in a standard such as ITU-T Rec.H.265. The encoded video sequence can conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file can select certain tools from all tools used in the video compression technique or standard as the only tools available under that configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be specified by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management sent in the encoded video sequence via signals.
[0087] Figure 6 A block diagram of a video encoder (603) according to one embodiment of the present disclosure is shown. The video encoder (603) can be used in a V-PCC encoder (300) for compressing point clouds. In one example, the video compression module (322), video compression module (323), and video compression module (332) are configured similarly to the encoder (603).
[0088] The video encoder (603) can receive images, such as filled geometric images, filled texture images, etc., and generate compressed images.
[0089] According to one embodiment, the video encoder (603) can encode and compress images of a source video sequence (images) in real time or under any other time constraints required by the application, into an encoded video sequence (compressed image). Implementing an appropriate encoding rate is a function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described below. For clarity, this coupling is not depicted. Parameters set by the controller (650) may include parameters related to rate control (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions related to the video encoder (603) optimized for a particular system design.
[0090] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an oversimplification, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and one(or more) reference images), and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (since any compression between the symbols and the encoded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference image memory (634). Since decoding of the symbol stream produces bit-accurate results independent of the decoder's location (local or remote), the contents of the reference image memory (634) are also bit-accurately corresponding between the local and remote encoders. In other words, the sample values of the reference image samples "seen" by the encoder's prediction portion are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This fundamental principle of image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.
[0091] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder such as a video decoder (510), the operation of which has already been described above. Figure 5 A detailed description has been provided. However, a brief reference is also provided. Figure 5 When symbols are available and the entropy encoder (645) and parser (520) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (510), including the parser (520), may not be fully implemented in the local decoder (633).
[0092] During operation, in some examples, the source encoder (630) may perform motion-compensated predictive coding, which predictively codes the input image with reference to one or more previously encoded images designated as "reference images" in the video sequence. In this way, the encoding engine (632) encodes the differences between pixel blocks of the input image and pixel blocks of one or more reference images that can be selected as one or more predictive references to the input image.
[0093] The local video decoder (633) can decode encoded video data of a picture that can be designated as a reference picture based on symbols created by the source encoder (630). The operation of the encoding engine (632) is preferably a lossy process. When the encoded video data can be decoded by the video decoder (630) Figure 6 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process, which can be performed by the video decoder on the reference image, and can store the reconstructed reference image in a reference image cache (634). In this way, the video encoder (603) can locally store copies of the reconstructed reference images that have the same content as the reconstructed reference images that will be obtained by the remote video decoder (without transmission errors).
[0094] The predictor (635) can perform a predictive search on the encoding engine (632). That is, for a new image to be encoded, the predictor (635) can search in the reference image memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate predictive references for the new image. The predictor (635) can operate pixel-by-pixel based on the sample blocks to find appropriate predictive references. In some cases, as determined by the search results obtained by the predictor (635), the input image may have predictive references extracted from multiple reference images stored in the reference image memory (634).
[0095] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.
[0096] The outputs of all the above functional units can be entropy encoded in the entropy encoder (645). The entropy encoder (645) translates the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, and arithmetic coding.
[0097] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a specific encoded image type to each encoded image, which can affect the encoding techniques applicable to each image. For example, an image can typically be designated as one of the following image types:
[0098] An intra-frame picture (I-picture) can be a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of those variations of I-pictures and their respective applications and characteristics.
[0099] A predicted image (P-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and reference index to predict sample values for each block.
[0100] A bidirectional prediction image (B-image) can be an image that can be encoded and decoded using either intra-frame or inter-frame prediction, which uses up to two motion vectors and a reference index to predict sample values for each block. Similarly, multiple prediction images can be used to reconstruct a single block using two or more reference images and associated metadata.
[0101] Source images are typically spatially subdivided into multiple sample blocks (e.g., blocks with 4x4, 8x8, 4x8, or 16x16 samples each) and encoded block by block. Blocks can be predictively coded by referencing other (already coded) blocks determined by the coding assignment applied to the corresponding image. For example, blocks of an I-image can be non-predictively coded, or they can be predictively coded (spatial or intra-frame prediction) by referencing already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded by referencing a previously coded reference image, either spatially or temporally. Blocks of a B-image can be predictively coded by referencing one or two previously coded reference images, either spatially or temporally.
[0102] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard (e.g., ITU-T REC.H.265). In this operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0103] Video can be in the form of multiple source images in a time series. Intra-frame image prediction (often abbreviated as intra-prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In one example, the specific image being encoded / decoded, referred to as the current image, is divided into blocks. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, that block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and in the case of multiple reference images, the motion vector can have a third dimension that identifies the reference images.
[0104] In some embodiments, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in the decoding order (but potentially in the past and future in the display order, respectively). A block in the current image can be encoded by a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be predicted by a combination of the first and second reference blocks.
[0105] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.
[0106] According to some embodiments of this disclosure, prediction is performed on a block-by-block basis, such as inter-frame picture prediction and intra-frame picture prediction. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression. The CTUs in the pictures have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Typically, a CTU comprises three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-divided into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one 64x64 pixel CU, or four 32x32 pixel CUs, or sixteen 16x16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as inter-frame prediction or intra-frame prediction. Depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes one luminance prediction block (PB) and two chrominance PBs. In one embodiment, prediction operations in encoding / decoding are performed on a per-prediction-block basis. Using the luminance prediction block as an example, the prediction block includes a matrix of pixel values (e.g., luminance values) such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0107] Figure 7 A block diagram of a G-PCC encoder (700) in some examples is shown. The G-PCC encoder (700) can be configured to receive point cloud data and compress the point cloud data to generate a bitstream carrying the compressed point cloud data. In one embodiment, the G-PCC encoder (700) may include a position quantization module (710), a duplicate point removal module (712), an octree encoding module (730), an attribute transfer module (720), a level of detail (LOD) generation module (740), an attribute prediction module (750), a residual quantization module (760), an arithmetic encoding module (770), an inverse residual quantization module (780), an addition module (781), and a memory (790) for storing the reconstructed attribute values.
[0108] As shown in the figure, an input point cloud (701) can be received at a G-PCC encoder (700). The position (e.g., 3D coordinates) of the point cloud (701) is provided to a quantization module (710). The quantization module (710) is configured to perform quantization on the coordinates to generate quantized positions. A duplicate point removal module (712) is configured to receive the quantized positions and perform a filtering process to identify and remove duplicate points. An octree encoding module (730) is configured to receive the filtered positions from the duplicate point removal module (712) and perform an octree-based encoding process to generate an octet code sequence describing a 3D mesh of voxels. The octet code is provided to an arithmetic encoding module (770).
[0109] The attribute transfer module (720) is configured to receive attributes of the input point cloud and, when multiple attribute values are associated with individual voxels, perform an attribute transfer process to determine the attribute value for each voxel. The attribute transfer process can be performed on reordered points output from the octree encoding module (730). The attributes after the transfer operation are provided to the attribute prediction module (750). The LOD generation module (740) is configured to operate on the reordered points output from the octree encoding module (730) and reorganize these points into different LODs. LOD information is supplied to the attribute prediction module (750).
[0110] The attribute prediction module (750) processes points according to the LOD-based order indicated by LOD information from the LOD generation module (740). The attribute prediction module (750) generates an attribute prediction for the current point based on the reconstructed attributes of the set of neighboring points stored in memory (790). The prediction residual can then be obtained based on the original attribute values received from the attribute transfer module (720) and the locally generated attribute prediction. When candidate indices are used in each attribute prediction process, the index corresponding to the selected prediction candidate can be provided to the arithmetic encoding module (770).
[0111] The residual quantization module (760) is configured to receive the prediction residuals from the attribute prediction module (750) and perform quantization to generate quantized residuals. The quantized residuals are then provided to the arithmetic coding module (770).
[0112] The inverse residual quantization module (780) is configured to receive the quantized residual from the residual quantization module (760) and generate the reconstructed prediction residual by performing the inverse operation of the quantization operation performed at the residual quantization module (760). The addition module (781) is configured to receive the reconstructed prediction residual from the inverse residual quantization module (780) and the corresponding attribute prediction from the attribute prediction module (750). By combining the reconstructed prediction residual and the attribute prediction, the reconstructed attribute value is generated and stored in the memory (790).
[0113] The arithmetic coding module (770) is configured to receive occupancy codes, candidate indices (if used), quantization residuals (if generated), and other information, and perform entropy coding to further compress the received values or information. As a result, a compressed bitstream (702) carrying the compressed information can be generated. The bitstream (702) can be transmitted or otherwise provided to a decoder that decodes the compressed bitstream, or it can be stored in a storage device.
[0114] Figure 8 A block diagram of a G-PCC decoder (800) according to one embodiment is shown. The G-PCC decoder (800) can be configured to receive a compressed bitstream and perform point cloud data decompression to decompress the bitstream to generate decoded point cloud data. In one embodiment, the G-PCC decoder (800) may include an arithmetic decoding module (810), an inverse residual quantization module (820), an octree decoding module (830), an LOD generation module (840), an attribute prediction module (850), and a memory (860) for storing the reconstructed attribute values.
[0115] As shown in the figure, a compressed bitstream (801) can be received at the arithmetic decoding module (810). The arithmetic decoding module (810) is configured to decode the compressed bitstream (801) to obtain the quantized residual (if generated) and occupancy code of the point cloud. The octree decoding module (830) is configured to determine the reconstructed position of the points in the point cloud based on the occupancy code. The LOD generation module (840) is configured to reorganize the points into different LODs based on the reconstructed positions and determine the LOD-based order. The inverse residual quantization module (820) is configured to generate the reconstructed residual based on the quantized residual received from the arithmetic decoding module (810).
[0116] The attribute prediction module (850) is configured to perform an attribute prediction process to determine the attribute predictions of points based on a LOD-based order. For example, the attribute predictions of the current point can be determined based on the reconstructed attribute values of the current point's neighboring points stored in memory (860). In some examples, the attribute predictions can be combined with the corresponding reconstructed residuals to generate the reconstructed attributes of the current point.
[0117] In one example, the sequence of reconstructed attributes generated from the attribute prediction module (850) and the reconstructed positions generated from the octree decoding module (830) together correspond to the decoded point cloud (802) output from the G-PCC decoder (800). Furthermore, the reconstructed attributes are also stored in memory (860) and can subsequently be used to derive attribute predictions for subsequent points.
[0118] In various embodiments, the encoder (300), decoder (400), encoder (700), and / or decoder (800) may be implemented in hardware, software, or a combination thereof. For example, the encoder (300), decoder (400), encoder (700), and / or decoder (800) may be implemented using processing circuitry such as one or more integrated circuits (ICs) that operate with or without software, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. In another example, the encoder (300), decoder (400), encoder (700), and / or decoder (800) may be implemented as software or firmware comprising instructions stored in a non-volatile (or non-transitory) computer-readable storage medium. When executed by processing circuitry such as one or more processors, these instructions cause the processing circuitry to perform the functions of the encoder (300), decoder (400), encoder (700), and / or decoder (800).
[0119] It should be noted that the attribute prediction module (750) and attribute prediction module (850) configured to implement the attribute prediction technology disclosed herein may include components that can have the same characteristics as those described herein. Figure 7 and Figure 8 Other decoders or encoders with similar or different structures as shown. Furthermore, in various examples, the encoder (700) and decoder (800) may be included in the same device or in different devices.
[0120] According to some aspects of this disclosure, mesh compression may use encoding tools different from PCC encoding tools, or may use PCC encoding tools, such as the PCC (e.g., G-PCC, V-PCC) encoders and the PCC (e.g., G-PCC, V-PCC) decoders described above.
[0121] An object's mesh (also known as a mesh model or mesh frame) can comprise polygons describing the object's surface. Each polygon can be defined by its vertices in 3D space and the edges connecting those vertices to the polygon. Information about how the vertices are connected (e.g., edge information) is called connectivity information. In some examples, the object's mesh is formed by connected triangles describing the object's surface. Two triangles sharing an edge are called two connected triangles. In other examples, the object's mesh is composed of connected quadrilaterals. Two quadrilaterals sharing an edge are called two connected quadrilaterals. It's important to note that the mesh can be formed by other suitable polygons.
[0122] In some examples, the mesh may also include properties associated with the vertices, such as color, normals, etc. By utilizing the mapping information of the mesh, which is parameterized as a 2D property graph, these properties can be associated with the mesh's surface. This mapping information is typically described by a set of parametric coordinates (called UV coordinates or texture coordinates) associated with the mesh vertices. The 2D property graph (called a texture graph in some examples) is used to store high-resolution property information, such as texture, normals, displacements, etc. This information can be used for various purposes, such as texture mapping and texture shading.
[0123] In some embodiments, the mesh may include components referred to as geometric information, connectivity information, mapping information, vertex attributes, and a property graph. In some examples, geometric information is described by a set of 3D positions associated with the vertices of the mesh. In one example, (x, y, z) coordinates may be used to describe the 3D position of a vertex and are also referred to as 3D coordinates. In some examples, connectivity information includes a set of vertex indices describing how vertices are connected to create a 3D surface. In some examples, mapping information describes how the mesh surface is mapped to a 2D region of a plane. In one example, mapping information is described together with connectivity information by a set of UV parameter / texture coordinates (u, v) associated with mesh vertices. In some examples, vertex attributes include scalar or vector attribute values associated with mesh vertices. In some examples, the property graph includes attributes associated with the mesh surface and stored as a 2D image / video. In one example, the mapping between the video (e.g., a 2D image / video) and the mesh surface is defined by mapping information.
[0124] According to one aspect of this disclosure, techniques known as UV mapping or mesh parameterization are used to map the surfaces of a mesh in a 3D domain to a 2D domain. In some examples, the mesh is divided into multiple patches in the 3D domain. A patch is a connected subset of the mesh, and the boundary of the patch is formed by boundary edges. A boundary edge of a patch is an edge belonging to only one polygon in that patch and is not shared by two adjacent polygons in that patch. In some examples, the vertices of the boundary edges in a patch are called the boundary vertices of the patch, while non-boundary vertices in the patch may be called the interior vertices of the patch.
[0125] In some examples, the mesh of an object is formed by connected triangles, and the mesh can be divided into multiple patches, each patch being a subset of connected triangles. A patch's boundary edge is an edge that belongs to only one triangle within the patch and is not shared by adjacent triangles within the patch. In some examples, the vertices of the boundary edges in a patch are called the patch's boundary vertices, while non-boundary vertices in a patch can be called the patch's interior vertices.
[0126] According to one aspect of this disclosure, in some examples, the sheets are parameterized as 2D shapes (also referred to as UV sheets). These 2D shapes can be packaged (e.g., oriented and placed) into maps, which in some examples are also referred to as atlases. In some examples, 2D image or video processing techniques can be used to further process the maps.
[0127] In one example, UV mapping generates a 2D UV atlas (also called a UV map) and one or more texture atlases (also called texture maps) corresponding to a 3D mesh. The UV atlas includes the assignment of 3D vertices of the 3D mesh to 2D points within a 2D domain (e.g., a rectangle). The UV atlas is a mapping between the coordinates of the 3D surface and the coordinates of the 2D domain. In one example, a point in the UV atlas at 2D coordinates (u, v) has a value formed by the coordinates (x, y, z) of the vertices in the 3D domain. In one example, the texture atlas includes color information for the 3D mesh. For example, a point in the texture atlas at 2D coordinates (u, v) (which has 3D values (x, y, z) in the UV atlas) has a color that specifies the color attribute of the point at (x, y, z) in the 3D domain. In some examples, the coordinates (x, y, z) in the 3D domain are referred to as 3D coordinates or xyz coordinates, and the 2D coordinates (u, v) are referred to as uv coordinates or UV coordinates.
[0128] According to some aspects of this disclosure, grid compression can be performed by representing a grid using one or more 2D maps (also referred to as 2D atlases in some examples) and then encoding the 2D maps using an image or video codec. Different techniques can be used to generate 2D maps.
[0129] Figure 9 A schematic diagram of the mapping from a 3D mesh (910) to a 2D atlas (920) is shown in some examples. Figure 9 In the example, the 3D mesh (910) comprises four vertices 1-4 forming four patches AD. Each patch has a set of vertices and associated attribute information. For example, patch A is formed by vertices 1, 2, and 3 connected to form a triangle; patch B is formed by vertices 1, 3, and 4 connected to form a triangle; patch C is formed by vertices 1, 2, and 4 connected to form a triangle; and patch D is formed by vertices 2, 3, and 4 connected to form a triangle. In some examples, vertices 1, 2, 3, and 4 may have their own attributes, and the triangles formed by vertices 1, 2, 3, and 4 may also have their own attributes.
[0130] In one example, pieces A, B, C, and D in 3D are mapped to a 2D domain, such as a 2D atlas (920), also known as a UV atlas (920) or map (920). For example, piece A is mapped to a 2D shape (also known as a UV piece) A' in map (920), piece B is mapped to a 2D shape (also known as a UV piece) B' in map (920), piece C is mapped to a 2D shape (also known as a UV piece) C' in map (920), and piece D is mapped to a 2D shape (also known as a UV piece) D' in map (920). In some examples, coordinates in the 3D domain are called (x, y, z) coordinates, and coordinates in the 2D domain, such as in map (920), are called UV coordinates. Vertices in the 3D mesh can have corresponding UV coordinates in map (920).
[0131] The map (920) can be a geometric map with geometric information, a texture map with color, normal, weave or other attribute information, or an occupancy map with occupancy information.
[0132] Although Figure 9 In the examples, each piece is represented by a triangle, but it's important to note that a piece can include any suitable number of connected subsets of vertices that are joined to form a mesh. In some examples, the vertices in a piece are connected into triangles. It's worth noting that other suitable shapes can be used to connect the vertices in a piece.
[0133] In one example, the geometric information of a vertex can be stored in a 2D geometry. For instance, the 2D geometry stores the (x, y, z) coordinates of a sample point at the corresponding point in the 2D geometry. For example, a point located at (u, v) in the 2D geometry has vector values with three components corresponding to the x, y, and z values of the corresponding sample point in the 3D mesh.
[0134] According to one aspect of this disclosure, areas in the map may not be fully occupied. For example, in Figure 9 In this context, the regions outside the 2D shapes A', B', C', and D' are undefined. Sample values from these regions can be discarded after decoding. In some cases, the occupancy map is used to store additional information for each pixel, such as storing binary values to identify whether a pixel belongs to a patch or is undefined.
[0135] According to one aspect of this disclosure, a dynamic mesh is a mesh in which at least one of its components (geometric information, connectivity information, mapping information, vertex attributes, and attribute graphs) changes over time. A dynamic mesh can be described by a series of meshes (also referred to as mesh frames). Dynamic meshes may require large amounts of data because they can include a large amount of information that changes over time. Mesh compression techniques can allow for the efficient storage and transmission of media content within a mesh representation.
[0136] In some examples, dynamic meshes can have constant connectivity information, time-varying geometry, and time-varying vertex properties. In other examples, dynamic meshes can have connectivity information that changes over time. In one example, digital content creation tools typically generate dynamic meshes with time-varying property graphs and time-varying connectivity information. In some examples, volumetric acquisition techniques are used to generate dynamic meshes. Volumetric acquisition techniques can generate dynamic meshes with time-varying connectivity information, especially under real-time constraints.
[0137] Several techniques are used for mesh compression. In some examples, UV atlas sampling and V-PCC can be used for mesh compression. For instance, UV atlases are sampled on a regular mesh to generate a geometric image with regular mesh samples. The connectivity of the regular mesh samples can be inferred. The regular mesh samples can be considered as points in a point cloud and therefore can be encoded using a PCC codec (e.g., a V-PCC codec).
[0138] According to one aspect of this disclosure, in order to effectively compress 3D mesh information, 2D maps such as geometry maps, texture maps (also referred to as attribute maps in some examples), occupancy maps, etc., can be downsampled before encoding.
[0139] Figure 10 A schematic diagram of downsampling in some examples is shown. Figure 10 In this process, the map (1020) is downsampled by a factor of 2 in both the horizontal and vertical directions, and a downsampled map (1030) is generated accordingly. The width (e.g., the number of pixels in the horizontal direction) of the downsampled map (1030) is half the width (e.g., the number of pixels in the horizontal direction) of the map (1020), and the height (e.g., the number of pixels in the vertical direction) of the downsampled map (1030) is half the height (e.g., the number of pixels in the vertical direction) of the map (1020).
[0140] exist Figure 10In the map (1020), there are 2D shapes (also known as UV sheets) A', B', C', and D', and the downsampled map (1030) includes sampled 2D shapes A”, B”, C”, and D” corresponding to the 2D shapes A', B', C', and D', respectively. In some examples, the downsampled map (1030) is then encoded by an image or video encoder on the grid encoder side.
[0141] In some examples, the downsampled image is decoded on the mesh decoder side. After decoding, the downsampled image is restored to its original resolution (e.g., the original number of pixels in the vertical direction and the original number of pixels in the horizontal direction) for use in reconstructing the 3D mesh.
[0142] Dynamic mesh sequences typically require large amounts of data because they can consist of a vast amount of information that changes over time. Sampling steps applied to 2D maps (e.g., UV atlases, attribute maps) can help reduce the bandwidth required to represent mesh information. However, sampling steps can also remove critical information during downsampling, such as some essential geometries of 3D meshes.
[0143] In some examples, adaptive sampling techniques can be used to process 2D atlases (also known as maps in 2D) without losing much important information. Adaptive sampling techniques can be used for both static mesh compression (where a single mesh frame or mesh content does not change over time) and dynamic mesh compression. Various adaptive sampling techniques can be applied individually or in any combination. In the following description, adaptive sampling methods are applied to 2D atlases (e.g., maps in 2D), which can be one or both of a geometry graph or an attribute (texture) graph.
[0144] Figure 11 A schematic diagram of a frame (1100) for mesh compression according to some embodiments of the present disclosure is shown. The frame (1100) includes a mesh encoder (1110) and a mesh decoder (1150). The mesh encoder (1110) receives an input mesh (1101) (a mesh frame in the case of dynamic mesh processing) and encodes the input mesh (1101) into a bitstream (1145), and the mesh decoder (1150) decodes the bitstream (1145) to generate a reconstructed mesh (1195) (a reconstructed mesh frame in the case of dynamic mesh processing).
[0145] The mesh encoder (1110) can be any suitable device, such as a computer, server computer, desktop computer, laptop computer, tablet computer, smartphone, gaming device, AR device, VR device, etc. The mesh decoder (1150) can be any suitable device, such as a computer, client computer, desktop computer, laptop computer, tablet computer, smartphone, gaming device, AR device, VR device, etc. The bitstream (1145) can be transmitted from the mesh encoder (1110) to the mesh decoder (1150) via any suitable communication network (not shown).
[0146] exist Figure 11 In the example, the mesh encoder (1110) includes a preprocessing module (1111), an adaptive sampling module (1120), a video encoder (1130), and an auxiliary data encoder (1140) coupled together. The video encoder (1130) is configured to encode image or video data (e.g., a 2D map in a representation of a 3D mesh).
[0147] exist Figure 11 In the example, the preprocessing module (1111) is configured to perform appropriate operations on the input mesh (1101) to generate a mesh (1105) with a UV atlas. For example, the preprocessing module (1111) can perform a series of operations, including tracing, remeshing, parameterization, and voxelization. Figure 11 In the examples, this series of operations is only an operation of the encoder, not part of the decoding process. In some examples, the mesh (1105) with a UV atlas includes 3D position information of the vertices, a UV atlas that maps the 3D position information to 2D, and other 2D attribute maps (e.g., 2D color maps, etc.).
[0148] It should be noted that in some examples, the input mesh (1101) is in the form of a mesh with a UV atlas, and then the preprocessing module (1111) can forward the input mesh (1101) into a mesh (1105) with a UV atlas.
[0149] An adaptive sampling module (1120) receives a grid (1105) with a UV atlas and performs adaptive sampling to generate an adaptively sampled map (1125). In some examples, the adaptive sampling module (1120) may use various techniques to detect characteristics in the map or different regions of the map, such as information density in the map, and determine different sampling rates for sampling the map or different regions of the map based on these characteristics. The 2D map can then be sampled according to the different sampling rates to generate the adaptively sampled map (1125). The adaptively sampled map (1125) may include a geometry map (also referred to as a geometric image in some examples), an occupancy map, other attribute maps (e.g., a color map), etc.
[0150] The video encoder (1130) can use image encoding and / or video encoding techniques to encode the adaptively sampled map (1125) into a bitstream (1145).
[0151] The adaptive sampling module (1120) also generates auxiliary data (1127) that indicates auxiliary information for adaptive sampling. The auxiliary data encoder (1140) receives the auxiliary data (1127) and encodes the auxiliary data (1127) into a bit stream (1145).
[0152] The operation of the adaptive sampling module (1120) and the auxiliary data encoder (1140) will be further described in this disclosure.
[0153] exist Figure 11 In the example, the bitstream (1145) is provided to the grid decoder (1150). The grid decoder (1150) includes, for example, Figure 11 The video decoder (1160), auxiliary data decoder (1170), and mesh reconstruction module (1190) are coupled together. In one example, the video decoder (1160) corresponds to the video encoder (1130) and can decode a portion of the bitstream (1145) encoded by the video encoder (1130) and generate a decoded map (1165). In some examples, the decoded map (1165) includes a decoded UV map, one or more decoded attribute maps, etc. In some examples, the decoded map (1165) includes a decoded occupancy map (e.g., an initial decoded map).
[0154] exist Figure 11 In the example, the auxiliary data decoder (1170) corresponds to the auxiliary data encoder (1140) and can decode a portion of the bitstream (1145) encoded by the auxiliary data encoder (1140) and generate decoded auxiliary data (1175).
[0155] exist Figure 11In the example, a decoded map (1165) and decoded auxiliary data (1175) are provided to the mesh reconstruction module (1190). The mesh reconstruction module (1190) generates a reconstructed mesh (1195) based on the decoded map (1165) and decoded auxiliary data (1175). In some examples, the mesh reconstruction module (1190) can determine the vertices and vertex information in the reconstructed mesh (1195), such as the individual 3D coordinates, UV coordinates, colors, etc. associated with the vertices. The operation of the auxiliary data decoder (1170) and the mesh reconstruction module (1190) will be further described in this disclosure.
[0156] It should be noted that the components in the trellis encoder (1110), such as the preprocessing module (1111), the adaptive sampling module (1120), the video encoder (1130), and the auxiliary data encoder (1140), can be implemented using various techniques. In one example, the components are implemented using integrated circuits. In another example, the components are implemented using software that can be executed by one or more processors.
[0157] It should be noted that the components in the mesh decoder (1150), such as the video decoder (1160), the auxiliary data decoder (1170), and the mesh reconstruction module (1190), can be implemented using various techniques. In one example, the components are implemented using integrated circuits. In another example, the components are implemented using software that can be executed by one or more processors.
[0158] In some embodiments, adaptive sampling can be based on map type. In some examples, the adaptive sampling module (1120) can apply different sampling rates to different types of maps. For example, different sampling rates can be applied to geometry maps and attribute maps. In one example, a grid is a model of objects with regular shapes and rich textures. For example, the objects have rectangular shapes and rich colors. Therefore, the information density of the geometry map is relatively low. In one example, the adaptive sampling module (1120) applies a first sampling rate of 2:1 to the geometry map (in both the vertical and horizontal directions) and a second sampling rate of 1:1 to the texture map (in both the vertical and horizontal directions).
[0159] In some examples, an A:B sampling rate in one direction indicates that B samples are generated from A pixels in the original map in that direction. For example, a 2:1 sampling rate in the horizontal direction indicates that one sample is generated for every two pixels in the original map in the horizontal direction. A 2:1 sampling rate in the vertical direction indicates that one sample is generated for every two pixels in the original map in the vertical direction.
[0160] In some examples, the term "sampling step size" is used. The sampling step size in one direction indicates the number of pixels between two adjacent sampling positions in that direction. For example, a sampling step size of 2 in the horizontal direction indicates two pixels between adjacent sampling positions in the horizontal direction; a sampling step size of 2 in the vertical direction indicates two pixels between adjacent sampling positions in the vertical direction. It should be noted that in this disclosure, the sampling rate is equivalent to the sampling step size. For example, a sampling rate of 2 (e.g., 2:1) corresponds to two pixels between adjacent sampling positions.
[0161] In some embodiments, sampling adaptation is based on sub-regions in the map. Different sampling rates can be applied to different parts of the map. In some examples, some pixel rows have less information to retain, so a larger sampling rate can be applied along these rows, resulting in fewer sample rows to be encoded. In some examples, some pixel columns have less information to retain, so a larger sampling rate can be applied along these columns, resulting in fewer sample columns to be encoded. For other regions, a smaller sampling rate is applied to minimize information loss after sampling.
[0162] Figure 12 A schematic diagram of adaptive sampling in some examples is shown. The map (1220) is divided into several block rows, each block row comprising a fixed number of sample (pixel) rows. Different sampling rates are applied to the block rows in the vertical direction to generate an adaptive sampled map (1230). For example, each block row is a CTU row (also referred to as a CTU line) and includes 64 rows of samples (also referred to as pixels). Figure 12 In the example, for block rows 0 and 6 in map (1220), a first sampling rate of 2:1 is applied in the vertical direction, resulting in 32 rows of samples for each block row in the adaptive sampled map (1230) after sampling. For block rows 1 to 5 in map (1220), a second sampling rate of 1:1 is applied in the vertical direction, resulting in 64 rows of samples for each block row in the adaptive sampled map (1230).
[0163] It is important to note that, in Figure 12 In the middle, a 1:1 sampling rate is applied in the horizontal direction.
[0164] In some examples, the adaptive sampled map (1230) is then encoded by an image or video encoder (such as a video encoder (1130)). On the decoder side, in one example, the adaptive sampled map (1230) is decoded. After decoding, the top 32 rows of samples are restored (upsampled) to the original resolution, such as 64 rows of samples; and the bottom 32 rows of samples are restored (upsampled) to the original resolution, such as 64 rows of samples.
[0165] In some other examples, the map to be encoded in a 2D representation of a 3D mesh can be divided into multiple sub-regions. Examples of such division within a map (e.g., an image) include slices, tiles, tile groups, coding tree units, etc. In some examples, different sampling rates can be applied to different sub-regions. In one example, different sampling rates associated with different sub-regions can be signaled in the bitstream carrying the 3D mesh. On the decoder side, after decoding the adaptively sampled map, each sub-region is restored to its original resolution based on the sampling rate associated with it.
[0166] In some examples, the process of adapting the sampled map to its original resolution is referred to as the inverse sampling process, which generates the restored map. After restoration from the inverse sampling process, the output of the restored map in 2D atlas form can be used for 3D mesh reconstruction.
[0167] Although Figure 12 The example in the text shows adaptive sampling of different block rows in the vertical direction, but similar adaptive sampling can be applied to different columns in the horizontal direction, or similar adaptive sampling can be applied in both the vertical and horizontal directions.
[0168] In some embodiments, sampling adaptation is patch-based. In some examples, different patches in the map may have different sampling rates.
[0169] Figure 13 A schematic diagram of adaptive sampling is shown in some examples. The map (1320), such as a high-resolution 2D atlas, comprises multiple 2D shapes, also referred to as UV sheets corresponding to patches in a 3D mesh, such as a first 2D shape A' and a second 2D shape B'. Figure 13 In the example, a first sampling rate of 2:1 is applied to the first 2D shape A' in both the vertical and horizontal directions to generate a first sampled 2D shape A”; and a second sampling rate of 1:1 is applied to the second 2D shape B' in both the vertical and horizontal directions to generate a second sampled 2D shape B”. The first sampled 2D shape A” and the second sampled 2D shape B” are placed in a new map called the adaptive sampled map (1330).
[0170] exist Figure 13In the example, the first sampled 2D shape A” is smaller than the first 2D shape A', and the second sampled 2D shape B” has the same size as the second 2D shape B'. The adaptive sampled map (1330) is encoded into a bitstream carrying a 3D mesh by an image or video encoder such as a video encoder (1130). In some examples, the sampling rate associated with the sampled 2D shape is encoded into a bitstream carrying a 3D mesh, for example by an auxiliary data encoder (1140).
[0171] In some examples, on the decoder side, an image / video decoder (such as a video decoder (1160)) decodes an initial map, such as an adaptive sampled map (1330), from the bitstream. Additionally, a sampling rate associated with the sampled 2D shape is decoded from the bitstream, for example, by an auxiliary data decoder (1170). Based on the sampling rate associated with the sampled 2D shape, the sampled 2D shape in the adaptive sampled map (1330) is restored to its original size (e.g., the same number of pixels in the vertical and horizontal directions) to generate a restored map. The restored map is then used for 3D mesh reconstruction.
[0172] According to one aspect of this disclosure, adaptive sampling information, such as sampling rates for different map types, sampling rates for different sub-regions, and sampling rates for different patches, is known on both the grid encoder and grid decoder sides. In some examples, the adaptive sampling information is appropriately encoded into a bitstream carrying a 3D grid. Therefore, the grid decoder and grid encoder can operate based on the same adaptive sampling information. The grid decoder can restore the map to the correct size.
[0173] According to one aspect of this disclosure, a mesh encoder, such as a mesh encoder (1110), can perform 2D atlas sampling (also known as UV atlas sampling). For example, an adaptive sampling module (1120) can receive a mesh (1105) having a UV atlas. Each vertex of the mesh (1105) having a UV atlas has a corresponding point in the UV atlas, and the position of the corresponding point in the UV atlas is specified by UV coordinates. In the UV atlas, the corresponding point may have a vector value including the 3D coordinates of the vertex in 3D space (e.g., (x, y, z)). Furthermore, the mesh (1105) having a UV atlas includes one or more attribute maps that store attribute values associated with vertices as attribute values of pixels at positions specified by UV coordinates in the one or more attribute maps. For example, a color map may store the color of a vertex as the color of a pixel at a position specified by UV coordinates in the color map.
[0174] The adaptive sampling module (1120) can apply adaptive sampling techniques to a mesh with a UV atlas (1105) and output an adaptive sampled map (1125) (also referred to as an adaptive sampled atlas in some examples), which may include, for example, a sampled geometry (also referred to as a sampled UV map or sampled UV atlas), one or more sampled attribute maps (also referred to as texture maps in some examples), etc. In some examples, the adaptive sampled map (1125) includes a sampled occupancy map.
[0175] According to one aspect of this disclosure, the same sampling rate configuration can be applied to various maps, such as geometric maps, attribute maps, occupancy maps, etc., to generate an adaptive sampled map. In some examples, the adaptive sampling module (1120) can generate an adaptive sampled map (1125) by sampling on a UV atlas (e.g., based on sampling locations on the UV atlas). The adaptive sampling module (1120) can determine the sampling locations on the UV atlas and then generate an adaptive sampled map based on the sampling locations on the UV atlas. For example, after determining the sampling locations on the UV atlas, for example, based on the sampling rate, the adaptive sampling module (1120) determines the location of the sampled points in each sampled map (1125) and then determines the value of each sampled point in the sampled map (1125).
[0176] In one example, when a sampling location on the UV atlas is inside a polygon defined by the vertices of the mesh, the sampling location is occupied, and the sampled point corresponding to the sampling location in the sampled-occupied map is set to occupied (e.g., the value of the sampled point in the sampled-occupied map is "1"). However, when the sampling location is not inside any polygon defined by the vertices of the mesh, the sampling location is not occupied, and the sampled point corresponding to the sampling location in the sampled-occupied map is set to unoccupied (e.g., the value of the sampled point in the sampled-occupied map is "0").
[0177] For each occupied sampling location on the UV atlas, the adaptive sampling module (1120) can determine the 3D (geometric) coordinates at that occupied sampling location and assign the determined 3D coordinates to the vector values of the corresponding sampled points in the sampled geometry (also known as the sampled UV atlas). Similarly, the adaptive sampling module (1120) can determine the attribute values (e.g., color, normal, etc.) at the occupied sampling location and assign the determined attribute values to the attribute values of the corresponding sampled points in the sampled attribute map. In some examples, the adaptive sampling module (1120) can determine the 3D coordinates and attributes of the occupied sampling location by interpolation from the associated polygon vertices.
[0178] In one example, the mesh is formed by triangles. The sampling position is inside the triangle defined by the three vertices of the mesh, therefore the sampling position is an occupied sampling position. The adaptive sampling module (1120) can determine the 3D (geometric) coordinates at the occupied sampling position based on, for example, the weighted average 3D coordinates of the 3D coordinates of the three vertices of the triangle. The adaptive sampling module (1120) can assign the weighted average 3D coordinates to the vector values of the corresponding sampled points in the sampled geometry. Similarly, the adaptive sampling module (1120) can determine the attribute values at the occupied sampling position based on, for example, the weighted average attribute values of the attributes of the three vertices of the triangle (e.g., weighted average color, weighted average normal, etc.). The adaptive sampling module (1120) can assign the weighted average attribute values to the attribute values of the corresponding sampled points in the sampled attribute graph.
[0179] In some examples, the sampling rate (SR) can be consistent across the entire 2D atlas (e.g., geometry, property graph, etc.), but the sampling rates for the u-axis and v-axis can differ. Using different sampling rates on the u-axis and v-axis makes anisotropic remeshing possible. (See reference...) Figure 12 and Figure 13 As described, in some examples, a 2D atlas can be divided into multiple regions, such as slices, tiles, or patches, and each region can have its own sampling rate. For example, a mesh consists of connected triangles, and the mesh can be divided into several patches, each patch comprising a subset of the entire mesh. Different sampling rates can be applied to the individual patches, for example, via an adaptive sampling module (1120).
[0180] According to one aspect of this disclosure, a grid encoder, such as a grid encoder (1110), can determine a suitable sampling rate for each region in a 2D atlas.
[0181] In some examples, the adaptive sampling module (1120) can determine the sampling rate distribution on a 2D atlas (e.g., a geometry graph, attribute graph, etc.). For example, the adaptive sampling module (1120) can determine a specific sampling rate (SR) for a region based on its characteristics. In one example, the specific sampling rate is determined based on the region's spectrum. For example, a textured region (or the entire 2D atlas) may have high spatial frequency components in its texture attribute values, and the adaptive sampling module (1120) can assign a sampling rate (e.g., a lower sampling rate, a lower sampling step size) suitable for the high spatial frequency components to the textured region. In another example, a region with high activity (or the entire 2D atlas) may include high spatial frequency components in its coordinates (e.g., 3D coordinates, UV coordinates), and the adaptive sampling module (1120) can assign a sampling rate (e.g., a lower sampling rate, a lower sampling step size) suitable for the high activity to the region. In another example, a smooth region (or the entire 2D atlas) may lack high spatial frequency components in its texture attribute values, and the adaptive sampling module (1120) can assign a sampling rate suitable for the smooth region (e.g., a higher sampling rate, a higher sampling step size). In another example, a region with low activity (or the entire 2D atlas) may lack high spatial frequency components in its coordinates (e.g., 3D coordinates, UV coordinates), and the adaptive sampling module (1120) can assign a sampling rate suitable for the low activity in that region (e.g., a higher sampling rate, a lower sampling step size).
[0182] According to one aspect of this disclosure, the sampling rate can be represented by an oversampling ratio (OR) parameter. The OR parameter is defined as the ratio between the number of sampled points in a region and the original number of vertices in that region. When the OR parameter of a region is greater than 1, the OR parameter indicates that the region is oversampled compared to the original number of vertices; when the OR parameter of a region is less than 1, the OR parameter indicates that the region is undersampled compared to the original number of vertices. For example, a region of a mesh consists of 1000 vertices, and when a specific sampling rate (SR) is applied to that region, 3000 sampled points will be obtained. Then, both the OR parameter and SR for that region are equal to 3, i.e., OR(SR) = 3.
[0183] In some embodiments, the adaptive sampling module (1120) can use an algorithm to determine the final sampling rate of a region to achieve an OR parameter that is closest to a predefined target OR parameter. For example, for a specific region i (e.g., i is a region index used to identify the specific region), a target OR is defined for that specific region i (by TOR). i (Representation) parameters. The adaptive sampling module (1120) can adjust the final sampling rate (determined by SR) for a specific region i. i(Indicates) is determined to generate OR (TOR) with respect to the target. i The sampling rate of the OR parameter that is closest to the OR parameter, such as that expressed by formula (1).
[0184] SR i =argmin SR |OR(SR)-TOR i | Formula (1)
[0185] In one example, the adaptive sampling module (1120) can try multiple sampling rates and select one of them that produces the OR parameter closest to the target OR parameter. In another example, the adaptive sampling module (1120) can use an algorithm such as a binary search algorithm to perform a search within a range of sampling rates to determine the final sampling rate that produces the OR parameter closest to the target OR parameter.
[0186] In some embodiments, the adaptive sampling module (1120) can use an algorithm to determine the final sampling rate of a specific region i (e.g., i is a region index used to identify the specific region), which can achieve a maximum OR parameter less than a predetermined threshold (denoted by Th0). In some examples, the algorithm starts with a relatively small base sampling rate (BSR) (e.g., 1:1) and uses an iterative process with iteration cycles to determine the final sampling rate. In each iteration cycle that tests the current BSR, the OR parameter of the region is determined. When the OR parameter is less than the threshold Th0, the current BSR is the final sampling rate of the specific region; and when the OR parameter is greater than the threshold Th0, a new BSR is calculated based on the current BSR. For example, a scaling factor F0 (e.g., greater than 1) is applied to the current BSR to determine the new BSR. The new BSR then becomes the current BSR, and the iteration process proceeds to the next iteration cycle. In some examples, the process for determining the final sampling rate of region (i) is defined as Equation (2):
[0187]
[0188] Among them SR i Here, Th0 is the final sampling rate, F0 is the threshold for the OR parameter, and F0 > 1 is the scaling factor used to increase the BSR. Therefore, when the OR parameter with BSR is less than the threshold Th0, the current region will only use BSR as the final sampling rate; otherwise, the final sampling rate will be changed according to the scaling factor F0. This process can be iterative until OR(SR) is reached. i Less than the threshold Th o .
[0189] In some embodiments, the adaptive sampling module (1120) can use an algorithm to determine the final sampling rate (SR) for a specific region i (e.g., i is a region index used to identify the specific region). i The final sampling rate can achieve an OR parameter within a specific range, such as less than a first predetermined threshold (denoted by Th0) and greater than a second predetermined threshold (denoted by Th1). In some examples, the algorithm starts with an arbitrary base sampling rate (BSR) and uses an iterative process with iteration cycles to determine the final sampling rate. In each iteration cycle testing the current BSR, the OR parameter for a specific region is determined. When the OR parameter is within a specific range, such as less than the first predetermined threshold Th0 and greater than the second predetermined threshold Th1, the current BSR is the final sampling rate for that region. However, when the OR parameter is greater than the first predefined threshold Th0, a first scaling factor F0 (e.g., greater than 1) is applied to the current BSR to determine a new BSR; when the OR parameter is less than the second predefined threshold Th1, a second scaling factor F1 (e.g., less than 1) is applied to the current BSR to determine a new BSR. The new BSR then becomes the current BSR, and the iteration process proceeds to the next iteration cycle.
[0190] In some examples, the process of determining the final sampling step size for region i is defined by formula (3):
[0191]
[0192] Wherein, BSR represents the basic sampling rate, Th0 represents the first predefined threshold, Th1 represents the second predefined threshold, F0>1 represents the first scaling factor used to increase BSR, and 0<F1<1 represents the second scaling factor used to decrease BSR.
[0193] It is important to note that scaling factors F0 and F1 can be different for each region. When the OR parameter with BSR is within a specific range defined by thresholds (e.g., Th0 and Th1), the current region can use BSR as the final sampling rate. When the OR parameter with BSR is equal to or greater than Th0, the sampling rate is lower. o When the final sampling rate increases by a scaling factor F0, and when the OR parameter with BSR is equal to or less than Th1, the final sampling rate decreases by a scaling factor F1. This process can be performed iteratively until OR(SR) is reached. i Within the range defined by Th0 and Th1.
[0194] According to some aspects of this disclosure, the adaptive sampling module (1120) can place regions with different sampling rates (or different sampling steps) in a single map.
[0195] It is important to note that when applying an adaptive sampling rate, the size of the sampled regions in the sampled map can vary proportionally compared to the original UV atlas or to the case of using a uniform sampling rate. The adaptive sampling module (1120) can place sampled regions with different sampling rates in the sampled map while ensuring that the sampled regions do not overlap, thus ensuring a one-to-one correspondence between each point in the sampled map and a specific region.
[0196] In some examples, the adaptive sampling module (1120) can determine a bounding box for each sampled region and then place the individual sampled regions according to that bounding box. In one example, for a specific region in the original UV atlas, u min It is the minimum u-coordinate of all vertices in a specific region, v min It is the minimum v-coordinate of all vertices in a specific region. Based on the minimum u-coordinate and the minimum v-coordinate, the adaptive sampling module (1120) can determine the bounding box for the sampled region corresponding to the specific region. For example, the top-left corner of the bounding box of the sampled region can be placed at position (u) in the sampled map. o ,v o At this location, it can be accessed via... and To calculate, where SR represents the sampling rate applied to a specific region when the same sampling rate is used in both the u and v directions. This represents the floor function that determines the smallest integer greater than the value C.
[0197] In some embodiments, the adaptive sampling module (1120) can place the sampled regions one after another into the sampled map in a certain order. To place the currently sampled region, the adaptive sampling module (1120) can first determine its position (u) based on the region's position. o ,v o The current sampled region is placed at the top left corner of the bounding box of the currently sampled region (e.g., the top left corner of the bounding box of the currently sampled region). When the adaptive sampling module (1120) detects that the current sampled region overlaps with an already placed sampled region, the adaptive sampling module (1120) can determine a new position to place the current sampled region to avoid overlapping with the previously placed sampled region.
[0198] In some examples, specific search windows and / or criteria can be defined, and the adaptive sampling module (1120) can find a new location for placing the currently sampled region based on the specific search window and / or criteria. n ,v n (This is represented by a symbol). It's important to note that in some examples, the new position (u) can be used. n ,v n ) and original position (uo ,v o The offset between the encoder and decoder (also known as the UV offset) is sent from the encoder side to the decoder side by a signal for reconstruction.
[0199] In some embodiments, the sampled regions are placed so that they not only do not overlap, but also have a certain amount of spacing between them. For example, each sampled region may need to be at least 10 pixels away from other sampled regions. It should be noted that the spacing between sampled regions can be defined using various techniques. In some examples, the minimum distance between sampled regions can be defined by a minimum horizontal distance l0 and a minimum vertical distance l1.
[0200] Some aspects of this disclosure also provide signaling techniques for adaptive sampling of grid compression.
[0201] In some embodiments, the sampling rate of different regions of the grid can be signaled in the bitstream carrying grid information. It is important to note that the sampling rate can be signaled at different levels within the bitstream. In one example, the sampling rate can be signaled in the sequence header of the entire grid sequence, including the grid frame sequence. In another example, the sampling rate can be signaled in the group header of a group of grid frames (i.e., conceptually similar to a group of pictures (GOP)). In another example, the sampling rate can be signaled in the frame header of each grid frame. In another example, the sampling rate of the stripes in the grid frame is signaled in the slice header of a slice. In another example, the sampling rate of the tiles in the grid frame is signaled in the tile header of a tile. In yet another example, the sampling rate of the pieces in the grid frame is signaled in the piece header of a patch.
[0202] Specifically, in some embodiments, signaled control flags can be used to indicate whether the adaptive sampling method is applied to different levels in the bitstream. In one example, the control flags are signaled in the sequence header of the entire grid sequence. In another example, the control flags are signaled in the group header of a set of grid frames. In yet another example, the control flags are signaled in the frame header of each grid frame. In another example, the control flags for the stripes in the grid frame are signaled in the strip header of the stripe. In yet another example, the control flags for the tiles in the grid frame are signaled in the tile header of the tile. In yet another example, the control flags for the pieces in the grid frame are signaled in the piece header of the piece. When the control flag for a level is true, adaptive sampling is enabled at that level, allowing an adaptive sampling rate to be applied. When the control flag for a level is false, adaptive sampling is disabled, applying a uniform sampling rate to that level.
[0203] In some examples, the control flag consists of 1 bit and can be encoded using various techniques. In one example, the control flag is encoded using entropy coding with fixed or updated probabilities (e.g., arithmetic coding and Huffman coding). In another example, the control flag is encoded using a lower-complexity encoding technique (known as bypass coding).
[0204] In some examples, the base sample rate can be signaled regardless of whether adaptive sampling is enabled. When adaptive sampling is enabled, the base sample rate can be used as a predictor, and each region can be signaled with the difference between the base sample rate and the base sample rate to indicate the actual sample rate for that region. When adaptive sampling is disabled, the base sample rate can be used as a uniform sample rate across the entire content at the appropriate level.
[0205] The basic sampling rate can also be signaled at different levels within the bitstream. In one example, the basic sampling rate can be signaled in the sequence header of the entire grid sequence. In another example, the basic sampling rate can be signaled in the group header of a set of grid frames. In yet another example, the basic sampling rate can be signaled in the frame header of each grid frame. In yet another example, the basic sampling rate of the stripes in a grid frame is signaled in the strip header of a stripe. In yet another example, the basic sampling rate of the tiles in a grid frame is signaled in the tile header of a tile. In yet another example, the basic sampling rate of the pieces in a grid frame is signaled in the piece header of a piece.
[0206] In some examples, the basic sampling rate can be binarized by a fixed-length or variable-length representation (e.g., a fixed k-bit representation and a k-order Exp-Golomb representation), and each bit can be encoded by entropy coding with fixed or updated probabilities (e.g., arithmetic coding and Huffman coding), or by bypass coding with lower complexity.
[0207] According to one aspect of this disclosure, when adaptive sampling is enabled, the sampling rate of regions in a grid frame is appropriately signaled. In some examples, the number of regions in the entire grid frame can be signaled or derived (e.g., as the number of CTU rows, the number of tiles, the number of slices, etc.).
[0208] According to one aspect of this disclosure, the sampling rate of a region can be signaled without prediction. In one example, the sampling rate of each region (or the entire 2D atlas) can be signaled directly without any prediction. In another example, the sampling rate of each region (or the entire 2D atlas) can be selected from a pre-established set of sampling rates known to both the encoder and decoder. Signaling a specific sampling rate can be performed by signaling an index of a specific sampling rate from a pre-established set of rates. For example, the pre-established set of sampling strides may include (every 2 pixels, every 4 pixels, every 8 pixels, etc.). Index 1 can be signaled to indicate a sampling rate of every 2 pixels (e.g., 2:1); index 2 can be signaled to indicate a sampling rate of every 4 pixels (e.g., 4:1); and index 3 can be signaled to indicate a sampling rate of every 8 pixels (e.g., 8:1).
[0209] According to another aspect of this disclosure, the sampling rate of a region can be predicted. It should be noted that any suitable prediction technique can be used.
[0210] In one example, the sampling rate for each region (or the entire 2D atlas) of a grid frame can be predicted based on a pre-established set of rates. In another example, the sampling rate for each region (or the entire 2D atlas) of a grid frame can be predicted based on sampling rates previously used in other already encoded regions of the same frame. In yet another example, the sampling rate for each region (or the entire 2D atlas) of a grid frame can be predicted based on sampling rates previously used in other already encoded grid frames.
[0211] According to another aspect of this disclosure, the sampling rate for each region (or the entire 2D atlas) can be determined in a manner that allows both prediction and direct signaling. In one example, the syntax can be organized to indicate whether the sampling rate will be predicted or directly signaled. When the syntax indicates predicting the sampling rate, for example, another syntax further signals which predictor to use to predict the sampling rate. When the syntax indicates directly signaling the sampling rate, for example, another syntax signals the value of the sampling rate.
[0212] In some examples, when the sampling rate is sent directly as a signal (either by sending the sampling rate as a signal or by sending an index pointing to the sampling rate as a signal), the sampling rate or the index pointing to the sampling rate can be binarized by a fixed-length or variable-length representation (e.g., a fixed k-bit representation and a k-order Exp-Golomb representation), and each bit can be encoded by entropy coding with fixed or updated probabilities (e.g., arithmetic coding and Huffman coding), or by a bypass coding with lower complexity.
[0213] In some examples, the prediction residual can be sent by signaling when the sampling rate is predicted. The prediction residual can be binarized using a fixed-length or variable-length representation (e.g., a fixed k-bit representation and a k-order Exp-Golomb representation), and each bit can be encoded using entropy coding with fixed or updated probabilities (e.g., arithmetic coding and Huffman coding), or using a bypass coding with lower complexity. For example, the sign bit of the prediction residual can be encoded using bypass coding, and the absolute value of the prediction residual can be encoded using entropy coding with updated probabilities.
[0214] According to some aspects of this disclosure, when adaptive sampling is enabled, the offset of the UV coordinates (also known as the UV offset) of each region in the grid frame can be encoded in the bitstream carrying the grid frame. u =u o -u n and offset v =v o -v n In some examples, the number of regions in the entire grid frame can be sent or derived using signals (e.g., the number of CTU rows, the number of tiles, the number of slices, etc.).
[0215] According to one aspect of this disclosure, the UV offset of a region can be transmitted by signaling without prediction. In one example, the UV offset of each region can be transmitted directly by signaling without any prediction.
[0216] According to another aspect of this disclosure, the UV offset of a region can be predicted. It should be noted that any suitable prediction technique can be used. In one example, the UV offset of each region is predicted based on a pre-established set of UV offsets. In another example, the UV offset of each region is predicted based on UV offsets previously used in other encoded regions of the same grid frame. In yet another example, the UV offset of each region is predicted based on UV offsets previously used in other encoded grid frames.
[0217] According to another aspect of this disclosure, the UV offset for each region can be signaled in a manner that allows both prediction and direct signaling. In some examples, syntax can be organized to indicate whether the UV offset will be predicted or directly signaled. When the syntax indicates that the UV offset is predicted, another syntax further signals which predictor is used to predict the UV offset. When the syntax indicates that the UV offset is directly signaled, another syntax signals the value of the UV offset.
[0218] In some examples, the UV offset is sent directly as a signal. The UV offset value can be binarized using a fixed-length or variable-length representation (e.g., a fixed k-bit representation and a k-order Exp-Golomb representation), and each bit can be encoded using entropy coding with fixed or updated probabilities (e.g., arithmetic coding and Huffman coding), or using a bypass coding with lower complexity. In one example, the sign bit of the UV offset can be encoded using bypass coding, and the absolute value of the UV offset can be encoded using entropy coding with updated probabilities.
[0219] In some examples, the sign bit of the UV offset can be inferred or predicted based on the sampling rate value. For instance, when the sampling rate of a region is greater than the base sampling rate, the sign bit of the UV offset can be inferred or predicted as positive; and when the sampling rate of the region is less than the base sampling rate, the sign bit of the UV offset can be inferred or predicted as negative. If the sampling rate of the region is equal to the base sampling rate, the UV offset can be inferred or predicted as zero.
[0220] In some examples, the UV offset can be predicted, and the prediction residual can be sent as a signal. For instance, the prediction residual can be binarized using a fixed-length or variable-length representation (e.g., a fixed k-bit representation and a k-order Exp-Golomb representation), and each bit can be encoded using entropy coding with fixed or updated probabilities (e.g., arithmetic coding and Huffman coding), or using a bypass coding with lower complexity. In one example, the sign bit of the prediction residual can be encoded using bypass coding, and the absolute value of the prediction residual can be encoded using entropy coding with updated probabilities.
[0221] Some aspects of this disclosure also provide mesh reconstruction techniques for the decoder side. These mesh reconstruction techniques can be used in mesh reconstruction modules, such as mesh reconstruction module (1190).
[0222] In some examples, the decoded map (1165) includes a decoded occupancy map, and the mesh reconstruction module (1190) can reconstruct the mesh frame based on the decoded map (1165) which includes the decoded occupancy map, the decoded geometry, and one or more decoded attribute maps. It should be noted that in some examples, the decoded occupancy map corresponds to a sampled occupancy map, which may be the result of adaptive sampling using the same sampling rate as the decoded geometry and the decoded attribute map; therefore, the decoded occupancy map includes sampled regions.
[0223] In some examples, the decoded map (1165) does not include an occupancy map, and the bitstream (1145) includes information identifying the boundary vertices of each region, as auxiliary data encoded, for example, by an auxiliary data encoder (1140). An auxiliary data decoder (1170) can decode the information identifying the boundary vertices of the regions. A mesh reconstruction module (1190) can infer regions for the inferred occupancy map based on the boundary vertices of the regions. It should be noted that in some examples, the inferred occupancy map has not yet been processed with adaptive sampling. The mesh reconstruction module (1190) can reconstruct a mesh frame based on the inferred occupancy map and the decoded map (1165), which includes a decoded geometry and one or more decoded attribute maps.
[0224] According to one aspect of this disclosure, the mesh reconstruction module (1190) can determine the UV coordinates of vertices in the mesh frame based on the decoded map (1165) and decoded auxiliary data (e.g., sampling rate, UV offset, information identifying boundary vertices, etc.), and determine the 3D coordinates and attribute values of vertices in the mesh frame.
[0225] In some embodiments, to determine the UV coordinates of vertices in a region, the region's sampling rate (SR) is determined based on the syntax values decoded from the bitstream (1145). In some examples, the region's UV offset, such as (offset), is determined based on the syntax values decoded from the bitstream (1145). u offset v ).
[0226] In one example, the decoded map (1165) includes a decoded occupancy map. The mesh reconstruction module (1190) can determine the UV coordinates of the vertices corresponding to each occupancy point in the sampled region of the decoded occupancy map. For example, for a sampled region with coordinates (u... i v i For each occupied point, the mesh reconstruction module (1190) can determine the UV coordinates (U coordinates) of the vertex corresponding to that occupied point according to formulas (4) and (5). i V i ):
[0227] U i =(u i +offset u )×SR formula (4)
[0228] V i =(v i +offset v )×SR formula (5)
[0229] In another example, the decoded map (1165) does not include the decoded occupancy map. The mesh reconstruction module (1190) can determine the inferred region in the inferred occupancy map based on the boundary vertices of the region. The UV coordinates of the vertices corresponding to the occupancy points in the inferred region can be directly inferred from the positions of the occupancy points in the region of the inferred occupancy map. For example, the occupancy points in the inferred occupancy map have specific UV coordinates (U... s V s The location defined by ) and specific UV coordinates (U s V s ) are the UV coordinates of the corresponding vertex in the mesh frame.
[0230] In some embodiments, for each occupancy point on the occupancy map (e.g., decoded occupancy map, inferred occupancy map), the mesh reconstruction module (1190) can recover the vertices on the mesh frame and determine the corresponding geometric values (e.g., 3D coordinates) and attribute values based on the corresponding positions in the decoded geometry map and the decoded attribute map.
[0231] In some embodiments, to derive the corresponding positions of vertices in a region in the geometry and attribute maps, the sampling rate (SR) of the region is determined based on the syntax values decoded from the bitstream (1145). In some examples, the UV offset of the region, such as (offset) is determined based on the syntax values decoded from the bitstream (1145). u offset v ).
[0232] In one example, the decoded map (1165) includes a decoded occupancy map. For the sampled regions of the decoded occupancy map, coordinates (u) are given. i ,v i Each occupied point can be directly obtained from (u) i ,v i This derives the corresponding locations in the decoded geometry and decoded attribute maps. For example, the sampling rate applied to the occupancy map to obtain the sampled occupancy map is the same as the sampling rate applied to the geometry (to obtain the sampled geometry) and the attribute map (to obtain the sampled attribute map). The decoded occupancy map corresponds to the sampled occupancy map, the decoded geometry map corresponds to the sampled geometry, and the decoded attribute map corresponds to the sampled attribute map. Therefore, an occupancy point in the decoded occupancy map can have a corresponding point with the same coordinates in both the decoded geometry and decoded attribute maps.
[0233] In another example, the decoded map (1165) does not include the decoded occupancy map. The mesh reconstruction module (1190) can determine the inferred region in the inferred occupancy map based on the boundary vertices of the region. For an inferred region with coordinates (U... i V iFor each occupied point, the corresponding position in the geometry and property graph can be derived from formulas (6) and (7):
[0234]
[0235]
[0236] In some embodiments, the mesh reconstruction module (1190) can infer connectivity information between vertices by inferring from occupied locations. In some embodiments, connectivity information can be explicitly sent as a signal in the bitstream (1145).
[0237] Figure 14 A flowchart of an overview process (1400) according to one embodiment of the present disclosure is shown. The process (1400) can be used during the mesh encoding process. In various embodiments, the process (1400) is executed by processing circuitry. In some embodiments, the process (1400) is implemented as software instructions, so that the processing circuitry executes the process (1400) when the software instructions are executed. The process begins at (S1401) and proceeds to (S1410).
[0238] At (S1410), a data structure of a mesh frame having polygons representing the surface of an object is received. The data structure of the mesh frame includes a UV atlas that associates the vertices of the mesh frame with UV coordinates in a UV atlas.
[0239] In (S1420), the sampling rate of each region of the grid frame is determined according to the characteristics of each region of the grid frame.
[0240] In (S1430), the sampling rate of each region of the grid frame is applied to the UV atlas to determine the sampling position on the UV atlas.
[0241] In (S1440), one or more sampled two-dimensional (2D) maps of grid frames are formed based on the sampling locations on the UV atlas.
[0242] In (S1450), the one or more sampled 2D maps are encoded into a bitstream.
[0243] To determine the individual sampling rates for each region of a grid frame, in some examples, a first sampling rate for the first region of the grid frame is determined based on a requirement that limits the first number of sampling locations in the first region. In one example, a first sampling rate is determined to achieve an oversampling ratio (OR) that is closest to the target value. OR is the ratio between the first number of sampling locations and the initial number of vertices in the first region of the grid frame.
[0244] In another example, the initial sampling rate is initialized to a relatively small value and adjusted until the oversampling rate (OR) is less than a first threshold. The OR is the ratio between the initial number of sampling locations and the initial number of vertices in the first region of the grid frame.
[0245] In another example, the first sampling rate is adjusted until the oversampling rate (OR) is less than a first threshold and greater than a second threshold. OR is the ratio between the first number of sampled locations and the number of vertices initially in the first region of the grid frame.
[0246] In some examples, the sampled 2D map of the one or more sampled 2D maps includes sampled regions with sampled points corresponding to sampled locations on the UV atlas. In one example, when the sampled location is within a polygon, that sampled location is determined to be occupied. Then, based on the properties of the polygon vertices, the properties of the sampled points in the sampled 2D map corresponding to the sampled location are determined.
[0247] In some examples, to form the one or more sampled 2D maps, sampled regions corresponding to each region of the grid can be determined based on their respective sampling rates. The sampled regions are then placed in a non-overlapping configuration to form the sampled map. In one example, the sampled regions are placed one after another. In another example, the bounding boxes of the sampled regions are determined, and the sampled regions are placed to avoid one or more corners of the bounding boxes overlapping with other already placed sampled regions.
[0248] For example, to place a sampling region in a non-overlapping configuration, after placing a subset of the sampled regions in a non-overlapping configuration, for the current sampled region, the initial placement position is determined based on the sampling rate of the current sampled region. Then, it is determined whether the current sampled region at the initial placement position will overlap with a subset of the sampled regions.
[0249] In some examples, an offset to the initial placement position is determined in response to the current sampled region overlapping with a subset of sampled regions at the initial placement position. This offset allows the current sampled region to not overlap with a subset of sampled regions.
[0250] In some examples, the non-overlapping configuration includes a minimum distance requirement between sampled regions.
[0251] In some embodiments, the sampling rates associated with each region are encoded using various techniques. In one example, the value of a first sampling rate for a first region is directly encoded into the bitstream. In another example, a first index is encoded into the bitstream, indicating the selection of a first sampling rate from a set of sampling rates. In yet another example, a syntax instructing a predictor is encoded into the bitstream, which predicts the first sampling rate from a pre-established set of sampling rates. In yet another example, a syntax instructing a predictor is encoded into the bitstream, which predicts the first sampling rate from previously used sampling rates for encoded regions of a grid frame. In yet another example, a syntax instructing a predictor is encoded into the bitstream, which predicts the first sampling rate from previously used sampling rates for encoded regions in another grid frame encoded prior to the grid frame.
[0252] In some embodiments, the encoder side may decide to signal or predict a first sampling rate associated with a first region, and encode a first syntax value indicating the decision into the bitstream. In one example, in response to the decision to signal the first sampling rate, the first sampling rate value is directly encoded into the bitstream. In another example, in response to the decision to signal the first sampling rate, an index is encoded into the bitstream, indicating the selection of the first sampling rate from a set of sampling rates.
[0253] In some examples, in response to this decision to predict a first sampling rate, a second syntax is encoded into the bitstream. This second syntax indicates the predictor used to predict the first sampling rate. Additionally, in one example, the prediction residual is encoded into the bitstream. This prediction residual is the difference between the first sampling rate and the predictor's sampling rate.
[0254] In some examples, the base sampling rate is encoded into the bitstream. Regardless of whether adaptive sampling is used, the base sampling rate can be encoded at any suitable level. In one example, the base sampling rate can be used as a predictor.
[0255] In some examples, control flags are encoded to indicate whether adaptive 2D atlas sampling on the grid frame is enabled or disabled.
[0256] In some examples, a first UV offset is determined that is associated with a first region of the grid frame. The first UV offset is applied to a first sampled region corresponding to the first region to avoid overlap with other sampled regions. One or more syntaxes that can indicate the first UV offset can be encoded into the bitstream. In one example, a syntax with a value for the first UV offset is directly encoded into the bitstream. In another example, a syntax indicating a predictor is encoded into the bitstream, which is used to predict the first UV offset based on a pre-established set of UV offsets. In yet another example, a syntax indicating a predictor is encoded into the bitstream, which is used to predict the first UV offset based on previously used UV offsets for encoded regions of the grid frame. In yet another example, a syntax indicating a predictor is encoded into the bitstream, which is used to predict the first UV offset based on previously used UV offsets for encoded regions in another grid frame encoded before the grid frame.
[0257] In some examples, the syntax of the predictor used to predict the first UV offset into the bitstream will be encoded into the bitstream, and the prediction residual, which is the difference between the first UV offset and the UV offset of the predictor, will be encoded into the bitstream.
[0258] Then, the process proceeds to (S1499) and terminates.
[0259] The process (1400) can be adjusted appropriately. One (or more) steps in the process (1400) can be modified and / or omitted. Another (or more) steps can be added. Any suitable implementation order can be used.
[0260] Figure 15 A flowchart of an overview process (1500) according to one embodiment of the present disclosure is shown. This process (1500) can be used during the decoding process of a mesh. In various embodiments, the process (1500) is executed by processing circuitry. In some embodiments, the process (1500) is implemented as software instructions, so that the processing circuitry executes the process (1500) when the software instructions are executed. The process begins at (S1501) and proceeds to (S1510).
[0261] At (S1510), multiple maps in 2D are decoded from the bitstream carrying grid frames. The grid frames use surfaces with polygonal representations of objects, and the multiple maps in 2D include at least a decoded geometry map and a decoded attribute map that have been sampled using an adaptive 2D atlas.
[0262] In (S1520), at least a first sampling rate and a second sampling rate are determined based on the syntax of the signals transmitted in the bitstream. During adaptive 2D atlas sampling, the first sampling rate is applied to a first region of the grid frame, and during adaptive 2D atlas sampling, the second sampling rate is applied to a second region of the grid frame. The first sampling rate and the second sampling rate are different.
[0263] In (S1530), based on the multiple maps, at least the first vertex of the grid frame is reconstructed according to the first sampling rate, and the second vertex of the grid frame is reconstructed according to the second sampling rate.
[0264] In some examples, the multiple maps include a decoded occupancy map sampled using an adaptive 2D atlas. To reconstruct at least a first vertex based on a first sampling rate, the initial UV coordinates of an occupant point in a first sampled region of the decoded occupancy map corresponding to a first region of a grid frame are determined, the occupant point corresponding to the first vertex. The recovered UV coordinates of the first vertex are then determined based on the initial UV coordinates and the first sampling rate. In one example, a first UV offset of the first sampled region is determined based on the syntax from the bitstream. The recovered UV coordinates of the first vertex are determined based on the initial UV coordinates, the first sampling rate, and the first UV offset. Furthermore, in one example, the recovered 3D coordinates of the first vertex are determined based on pixels at the initial UV coordinates in the decoded geometry map, and the recovered attribute values of the first vertex are determined based on pixels at the initial UV coordinates in the decoded attribute map.
[0265] In some embodiments, the plurality of maps does not include an occupancy map. To reconstruct at least a first vertex according to a first sampling rate, information indicating a first boundary vertex of a first region is decoded from the bitstream. A first occupancy region corresponding to the first region is inferred from the first boundary vertex in the occupancy map. The UV coordinates of occupancy points in the first occupancy region are obtained, which may correspond to the first vertex. The UV coordinates are converted to sampled UV coordinates at least according to the first sampling rate. The first vertex is reconstructed based on the sampled UV coordinates from the plurality of maps.
[0266] In some examples, in order to reconstruct the first vertex, the recovered 3D coordinates of the first vertex are determined based on the pixels at the sampled UV coordinates in the decoded geometry, and the recovered attribute values of the first vertex are determined based on the pixels at the sampled UV coordinates in the decoded attribute map.
[0267] In some examples, in order to convert UV coordinates to sampled UV coordinates, a first UV offset associated with a first region is decoded from the bitstream, and the UV coordinates are converted to sampled UV coordinates based on a first sampling rate and a first UV offset.
[0268] In some embodiments, at least a first sampling rate and a second sampling rate can be determined using various techniques. In one example, the values of at least the first sampling rate and the second sampling rate are directly decoded from the bitstream. In another example, at least a first index and a second index are decoded from the bitstream, the first index indicating the selection of a first sampling rate from a set of sampling rates, and the second index indicating the selection of a second sampling rate from that set of sampling rates. In yet another example, the first sampling rate is predicted based on a pre-established set of rates. In yet another example, the first sampling rate is predicted based on the previously used sampling rate of a decoded region of a grid frame. In yet another example, the first sampling rate is predicted based on the previously used sampling rate of a decoded region in another grid frame decoded prior to the grid frame.
[0269] In some embodiments, to determine the first sampling rate, a first syntax value indicating whether the first sampling rate is signaled or predicted is decoded from the bitstream. In one example, in response to the first syntax value indicating that the first sampling rate is signaled, the value of the first sampling rate is decoded directly from the bitstream. In another example, in response to the first syntax value indicating that the first sampling rate is signaled, an index is decoded from the bitstream. This index indicates the selection of the first sampling rate from the set of sampling rates.
[0270] In some examples, in response to a first syntax value indicating a predicted first sampling rate, a second syntax is decoded from the bitstream, which indicates a predictor used to predict the first sampling rate. Furthermore, in one example, a prediction residual is determined based on the syntax value decoded from the bitstream; and the first sampling rate is determined based on the predictor and the prediction residual.
[0271] In some examples, the base sample rate is decoded from the bitstream, and at least a first sample rate and a second sample rate can be determined based on the base sample rate. For example, the base sample rate is used as a predictor.
[0272] In some examples, control flags indicating the need to enable adaptive 2D atlas sampling on the grid frame are decoded from the bitstream. Then, for example, multiple regions in the grid frame and the sampling rate for each of these regions are determined based on the syntax in the bitstream.
[0273] In some examples, a first UV offset associated with a first region is determined from the bitstream, and the first vertex of a grid frame is reconstructed based on multiple maps, a first sampling rate, and the first UV offset. In one example, to determine the first UV offset associated with the first region, the value of the first UV offset is decoded directly from the bitstream. In another example, the first UV offset is predicted based on a pre-established set of UV offsets. In yet another example, the first UV offset is predicted based on previously used UV offsets of decoded regions of the grid frame. In yet another example, the first UV offset is predicted based on previously used UV offsets of decoded regions in another grid frame decoded prior to this grid frame.
[0274] In some embodiments, to determine a first UV offset associated with a first region, a first syntax value indicating whether the first UV offset is signaled or predicted is decoded from the bitstream. In one example, in response to a first syntax value indicating that the first UV offset is signaled, the value of the first UV offset is decoded directly from the bitstream. In one example, the sign of the first UV offset is inferred based on a comparison of a first sampling rate with a fundamental sampling rate.
[0275] In another example, in response to a first syntax value indicating a predicted first UV offset, a second syntax is decoded from the bitstream, which indicates a predictor used to predict the first UV offset. Furthermore, a prediction residual is determined based on the syntax value decoded from the bitstream, and the first UV offset is determined based on the predictor and the prediction residual.
[0276] Then, the process proceeds to (S1599) and terminates.
[0277] The process (1500) can be adjusted appropriately. One (or more) steps in the process (1500) can be modified and / or omitted. One (or more) additional steps can be added. Any suitable implementation order can be used.
[0278] The techniques disclosed in this disclosure can be used individually or in any combination in any order. Furthermore, each of the techniques (e.g., methods, embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In some examples, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0279] The above-described technology can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 16 A computer system (1600) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0280] Computer software can be coded using any suitable machine code or computer language, which can be assembled, compiled, linked or similar mechanisms to create code containing instructions that can be executed directly or through interpretation, microcode execution or other means by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0281] The instructions can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0282] Figure 16 The components of the computer system (1600) shown are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement on any one or combination of the components illustrated in the exemplary embodiments of the computer system (1600).
[0283] The computer system (1600) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from one or more human users through, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as sounds, clapping), visual input (such as gestures), and olfactory input (not depicted). The human-computer interface device may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as speech, music, ambient sounds), images (such as scanned images from a still image camera, photographic images), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0284] The input human-machine interface device may include at least one of the following (only one of each is shown): keyboard (1601), mouse (1602), touchpad (1603), touch screen (1610), data glove (not shown), joystick (1605), microphone (1606), scanner (1607), and camera (1608).
[0285] The computer system (1600) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (1610), data gloves (not shown), or joystick (1605), but may also include tactile feedback devices that are not used as input devices), audio output devices (such as speakers (1609), headphones (not depicted), visual output devices (such as screens (1610), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability, some of which are capable of outputting two-dimensional or more than three-dimensional visual output in a manner such as stereoscopic output; virtual reality glasses (not depicted), holographic displays, and smoke canisters (not depicted)), and printers (not depicted).
[0286] The computer system (1600) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1620) having CD / DVD or similar media (1621), thumb drives (1622), removable hard disk drives or solid-state drives (1623), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0287] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0288] The computer system (1600) may also include an interface (1654) to one or more communication networks (1655). The network may be, for example, wireless, wired, or optical. The network may also be a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a vehicle network and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include LANs such as Ethernet, wireless LANs, and cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicle networks and industrial networks including CAN buses, etc. Some networks typically require external network interface adapters to connect to certain general-purpose data ports or peripheral buses (1649) (such as, for example, a USB port on the computer system (1600)); other networks are typically integrated into the core of the computer system (1600) via system buses (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). The computer system (1600) can use any of these networks to communicate with other entities. This communication can be unidirectional, receive-only (e.g., broadcasting TV), transmit-only (e.g., to the CAN bus of some CAN bus device), or bidirectional, such as to other computer systems using local or wide area digital networks. As mentioned above, certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.
[0289] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be connected to the kernel (1640) of the computer system (1600).
[0290] The kernel (1640) may include one or more central processing units (CPU) (1641), graphics processing units (GPUs) (1642), dedicated programmable processing units (1643) in the form of field-programmable gate arrays (FPGAs), hardware accelerators (1644) for certain tasks, graphics adapters (1650), etc. These devices, along with read-only memory (ROM) (1645), random access memory (1646), and internal mass storage (1647) such as internal non-user-accessible hard disk drives, SSDs, etc., may be connected via a system bus (1648). In some computer systems, the system bus (1648) may be accessed as one or more physical connectors to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the kernel's system bus (1648) or connected via a peripheral bus (1649). In one example, a screen (1610) may be connected to a graphics adapter (1650). Peripheral bus architectures include PCI, USB, etc.
[0291] The CPU (1641), GPU (1642), FPGA (1643), and accelerator (1644) can execute certain instructions that, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (1645) or RAM (1646). Temporary data can also be stored in RAM (1646), while permanent data can be stored, for example, in internal mass storage (1647). Fast storage and retrieval of any memory device can be enabled by using a cache memory, which can be closely associated with one or more CPUs (1641), GPUs (1642), mass storage (1647), ROM (1645), RAM (1646), etc.
[0292] Computer-readable media may contain computer code for performing various computer-implemented operations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type well-known and available to those skilled in the art of computer software.
[0293] By way of example and not limitation, a computer system (1600) having an architecture, particularly a kernel (1640), can provide the functionality of executing software embodied in one or more tangible computer-readable media as a processor (including CPU, GPU, FPGA, accelerator, etc.). Such a computer-readable medium can be a medium associated with a user-accessible mass storage as described above, as well as some memory of the kernel (1640) having a non-transitory nature, such as the kernel's internal mass storage (1647) or ROM (1645). Software implementing various embodiments of this disclosure can be stored in such a device and executed by the kernel (1640). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the kernel (1640), particularly the processors therein (including CPU, GPU, FPGA, etc.), to execute a specific process or a specific portion of a specific process described herein, including defining data structures stored in RAM (1646) and modifying these data structures according to the software-defined process. In addition to or as an alternative, the computer system may provide functionality as a result of hard-wired logic or otherwise embodied in circuitry (e.g., an accelerator (1644)) that may replace or operate with software to perform the specific process or a specific portion of the specific process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0294] While this disclosure describes several exemplary embodiments, variations, arrangements, and various alternative equivalents thereof all fall within the scope of this disclosure. Therefore, those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.
Claims
1. A method for mesh decompression, characterized in that, The method includes: Decode multiple maps in two dimensions (2D) from a bitstream carrying grid frames, the grid frames using polygons to represent the surfaces of objects, the multiple maps including at least a decoded geometry map and a decoded attribute map with adaptive 2D atlas sampling applied; At least a first sampling rate and a second sampling rate are determined based on the syntax of signal transmission in the bitstream, the first sampling rate being applied to a first region of the grid frame during the adaptive 2D atlas sampling, and the second sampling rate being applied to a second region of the grid frame during the adaptive 2D atlas sampling, wherein the first sampling rate and the second sampling rate are different; and Based on the plurality of maps, at least a first vertex of the grid frame is reconstructed according to the first sampling rate, and a second vertex of the grid frame is reconstructed according to the second sampling rate; wherein determining at least the first sampling rate and the second sampling rate further includes at least one of the following: Decode at least the values of the first sampling rate and the second sampling rate directly from the bitstream; Decode at least a first index and a second index from the bitstream, wherein the first index indicates the selection of the first sampling rate from the set of sampling rates, and the second index indicates the selection of the second sampling rate from the set of sampling rates; The first sampling rate is predicted based on a pre-established set of rates; The first sampling rate is predicted based on the previously used sampling rate of the decoded region of the grid frame; and The first sampling rate is predicted based on the previously used sampling rate of a decoded region in another grid frame that was decoded before the first grid frame.
2. The method according to claim 1, characterized in that, The plurality of maps includes a decoded occupancy map that has been sampled using the adaptive 2D atlas, and the reconstructing of at least the first vertex of the grid frame based on the first sampling rate includes: Obtain the initial UV coordinates of an occupied point in a first sampled region of the decoded occupied map corresponding to the first region of the grid frame, the occupied point corresponding to the first vertex; and The recovered UV coordinates of the first vertex are determined based on the initial UV coordinates and the first sampling rate.
3. The method according to claim 2, characterized in that, The method further includes: Decode the first UV offset of the first sampled region from the bitstream; and The recovered UV coordinates of the first vertex are determined based on the initial UV coordinates, the first sampling rate, and the first UV offset.
4. The method according to claim 2, characterized in that, The method further includes: The recovered 3D coordinates of the first vertex are determined based on the pixels at the initial UV coordinates in the decoded geometry; and The recovered attribute value of the first vertex is determined based on the pixel at the initial UV coordinates in the decoded attribute map.
5. The method according to claim 1, characterized in that, The plurality of maps do not include an occupancy map, and the reconstructing of at least the first vertex of the grid frame based on the first sampling rate includes: Decode information indicating the first boundary vertex of the first region from the bitstream; Based on the first boundary vertex, infer the first occupied region in the occupied map corresponding to the first region; Obtain the UV coordinates of the occupied point in the first occupied region, wherein the occupied point corresponds to the first vertex; At least according to the first sampling rate, the UV coordinates are converted to sampled UV coordinates; and Based on the multiple maps, the first vertex is reconstructed according to the sampled UV coordinates.
6. The method according to claim 5, characterized in that, The step of reconstructing the first vertex based on the sampled UV coordinates includes: The recovered 3D coordinates of the first vertex are determined based on the pixels at the sampled UV coordinates in the decoded geometry; and The recovered attribute value of the first vertex is determined based on the pixel at the sampled UV coordinates in the decoded attribute map.
7. The method according to claim 5, characterized in that, The step of converting the UV coordinates to sampled UV coordinates at least according to the first sampling rate includes: Decode the first UV offset associated with the first region from the bitstream; and The UV coordinates are converted into the sampled UV coordinates based on the first sampling rate and the first UV offset.
8. The method according to claim 1, characterized in that, Determining at least a first sampling rate includes: Decode the first syntax value, which indicates whether the first sampling rate is sent by signaling or predicted.
9. The method according to claim 8, characterized in that, In response to the first syntax value indicating that the first sampling rate is signaled, the method includes at least one of the following: Decode the value of the first sampling rate directly from the bitstream; and Decode an index from the bitstream, the index indicating the selection of the first sampling rate from the set of sampling rates.
10. The method according to claim 8, characterized in that, In response to the first syntax value indicating the prediction of the first sampling rate, the method includes: Decode a second syntax from the bitstream, the second syntax indicating a predictor for predicting the first sampling rate.
11. The method according to claim 10, characterized in that, The method further includes: The prediction residual is determined based on the syntax values decoded from the bitstream; and The first sampling rate is determined based on the predictor and the prediction residual.
12. The method according to claim 1, characterized in that, Determining at least a first sampling rate and a second sampling rate includes: Decode the basic sampling rate from the bitstream; and At least the first sampling rate and the second sampling rate are determined based on the basic sampling rate.
13. The method according to claim 1, characterized in that, The method further includes: Decode the control flag, which indicates that adaptive 2D atlas sampling is enabled; Determine multiple regions within the grid frame; and Determine the sampling rate for each of the multiple regions.
14. The method according to claim 1, characterized in that, The method further includes: Determine the first UV offset associated with the first region from the bitstream; and Based on the multiple maps, the first vertex of the grid frame is reconstructed according to the first sampling rate and the first UV offset.
15. The method according to claim 14, characterized in that, The determination of the first UV offset associated with the first region further includes at least one of the following: The value of the first UV offset is decoded directly from the bitstream; The first UV offset is predicted based on a pre-established set of UV offsets; The first UV offset is predicted based on the previously used UV offset of the decoded region of the grid frame; as well as The first UV offset is predicted based on the previously used UV offset of a decoded region in another grid frame that was decoded before the grid frame.
16. The method according to claim 14, characterized in that, Determining the first UV offset associated with the first region includes: Decode the first syntax value, which indicates whether the first UV offset is sent by signaling or predicted.
17. The method according to claim 16, characterized in that, In response to the first syntax value indicating that the first UV offset is sent by a signal, the method includes: Decode the value of the first UV offset directly from the bitstream; and The sign of the first UV offset is inferred based on a comparison between the first sampling rate and the basic sampling rate.
18. The method according to claim 16, characterized in that, In response to the first syntax value indicating the prediction of the first UV offset, the method includes: Decode the second syntax from the bitstream, the second syntax indicating a predictor for predicting the first UV offset; The prediction residual is determined based on the syntax values decoded from the bitstream; and The first UV offset is determined based on the predictor and the prediction residual.
19. An apparatus for mesh decompression, characterized in that, The system includes processing circuitry configured to decode multiple maps in two dimensions (2D) from a bitstream carrying grid frames, the grid frames representing the surfaces of objects using polygons, the multiple maps including at least a decoded geometry map and a decoded attribute map with adaptive 2D atlas sampling applied. At least a first sampling rate and a second sampling rate are determined based on the syntax of signal transmission in the bitstream, wherein the first sampling rate is applied to a first region of the grid frame during the adaptive 2D atlas sampling, and the second sampling rate is applied to a second region of the grid frame during the adaptive 2D atlas sampling, wherein the first sampling rate and the second sampling rate are different. as well as Based on the plurality of maps, at least the first vertex of the grid frame is reconstructed according to the first sampling rate, and the second vertex of the grid frame is reconstructed according to the second sampling rate; wherein determining at least the first sampling rate and the second sampling rate further includes at least one of the following: Decode at least the values of the first sampling rate and the second sampling rate directly from the bitstream; Decode at least a first index and a second index from the bitstream, wherein the first index indicates the selection of the first sampling rate from the set of sampling rates, and the second index indicates the selection of the second sampling rate from the set of sampling rates; The first sampling rate is predicted based on a pre-established set of rates; The first sampling rate is predicted based on the previously used sampling rate of the decoded region of the grid frame; and The first sampling rate is predicted based on the previously used sampling rate of a decoded region in another grid frame that was decoded before the first grid frame.
20. The apparatus according to claim 19, characterized in that, The plurality of maps include a decoded occupancy map that has been sampled using the adaptive 2D atlas, and the processing circuitry is further configured to obtain the initial UV coordinates of an occupied point in a first sampled region of the decoded occupancy map corresponding to the first region of the grid frame, the occupied point corresponding to the first vertex; as well as The recovered UV coordinates of the first vertex are determined based on the initial UV coordinates and the first sampling rate.
21. The apparatus according to claim 20, characterized in that, The processing circuit is also configured to decode the first UV offset of the first sampled region from the bitstream; as well as The recovered UV coordinates of the first vertex are determined based on the initial UV coordinates, the first sampling rate, and the first UV offset.
22. A non-transitory computer-readable medium storing instructions, characterized in that, When the instructions are executed by a computer, they implement the method for grid decompression as described in any one of claims 1 to 18.
Citation Information
Patent Citations
Point cloud and mesh compression using image / video codecs
CN110892453A
Method and apparatus for encoding and decoding video signal using adaptive sampling
US20160286219A1